Skip to content

How Tinnitus Is Measured: What THI and TFI Scores Actually Mean

Tinnitus Clarified Editorial Team5 min readUpdated September 5, 2026

Tinnitus has no blood test, no scan finding, and no meter reading. Nobody can measure it from outside your head. So when a trial reports that a treatment worked, what it is reporting is a change in what people wrote on a questionnaire — and knowing which questionnaire, and what a change on it means, is most of what you need to read those trials properly.

The two that matter

The Tinnitus Handicap Inventory (THI), published in 1996, was built to be brief, easy to administer and interpret, broad in scope, and psychometrically robust. Its development ran a 45-item version past 84 patients, then a 25-item version past 66 more, checking it against an existing handicap questionnaire and against measures of depression, somatic perception, and rated annoyance, sleep disruption and concentration.

The Tinnitus Functional Index (TFI), published in 2012, was built for a different reason: to measure change. Its authors state the problem directly — effective treatments were urgently needed and evaluating them was hampered by the lack of a measure validated both for assessing severity at intake and for detecting treatment-related change.

The way they built it is worth knowing, because it explains why the TFI covers what it covers. A panel of 17 expert judges went through 175 items drawn from nine existing tinnitus questionnaires, identified 13 separate domains of tinnitus distress, and selected the 70 items most likely to respond to treatment. That was cut to 43 for the first prototype, tested at five clinics on 326 patients, and refined into the final instrument.

What the numbers behind the TFI look like

  • Eight factors. Every version of the instrument produced the same eight underlying dimensions of tinnitus severity and impact.
  • Internal consistency: Cronbach's alpha 0.97. Very high — the items measure a coherent thing.
  • Test–retest reliability: 0.78. Good, not perfect; the same person on two days will not give the identical score.
  • Agreement with the THI: r = 0.86. The two instruments largely agree about who is worse off.
  • Separation from depression: r = 0.56 against the Beck Depression Inventory–Primary Care. High enough to show the constructs are related, low enough to show it is not simply measuring depression.

The number to remember: 13

The TFI's developers evaluated how much a score has to move before the movement means something, and settled on a 13-point reduction as a preliminary criterion for meaningful improvement. They wrote "preliminary" and it should be quoted that way.

That single figure changes how a lot of tinnitus research reads. A trial can report a statistically significant improvement of five or six TFI points — genuinely unlikely to be chance — that sits well under the threshold its own instrument uses for clinical meaning. Statistical significance and clinical meaning are different questions, and in a field with small effects the gap between them is where most of the overclaiming happens.

When you next read that something "significantly improved tinnitus", the useful question is: by how many points, against a threshold of thirteen?

What these instruments do not measure

They do not measure loudness. They do not measure pitch. They do not measure whether the sound is still there.

They measure sleep, concentration, mood, emotional distress, sense of control, and how much the tinnitus intrudes on ordinary life. That is a deliberate choice and the right one, because those are the things that make tinnitus a problem. But it has a consequence people are rarely told:

A treatment can produce a large, real improvement on these scales without changing the sound at all.

That is exactly what cognitive behavioural therapy does, and it is why CBT has the strongest evidence base of anything on the treatment comparison while making no claim to reduce volume. It is not a loophole. It is the outcome the instruments were designed to capture, and the outcome that most changes how people live. It is only misleading when a product implies the other thing.

The same logic runs the other way. If you feel your tinnitus is no better but your life is, the questionnaires would agree with you, and would call that a good result.

Reading a tinnitus trial with this in hand

  1. Which instrument? TFI, THI, or a visual analogue scale. The TFI is the most responsive of the three, by its own validation.
  2. How big was the change, in points? Not the p-value. The points.
  3. Against what threshold? Thirteen for the TFI. The THI has no single agreed equivalent, which is part of why the TFI was built.
  4. Compared with what? A change from baseline in one group is not a treatment effect. Tinnitus distress falls on its own, which is what a control arm is for.
  5. Measured when? The TFI validation looked at 3 and 6 months. Effects at two weeks are a different claim.

How this site rates evidence sets out the rest of the standard, and finding and reading trials covers registry entries. This page is the piece in the middle: what the outcome actually is, in the studies you are reading about.

About the self-check on this site

The impact self-check here uses its own original questions and scoring, and says so on its own page. It is not the THI, the TFI, or a version of either.

That distinction is deliberate. Reproducing a validated instrument outside the context it was validated in produces a number that looks official and means less than it appears to — and the instruments are the property of the people who built them. What the self-check is for is giving you concrete language to bring to an appointment, which is a smaller and more honest job.

Sources

  1. Meikle et al., 2012 — The Tinnitus Functional Index: development of a new clinical measure for chronic, intrusive tinnitus, Ear and Hearing, PubMed
  2. Newman, Jacobson & Spitzer, 1996 — Development of the Tinnitus Handicap Inventory, Archives of Otolaryngology–Head & Neck Surgery, PubMed

Frequently asked questions

How do researchers measure tinnitus if there is no test for it?+

With questionnaires that ask what the tinnitus does to you, not what it sounds like. The two standard ones are the Tinnitus Handicap Inventory, published in 1996, and the Tinnitus Functional Index, published in 2012. Almost every trial you will read about used one of them as its outcome, which means the result you are being told about is a change in a self-reported score.

What does a change in TFI score have to mean to matter?+

Thirteen points. The team that developed the Tinnitus Functional Index set a 13-point reduction as their preliminary criterion for a meaningful improvement, and they called it preliminary in the paper. It is the single most useful number for reading a tinnitus trial: a study reporting a statistically significant drop of six points has found something real and smaller than the threshold its own instrument uses.

Do these questionnaires measure how loud my tinnitus is?+

No, and this is the most common misreading of them. They measure distress, intrusion and functional impact — sleep, concentration, mood, sense of control. Someone whose tinnitus is unchanged in volume but who has stopped fearing it will score dramatically better. That is a real improvement and worth having; it is not the same claim as the sound getting quieter.

Is the Tinnitus Functional Index better than the Tinnitus Handicap Inventory?+

It was built to be more responsive to change, and its development paper reports that it succeeded: effect sizes for detecting improvement were generally larger than those for the THI or a visual analogue scale. The two agree closely on severity, correlating at r = 0.86. The THI remains widely used and is shorter to complete, so which one a study picked usually reflects when it was designed.

Can I score myself?+

You can use the instruments if you find them, but a number without a clinician to interpret it is of limited use, and this site does not reproduce them. The self-check here uses its own original questions and says so on the page, because presenting a home-made questionnaire as a validated instrument would misrepresent both.