Skip to content

The Patterns That Make a Weak Tinnitus Result Look Strong

Nine ways a tinnitus study can produce an impressive number without establishing anything — each one with a real, cited example from this site, so you can check the pattern against a case rather than take it on trust.

By Tinnitus Clarified TeamUpdated 8 min read

Key takeaways

  • Ask what the treatment was compared with: nothing, a waiting list, or a convincing sham. Only a sham controls for attention and expectation.
  • Look for the difference between groups, not improvement within one. In a 2025 drug trial both drug arms improved on their own, but the arms did not differ (p = .265).
  • Check the size of the change in points, however small the p-value. On the THI, a 2025 study estimated the smallest change people notice at about 8 to 12 points, with a point estimate of 11.
  • Distrust effects far larger than the field produces, intervals that are suspiciously narrow, and 100% success rates from small case series.

A weak tinnitus result can look strong because of how the study was built: a within-group improvement, a waiting-list or no control, a change too small to feel, or an effect too large to be believable.

Most of Tinnitus Clarified is about what the evidence says. This page is about how to read it, because the same handful of patterns appear again and again — in journal abstracts, in press releases, and in the marketing built on top of both.

Every pattern below has a real example, cited, from another Tinnitus Clarified article. The point is not to make you sceptical of everything. It is to give you a specific thing to check, with a case to check it against.

1. The improvement is within the group, not between the groups

The pattern. A study reports that the treated group improved significantly from where it started. It does not report — or buries — whether the treated group improved more than the control group did.

Why it matters. People improve for many reasons that have nothing to do with treatment. They seek help when symptoms are at their worst, symptoms fluctuate, and expecting to improve is itself part of improving. Only a between-group comparison separates the treatment from all of that.

The example. A 2025 placebo-controlled trial of two drug combinations found nortriptyline–topiramate improving significantly against its own baseline (p < .001) and verapamil–paroxetine likewise (p = .004), while placebo did not reach significance (p = .086). It reads like a result. Between the arms there was nothing: ANOVA p = .265. The placebo group had dropped six points too. Drugs for tinnitus sets it out in full.

What to ask. "Compared with what, and by how much more?"

2. The comparison is a waiting list rather than a sham

The pattern. The control group receives nothing and waits.

Why it matters. A waiting list controls for time passing and for nothing else — not the clinician's attention, not the expectation of being treated, not the effort of attending. In an unblinded trial of a symptom only the patient can report, those are the whole problem.

The example. A 2025 multicentre trial found electroacupuncture beating its third arm on loudness at five and ten weeks with tight intervals. The third arm was a waiting list. Across this literature, acupuncture performs well against doing nothing and has never performed well against convincingly faked needling.

What to ask. "Could the control group tell they were the control group?"

3. There is no control arm at all

The pattern. Everyone in the study receives the treatment. The report compares doses, or before against after.

Why it matters. In a condition that improves on its own, an uncontrolled study measures the natural course and the treatment together, and cannot separate them.

The example. A 2026 study followed 195 people given ginkgo after sudden hearing loss. Handicap scores fell from 48.99 to 10.54 over eighteen months, with over 90% meeting the improvement criterion. Both arms received ginkgo — it compared 240 mg against 120 mg. Tinnitus after sudden hearing loss improves substantially anyway, so the study says nothing about the drug. Ginkgo biloba covers it.

What to ask. "Was there anyone who did not get it?"

4. The effect is real and too small to feel

The pattern. An overwhelming p-value attached to a change smaller than the amount a person notices.

Why it matters. A p-value answers whether an effect is likely to be zero. It is silent on whether the effect is big enough to matter, and with enough participants a trivial difference becomes statistically certain.

The example. A 2026 meta-analysis of stellate ganglion block reported a 5.73-point improvement on the Tinnitus Handicap Inventory at p < .00001. A 2025 study estimated the smallest change people notice on that scale at about 8 to 12 points, with a point estimate of 11 — so the result is very unlikely to be chance and still below the threshold for being felt. How tinnitus is measured sets out where those thresholds come from and why there is a range rather than a number.

What to ask. "How many points, and how many points does it take to notice?"

5. The effect size is too large for the field

The pattern. A number far outside what everything else in the area produces.

Why it matters. Tinnitus treatments produce small to moderate effects. When something reports an enormous one, the explanation is usually the trials rather than the treatment — weak blinding, small samples, selective publication.

The example. A 2025 network meta-analysis ranked a supplement first with a standardised mean difference of −3.11, roughly six times what transcranial stimulation achieves. In the same analysis, only 22% of trials were at low risk of bias and Egger's test detected publication bias. An MRI review of hyperacusis reported effect sizes above 5.0, which would be among the largest ever recorded for anything.

What to ask. "How does this compare with everything else in the field?"

6. The confidence interval is impossibly narrow — or crosses zero

The pattern. Either an interval far too tight for a subjective outcome measured across separate trials, or one that quietly includes no effect.

Why it matters. The interval tells you how well the number is pinned down. Tinnitus questionnaires are variable, so honest pooled estimates have wide intervals. Suspiciously tight ones imply the trials agreed almost perfectly, which tinnitus trials do not.

The examples, in both directions. The stellate ganglion block meta-analysis reported −5.73 with an interval of −6.10 to −5.36 — under three-quarters of a point wide, across eleven separate trials. In the other direction, a notched music meta-analysis reported −8.62 with an interval running to −0.23, an upper bound of effectively nothing, at p = 0.044.

What to ask. "What is the range, and does it include zero?"

7. The success rate is 100% and the studies are small

The pattern. A pooled figure of complete success, assembled from case reports and small series.

Why it matters. Nobody writes up the procedure that did not work. A literature of case reports measures what gets published as much as what happens.

The example. A 2026 review of endovascular treatment for sigmoid sinus diverticulum found complete or near-complete resolution in all patients — from 17 articles describing 26 patients. The surgical comparison, resting on four times as many patients, reported 77.6%. The gap is not a measured difference between treatments.

What to ask. "How many patients, across how many papers?"

8. The uncontrolled studies show a bigger effect than the controlled ones

The pattern. A review reports both, and the difference between them is larger than the treatment effect.

Why it matters. This is the clearest demonstration available that uncontrolled designs inflate results, and it is unusually visible when one paper contains both.

The example. A 2026 review of single-session tinnitus counselling found a small-to-medium effect in the one randomised trial and large effects in the pretest–posttest studies — the same intervention, the same outcomes. The review states it plainly: moderate benefits in controlled conditions, larger effects when evaluated uncontrolled. Progressive tinnitus management covers it.

What to ask. "Does this review separate the controlled studies from the rest?"

9. The weakest studies are driving the result

The pattern. A review reports what predicted effect size, and study quality is one of the predictors.

Why it matters. This is the previous pattern, quantified. When risk of bias predicts the size of the effect, the headline figure is inflated by exactly the studies least able to support it.

The example. A 2025 meta-analysis of e-health interventions reported a large effect on tinnitus distress, d = 0.83. It then tested what moderated that figure, and risk of bias was one of the two moderators it identified. Internet-based CBT carries the full picture, which remains positive — just smaller than the headline.

What to ask. "Did the review check whether the weaker studies found bigger effects?"

A shorter version

If you only keep three questions, keep these:

  1. Compared with what? Nothing, a waiting list, or a convincing sham.
  2. Between groups or within one? Only the first tells you anything.
  3. How many points, and is that enough to notice? Roughly 8–12 on the THI.

A claim that survives all three is worth taking seriously. Most of what is marketed for tinnitus does not survive the first.

What this is not

It is not a reason to dismiss everything. Several treatments on this site clear all three questions — CBT most clearly, and the 2026 Nature Reviews Disease Primers volume names counselling and CBT as first-line independently of anything argued here. Scepticism that rejects the good evidence along with the bad is not scepticism, it is just a different way of being wrong.

And it is not a reason to distrust your own experience. These patterns are about what a study establishes for a population. If something helps you and carries no risk, the trial literature is not the authority on your own week. What it is good for is deciding what to spend money and hope on before you know.

If a treatment has no studies at all rather than weak ones, checking that yourself is a different exercise with its own page.

Frequently asked questions

How can I tell whether a tinnitus study shows anything?

Ask three questions before looking at the headline number. What was it compared against — nothing, a waiting list, or a convincing sham? Was the difference measured between the groups, or just within the treated group before and after? And is the size of the change big enough for a person to notice, which on the Tinnitus Handicap Inventory means somewhere around 8 to 12 points. A result can be statistically overwhelming and fail all three.

What is the difference between within-group and between-group significance?

Within-group means the treated group improved compared with how it started. Between-group means the treated group improved more than the control group did. Only the second tells you the treatment did anything, because people improve for many reasons — they sought help at their worst, symptoms fluctuate, and expecting to improve is itself part of improving. A 2025 drug trial found both active arms improving significantly against their own baselines while placebo did not, and no significant difference between the arms at all, with an ANOVA p-value of .265.

Why does a waiting-list control not count as a proper comparison?

Because it controls only for the passage of time. It does not control for the attention of a clinician, the expectation created by starting a treatment, or the effort of turning up — and in a subjective symptom reported by unblinded people, those are exactly the things a sham comparison is designed to strip out. Acupuncture for tinnitus performs well against waiting lists and has never performed well against convincingly faked needling, and both of those statements have been true for over a decade.

Can a statistically significant result still be too small to matter?

Routinely, and it is one of the commonest gaps between a headline and a life. A p-value tells you an effect is unlikely to be zero. It says nothing about whether the effect is big enough to feel. A 2026 meta-analysis of stellate ganglion block reported a p-value below .00001 for a 5.73-point improvement on a scale where the smallest change people notice is estimated at about 8 to 12 points.

Is a very large effect size a good sign?

Usually the opposite, in this field. Tinnitus treatments produce small to moderate effects, so a number far outside that range is more often a fact about the studies than about the treatment. A supplement network meta-analysis reported a standardised mean difference of −3.11, roughly six times what transcranial stimulation achieves, in trials where only 22% were at low risk of bias and publication bias was detected. An MRI review of hyperacusis reported effect sizes above 5.0.

Sources

4 named sources

Show the list
  1. Abouzari, Tawk et al., 2025Clinical trial

    Efficacy of Nortriptyline-Topiramate and Verapamil-Paroxetine in Tinnitus Management: A Randomized Placebo-Controlled Trial, Otolaryngology–Head and Neck Surgery, PubMed (opens in a new tab)
  2. Sattel, Brueggemann et al., 2025Systematic review

    Short- and Long-Term Outcomes of e-Health and Internet-Based Psychological Interventions for Chronic Tinnitus: A Systematic Review and Meta-Analysis, Telemedicine and e-Health, PubMed (opens in a new tab)
  3. Engelke, Basso et al., 2025Journal article

    Estimation of Minimal Clinically Important Difference for Tinnitus Handicap Inventory and Tinnitus Functional Index, Otolaryngology–Head and Neck Surgery, PubMed (opens in a new tab)
  4. Pandey, Knoetze et al., 2026Systematic review

    Efficacy of Single-Session Intervention of Tinnitus Educational Counseling: A Systematic Review, American Journal of Audiology, PubMed (opens in a new tab)

On this page