A weak tinnitus result can look strong because of how the study was built: a within-group improvement, a waiting-list or no control, a change too small to feel, or an effect too large to be believable.
Most of Tinnitus Clarified is about what the evidence says. This page is about how to read it, because the same handful of patterns appear again and again — in journal abstracts, in press releases, and in the marketing built on top of both.
Every pattern below has a real example, cited, from another Tinnitus Clarified article. The point is not to make you sceptical of everything. It is to give you a specific thing to check, with a case to check it against.
1. The improvement is within the group, not between the groups
The pattern. A study reports that the treated group improved significantly from where it started. It does not report — or buries — whether the treated group improved more than the control group did.
Why it matters. People improve for many reasons that have nothing to do with treatment. They seek help when symptoms are at their worst, symptoms fluctuate, and expecting to improve is itself part of improving. Only a between-group comparison separates the treatment from all of that.
The example. A 2025 placebo-controlled trial of two drug combinations found nortriptyline–topiramate improving significantly against its own baseline (p < .001) and verapamil–paroxetine likewise (p = .004), while placebo did not reach significance (p = .086). It reads like a result. Between the arms there was nothing: ANOVA p = .265. The placebo group had dropped six points too. Drugs for tinnitus sets it out in full.
What to ask. "Compared with what, and by how much more?"
2. The comparison is a waiting list rather than a sham
The pattern. The control group receives nothing and waits.
Why it matters. A waiting list controls for time passing and for nothing else — not the clinician's attention, not the expectation of being treated, not the effort of attending. In an unblinded trial of a symptom only the patient can report, those are the whole problem.
The example. A 2025 multicentre trial found electroacupuncture beating its third arm on loudness at five and ten weeks with tight intervals. The third arm was a waiting list. Across this literature, acupuncture performs well against doing nothing and has never performed well against convincingly faked needling.
What to ask. "Could the control group tell they were the control group?"
3. There is no control arm at all
The pattern. Everyone in the study receives the treatment. The report compares doses, or before against after.
Why it matters. In a condition that improves on its own, an uncontrolled study measures the natural course and the treatment together, and cannot separate them.
The example. A 2026 study followed 195 people given ginkgo after sudden hearing loss. Handicap scores fell from 48.99 to 10.54 over eighteen months, with over 90% meeting the improvement criterion. Both arms received ginkgo — it compared 240 mg against 120 mg. Tinnitus after sudden hearing loss improves substantially anyway, so the study says nothing about the drug. Ginkgo biloba covers it.
What to ask. "Was there anyone who did not get it?"
4. The effect is real and too small to feel
The pattern. An overwhelming p-value attached to a change smaller than the amount a person notices.
Why it matters. A p-value answers whether an effect is likely to be zero. It is silent on whether the effect is big enough to matter, and with enough participants a trivial difference becomes statistically certain.
The example. A 2026 meta-analysis of stellate ganglion block reported a 5.73-point improvement on the Tinnitus Handicap Inventory at p < .00001. A 2025 study estimated the smallest change people notice on that scale at about 8 to 12 points, with a point estimate of 11 — so the result is very unlikely to be chance and still below the threshold for being felt. How tinnitus is measured sets out where those thresholds come from and why there is a range rather than a number.
What to ask. "How many points, and how many points does it take to notice?"
5. The effect size is too large for the field
The pattern. A number far outside what everything else in the area produces.
Why it matters. Tinnitus treatments produce small to moderate effects. When something reports an enormous one, the explanation is usually the trials rather than the treatment — weak blinding, small samples, selective publication.
The example. A 2025 network meta-analysis ranked a supplement first with a standardised mean difference of −3.11, roughly six times what transcranial stimulation achieves. In the same analysis, only 22% of trials were at low risk of bias and Egger's test detected publication bias. An MRI review of hyperacusis reported effect sizes above 5.0, which would be among the largest ever recorded for anything.
What to ask. "How does this compare with everything else in the field?"
6. The confidence interval is impossibly narrow — or crosses zero
The pattern. Either an interval far too tight for a subjective outcome measured across separate trials, or one that quietly includes no effect.
Why it matters. The interval tells you how well the number is pinned down. Tinnitus questionnaires are variable, so honest pooled estimates have wide intervals. Suspiciously tight ones imply the trials agreed almost perfectly, which tinnitus trials do not.
The examples, in both directions. The stellate ganglion block meta-analysis reported −5.73 with an interval of −6.10 to −5.36 — under three-quarters of a point wide, across eleven separate trials. In the other direction, a notched music meta-analysis reported −8.62 with an interval running to −0.23, an upper bound of effectively nothing, at p = 0.044.
What to ask. "What is the range, and does it include zero?"
7. The success rate is 100% and the studies are small
The pattern. A pooled figure of complete success, assembled from case reports and small series.
Why it matters. Nobody writes up the procedure that did not work. A literature of case reports measures what gets published as much as what happens.
The example. A 2026 review of endovascular treatment for sigmoid sinus diverticulum found complete or near-complete resolution in all patients — from 17 articles describing 26 patients. The surgical comparison, resting on four times as many patients, reported 77.6%. The gap is not a measured difference between treatments.
What to ask. "How many patients, across how many papers?"
8. The uncontrolled studies show a bigger effect than the controlled ones
The pattern. A review reports both, and the difference between them is larger than the treatment effect.
Why it matters. This is the clearest demonstration available that uncontrolled designs inflate results, and it is unusually visible when one paper contains both.
The example. A 2026 review of single-session tinnitus counselling found a small-to-medium effect in the one randomised trial and large effects in the pretest–posttest studies — the same intervention, the same outcomes. The review states it plainly: moderate benefits in controlled conditions, larger effects when evaluated uncontrolled. Progressive tinnitus management covers it.
What to ask. "Does this review separate the controlled studies from the rest?"
9. The weakest studies are driving the result
The pattern. A review reports what predicted effect size, and study quality is one of the predictors.
Why it matters. This is the previous pattern, quantified. When risk of bias predicts the size of the effect, the headline figure is inflated by exactly the studies least able to support it.
The example. A 2025 meta-analysis of e-health interventions reported a large effect on tinnitus distress, d = 0.83. It then tested what moderated that figure, and risk of bias was one of the two moderators it identified. Internet-based CBT carries the full picture, which remains positive — just smaller than the headline.
What to ask. "Did the review check whether the weaker studies found bigger effects?"
A shorter version
If you only keep three questions, keep these:
- Compared with what? Nothing, a waiting list, or a convincing sham.
- Between groups or within one? Only the first tells you anything.
- How many points, and is that enough to notice? Roughly 8–12 on the THI.
A claim that survives all three is worth taking seriously. Most of what is marketed for tinnitus does not survive the first.
What this is not
It is not a reason to dismiss everything. Several treatments on this site clear all three questions — CBT most clearly, and the 2026 Nature Reviews Disease Primers volume names counselling and CBT as first-line independently of anything argued here. Scepticism that rejects the good evidence along with the bad is not scepticism, it is just a different way of being wrong.
And it is not a reason to distrust your own experience. These patterns are about what a study establishes for a population. If something helps you and carries no risk, the trial literature is not the authority on your own week. What it is good for is deciding what to spend money and hope on before you know.
If a treatment has no studies at all rather than weak ones, checking that yourself is a different exercise with its own page.