How the verdict is built

linked from every drug and comparison page
01 · DIVIDE

By route, not one big list

Oral and injectable are graded as separate divisions, so a best-in-class pill is never ranked against an injection. Not because the weight-loss numbers cannot be compared across routes, they can, and we compare them directly on the oral vs injectable page. It is because a single cross-route grade would smuggle in a verdict on the needle-versus-pill trade-off, and that trade-off is yours to price, not ours.

02 · STANDARDIZE

The same question for everyone

We score placebo-adjusted loss using each trial’s treatment-policy or treatment-regimen estimate, at the highest standard maintenance dose reached through the usual labeled escalation (or the pivotal trial dose while a drug is still investigational), near a 68-week window. These estimates incorporate treatment discontinuation and other changes specified in the trial’s analysis plan. Off-window numbers are shown and flagged, never extrapolated. Zepbound 15 mg is a standard maintenance dose, so it is scored; Wegovy HD has an added clinical gate, so it appears unranked beneath the 2.4 mg semaglutide verdict. Approval can lower a drug’s score when the approved dose is below the trial ceiling.

03 · SPLIT

Efficacy and tolerability apart

Two scores, never blended. Convenience stays a labeled profile, because the needle-versus-pill trade-off is your value judgment, not a fact we bake in.

04 · GRADE THE EVIDENCE

Confidence caps the verdict

Four levels, top to bottom. A regulator clearance is the ceiling: those read Approved. A Phase 3 result, topline or published, earns Certified and a place on the board. Below the line, Projected (only Phase 2) and NR (too early to rate) stay on the watchlist, never mixed into the ranking: a hyped number cannot be crowned until a Phase 3 readout lands.

The two ideas that do the heavy lifting

Placebo-adjusted

Credit the drug, not the diet

Every obesity trial runs a diet-and-lifestyle program in both arms, so the placebo group already loses a few percent. We report the gap above placebo, not the raw number: a drug whose arm lost 15% against a placebo arm that lost 3% really delivered 12. Absolute numbers flatter drugs tested against weaker placebo arms; the gap is the fair comparison.

Two estimands, two questions

Same population, different treatment scenarios

Obesity trials often report two estimates that target the same randomized trial population.

The treatment-policy or treatment-regimen estimand asks what happened after participants were assigned treatment, regardless of whether they later stopped it or had other specified treatment changes. Follow-up measurements after discontinuation are used when available, and missing outcomes are estimated.

The trial-product or efficacy estimand asks a hypothetical question: what would the average effect have been if participants had continued assigned treatment as intended and avoided specified rescue or prohibited weight-management treatment? Participants who stopped are still part of the target population, but their outcomes under that continued-treatment scenario must be estimated.

In obesity trials, the hypothetical estimate is often larger. The estimands answer different questions. We rank on the treatment-policy or treatment-regimen estimate because it incorporates discontinuation and other specified treatment changes. Sponsor terminology and exact methods vary, so we verify each trial’s protocol and statistical analysis plan.

Worked example: survodutide

How a 16.6% headline becomes a Lagging verdict: the estimand and the tolerability cost, one step at a time.

16.6%
The headlineabsolute loss, efficacy estimand
13.4
Subtract placebo16.6% − 3.2%, same estimand
7.6
Switch questions13.0% − 5.4%: treatment-regimen estimate incorporates treatment changes
Heavy
The tolerability costabout 20% stopped for GI (+17.3)
C
Verdict: Lagging, Certified. The evidence is solid Phase 3: we do not doubt the trial. But the treatment-regimen question yields a 7.6-point placebo-adjusted result, and the GI cost is substantial. That makes the 16.6% efficacy-estimand headline a more modest, hard-to-tolerate result for our scoring question. It is why the choice of estimand and the separate tolerability score matter.

← Back to the board