Independent · non-commercial · publishes on a quarterly cycle|Current cycle 2026 Q3
Compound Evidence InstituteEvidence synthesis · established 2023Graded assessments of compounds, trials, methods and supply
Document set current to 30 July 2026
Glossary category

Statistics and inference

Terms describing how effects are estimated and how uncertainty is expressed.

Terms

22 documents

Table 1. Terms in the statistics and inference category.

TermDefinition
Absolute risk differenceRisk differenceThe difference in event probability between arms. Depends on baseline risk and is the measure from which a number needed to treat is derived.
Confidence intervalCIA range constructed so that, over repeated experiments, the stated proportion of such ranges would contain the true value. The Institute reports the interval alongside every point estimate and treats its width as the primary expression of precision.
Effect sizeThe magnitude of a difference, expressed on the scale of the outcome or standardised. A significant result with an effect size below the minimal clinically important difference is a detected difference that does not matter.
Fixed-effect modelA pooling model assuming a single true effect common to all contributing studies. Appropriate only where that assumption is defensible.
Forest plotA figure showing each contributing estimate with its interval against a line of no effect, with the pooled estimate where one is presented. The standard figure of the Institute’s synthesis series.
Funnel plotA scatter of effect estimate against study precision, inspected for asymmetry as an indication of small-study effects. Uninformative below approximately ten contributing studies.
Hazard ratioHRThe ratio of the instantaneous event rates in two arms, assumed constant over follow-up. A hazard ratio without an absolute risk difference does not convey how much benefit a given population would obtain.
HeterogeneityVariation in effect estimates across studies beyond that expected from chance. Quantified imperfectly by the I-squared statistic and assessed by inspection of the estimates and their intervals.
I-squaredThe proportion of observed variation across studies attributable to heterogeneity rather than to chance. Does not indicate whether the heterogeneity matters, and takes low values when studies are small regardless of true variation.
Minimal clinically important differenceMCIDThe smallest change in an outcome that a patient would identify as meaningful. Frequently contested and frequently unavailable, and the Institute names the source of any value it uses.
Missing dataObservations not collected. The assumption under which they are handled is rarely testable and is the most common unstated assumption in a trial report.
Multiple imputationA method for handling missing data by generating several completed datasets under a stated model and combining the results. Depends entirely on the plausibility of the model.
MultiplicityThe inflation of false-positive rates that follows from testing many hypotheses. Controlled by a pre-specified testing hierarchy, whose position each reported endpoint occupies.
Number needed to treatNNTThe reciprocal of the absolute risk difference: the number of people who must be treated for the stated duration for one additional person to benefit. Rises sharply as baseline risk falls.
Odds ratioORThe ratio of the odds of an event in two arms. Approximates the risk ratio only when the event is uncommon, and departs from it materially for common outcomes.
Optimal information sizeOISThe number of participants or events a body of evidence would require to be as informative as an adequately powered single trial. Used by the Institute in assessing imprecision.
Point estimateThe single value that best summarises the observed effect. Never reported by the Institute without its interval.
Proportional hazardsThe assumption that the ratio of hazards between arms is constant over time. Where the assumption fails, a single hazard ratio summarises a time-varying effect and can mislead.
Random-effects modelA pooling model assuming that true effects vary across studies and estimating their distribution. Produces wider intervals than a fixed-effect model and gives relatively more weight to small studies.
Risk ratioRelative riskThe ratio of event probabilities in two arms. More directly interpretable than an odds ratio for common outcomes.
Statistical significanceP valueA property of a result relative to a threshold, not a measure of effect size or of importance. The Institute reports effect sizes alongside any statement of significance as a standing drafting rule.
Testing hierarchyHierarchical testingA pre-specified order in which endpoints are tested, in which testing stops at the first failure. An endpoint below the point of failure is not confirmatory regardless of its nominal significance.
Nothing published by the Institute is medical advice, a diagnosis, a prescription, a treatment recommendation or a purchasing recommendation. Compounds supplied for research use are not approved for human or veterinary use in any jurisdiction, and a favourable analytical assessment of a supplier is not a statement that any product is safe or effective. The Institute publishes certainty ratings and never recommendations. No telephone number, messaging handle or ordering channel appears anywhere on this site.