Glossary category
Statistics and inference
Terms describing how effects are estimated and how uncertainty is expressed.
Terms
22 documents
Table 1. Terms in the statistics and inference category.
| Term | Definition |
|---|---|
| Absolute risk differenceRisk difference | The difference in event probability between arms. Depends on baseline risk and is the measure from which a number needed to treat is derived. |
| Confidence intervalCI | A range constructed so that, over repeated experiments, the stated proportion of such ranges would contain the true value. The Institute reports the interval alongside every point estimate and treats its width as the primary expression of precision. |
| Effect size | The magnitude of a difference, expressed on the scale of the outcome or standardised. A significant result with an effect size below the minimal clinically important difference is a detected difference that does not matter. |
| Fixed-effect model | A pooling model assuming a single true effect common to all contributing studies. Appropriate only where that assumption is defensible. |
| Forest plot | A figure showing each contributing estimate with its interval against a line of no effect, with the pooled estimate where one is presented. The standard figure of the Institute’s synthesis series. |
| Funnel plot | A scatter of effect estimate against study precision, inspected for asymmetry as an indication of small-study effects. Uninformative below approximately ten contributing studies. |
| Hazard ratioHR | The ratio of the instantaneous event rates in two arms, assumed constant over follow-up. A hazard ratio without an absolute risk difference does not convey how much benefit a given population would obtain. |
| Heterogeneity | Variation in effect estimates across studies beyond that expected from chance. Quantified imperfectly by the I-squared statistic and assessed by inspection of the estimates and their intervals. |
| I-squared | The proportion of observed variation across studies attributable to heterogeneity rather than to chance. Does not indicate whether the heterogeneity matters, and takes low values when studies are small regardless of true variation. |
| Minimal clinically important differenceMCID | The smallest change in an outcome that a patient would identify as meaningful. Frequently contested and frequently unavailable, and the Institute names the source of any value it uses. |
| Missing data | Observations not collected. The assumption under which they are handled is rarely testable and is the most common unstated assumption in a trial report. |
| Multiple imputation | A method for handling missing data by generating several completed datasets under a stated model and combining the results. Depends entirely on the plausibility of the model. |
| Multiplicity | The inflation of false-positive rates that follows from testing many hypotheses. Controlled by a pre-specified testing hierarchy, whose position each reported endpoint occupies. |
| Number needed to treatNNT | The reciprocal of the absolute risk difference: the number of people who must be treated for the stated duration for one additional person to benefit. Rises sharply as baseline risk falls. |
| Odds ratioOR | The ratio of the odds of an event in two arms. Approximates the risk ratio only when the event is uncommon, and departs from it materially for common outcomes. |
| Optimal information sizeOIS | The number of participants or events a body of evidence would require to be as informative as an adequately powered single trial. Used by the Institute in assessing imprecision. |
| Point estimate | The single value that best summarises the observed effect. Never reported by the Institute without its interval. |
| Proportional hazards | The assumption that the ratio of hazards between arms is constant over time. Where the assumption fails, a single hazard ratio summarises a time-varying effect and can mislead. |
| Random-effects model | A pooling model assuming that true effects vary across studies and estimating their distribution. Produces wider intervals than a fixed-effect model and gives relatively more weight to small studies. |
| Risk ratioRelative risk | The ratio of event probabilities in two arms. More directly interpretable than an odds ratio for common outcomes. |
| Statistical significanceP value | A property of a result relative to a threshold, not a measure of effect size or of importance. The Institute reports effect sizes alongside any statement of significance as a standing drafting rule. |
| Testing hierarchyHierarchical testing | A pre-specified order in which endpoints are tested, in which testing stops at the first failure. An endpoint below the point of failure is not confirmatory regardless of its nominal significance. |