Methodology & standards

Overall structure

The chart below traces every input that goes into an evidence grade. Every box is a clickable link to the section that explains it — start with the badge at the bottom and work backwards to see how it was derived, or click any input box to read what it contributes. The text contents list is below.

Locked PICOPopulationInterventionComparatorOutcomeStudy selectionIdentification → Screening → Eligibility → IncludedEligibility criteria are read from the locked PICOPer-study RoBRoB2 — 5 domainsper studyworst-domain overallPer-study statsn, mean, SD (or events), scale, dataTypeHedges' g and variance per studyIndirectnessPICO match(P / I / C / O)manual judgementRisk of biasweight-based:> 30% weight athigh RoB → −1(Cochrane §14.2)InconsistencyI², τ², Cochran Q,prediction intervalI² > 50 % + PIcrosses null → −1Imprecision95 % CI, p-value,OISCI crosses null → −1(p-value enters here)Publication biasEgger's p, kk ≥ 10 + Eggerp < 0.10 → −1(else 'unknown')Effect sizepooled Hedges' g,95 % CI, kcompared to scaleMCID (if known)Certainty letter (GRADE)HIGH → MODERATE → LOW → VERY LOW = A · B · C · Dauto-computed; reviewer can overrideMagnitude number1 large2 moderate3 smallnoneEvidence grade — letter + numbere.g. B2 = 'moderate certainty, moderate effect'

Contents

Locked PICO

Fix the question, PICO and eligibility before searching, then lock it. Everything downstream inherits this scope.

PICO is the standard 4-part framework for defining a clinical question:

Population — who we're studying (e.g. “adults with stress or anxiety”).
Intervention — what's being given (e.g. “oral ashwagandha root extract, any standardised dose”).
Comparator — vs what (e.g. “placebo”).
Outcome — measured how (e.g. “perceived stress on PSS or a validated anxiety scale”).

“Locked” means we commit to these definitions before running searches or screening, and they can't be silently changed afterwards. The PICO + inclusion / exclusion criteria are stored as JSON on the MA row with a lockedAt timestamp; the build workflow won't move past Stage 1 until you click “lock protocol.”

This is the same idea PROSPERO registration enforces in academic practice — a timestamped commitment to the question and methods before looking at data. It prevents three classic failure modes:

Cherry-picking studies after the fact — narrowing the population once you see which trials helped your hypothesis.
Post-hoc outcome switching — “the trials don't all use PSS, let's switch our outcome to HAM-A” (a HARKing risk).
Scope creep — quietly adding eligible designs or comparators because the original ones gave too few trials.

Downstream, the locked PICO drives the Indirectness GRADE judgement (do the included studies actually match what we asked for?), the screening eligibility criteria, and the PRISMA item 24a (Registration) compliance status.

Study selection

Search every database, screen candidates against eligibility with reasons, and collect the full text of the included set (PRISMA flow).

Selection is the first place a synthesis can go wrong. We follow PRISMA 2020 — the international reporting standard used by Cochrane and the major medical journals — so the path from “everything the search returned” to “studies we pooled” is fully visible. Each record carries the decision (include / exclude), the reason for the decision, and a verbatim abstract excerpt under fair dealing that supports it.

The funnel has four PRISMA stages, with counts and reasons at each:

Identification — records returned by the registered searches, with duplicates removed.
Screening — title-and-abstract triage against the locked PICO.
Eligibility — full-text assessment of records that passed screening.
Included — the studies that survive into the pooled analysis. They become the input population for the per-study RoB and per-study statistics downstream.

Selection sits between the locked PICO (which defines what we're looking for) and the per-study inputs (the extractions and judgements we make on each included paper). On a meta-analysis page the full record-by-record funnel and PRISMA-style flow diagram appear in the study selection section.

Per-study RoB

For each RCT, assess the 5 RoB2 domains from objective trial facts (attrition numbers, blinding, allocation concealment, pre-registration…) — each signalling question answered with a snippet-backed fact; the domain rating auto-derives from the answers. A RoB observation is recorded as the relevant domain’s Fact (with its snippet) in this table — its native format — never as a standalone note; if it changes an answer, re-derive the rating. Record any author-stated limitations separately; they are displayed but not scored.

For each randomised trial we apply RoB2 (Sterne et al., BMJ 2019), the current Cochrane risk-of-bias tool. Each study is assessed across five domains and rated low / some concerns /high on each. The overall rating is the worst-domain rating (Cochrane's worst-domain rule).

Facts, not author self-criticism

A core principle of our assessment: ratings are derived from objective trial facts, not from the authors' own statements about limitations. An author who writes “a limitation is that blinding was not possible” and one who says nothing about blinding are rated identically if the underlying trial had the same design. A self-critical author is not penalised for their candour; an over-confident one gets no credit for staying silent. What the trial did — attrition numbers, allocation concealment procedures, pre-registration status, blinding arrangements — is the input. What the authors said about what they did is not.

Author-stated limitations are separately captured in a limitations field and displayed alongside the RoB summary on the MA page. They inform readers but are never fed into the domain ratings.

The five RoB2 domains

Each domain is assessed by answering its Cochrane signalling questions (Yes / Probably yes / Probably no / No / No information). Each answer is backed by an extracted fact — a concrete, verifiable detail — linked to a verbatim snippet from the paper. The per-domain rating (low / some concerns / high) is then auto-derived by the RoB2 decision rules; it is not typed by hand.

D1 · Randomisation process — was the allocation sequence random? Was it concealed until assignment? Did baseline differences suggest a problem? Fact sources: methods section, CONSORT flow diagram, supplementary protocol.
D2 · Deviations from intended interventions — were participants and carers/personnel blinded? Were there trial-context deviations from the intended intervention? Was an appropriate intention-to-treat analysis used? Fact sources: blinding method description, protocol deviations, analysis population.
D3 · Missing outcome data — were outcome data available for all or nearly all randomised participants? If not, is there evidence the result was unaffected? Could dropout have depended on the true outcome value? Fact sources: attrition numbers per arm, CONSORT flow, imputation method.
D4 · Measurement of the outcome — was the outcome measurement method appropriate? Could it have differed between groups? Were outcome assessors blind, or was the outcome objective (e.g. automated device)? Fact sources: instrument description, assessor-blinding statement, outcome-acquisition method.
D5 · Selection of the reported result — was the result analysed in line with a pre-specified plan, protocol, or registration? Were multiple measurements or analyses available from which the reported one was selected? Fact sources: trial registration (ClinicalTrials.gov, CTRI, etc.), pre-specified statistical analysis plan, correspondence between registered and reported outcomes.

Every per-domain judgement is backed by a verbatim quote (≤50-word fair-dealing extract) stored as a Snippet row, so the rationale is auditable. The rendered MA page shows each domain rating with an inline citation link to that snippet. Internal storage uses the Cochrane terms (low / some_concerns /high); we render them in plain language (low / minor concerns / major concerns) for non-specialist readers.

This card is an input — the per-study judgements are then aggregated into the Risk of bias synthesis card. Studies we couldn't obtain are marked not assessable rather than assumed clean.

Per-study stats

Extract n, mean, SD (or events) and the scale per included study; the per-study effect (Hedges’ g) is computed from this. Any extraction caveat (which arm/subgroup was pooled, combination products, implausibly tight SDs) is recorded as a note under that study’s row here — in the section’s native format — never in a standalone notes box.

Each included paper is read end-to-end and the primary-outcome data are extracted into a structured row on the Study:

Scale — the measurement instrument (e.g. PSS-10, HAM-A, DASS-21 stress subscale).
Data type — “endpoint” or “change from baseline” (read the table heading carefully — many tables labelled “change” actually report endpoint values).
Timepoint — e.g. day 60, week 8.
n, mean, SD (or event counts) for the ashwagandha arm and the placebo arm.
Sanity flags — recorded automatically when something looks wrong: SEM_not_SD (SE × √n conversion applied), baseline_only, combination_product, imputed_sd, crossover_design, endpoint_only_n_unclear.
Verbatim quote + table reference — fair-dealing extract that supports the numbers, plus the source location (e.g. “Table 2”).

From these we compute Hedges' g (small-sample-corrected SMD) and its variance per study, using the small-sample correction factor J = 1 − 3 / (4(n₁+n₂) − 9). When the paper reports SE rather than SD we convert via SD = SE × √n and flag the conversion.

This card is an input. The per-study g values flow downstream into Inconsistency (do they agree?), Imprecision (how precise is the pool?), Publication bias (funnel asymmetry?), and Effect size (the pooled estimate itself).

Indirectness

Judge how directly the population, intervention, comparator and outcome match the question. Human assessment, preserved across recompute.

Indirectness asks: do the included studies actually match the question we said we were answering? That comparison only makes sense against a fixed reference — the locked PICO. Each study is checked against the four PICO dimensions:

Population mismatch — e.g. a trial in athletes when our PICO specified stressed adults.
Intervention mismatch — e.g. a multi-herb combination product when our PICO specified pure ashwagandha.
Comparator mismatch — e.g. active comparator when our PICO specified placebo.
Outcome mismatch — e.g. salivary cortisol when our PICO specified a validated psychometric scale.

This is the GRADE domain that is hardest to automate — there is no statistic that captures “does the evidence match the question.” So we never let the analysis decide it. Instead the analysis only flagsit: it marks indirectness pending (an outstanding review item, surfaced in the build workflow's Analyse stage) and notes any trigger it can detect — e.g. a combination product in the pool. A reviewer then resolves it against the locked PICO, recording a rating and a note. That judgement is stored separately (manualDomains) so it survives every recompute, and a serious call counts toward the certainty letter. Until resolved, the SR page shows a pending review state rather than a fabricated default.

Risk of bias (synthesis)

Downgrade for study limitations weighted by each study’s share of the pooled information.

The synthesis card aggregates the per-study RoBjudgements into a single GRADE domain rating. The threshold is the most subtle GRADE judgement and mis-calibrating it is a common cause of overly pessimistic certainty ratings.

Cochrane Handbook §14.2 wording: downgrade only when “most information is from results NOT at low risk of bias.”The unit of analysis is the body of evidence contributing to the pool, judged by weight not study count. We now read this across all RoB tiers rather than the high tier alone. Each study contributes its pooled inverse-variance weight times a bias fraction, and we take the weighted mean — the bias share:

low → bias fraction 0 (clean evidence).
some concerns / moderate → bias fraction 0.5.
high / serious / critical → bias fraction 1.0.

Downgrade rule: serious (−1 to the certainty letter) when the weighted-mean bias share exceeds 50 %. A pool that is entirely some-concerns sits at exactly 0.50 and is not downgraded (the borderline is treated as discharged). The distribution is still reported (“1 of 14 studies at high RoB contributing 5 % of pooled weight”) so it stays visible even where no downgrade applies.

The weighted-across-tiers reading replaces the old high-RoB-only > 30 % rule, which missed pools with no clean evidence — e.g. 72 % some-concerns + 28 % high carries a bias share of 0.72, well over the threshold, yet contributes 0 % high-RoB weight under the old high-only cutoff. It still avoids the opposite pitfall — a single high-RoB study with low weight does not on its own tip a low-risk pool over 50 % — which matches Cochrane practice and published ashwagandha reviews (Akhgarjand 2022, low certainty). Count-based rules that downgrade for any one high-RoB study remain something we deliberately avoid.

Low-RoB sensitivity re-pool. As a GRADE robustness check, when ≥ 2 studies are at low risk of biasthe engine re-pools just those low-RoB studies — same Paule–Mandel random-effects / HKSJ machinery as the full pool. If the low-RoB-only estimate agrees with the full pool (same direction, overlapping CIs) the risk-of-bias downgrade is dischargeable — the bias didn't move the answer; if it diverges, the bias mattered. The re-pool is drawn as a labelled row in the forest plot beside the full-pool diamond.

Inconsistency

Downgrade for unexplained heterogeneity (I² and a prediction interval that crosses the null).

Inconsistency asks how much the studies disagree once we've accounted for chance. The four signals:

— percentage of variability across studies that is due to true differences (not chance). 0 % = none, > 75 % = substantial per the Cochrane Handbook §10 rough guides.
τ² — variance of the true effects between studies, now estimated by Paule–Mandel. The raw quantity I² is derived from (I² and Q are unchanged — both are fixed-effect based).
Cochran's Q (with p-value) — heterogeneity test; underpowered when k is small, so non-significant Q does not imply homogeneity.
Prediction interval — the range a future trial's true effect would plausibly fall in, built on the Paule–Mandel τ². Reported whenever k ≥ 3. Far more honest than the CI when heterogeneity is high.

Downgrade rule: serious (−1) only when I² > 50 % and the prediction interval crosses the null. High I² whose prediction interval still excludes the null is heterogeneity that does not change the conclusion — we are still confident of the effect's direction — so it draws no downgrade. The prediction interval is undefined for k < 3; we treat that case as crossing (the conservative reading), so a small, heterogeneous pool is not let off on a missing interval.

What the textbook says vs. what we do. Cochrane (Handbook §10.10) treats inconsistency as a holistic judgement — weighing how far the point estimates spread, whether the confidence intervals overlap, I², and the χ² test — with deliberately no fixed I² cut-off. Crucially, it says heterogeneity that is explained by a moderator (dose, duration, population) should be handled by a subgroup analysis or meta-regression, not a downgrade. We make that judgement reproducible by operationalising it as the two conditions above (I² > 50 % and the prediction interval crosses the null), capped at one level. A third condition — whether the heterogeneity is explained by a moderator — is on our roadmap (dose-response / subgroup analysis); until it lands it is treated as not assessed and we take the conservative reading.

A consistent direction with wide magnitude variation (e.g. ash effects from g = −0.09 to g = −4.0, all favouring treatment) is the textbook case for −1 once the prediction interval crosses the null: we're confident there's an effect, but we don't know its size.

Imprecision

Downgrade if the CI crosses the null or the sample falls short of the optimal information size.

Imprecision asks how precise the pooled estimate is. The inputs are the 95 % confidence interval, implicitly the p-value, and the optimal information size (OIS) — does the evidence base have enough participants to detect a clinically meaningful effect?

The pooled CI now uses Hartung–Knapp–Sidik–Jonkman (HKSJ). Rather than a normal-z interval around the DerSimonian–Laird estimate, we build the interval on a tk−1 quantile with a variance-inflation factor truncated at 1 (so the HKSJ interval is never narrower than the classic one). This is the modern default for few-study meta-analyses, where DerSimonian–Laird + normal-z understates the CI and overstates precision. It is gated to k ≥ 3: at k = 2 the t1 = 12.71 quantile is too extreme, so the engine falls back to DL + normal-z.

Downgrade rule: serious when the 95 % CI crosses the null, or when the CI excludes the null but the pooled sample falls below the optimal information size. The OIS check is now automated (it used to be applied by hand). For a continuous outcome the OIS is the total N needed to detect an effect δ (default 0.2 SMD) at α = 0.05 with 80 % power — 31.36 / δ² (≈ 784 participants for δ = 0.2). Ratio-outcome OIS is deferred, so for binary/ratio outcomes the domain currently downgrades on the crosses-null condition only.

This is where the p-value enters the badge

The p-value isn't a direct input to either part of the evidence grade — it's implicit in this domain:

p < 0.05 ⇔ the 95 % HKSJ CI excludes the null ⇒ imprecision not seriousprovided the pooled sample also clears the OIS; a CI that excludes the null on too few participants still downgrades.
p ≥ 0.05 ⇔ the CI crosses the null ⇒ imprecision serious, −1 to the certainty letter; AND the Magnitude number is suppressed because the effect size is then indeterminate.

We deliberately don't render the raw p-value on the badge because two estimates with p = 0.04 and p = 0.001 can have very different certainty letters depending on heterogeneity and publication bias — the letter is the more honest summary.

Publication bias

Downgrade for funnel asymmetry / small-study effects (Egger, Begg) when there are enough studies to assess it.

Publication bias asks whether the published literature systematically over-states the true effect. The classic mechanism: small null trials never get written up, while small positive trials do — so the published set is asymmetric and the pooled estimate is inflated.

Funnel plot inspection — visual asymmetry of effect-size vs precision; meaningful only with k ≥ 10 (Cochrane minimum).
Egger's regression test — formal asymmetry test; the standard pair with the funnel plot, and the test that drives the downgrade.
Begg's rank-correlation test — now computed and reported alongside Egger's at k ≥ 10 as a robustness check (it is more robust under high heterogeneity); it does not itself trigger the downgrade.
Trim-and-fill — estimates the “missing” studies and an adjusted pool; reported only when asymmetry is detected.

Downgrade rule: serious when k ≥ 10 and Egger's p < 0.10. k < 10 is reported as unknown rather than downgraded — too few studies to detect asymmetry reliably. Begg's p is shown beside Egger's (k ≥ 10) so the two agree-or-disagree at a glance, but Egger's remains the test of record. GRADE caps this domain at −1.

Dose–response (biological gradient)

A genuine dose-response gradient — larger doses producing larger effects — is one of Bradford Hill's criteria for causality and one of GRADE's three upgrade factors for observational evidence. We test for it directly: each trial's dose is parsed to a comparable mg/day figure (handling multi-arm regimens and twice-daily dosing) and the per-study effect is regressed on dose with a precision (random-effects) weighting.

Slope — effect change per +100 mg/day. A slope toward more benefit at higher doses is the biological gradient.
p-value & R² — whether the trend is more than noise, and how much of the between-study spread dose explains.
Needs ≥ 4 trials with a parseable dose; otherwise the domain is left unassessed.

We use it two ways. For observational pools a beneficial gradient rates the certainty up one level. For randomised pools (already starting HIGH) it is contextual — but it also answers whether dose explains the heterogeneity: if the trend runs the wrong way (an inverse gradient, where the smallest-dose trials report the largest effects), that is biologically implausible, flags those trials as outliers, and means the inconsistency is not explained by dose — so the downgrade stands.

Effect size

Pool the per-study effects (random-effects Hedges’ g, Paule–Mandel τ², HKSJ CIs) into the summary estimate, I² and prediction interval. This is the hub every GRADE domain reads from.

This card is the pooled estimate itself — the number you'd quote if asked “by how much does this intervention work?” — distinct from the downstream Magnitude number, which is the 1 / 2 / 3 label we derive from this pool.

Pooled Hedges' g — Paule–Mandel random-effects (DerSimonian–Laird's successor) inverse-variance-weighted mean of the per-study g values. The between-study variance τ² is estimated iteratively (Paule–Mandel is distribution-free and does not underestimate τ² for few studies the way DerSimonian–Laird does); the random-effects weights and prediction interval inherit this τ².
95 % CI — uncertainty in the pooled estimate, computed by HKSJ (tk−1, variance-inflation truncated at 1) for k ≥ 3, falling back to DL + normal-z at k = 2.
k — number of studies contributing to the pool; N — total participants.
MCID comparison — when a minimum clinically important difference is known for the scale (e.g. ~5 points on PSS-10, ~3 on HAM-A), the pooled effect is interpreted against it.

The effect-size card is fed by per-study statisticsand itself feeds Magnitude. See the Statistical methods and Analytical choices sections below for the full menu of pooling options.

Certainty letter

The certainty letter is the GRADE rating — the framework used by Cochrane, the WHO and most clinical guideline bodies. GRADE is design-stratified: a body of randomised trials starts at HIGH (= A) and can only be downgraded — one level per serious domain — because HIGH is the ceiling. A body of observational (non-randomised) evidence starts at LOW and can instead be rated up; see “Upgrades vs downgrades” below. When an analysis contains both, we pool and grade each stratum separately and present them side by side — so each design carries its own letter.

A — HIGH: we are very confident that the true effect lies close to our pooled estimate.
B — MODERATE: we are moderately confident; the true effect is likely to lie close to the pool but there is a possibility it is substantially different.
C — LOW: our confidence in the effect estimate is limited; the true effect may be substantially different.
D — VERY LOW: we have very little confidence in the effect estimate; the true effect is likely to be substantially different.

Separate, don't merge. When randomised and non-randomised studies both bear on a question, we pool them in separate strata — a primary RCT pool and an observational (NRSI) pool — and compare them, rather than blending them into one number. The reason is mechanical: a meta-analysis weights each study by its precision, which scales with sample size, not with freedom from bias. Observational datasets are often enormous, so in a merged pool they dominate — and their confounding is systematic error that does not shrink with n. Merging therefore narrows the confidence interval around a possibly-biased estimate, manufacturing false precision. Keeping the strata apart lets each be graded on its own terms (RCTs start HIGH, observational LOW) and turns the comparison into a finding: when the two agree, that is corroboration across independent designs; when they diverge, that is a signature of confounding, and the RCT pool is the more trustworthy estimate. Case series and single-arm before-after studies have no concurrent comparator, so they yield no comparative effect at all — they are catalogued (and, where useful, summarised as a pooled proportion), never placed in a comparative forest plot. Lower-tier evidence still earns its place — harms, rare events, questions where RCTs are impossible — as a labelled, lower-certainty strand.

Upgrades vs downgrades. The five domains above only downgrade. GRADE also defines three upgradefactors — a large effect, a dose-response gradient, and “plausible confounding would only shrink the effect” — but these apply to observational evidence (which starts at LOW); they exist to rescue uncontrolled evidence that turns out to be compelling. They are not a counterweight to downgrades: rating up is only considered when there is no serious reason to rate down, so a −1 for risk of bias is not cancelled by a +1 for dose-response. It is precedence, not arithmetic netting. Upgrades apply only to the observational stratum (which starts at LOW); the RCT stratum, starting at HIGH, is downgrade-only. We auto-apply the large-effect upgrade (a pooled RR/OR ≤0.5 or ≥2 with a CI excluding the null rates up one level; ≤0.2 or ≥5 rates up two), and surface dose-response and opposing-confounding for manual judgement since they need data we do not yet capture. Risk of bias is assessed with RoB2 for RCTs and ROBINS-I for non-randomised studies (its “serious”/“critical” ratings carry the downgrade).

The auto-algorithm computes the rating from the five GRADE domain cards above (see Risk of bias, Inconsistency, Indirectness, Imprecision, Publication bias) using Cochrane-calibrated thresholds documented in “Frameworks we follow.” A reviewer can override the letter at the Review stage; the stored letter and any override carry a rationale and timestamp.

Consilience continuous certainty. Alongside the canonical discrete GRADE letter — never replacing it — we run a continuous certainty engine. Each of the five domains emits a fractional downgrade in [0, 2] rather than a 0/−1/−2 step: risk of bias ramps up from the 50 % bias-share zero-point, inconsistency from I², imprecision from crosses-null plus OIS shortfall, publication bias from Egger's p, and indirectness from the manual judgement. The five partials are summed and rounded once onto the design-anchored start (RCTs HIGH), floored at very low. This addresses GRADE's “death by a thousand cuts”: several domains that are each borderline-but-not-serious can together cost a level, which discrete GRADE — rounding each domain independently — silently rounds away. The MA page shows this as a “Consilience continuous certainty” panel with the per-domain partials, beside the canonical letter; where the two diverge that gap is surfaced, not hidden.

Magnitude number

The magnitude number is our addition on top of GRADE: it answers “how big is the effect?” on a 1 / 2 / 3 scale, derived from the pooled effect-size card.

1 — large: SMD ≥ 0.8 (Cohen); RR ≤ 0.5 or ≥ 2.0; or MD ≥ ~1× MCID on a validated scale.
2 — moderate: SMD 0.5–0.8; RR 0.5–0.67 or 1.5–2.0; or MD ~½–1× MCID.
3 — small: SMD 0.2–0.5; RR 0.67–0.83 or 1.2–1.5; or MD < ½× MCID.
(blank): the 95 % CI crosses the null — the magnitude is indeterminate and we deliberately suppress the number rather than show a misleading classification.

Showing magnitude separately from certainty matters because “how sure?” and “how big?” are independent questions. A large effect can be uncertain; a small effect can be nailed down. GRADE alone conflates these by reporting only the letter.

Evidence grade — letter + number

Sum the five domain downgrades into the certainty letter, map the pooled effect to a 1–3 magnitude number, and combine them into the evidence grade.

The evidence grade is the final two-part badge that combines the certainty letter and the magnitude number into a single token.

B2 = moderate certainty in a moderate effect.
B1 = a large effect we are moderately certain about.
A3 = a small effect we are highly certain about (probably more reliable than B1 for a clinical decision).
A alone (no number) = high certainty but the effect itself is indeterminate (CI crosses null).

The badge is auto-computed, but a reviewer can override either the letter or the number with a recorded rationale at the Review stage of the build workflow. The override is timestamped and visible in the History.

How the frameworks fit together

Evidence synthesis borrows from several frameworks that are easy to confuse because they overlap. They are not competitors — each governs a different stage of the same pipeline. Here is the order in which they act and what each one decides.

  1. PRISMAhow the studies are found and reported. A reporting standard for the search → screen → include funnel. It governs study selection, not the statistics. Answers: “is the search transparent and reproducible?”
  2. Study design classwhich pool a study enters. RCT vs non-randomised (cohort, case-control) vs uncontrolled (case series). This routes each study to the right stratum and the right risk-of-bias tool. Design is a starting prior for trustworthiness, not a verdict.
  3. RoB2 / ROBINS-Ihow much to trust each individual study. RoB2 (5 domains) for RCTs, ROBINS-I (7 domains, incl. confounding) for non-randomised studies. Per-study; worst domain sets the overall. Feeds the GRADE risk-of-bias domain.
  4. Effect size + poolingwhat the combined number is. Each study's effect (SMD/MD, or log RR/OR) is combined by inverse-variance random-effectsmeta-analysis, separately within each design stratum. Answers: “how big is the effect, and how consistent?”
  5. GRADEhow much to trust the combined number. The capstone. Takes the pooled estimate + heterogeneity + RoB + precision + publication bias and emits a single certainty letter per stratum (RCTs start HIGH, observational LOW). It consumes the outputs of all the steps above.
  6. Bradford Hillcausation, beyond a single estimate. Nine viewpoints (strength, consistency, biological gradient/dose-response, temporality, plausibility…) for judging whether an association is causal. Two of them double as GRADE upgrade factors for observational evidence (large effect = strength; dose-response = biological gradient).

The one-line map: PRISMA picks the studies → design class sorts them → RoB2/ROBINS-I weighs each one → pooling combines them per stratum → GRADE rates the combination → Bradford Hill asks whether it's causal. The first five are sequential stages of our pipeline; Bradford Hill is the interpretive lens laid over the result.

A frequent confusion: RoB and GRADE both concern “bias/quality,” but RoB rates one study while GRADE rates the body of evidence for one outcome. Likewise PRISMA (reporting) is orthogonal to GRADE (certainty) — a perfectly PRISMA-compliant review can still be GRADE VERY LOW, and vice versa.

Gold standards

A modern systematic review is governed by four overlapping standards. Our methodology aligns with all four and the choices we made on each are described below.

FrameworkWhat it specifiesHow we use it
PRISMA 2020
Page MJ et al., BMJ 2021;372:n71
27-item reporting checklist + 4-stage flow diagram.Both rendered in full on every MA. Each item's compliance status (done / partial / to-do / n/a) is computed deterministically from MA state. The act of building the MA populates the checklist.
Cochrane Handbook
Higgins JPT et al., current edition
Methodological standard — search, eligibility, extraction, RoB, synthesis, certainty.PICO-locked eligibility, per-paper RoB2, Paule–Mandel random-effects pooling with HKSJ CIs (k ≥ 3), prediction intervals when k ≥ 3, Egger's only when k ≥ 10, weight-based RoB downgrade for GRADE (§14.2).
GRADE
Guyatt GH et al., BMJ 2008;336:924
Certainty framework — A / B / C / D, downgraded across 5 domains.Auto-algorithm with calibrated thresholds; reviewer can override at the Review stage. Per-domain rating + note + verbatim quote on the MA.
RoB2
Sterne JAC et al., BMJ 2019;366:l4898
Per-RCT risk-of-bias tool — 5 domains, worst-domain overall.Every in-pool study gets 5-domain rating + per-domain rationale + verbatim quote (citable Snippet rows).

Adjacent standards we touch:

AMSTAR-2 (Shea BJ et al., BMJ 2017;358:j4008) — quality of systematic reviews. We satisfy its 16 critical items via PRISMA compliance + locked protocol + RoB2 + sensitivity analyses.
CONSORT 2010 — for the underlying trials. CONSORT-style deficits in the included RCTs surface as RoB downgrades.
SPIRIT 2013 — trial-protocol reporting. We don't produce trials, but we use SPIRIT items as a quality lens when judging whether an included trial's protocol was robust.
MOOSE (Stroup DF et al., JAMA 2000;283:2008) — observational-study meta-analysis reporting. Applied when an MA pools observational designs alongside RCTs.
ICH E9 — statistical principles for clinical trials. Used as a lens for handling missing data, ITT, and primary-outcome pre-specification at the screening stage.
EQUATOR Network — clearing-house for reporting guidelines. We cross-reference EQUATOR when adding a new framework so we don't miss an applicable standard.

See “Improvements over gold standard” below for the specific ways our methodology goes beyond or deliberately diverges from these frameworks.

Improvements over gold standard

The frameworks above are reporting and quality standards designed for static, published reviews. Building a live evidence-synthesis platform gave us scope to push past the standards where the gain to readers is real and to deliberately substitute lower-cost alternatives where the standard's ceremony exceeds its value.

Where we go beyond the standards

Fair-dealing snippet enforcement on every per-study claim. PRISMA requires extracted data to be cited but is vague about quoting; we require a verbatim ≤50-word extract under fair dealing for every per-study claim (extracted data, RoB judgement, indirectness rating). Each in-text [Nx] resolves to the source paper's exact words, so nothing on a published MA is an unsupported claim.
Live recompute. Cochrane reviews are frozen on publication. Our pooled estimate, GRADE, leave-one-out and prediction interval recompute every time a study row changes; the MA you read today reflects today's evidence base, not the cut-off date of a published paper.
Out-of-pool transparency. Studies that passed PRISMA inclusion but couldn't be pooled (no SD reported, safety-arm only, no psychometric outcome) appear in References with an orange pill and a “Why not pooled” item carrying the prose notes plus the verbatim quote backing the exclusion. Cochrane usually omits these or stashes them in an appendix.
PRISMA flow diagram as runtime UI — entries link to the corresponding evidence in References rather than sitting as a static figure.
Two-part evidence grade. The certainty letter (pure GRADE) and the effect-size number (1/2/3 from MCID/Cohen) are shown separately, so “how sure” and “how big” are independently visible. GRADE alone conflates these by reporting only the letter.
Methodology choices menu. The analytical surface area (effect-size measures, pooling models, heterogeneity quantification, sensitivity analyses, publication-bias methods) is listed in full on this page; per-MA the actual choices made are spelled out under the “Methodology” section with reasoning.
Visual flow chart + clickable navigation. The chart at the top of this page doubles as a table of contents — every input/synthesis/output box links to its explanation section. Standard methodology documents bury the same content in prose.

Where we deliberately diverge

One reviewer + LLM instead of two human reviewers for screening and extraction. The LLM is fast and consistent; a human adjudicates the disagreements and writes the final RoB judgements. This is an explicit cost / coverage trade-off, documented in PRISMA item 8 compliance note, not a methodological lapse.
No PROSPERO registration for the protocol. The locked-protocol-with-timestamp + immutable JSON in our database serves the same audit purpose at lower cost. PRISMA item 24a is reported as “partial” with this note rather than “done.”
Weight-based bias-share RoB downgrade (calibrated against Cochrane Handbook §14.2). We weight a bias fraction (low 0, some-concerns 0.5, high 1.0) by pooled inverse-variance weight across all tiers and downgrade when the bias share exceeds 50 %, rather than counting high-RoB studies. This catches a pool with no clean evidence (e.g. 72 % some-concerns + 28 % high) that a high-RoB-only rule misses, while still not downgrading on one low-weight high-RoB study. Many off-the-shelf GRADE implementations use a count-based rule and downgrade for any single high-RoB study; we don't.
I² > 50 % + PI-crosses-null required for an Inconsistency downgrade. Cochrane Handbook §10 deliberately gives no I² cutoff; we set one and pair it with the prediction-interval check because direction-consistent heterogeneity (all studies favour treatment, magnitudes vary) should not cost a level on its own.

Our approach

The goal is to produce high-quality, transparent meta-analyses — our own synthesis of the primary evidence, graded and fully traceable to source. Where a published meta-analysis already exists on the same question, we use it as a comparison and a check on our own quality: if our independent result and a well-conducted prior review disagree, that is a signal to find out why before trusting either. Throughout we favour skeptical accuracy over optimism — where the evidence is thin or fragile we say so, rather than presenting a single confident number.

Statistical methods

We pool with the standard tools of a modern systematic review:

Random-effects pooling — Paule–Mandel random-effects (DerSimonian–Laird's successor), with the pooled CI by HKSJ for k ≥ 3; assumes the true effect varies between studies.
Hedges' g for standardised mean differences — a small-sample-corrected effect size.
Heterogeneity — I², τ² (Paule–Mandel) and Cochran's Q quantify how much the studies disagree.
Prediction interval — the range a future trial's result would plausibly fall in.
Leave-one-out — re-pool dropping each study in turn.
Egger's test — checks funnel-plot asymmetry for publication bias (only meaningful with ≥10 studies).

Analytical choices — when each applies

A systematic review involves a chain of analytical choices. Each MA records the specific choices it made; below is the menu of options we draw from and the circumstances that typically lead to each.

Effect-size measure

Hedges' g (small-sample-corrected SMD) — default for continuous outcomes on different scales.
Cohen's d — uncorrected SMD; used only when sample sizes are large.
Glass's Δ — SMD standardised by the control-arm SD; appropriate when arm variances differ markedly.
Mean difference (MD) — when every trial measures on the same clinically meaningful scale.
Risk Ratio (RR) — binary outcomes from prospective designs.
Odds Ratio (OR) — binary outcomes; case-control designs and rare events.
Peto OR — variant of OR for very rare events.
Hazard Ratio (HR) — time-to-event outcomes.

Pooling model

Random-effects (Paule–Mandel) — default; iterative, distribution-free τ², DerSimonian–Laird's successor.
HKSJ pooled CI — Hartung–Knapp–Sidik–Jonkman tk−1 interval (variance-inflation truncated at 1), applied by default for k ≥ 3; falls back to DL + normal-z at k = 2.
Random-effects (DerSimonian–Laird) — the classic moment estimator; retained for the k = 2 fallback and for comparison.
Random-effects (REML) — alternative when k is small (≤10).
Fixed-effect (inverse-variance) — assumes a single true effect; only when trials are very similar.
Mantel–Haenszel — sparse binary data (zero cells).

Publication-bias assessment

Funnel plot — visual asymmetry (k ≥ 10 minimum).
Egger's regression test — formal asymmetry test.
Trim-and-fill — adjusted pool when asymmetry is detected.
Begg's rank correlation — robust alternative to Egger's.

Referencing & fair dealing

Every paper — the included studies and the reference meta-analysis — is a numbered reference. Extracted data and any subjective judgement (risk of bias, indirectness) are backed by a short verbatim extractfrom the source, quoted under fair dealing (typically ≤ ~50 words, with its location).

In-text markers link to the evidence: [N] jumps to a paper, and [Nx] opens the exact extract that supports a statement. Nothing on the page is an unsupported claim — you can always trace a number or a rating back to the words it came from.

Data provenance & quality checks

We are explicit about where each number comes from. Data extracted from the original paper is shown in that paper's entry; data we could only take from the reference meta-analysis' table is labelled as such. An analysis is only marked replicable once every study has been independently extracted.

During extraction we run automated sanity checks — the common ways a number goes wrong, such as a standard error mistaken for a standard deviation, a crossover or combination design, or a result that counts only trial completers. These act like unit tests on the data: a warning means we adjusted or double-checked something, and the raw check is always available behind the plain-language summary.