Methodology & standards
Overall structure
The chart below traces every input that goes into an evidence grade. Every box is a clickable link to the section that explains it — start with the badge at the bottom and work backwards to see how it was derived, or click any input box to read what it contributes. The text contents list is below.
Contents
Question
Locked PICOSelection
Study selectionEffect-size synthesis
Effect sizeCombined
Evidence grade — letter + numberLocked PICO
Fix the question, PICO and eligibility before searching, then lock it. Everything downstream inherits this scope.
PICO is the standard 4-part framework for defining a clinical question:
“Locked” means we commit to these definitions before running searches or screening, and they can't be silently changed afterwards. The PICO + inclusion / exclusion criteria are stored as JSON on the MA row with a lockedAt timestamp; the build workflow won't move past Stage 1 until you click “lock protocol.”
This is the same idea PROSPERO registration enforces in academic practice — a timestamped commitment to the question and methods before looking at data. It prevents three classic failure modes:
Downstream, the locked PICO drives the Indirectness GRADE judgement (do the included studies actually match what we asked for?), the screening eligibility criteria, and the PRISMA item 24a (Registration) compliance status.
Study selection
Search every database, screen candidates against eligibility with reasons, and collect the full text of the included set (PRISMA flow).
Selection is the first place a synthesis can go wrong. We follow PRISMA 2020 — the international reporting standard used by Cochrane and the major medical journals — so the path from “everything the search returned” to “studies we pooled” is fully visible. Each record carries the decision (include / exclude), the reason for the decision, and a verbatim abstract excerpt under fair dealing that supports it.
The funnel has four PRISMA stages, with counts and reasons at each:
Selection sits between the locked PICO (which defines what we're looking for) and the per-study inputs (the extractions and judgements we make on each included paper). On a meta-analysis page the full record-by-record funnel and PRISMA-style flow diagram appear in the study selection section.
Per-study RoB
For each RCT, assess the 5 RoB2 domains from objective trial facts (attrition numbers, blinding, allocation concealment, pre-registration…) — each signalling question answered with a snippet-backed fact; the domain rating auto-derives from the answers. A RoB observation is recorded as the relevant domain’s Fact (with its snippet) in this table — its native format — never as a standalone note; if it changes an answer, re-derive the rating. Record any author-stated limitations separately; they are displayed but not scored.
For each randomised trial we apply RoB2 (Sterne et al., BMJ 2019), the current Cochrane risk-of-bias tool. Each study is assessed across five domains and rated low / some concerns /high on each. The overall rating is the worst-domain rating (Cochrane's worst-domain rule).
Facts, not author self-criticism
A core principle of our assessment: ratings are derived from objective trial facts, not from the authors' own statements about limitations. An author who writes “a limitation is that blinding was not possible” and one who says nothing about blinding are rated identically if the underlying trial had the same design. A self-critical author is not penalised for their candour; an over-confident one gets no credit for staying silent. What the trial did — attrition numbers, allocation concealment procedures, pre-registration status, blinding arrangements — is the input. What the authors said about what they did is not.
Author-stated limitations are separately captured in a limitations field and displayed alongside the RoB summary on the MA page. They inform readers but are never fed into the domain ratings.
The five RoB2 domains
Each domain is assessed by answering its Cochrane signalling questions (Yes / Probably yes / Probably no / No / No information). Each answer is backed by an extracted fact — a concrete, verifiable detail — linked to a verbatim snippet from the paper. The per-domain rating (low / some concerns / high) is then auto-derived by the RoB2 decision rules; it is not typed by hand.
Every per-domain judgement is backed by a verbatim quote (≤50-word fair-dealing extract) stored as a Snippet row, so the rationale is auditable. The rendered MA page shows each domain rating with an inline citation link to that snippet. Internal storage uses the Cochrane terms (low / some_concerns /high); we render them in plain language (low / minor concerns / major concerns) for non-specialist readers.
This card is an input — the per-study judgements are then aggregated into the Risk of bias synthesis card. Studies we couldn't obtain are marked not assessable rather than assumed clean.
Per-study stats
Extract n, mean, SD (or events) and the scale per included study; the per-study effect (Hedges’ g) is computed from this. Any extraction caveat (which arm/subgroup was pooled, combination products, implausibly tight SDs) is recorded as a note under that study’s row here — in the section’s native format — never in a standalone notes box.
Each included paper is read end-to-end and the primary-outcome data are extracted into a structured row on the Study:
SEM_not_SD (SE × √n conversion applied), baseline_only, combination_product, imputed_sd, crossover_design, endpoint_only_n_unclear.From these we compute Hedges' g (small-sample-corrected SMD) and its variance per study, using the small-sample correction factor J = 1 − 3 / (4(n₁+n₂) − 9). When the paper reports SE rather than SD we convert via SD = SE × √n and flag the conversion.
This card is an input. The per-study g values flow downstream into Inconsistency (do they agree?), Imprecision (how precise is the pool?), Publication bias (funnel asymmetry?), and Effect size (the pooled estimate itself).
Indirectness
Judge how directly the population, intervention, comparator and outcome match the question. Human assessment, preserved across recompute.
Indirectness asks: do the included studies actually match the question we said we were answering? That comparison only makes sense against a fixed reference — the locked PICO. Each study is checked against the four PICO dimensions:
This is the GRADE domain that is hardest to automate — there is no statistic that captures “does the evidence match the question.” So we never let the analysis decide it. Instead the analysis only flagsit: it marks indirectness pending (an outstanding review item, surfaced in the build workflow's Analyse stage) and notes any trigger it can detect — e.g. a combination product in the pool. A reviewer then resolves it against the locked PICO, recording a rating and a note. That judgement is stored separately (manualDomains) so it survives every recompute, and a serious call counts toward the certainty letter. Until resolved, the SR page shows a pending review state rather than a fabricated default.
Risk of bias (synthesis)
Downgrade for study limitations weighted by each study’s share of the pooled information.
The synthesis card aggregates the per-study RoBjudgements into a single GRADE domain rating. The threshold is the most subtle GRADE judgement and mis-calibrating it is a common cause of overly pessimistic certainty ratings.
Cochrane Handbook §14.2 wording: downgrade only when “most information is from results NOT at low risk of bias.”The unit of analysis is the body of evidence contributing to the pool, judged by weight not study count. We now read this across all RoB tiers rather than the high tier alone. Each study contributes its pooled inverse-variance weight times a bias fraction, and we take the weighted mean — the bias share:
Downgrade rule: serious (−1 to the certainty letter) when the weighted-mean bias share exceeds 50 %. A pool that is entirely some-concerns sits at exactly 0.50 and is not downgraded (the borderline is treated as discharged). The distribution is still reported (“1 of 14 studies at high RoB contributing 5 % of pooled weight”) so it stays visible even where no downgrade applies.
The weighted-across-tiers reading replaces the old high-RoB-only > 30 % rule, which missed pools with no clean evidence — e.g. 72 % some-concerns + 28 % high carries a bias share of 0.72, well over the threshold, yet contributes 0 % high-RoB weight under the old high-only cutoff. It still avoids the opposite pitfall — a single high-RoB study with low weight does not on its own tip a low-risk pool over 50 % — which matches Cochrane practice and published ashwagandha reviews (Akhgarjand 2022, low certainty). Count-based rules that downgrade for any one high-RoB study remain something we deliberately avoid.
Low-RoB sensitivity re-pool. As a GRADE robustness check, when ≥ 2 studies are at low risk of biasthe engine re-pools just those low-RoB studies — same Paule–Mandel random-effects / HKSJ machinery as the full pool. If the low-RoB-only estimate agrees with the full pool (same direction, overlapping CIs) the risk-of-bias downgrade is dischargeable — the bias didn't move the answer; if it diverges, the bias mattered. The re-pool is drawn as a labelled row in the forest plot beside the full-pool diamond.
Inconsistency
Downgrade for unexplained heterogeneity (I² and a prediction interval that crosses the null).
Inconsistency asks how much the studies disagree once we've accounted for chance. The four signals:
Downgrade rule: serious (−1) only when I² > 50 % and the prediction interval crosses the null. High I² whose prediction interval still excludes the null is heterogeneity that does not change the conclusion — we are still confident of the effect's direction — so it draws no downgrade. The prediction interval is undefined for k < 3; we treat that case as crossing (the conservative reading), so a small, heterogeneous pool is not let off on a missing interval.
A consistent direction with wide magnitude variation (e.g. ash effects from g = −0.09 to g = −4.0, all favouring treatment) is the textbook case for −1 once the prediction interval crosses the null: we're confident there's an effect, but we don't know its size.
Imprecision
Downgrade if the CI crosses the null or the sample falls short of the optimal information size.
Imprecision asks how precise the pooled estimate is. The inputs are the 95 % confidence interval, implicitly the p-value, and the optimal information size (OIS) — does the evidence base have enough participants to detect a clinically meaningful effect?
The pooled CI now uses Hartung–Knapp–Sidik–Jonkman (HKSJ). Rather than a normal-z interval around the DerSimonian–Laird estimate, we build the interval on a tk−1 quantile with a variance-inflation factor truncated at 1 (so the HKSJ interval is never narrower than the classic one). This is the modern default for few-study meta-analyses, where DerSimonian–Laird + normal-z understates the CI and overstates precision. It is gated to k ≥ 3: at k = 2 the t1 = 12.71 quantile is too extreme, so the engine falls back to DL + normal-z.
Downgrade rule: serious when the 95 % CI crosses the null, or when the CI excludes the null but the pooled sample falls below the optimal information size. The OIS check is now automated (it used to be applied by hand). For a continuous outcome the OIS is the total N needed to detect an effect δ (default 0.2 SMD) at α = 0.05 with 80 % power — 31.36 / δ² (≈ 784 participants for δ = 0.2). Ratio-outcome OIS is deferred, so for binary/ratio outcomes the domain currently downgrades on the crosses-null condition only.
This is where the p-value enters the badge
The p-value isn't a direct input to either part of the evidence grade — it's implicit in this domain:
We deliberately don't render the raw p-value on the badge because two estimates with p = 0.04 and p = 0.001 can have very different certainty letters depending on heterogeneity and publication bias — the letter is the more honest summary.
Publication bias
Downgrade for funnel asymmetry / small-study effects (Egger, Begg) when there are enough studies to assess it.
Publication bias asks whether the published literature systematically over-states the true effect. The classic mechanism: small null trials never get written up, while small positive trials do — so the published set is asymmetric and the pooled estimate is inflated.
Downgrade rule: serious when k ≥ 10 and Egger's p < 0.10. k < 10 is reported as unknown rather than downgraded — too few studies to detect asymmetry reliably. Begg's p is shown beside Egger's (k ≥ 10) so the two agree-or-disagree at a glance, but Egger's remains the test of record. GRADE caps this domain at −1.
Dose–response (biological gradient)
A genuine dose-response gradient — larger doses producing larger effects — is one of Bradford Hill's criteria for causality and one of GRADE's three upgrade factors for observational evidence. We test for it directly: each trial's dose is parsed to a comparable mg/day figure (handling multi-arm regimens and twice-daily dosing) and the per-study effect is regressed on dose with a precision (random-effects) weighting.
We use it two ways. For observational pools a beneficial gradient rates the certainty up one level. For randomised pools (already starting HIGH) it is contextual — but it also answers whether dose explains the heterogeneity: if the trend runs the wrong way (an inverse gradient, where the smallest-dose trials report the largest effects), that is biologically implausible, flags those trials as outliers, and means the inconsistency is not explained by dose — so the downgrade stands.
Effect size
Pool the per-study effects (random-effects Hedges’ g, Paule–Mandel τ², HKSJ CIs) into the summary estimate, I² and prediction interval. This is the hub every GRADE domain reads from.
This card is the pooled estimate itself — the number you'd quote if asked “by how much does this intervention work?” — distinct from the downstream Magnitude number, which is the 1 / 2 / 3 label we derive from this pool.
The effect-size card is fed by per-study statisticsand itself feeds Magnitude. See the Statistical methods and Analytical choices sections below for the full menu of pooling options.
Certainty letter
The certainty letter is the GRADE rating — the framework used by Cochrane, the WHO and most clinical guideline bodies. GRADE is design-stratified: a body of randomised trials starts at HIGH (= A) and can only be downgraded — one level per serious domain — because HIGH is the ceiling. A body of observational (non-randomised) evidence starts at LOW and can instead be rated up; see “Upgrades vs downgrades” below. When an analysis contains both, we pool and grade each stratum separately and present them side by side — so each design carries its own letter.
Separate, don't merge. When randomised and non-randomised studies both bear on a question, we pool them in separate strata — a primary RCT pool and an observational (NRSI) pool — and compare them, rather than blending them into one number. The reason is mechanical: a meta-analysis weights each study by its precision, which scales with sample size, not with freedom from bias. Observational datasets are often enormous, so in a merged pool they dominate — and their confounding is systematic error that does not shrink with n. Merging therefore narrows the confidence interval around a possibly-biased estimate, manufacturing false precision. Keeping the strata apart lets each be graded on its own terms (RCTs start HIGH, observational LOW) and turns the comparison into a finding: when the two agree, that is corroboration across independent designs; when they diverge, that is a signature of confounding, and the RCT pool is the more trustworthy estimate. Case series and single-arm before-after studies have no concurrent comparator, so they yield no comparative effect at all — they are catalogued (and, where useful, summarised as a pooled proportion), never placed in a comparative forest plot. Lower-tier evidence still earns its place — harms, rare events, questions where RCTs are impossible — as a labelled, lower-certainty strand.
Upgrades vs downgrades. The five domains above only downgrade. GRADE also defines three upgradefactors — a large effect, a dose-response gradient, and “plausible confounding would only shrink the effect” — but these apply to observational evidence (which starts at LOW); they exist to rescue uncontrolled evidence that turns out to be compelling. They are not a counterweight to downgrades: rating up is only considered when there is no serious reason to rate down, so a −1 for risk of bias is not cancelled by a +1 for dose-response. It is precedence, not arithmetic netting. Upgrades apply only to the observational stratum (which starts at LOW); the RCT stratum, starting at HIGH, is downgrade-only. We auto-apply the large-effect upgrade (a pooled RR/OR ≤0.5 or ≥2 with a CI excluding the null rates up one level; ≤0.2 or ≥5 rates up two), and surface dose-response and opposing-confounding for manual judgement since they need data we do not yet capture. Risk of bias is assessed with RoB2 for RCTs and ROBINS-I for non-randomised studies (its “serious”/“critical” ratings carry the downgrade).
The auto-algorithm computes the rating from the five GRADE domain cards above (see Risk of bias, Inconsistency, Indirectness, Imprecision, Publication bias) using Cochrane-calibrated thresholds documented in “Frameworks we follow.” A reviewer can override the letter at the Review stage; the stored letter and any override carry a rationale and timestamp.
Consilience continuous certainty. Alongside the canonical discrete GRADE letter — never replacing it — we run a continuous certainty engine. Each of the five domains emits a fractional downgrade in [0, 2] rather than a 0/−1/−2 step: risk of bias ramps up from the 50 % bias-share zero-point, inconsistency from I², imprecision from crosses-null plus OIS shortfall, publication bias from Egger's p, and indirectness from the manual judgement. The five partials are summed and rounded once onto the design-anchored start (RCTs HIGH), floored at very low. This addresses GRADE's “death by a thousand cuts”: several domains that are each borderline-but-not-serious can together cost a level, which discrete GRADE — rounding each domain independently — silently rounds away. The MA page shows this as a “Consilience continuous certainty” panel with the per-domain partials, beside the canonical letter; where the two diverge that gap is surfaced, not hidden.
Magnitude number
The magnitude number is our addition on top of GRADE: it answers “how big is the effect?” on a 1 / 2 / 3 scale, derived from the pooled effect-size card.
Showing magnitude separately from certainty matters because “how sure?” and “how big?” are independent questions. A large effect can be uncertain; a small effect can be nailed down. GRADE alone conflates these by reporting only the letter.
Evidence grade — letter + number
Sum the five domain downgrades into the certainty letter, map the pooled effect to a 1–3 magnitude number, and combine them into the evidence grade.
The evidence grade is the final two-part badge that combines the certainty letter and the magnitude number into a single token.
The badge is auto-computed, but a reviewer can override either the letter or the number with a recorded rationale at the Review stage of the build workflow. The override is timestamped and visible in the History.
How the frameworks fit together
Evidence synthesis borrows from several frameworks that are easy to confuse because they overlap. They are not competitors — each governs a different stage of the same pipeline. Here is the order in which they act and what each one decides.
The one-line map: PRISMA picks the studies → design class sorts them → RoB2/ROBINS-I weighs each one → pooling combines them per stratum → GRADE rates the combination → Bradford Hill asks whether it's causal. The first five are sequential stages of our pipeline; Bradford Hill is the interpretive lens laid over the result.
A frequent confusion: RoB and GRADE both concern “bias/quality,” but RoB rates one study while GRADE rates the body of evidence for one outcome. Likewise PRISMA (reporting) is orthogonal to GRADE (certainty) — a perfectly PRISMA-compliant review can still be GRADE VERY LOW, and vice versa.
Gold standards
A modern systematic review is governed by four overlapping standards. Our methodology aligns with all four and the choices we made on each are described below.
Adjacent standards we touch:
See “Improvements over gold standard” below for the specific ways our methodology goes beyond or deliberately diverges from these frameworks.
Improvements over gold standard
The frameworks above are reporting and quality standards designed for static, published reviews. Building a live evidence-synthesis platform gave us scope to push past the standards where the gain to readers is real and to deliberately substitute lower-cost alternatives where the standard's ceremony exceeds its value.
Where we go beyond the standards
[Nx] resolves to the source paper's exact words, so nothing on a published MA is an unsupported claim.Where we deliberately diverge
Our approach
The goal is to produce high-quality, transparent meta-analyses — our own synthesis of the primary evidence, graded and fully traceable to source. Where a published meta-analysis already exists on the same question, we use it as a comparison and a check on our own quality: if our independent result and a well-conducted prior review disagree, that is a signal to find out why before trusting either. Throughout we favour skeptical accuracy over optimism — where the evidence is thin or fragile we say so, rather than presenting a single confident number.
Statistical methods
We pool with the standard tools of a modern systematic review:
Analytical choices — when each applies
A systematic review involves a chain of analytical choices. Each MA records the specific choices it made; below is the menu of options we draw from and the circumstances that typically lead to each.
Effect-size measure
Pooling model
Publication-bias assessment
Referencing & fair dealing
Every paper — the included studies and the reference meta-analysis — is a numbered reference. Extracted data and any subjective judgement (risk of bias, indirectness) are backed by a short verbatim extractfrom the source, quoted under fair dealing (typically ≤ ~50 words, with its location).
In-text markers link to the evidence: [N] jumps to a paper, and [Nx] opens the exact extract that supports a statement. Nothing on the page is an unsupported claim — you can always trace a number or a rating back to the words it came from.
Data provenance & quality checks
We are explicit about where each number comes from. Data extracted from the original paper is shown in that paper's entry; data we could only take from the reference meta-analysis' table is labelled as such. An analysis is only marked replicable once every study has been independently extracted.
During extraction we run automated sanity checks — the common ways a number goes wrong, such as a standard error mistaken for a standard deviation, a crossover or combination design, or a result that counts only trial completers. These act like unit tests on the data: a warning means we adjusted or double-checked something, and the raw check is always available behind the plain-language summary.