If you read evidence for a living, here's what this is and isn't. Every claim is an atomic Subject–Predicate–Object assertion; every study attached is a single graded appraisal — support, contradict, null, or mixed — weighted by grade and quality, never by count. The headline number is a weighted mean of signed stances on [−1, +1]; it measures how much the evidence agrees on a direction, and deliberately nothing else. Read it alongside the two figures next to it: the evidence ceiling (a strong-support built only on mice says so) and the independent-group count (two or more, or it's flagged insufficient). Well-powered nulls count against directional claims; retractions stay on the record at zero weight rather than being deleted; industry funding and single-network sourcing are flagged. None of this is a verdict on truth — it's an auditable snapshot of where the evidence sits today, every source exposed so you can check the arithmetic and watch the grade move when better studies land. Consensus is a question we keep re-asking, not an answer we defend.
The full model
Each claim is a graded Subject–Predicate–Object triplet. Each source attached to it is one appraisal — supports, contradicts, tested-null, or mixed — weighted by study design and quality, never by count. Each appraisal i carries a weight:
k is the number of sign-bearing appraisals sharing a research group (lab, cohort, or author network), so a cluster of k papers from one group carries about √k votes rather than k. Mixed and tested-null appraisals cast no directional vote and are not discounted. For unclustered sources k = 1 — see the honesty rules below for how often that label is actually filled in.
| Study grade | weight |
| mechanism / in-vitro | 1 |
| animal | 2 |
| observational | 3 |
| RCT / n-of-1 | 5 |
| meta-analysis / review | 8 |
| Quality | mult |
| low | 0.5 |
| moderate | 1.0 |
| high | 1.5 |
How the score is computed
The stance sign is +1 supports, −1 contradicts, 0 for tested-null and mixed. Nulls and mixed contribute 0 to the numerator but full weight to the denominator — a pile of nulls correctly drags a claim toward equipoise. The consensus score is the weighted mean of signed stances:
Score → state
| score | consensus_state |
| ≥ +0.6 | strong-support |
| +0.2 … +0.6 | leans-support |
| −0.2 … +0.2 | contested |
| −0.6 … −0.2 | leans-against |
| ≤ −0.6 | refuted |
Independence override. Regardless of score, a claim is forced to insufficient when total weight is tiny, or when fewer than two independent research groups back it (appraisals sharing a dataset count as one). A single meta-analysis is the exception — it pools many groups on its own.
Evidence ceiling. Reported alongside the score, never folded into it: the highest study grade behind the claim. A strong-support with an animal ceiling reads "consistently shown in mice," never "proven in humans." The score measures direction-agreement, not evidence grade.
The honesty rules
These are pulled verbatim from the engine's export, in its order — including the ones that are unflattering. If a rule is missing here, it is missing in the engine.
This guard is not mostly-right-with-a-few-bad-labels; it is mostly not applied.
- Null results for directional claims are graded contradicts, not neutral.
- Retracted studies are kept visible under the claim (struck-through) but carry zero weight — they don't move the consensus score or count as independent evidence; flagged automatically via Crossref/Retraction-Watch.
- Studies from the same research group are discounted: a cluster of k papers from one lab, cohort or author network counts as roughly √k independent ones, not k. Two deliberate exceptions: a study graded mixed or tested-null casts no directional vote, so it is not discounted; and a claim backed by a single group is forced to insufficient regardless of its score.
- A claim backed by only one research group is flagged 'insufficient' outright, however many papers that group published.
- Reviews replace what they pool, but only where checked. Matching a review to the trials it pools has to be done by hand, because a review usually pools other labs' work and no automatic rule can see the overlap. Done for 31 of the 370 claims where a review sits alongside trials it may contain; on the other 339 that evidence is still counted twice. ~45% of the reviews we try to open are paywalled and the unreachable ones cluster where overlap is most likely, so treat our count as a floor.
- The same-group discount relies on a label we type by hand, and most sources do not carry a meaningful one. Measured 2026-08-14: 6,987 distinct group tags across 7,458 sources, because the tag was filled in per STUDY rather than per LAB — so the discount fires on only 131 of 645 claims and touches 6% of directional studies. The other 94% are treated as independent because nobody grouped them. A missing label does not error; it makes a study look like its own independent group, voting at full strength. 23 of 27 hand-checked cases were confirmed dependent (one cohort under three tags, one lab under four, four tags literally prefixed 'independent-' sharing two authors) and 13 verdicts moved; 2,280 candidates remain unread. Blind spot: we store only each paper's first author, so we cannot match on the senior author, the strongest same-lab signal. An unclustered claim is not proof of independent replication.
- The same label can err the other way: a tag named after a PUBLISHER had grouped four unrelated reviews, wrongly discounting independent work. Group by team, cohort or declared network — never by journal or publisher.
- Known and unfixed: checked against PubMed publication types, 16% of the sources filed as individual studies are indexed there as reviews or meta-analyses. A review counted as a primary study is the double-count above in disguise, and the review-replacement rule cannot see it because nothing marks it as a review. Measured, not yet corrected.
- Some counts here were previously wrong in both directions: an audit of our own files in August 2026 found 132 studies voting twice on the same claim, and 29 studies across 17 claims not counted at all. Both fixed; a file the engine cannot read is now a hard error rather than a silent omission.
- Industry conflicts of interest are surfaced, not hidden.
How we count evidence: independence, not vote-counting
The score is a weighted mean, not a tally. Every study enters at its design-tier weight times a quality modifier (the tables above), so ten mouse studies never outweigh one large trial. Two rules then govern which studies count, and how much.
- Independence, not repetition. Studies from the same research group are discounted: a cluster of k papers from one lab, cohort or author network counts as roughly √k independent ones, not k. Two deliberate exceptions: a study graded mixed or tested-null casts no directional vote, so it is not discounted; and a claim backed by a single group is forced to insufficient regardless of its score. Ten analyses of a single biobank are closer to three independent confirmations than to ten, and we don't let them masquerade as ten. Studies from the same research group are discounted: a cluster of k papers from one lab, cohort or author network counts as roughly √k independent ones, not k. Two deliberate exceptions: a study graded mixed or tested-null casts no directional vote, so it is not discounted; and a claim backed by a single group is forced to insufficient regardless of its score. The claim page shows the independent-group count, not the raw paper count.
- Every genuinely independent study counts. We don't sample, cap, or stop once a claim "looks settled." A well-replicated finding should accumulate the weight of its replications; that's how a claim earns a strong rating. The only sources we leave out of the score are strict redundancies: a duplicate of a paper already counted, another paper from a cohort already represented, or a review that merely restates primary studies we've graded individually.
The upshot: a claim reads strong because independent groups keep finding the same thing, never because we hand-picked the supportive ones.
The double-count, said plainly: fixed where checked, and mostly unchecked
A review that pools twenty trials should replace those trials in our count. For a long time it sat alongside them instead, double-counting evidence on roughly 370 claims. That is now being repaired, and the repair's own coverage is the number to hold us to:
Reviews replace what they pool, but only where checked. Matching a review to the trials it pools has to be done by hand, because a review usually pools other labs' work and no automatic rule can see the overlap. Done for 31 of the 370 claims where a review sits alongside trials it may contain; on the other 339 that evidence is still counted twice. ~45% of the reviews we try to open are paywalled and the unreachable ones cluster where overlap is most likely, so treat our count as a floor.
Superseded studies stay visible on the claim page — struck through, zero weight, each linking to the review that absorbed it — so you can see the double-count being removed rather than take our word for it.
There is no sound automatic fix. A review usually pools other labs' work, and our grouping tag cannot see that overlap. We remove a trial only when it appears in the review's own included-studies list, verified by reading the full text, the PRISMA diagram or the forest plot. We never assume an overlap we have not read. That standard is right, and it is slow. The gap is not reviews we cannot parse. It is reviews nobody has sat down with yet — and the paywalled ones cluster exactly where overlap is most likely.
The distortion runs in whichever direction the duplicated evidence points. Usually that means a claim reads more confident than it should. Sometimes it means the opposite: when a null review sits on top of positive trials, the double-count suppresses the verdict instead.
We measured how much it could matter. Of those 370 claims, 85 would cross into a different verdict band if, wherever a pooled review exists, we dropped every primary study on that claim outright. That cuts deliberately too much: we do not know which primaries any given review actually pooled, which is the whole reason the rule needs reading by hand. So 85 is a ceiling, not a prediction: the most this correction could ever reach, not a count of verdicts we believe are wrong. Of the 85, 65 would move toward less support and 20 toward more. The tendency is toward over-confidence, and close to one in four runs the other way.
We would rather you knew the weakness than assume it away.
Studies we saw but did not grade
For every claim, alongside the graded evidence we publish the sources we encountered but did not fold into the score — each with a reason. It's there so you can see what was set aside and check our work.
The reason is always a structural fact about the source, never a verdict on what it found. A study is never excluded for being inconvenient or unconvincing; there is deliberately no "weak" or "unconvincing" exclusion reason. A disconfirming study that meets the bar is graded contradicts, full stop. The reasons are a fixed, closed set:
| reason | meaning |
off-topic | doesn't actually test this claim's specific endpoint |
duplicate-cohort / non-independent | same cohort, dataset, or lab as a study already counted |
subsumed-in-ma | a primary study already pooled inside a meta-analysis we grade (a manual exclusion: applied only when the primary is shown in the review's own included-studies list, PRISMA diagram or forest plot, never assumed) |
retracted-or-predatory | retracted, or from a source failing basic integrity standards |
methodology-floor | a design that can't isolate the claim: a confounded combination product, an unvalidated model |
For pickled vegetables & esophageal cancer, three primary case-control studies are listed subsumed-in-ma — already inside the Islami 2009 meta-analysis we grade, so grading them again would double-count. For resveratrol & lifespan, a resveratrol-plus-nanodiamond study in crickets is listed methodology-floor: the combination and the model can't isolate resveratrol's effect on lifespan.
"Last reviewed" dates
Some claims carry a Last reviewed date — the last time we actively re-searched the current literature for that claim and re-graded it by hand. It's distinct from the score's recompute timestamp, which updates mechanically whenever the graph changes. Absent means not yet re-surveyed under the ongoing review campaign. It's there so you can judge how current a verdict is, independent of where the score happens to sit.
Two memories, under the hood
ConsensusLab runs on an AI agent with a science-research toolchain and two very different memories — don't conflate them:
- The public consensus vault (what you see): a typed graph of claim–evidence SPO triplets — you can explore it live. Its design follows Agents-K1: Towards Agent-native Knowledge Orchestration (arXiv:2606.13669) — the thesis that the bottleneck for agent knowledge isn't retrieval but structure: capture entities, claims, evidence, and typed relations, not papers flattened to abstracts and citation edges. Explore it live →
- The agent's own operational memory (private, un-graded): a flat directory of ~17 small typed Markdown notes plus a one-line index — four types (project / feedback / reference / user), YAML frontmatter,
[[wikilink]] cross-references. Optimized for "what do I need to remember to work well," not for truth-tracking. The opposite of the vault on every axis.
Grading runs on the open-source Google DeepMind "Science Skills" (Apache 2.0) — structured agent skills for scientific-database lookups: PubMed / Europe PMC / OpenAlex, ClinicalTrials.gov, ClinVar / gnomAD / Ensembl, PubChem / ChEMBL, Open Targets / Reactome / UniProt, and more — with the honest weighting layered on top.
Disagree with a grade? Open any claim and use "Challenge the grade" — it goes to a real review queue.
Educational only, not medical advice.