Skip to how evidence gets into the library

One rule, applied the same way every time.

Every finding follows the same path from a saved source to a public page. Reported details, names we line up, ratings we work out and gaps we cannot fill stay visibly separate all the way through.

This is a library, not advice. No recommendation, no dose, no diagnosis, no ranking. We never ask about your symptoms, and nothing here is tailored to you.

The whole path

From a source to something you can explore

The same checked information powers the question pages, the full table and the concept map. Those are different views of one body of work, not separate answers.

  1. Read and record. A source is saved where possible, then its stated people, comparison, measure and numbers are taken down without filling omissions.
  2. Line up the names. Different written forms are connected to one shared vocabulary, while any name that cannot be resolved stays outside the view.
  3. Build connected facts. Eligible rows become fact records; related facts form pages and links that can always be rebuilt from those records.
  4. Publish only checked views. Public pages withhold source extracts and working references. Internal screening and review material stays behind access controls instead of leaking through the reading room.

The system behind the library

Five parts of the system, and where to watch each one

What you can read here is the published end of something larger. These five parts describe the rest of it in the same plain terms the findings get, and each one names a page on this website where you can watch it working instead of taking our word.

Reading a source: machines suggest, fixed code decides

A language model — a program that writes text — is allowed to suggest: which papers might be worth reading, which written name might be the same thing as another. It is never allowed to decide. Every suggestion is set aside and checked by fixed code, and nothing a model writes reaches a page as a result of its own. Put the same inputs through the deciding code twice and the same thing comes out both times, with nothing fetched from the network in between.

A source that gets stopped on the way through is a real outcome, not a missing one. It is recorded as stopped, with the reason, in the same way an unreported number is recorded as unreported instead of being counted as a zero.

Rating: the argument, not the letter

A letter on its own is not an argument. Being told a finding is low certainty gives a reader nothing to agree or disagree with. So every check we run records one of three answers, never two: it fired and stepped the rating down, it was checked and did not apply, or it could not be decided at all.

That third answer is the one most systems drop. A check that asks whether a range of results is implausibly wide, and is handed no range, has not passed it — it has not been decided, and writing it down as passed would be asserting something that is not true. Most steps down in this library are of that kind: taken because a number was never reported, not because anyone judged a study weak.

One set of facts, many ways in

A substance, a thing measured, a health topic, a paper: each is a different door into the same set of checked facts. A page here is a walk across those facts from whichever side you came in, not a separately written article. That is why the same statement carries the same note of where it came from no matter which page you meet it on, and why a gap shows up as the same named gap from every side.

The complete internal pages are not published, and it is worth saying plainly why. They carry internal review identifiers and working references that belong to the people doing the reviewing, so they stay behind the operator sign-in. This website builds a smaller map of names from those same facts and never ships the full pages. Nothing in the public map depends on anything withheld.

The screening check: four questions, and working you can re-check

Separately from this library, the system can screen a named thing against what has been reviewed. It asks four fixed questions — is there a known clash, is there a recorded reason to consider it at all, is the detail the check needs actually present, and is there a reviewed basis. Each question records its own answer, and those answers are more than a yes or a no: one says the question was never reached because an earlier one had already settled the matter. The four together produce one answer for the screen as a whole: held, needs more context, or allowed. It never advises and never doses; held and needs-more-context are ordinary answers, not errors.

Every answer comes with its working attached, and a short fingerprint of that working. The fingerprint can be checked offline with a separate tool that ships with the system’s own package. To be exact about what that proves: it confirms the shape of the recorded working and that the fingerprint matches it. It does not re-run the check or work the answer out again.

Publishing: a signature the software cannot produce

Nothing reaches the published files because a program decided it was ready. Any approval a program can produce for itself is not a control at all: the same process that prepared the change could write the approving value in one line. So the gate is one thing the software cannot make — a signature from a key that lives on a person’s own machine, where no part of the system can read it. The software can check that signature. It can never mint one.

A person signing is the decision, and it is recorded as such. Being exact about where that stands as of August 2026: the gate is in place and has been used. In late July 2026 the project’s owner signed six approvals that re-checked findings which were already published — so sixteen of the thirty-one findings the answers draw on now carry a person’s signature over their exact saved source records, and the software re-derives those ratings from the signed records and reaches the same answer. What has not happened yet is a brand-new finding entering the published set through the gate.

The reading-room findings, re-read by an AI agent

The several hundred reading-room findings were all published before that gate existed. Reading them one by one is more work than the owner has hours for, so the owner did something different and is telling you exactly what it was: an AI agent was given each finding and the saved copy of its source, and asked four questions. Is the quoted sentence really the source’s sentence about this ingredient and this outcome? Is every number the number the source gives for this comparison, and not one that happens to appear elsewhere in the text? Is every unit the unit the source states? And does the direction describe the thing the finding says it measured?

No person read these. The owner authorised the agent to sign each batch of about ten with the owner’s own key. That signature is worth less than the one described above, and we would rather say so than let the word carry more than it earned. What it proves is that the batch came from the party the owner authorised, and that neither the verdicts nor the finding’s own text has changed by so much as a byte since — every rebuild rechecks both, and a finding edited after its review silently loses the label rather than keeping a claim that no longer describes it. What it does not prove is that a second party agreed. Nothing here samples the agent against a human reader.

Findings the agent could not confirm are published saying so, beside the ones it confirmed, with both counts shown together. That is deliberate. A review that only ever added a mark to the findings it liked would be publishing its own success rate and calling it a fact about the evidence.

The version code the published files carry is a separate thing, and worth not confusing with this one. It is worked out from the set of facts itself: it names exactly which set you are reading and changes whenever they do. It says nothing about who approved them.

How we work

Which words came from the paper, and which came from us:

  • Reported by the source: the number, the range, how many people took part, the kind of study, who was studied, what it was compared with, how much they took and for how long — exactly as the source stated them.
  • Standardised by us: the ingredient name, the name of the thing measured, and the unit, lined up so findings can sit in one table. The name and the unit may change; the value is unchanged.
  • Rated by us: the certainty, the rule that produced it, and every reason it stepped down. This part is our work, and we label it as ours.

Six rules govern every finding:

  1. Trace every number to a source. Every finding names and links the source it came from, and every total on these pages is worked out from the library itself. No number here is typed in by hand.
  2. Keep the three layers apart. What the source reported, what we standardised and what we worked out are labelled separately on every finding, never blurred into one voice.
  3. Say out loud how unsure we are. Every finding carries a certainty level and the reasons behind it, written down rather than smoothed over.
  4. Leave a gap as a gap. The finding names what was missing from the reading: “Not in the abstract”, “Not in this document” or “Not recorded in this reading”. Never a zero, never a blank, never a guess. A real zero is shown as a zero.
  5. Record what happened to each source later. We read corrections and updates from the saved copy and show what we find: the notice where there is one, “We checked the saved copy: no notice” where we looked and found none, and “No correction check recorded yet” where nothing has been checked. We never call unchecked clear.
  6. Stay a library, not advice. No recommendation, no dose, no diagnosis and no ranking. We describe evidence; we do not tell anyone what to do.

Exactly what we are claiming

What “reviewed” means here. 515 findings were rated by the fixed rule with nobody signing off, 28 had a person go through the criteria one by one, and on 3 a person read the claim the rating was made from and wrote their own reason for each step — the only ratings here whose reasons are somebody’s words rather than the rule’s. That is the whole of it: we never claim all 546 findings were human-reviewed. On the 28 a person went through, what they wrote against each step was not kept in a form we can show, and those findings say so where the reason would be — a person having looked is not the same as a reason you can read.

Eight steps, always in this order

How evidence gets into the library

  1. Identify an in-scope source: a study or review that has something to say about one of the questions on our list.
  2. Save a copy of the source where we can, so every later check reads the same text.
  3. Take down only what is reported: the numbers and descriptions the source itself states, with nothing inferred.
  4. Leave the gaps as gaps: anything the text we read does not state stays missing, with the abstract, document or full-paper reading named. We never fill it in.
  5. Line up names and units so findings can sit in one table: ingredient names, the names of things measured, and units. The reported value is never changed.
  6. Apply the certainty rule and write down what came out and every reason, exactly as set out below.
  7. Record what happened to the source later: any correction or update, read from our saved copy, never guessed. “Not recorded” means nothing was checked.
  8. Publish only what we are allowed to, holding back extracts and internal references, and stamp it with a code worked out from the content itself. The code on the library page tells you exactly what was published.

The certainty rule

Our certainty ladder follows GRADE, the standard method for rating how certain evidence is, applied as one fixed rule. Where a finding starts depends on the kind of study; from there it can only go down.

Where a finding starts

  • High certainty, level 4 of 4systematic reviews — studies that pool all the trials on a question — meta-analyses and randomised trials start at high certainty
  • Low certainty, level 2 of 4cohort and case-control studies start at low
  • Very low certainty, level 1 of 4case series and mechanism evidence start at very low
  • Low certainty, level 2 of 4any other design starts at low

Certainty never steps up.

Why a finding steps down

From wherever it starts, a finding drops one level for each kind of problem the source itself shows:

  • Risk of bias — how likely a study’s design distorted its result: nobody appraised the quality, or the source itself reports serious limitations.
  • Imprecision: the range is missing, or it crosses the no-difference point, or it is very wide, or too few people took part.
  • Inconsistency: the pooled studies disagreed a lot — heterogeneity, how much the pooled studies disagreed, scoring above 60 out of 100 — or the direction they reported is mixed.
  • Indirectness: the evidence comes from something other than people, or what was measured is a stand-in for what you would actually care about.

Every step down is shown on the finding itself, with the reason where one was recorded — on the ratings a person went through, it says so instead of inventing one — and every criterion that took no step is named there too.

What a low rating is usually measuring

Most readings here are published abstracts. 2 signed public-reading findings use the full paper, and some use other published documents. The reasons above can fire when something is simply not in what we read: no appraisal of the studies behind it, no count of the people pooled, no range around the number. Of the 1,019 steps down taken across the whole library, 751 are of that kind.

So a Low rating here usually means we could not see enough to be sure, and only sometimes means the study itself was weak. The two are different, and every finding says which of them each of its steps was. A rating is never a verdict on an ingredient, and it is not a verdict on a paper either.

One rating, start to finish

Not an example we picked. 219 of the 546 findings here take the same path through the ladder — more than any other — and this is the first of them in order of its own record identifier: Ashwagandha and Anxiety.

What the source is
Effects of Ashwagandha (Withania Somnifera) on stress and anxiety: A systematic review and meta-analysis. Explore (NY). 2024;20(6):103062. doi:10.1016/j.explore.2024.103062. PMID:39348746.
Number reported
-2.19 points on the Hamilton Anxiety Scale, mean difference vs placebo
Likely range
-3.83 to -0.55
How many people took part
Not in the abstract
How this rating was reachedstarts at High, 2 steps down: risk of bias and imprecision

A fixed rule worked this out from the source’s own fields. Nobody checked this rating by hand.

  1. Starts at High
  2. Risk of bias brought it down to Moderatethe abstract reports no risk-of-bias or quality appraisal, so study limitations cannot be ruled outNot in the abstract
  3. Imprecision brought it down to Lowthe abstract does not state how many participants were pooledNot in the abstract
  4. Ends at Low2 steps down.
Checked, nothing to step it down for
indirectness
Inconsistency
No disagreement reported, and no count of studies behind it

Imprecision stops at the first check that applies; others may apply too.

The finding itself, where the same working sits under the record it belongs to, and the source record.