Every certainty rating Evidara shows you comes from a deterministic scaffold, not a vibe. This page explains exactly how it works — the rules, the thresholds, and the places it deliberately declines to guess. Including the parts a pitch page would normally leave out, like why certainty so often lands at Moderate.
The certainty rules are deterministic — the same cited evidence always produces the same rating. The evidence set is not frozen: live retrieval means re-running a query later can surface new literature, and LLM-synthesized narrative is not reproduced verbatim between runs.
Every graded outcome lands on one of five GRADE certainty levels. There's a level reserved specifically for "we didn't have enough evidence to judge" — that's a different statement from "we judged it and it was weak," and Evidara keeps the two separate rather than collapsing them.
GRADE defines five reasons to downgrade certainty. Evidara automates four of them from what a study's abstract actually states. The fifth is named on purpose: it's a judgment call we route to a human rather than approximate.
Which risk-of-bias tool applies depends on the study design being assessed. Whichever tool runs, the same aggregation rule applies: the worst domain that could actually be assessed sets the overall verdict — and a domain marked "not assessed" is left out of that comparison entirely, never quietly counted as low risk.
Two structural rules combine to make Moderate the most common outcome, even for evidence that includes strong trial data. First: when an answer draws on both randomized and observational sources, the overall certainty follows the weaker pool, not the stronger one. Second: the observational ceiling described above means any pool of real-world evidence tops out at Moderate regardless of effect size. Put those together, and a mixed evidence base rarely reaches High — not because the rules are broken, but because that's what a mixed evidence base honestly supports.
A rating is only as trustworthy as the citations under it. Evidara runs two separate, live verification paths — one over generated content, one over chat and retrieval — because they catch different failure modes.
GRADE was designed for panels of methodologists working from full-text access to every included study. Evidara's implementation approximates that logic mechanically, from what's stated in an abstract — it is not a substitute for that panel, and it doesn't claim to be. Every automated rating is marked as automated. Where a judgment genuinely requires a person — indirectness, full-text-only risk-of-bias items, the ten AMSTAR-2 items an abstract can't answer — Evidara says so and stops, rather than filling the gap with a guess.
The certainty rules themselves are deterministic — the same cited evidence always produces the same rating. The evidence set is not frozen: live retrieval means re-running a query later can surface new literature, and LLM-synthesized narrative is not reproduced verbatim between runs.
Evidara is built for evidence synthesis from published literature. That's a specific job, not every job — five places where a different tool is the right call:
The clearest way to understand a grading system is to watch it grade something you actually care about.