Pricing Platform Skills Contact
Methodology

How Evidara grades evidence.

Every certainty rating Evidara shows you comes from a deterministic scaffold, not a vibe. This page explains exactly how it works — the rules, the thresholds, and the places it deliberately declines to guess. Including the parts a pitch page would normally leave out, like why certainty so often lands at Moderate.

The certainty rules are deterministic — the same cited evidence always produces the same rating. The evidence set is not frozen: live retrieval means re-running a query later can surface new literature, and LLM-synthesized narrative is not reproduced verbatim between runs.

Five outcomes, not a percentage

Certainty is a rating,
not a confidence score.

Every graded outcome lands on one of five GRADE certainty levels. There's a level reserved specifically for "we didn't have enough evidence to judge" — that's a different statement from "we judged it and it was weak," and Evidara keeps the two separate rather than collapsing them.

⊕⊕⊕⊕
High
Confident the true effect is close to the estimate. Reserved for evidence that started strong and survived every downgrade check.
⊕⊕⊕◯
Moderate
Likely close to the true effect, with meaningful uncertainty. This is the ceiling for observational evidence, and the most common landing point overall — more on why below.
⊕⊕◯◯
Low
The true effect may differ substantially from the estimate. Typically observational evidence that picked up an additional downgrade.
⊕◯◯◯
Very Low
Very little confidence in the estimate. The floor — certainty never drops further, no matter how many downgrade factors apply.
◯◯◯◯
Insufficient
Not a quality judgment — a coverage one. Fewer than two gradeable sources were found. This is deliberately distinct from Very Low: "not enough to judge" and "judged and found weak" are different claims, and only one of them is a statement about quality.
Where a rating starts

Study design sets the floor —
and, for one design, the ceiling.

Starting point

Randomized trials and reviews start High

Any body of evidence containing a randomized controlled trial, systematic review, or meta-analysis starts at the top of the scale. From there, it can only move down — through the downgrade checks below.
Starting point

Observational evidence starts Low

Cohort studies, case-control studies, and other non-randomized designs start two levels down. Unlike randomized evidence, this starting point can move up — but only under specific conditions.
Structural ceiling

The observational upgrade stops at Moderate

A large detected effect (an odds, risk, or hazard ratio above 2.0 or below 0.5) can upgrade observational evidence by one level. It cannot upgrade it past Moderate — never to High, regardless of effect size. This is a hard ceiling, not a rounding artifact.
Five checks, run mechanically

What pulls a rating down —
and what we won't automate.

GRADE defines five reasons to downgrade certainty. Evidara automates four of them from what a study's abstract actually states. The fifth is named on purpose: it's a judgment call we route to a human rather than approximate.

01 — Risk of Bias
Downgrades when bias signals cluster
Certainty drops one level once at least 3 studies have been assessed for bias, and either 25% or more come back high-risk, or 50% or more carry any concern. Below that 3-study minimum, no downgrade applies — an unassessed body isn't quietly treated as low-risk.
02 — Inconsistency
Downgrades when designs don't match
If the cited evidence mixes study designs — or includes any source whose design couldn't be classified at all — certainty drops one level. Heterogeneity can't be computed from abstract text alone, so a mixed body is treated conservatively rather than assumed consistent.
03 — Imprecision
Downgrades on small or underreported samples
Certainty drops one level if the combined evidence covers fewer than 100 participants. Only when that check doesn't fire does a second check apply: no cited source reporting a confidence interval also costs one level. The two never stack — imprecision costs at most one level, decided against what's actually stated in the abstracts, not assumed.
04 — Publication Bias
Downgrades when a funnel plot isn't possible
A small evidence base — five or fewer studies, five or fewer of them randomized — can't support a meaningful publication-bias check. Rather than skip the question, Evidara downgrades for it.
05 — Indirectness
Never auto-applied
Judging whether a study's population or outcome actually matches the question asked requires a reviewer's read, not a keyword match. Evidara names this as a required GRADE domain and flags it for review — it does not score it automatically.
Flagged for human review, not scored
Risk of bias, by design

Three tools, one rule for
combining them.

Which risk-of-bias tool applies depends on the study design being assessed. Whichever tool runs, the same aggregation rule applies: the worst domain that could actually be assessed sets the overall verdict — and a domain marked "not assessed" is left out of that comparison entirely, never quietly counted as low risk.

Randomized trials

RoB 2

Five domains, judged from what an abstract states about trial conduct. A registered trial ID, for instance, supports a low-risk call on the reported-result domain — its absence doesn't count against the study, it just leaves that domain unassessed.
Randomization process
Deviations from the intended intervention
Missing outcome data
Measurement of the outcome
Selection of the reported result
Observational studies

ROBINS-I

Seven domains for cohort, case-control, and other non-randomized designs. Confounding is treated differently here than in RoB 2: it defaults to high risk unless the abstract explicitly states an adjustment — observational evidence carries that risk until told otherwise, not the reverse. Two of the seven — marked below — aren't judged from abstract signals at all; they're always left not-assessed rather than guessed.
Confounding
Selection of participants
Classification of interventions*
Deviations from intended interventions*
Missing data
Measurement of outcomes
Selection of the reported result
* always not-assessed — no abstract signal drives these two currently
Systematic reviews

AMSTAR-2 (partial)

Systematic reviews and meta-analyses are checked against 6 of AMSTAR-2's 16 items — the ones an abstract can actually speak to, like whether a protocol was pre-registered or risk of bias in the included studies was assessed. The other 10 require the review's full methods section and are left unassessed rather than guessed.
The honest part

Why certainty so often lands at Moderate, not High

Two structural rules combine to make Moderate the most common outcome, even for evidence that includes strong trial data. First: when an answer draws on both randomized and observational sources, the overall certainty follows the weaker pool, not the stronger one. Second: the observational ceiling described above means any pool of real-world evidence tops out at Moderate regardless of effect size. Put those together, and a mixed evidence base rarely reaches High — not because the rules are broken, but because that's what a mixed evidence base honestly supports.

Illustrative — not a live query result. A body of evidence with 3 randomized trials (High-starting) and 2 observational cohort studies (Low-starting, Moderate ceiling), cited together in one answer.
Certainty follows the weaker pool across the combined answer, not the stronger one
The observational pool's ceiling is Moderate — a large effect can move it up one level, never past Moderate
Combined answer-level certainty: Moderate — even though 3 of the 5 cited sources are randomized trials
Illustrative example, not a real run · reflects the rating logic as implemented
Every citation, checked live

Two verification systems,
both checking against the source.

A rating is only as trustworthy as the citations under it. Evidara runs two separate, live verification paths — one over generated content, one over chat and retrieval — because they catch different failure modes.

Generated content

Agent-output validation

PMIDs extracted from generated text are checked live against NCBI as the output is produced.
An unverifiable citation is never silently dropped — it's tagged inline, right where it appears
If the live NCBI check itself can't complete — a network timeout, for instance — a citation defaults to valid rather than getting flagged. That's the opposite failure direction from the system below, on purpose: this check runs inline during generation and is built to not block output on an outage
Chat & retrieval

Reference verification

PMIDs and NCT numbers are checked live against NCBI and ClinicalTrials.gov, and cached only once resolved.
A failed check is never cached, so a temporary outage can't lock in a wrong answer
An identifier is marked verified because the API confirmed it — never because a check couldn't be completed
Citations the model writes itself are checked against the sources actually supplied to it; anything outside that set is flagged to you, not rendered as fact
Scope & limitations

This is a scaffold,
not a verdict.

GRADE was designed for panels of methodologists working from full-text access to every included study. Evidara's implementation approximates that logic mechanically, from what's stated in an abstract — it is not a substitute for that panel, and it doesn't claim to be. Every automated rating is marked as automated. Where a judgment genuinely requires a person — indirectness, full-text-only risk-of-bias items, the ten AMSTAR-2 items an abstract can't answer — Evidara says so and stops, rather than filling the gap with a guess.

The certainty rules themselves are deterministic — the same cited evidence always produces the same rating. The evidence set is not frozen: live retrieval means re-running a query later can surface new literature, and LLM-synthesized narrative is not reproduced verbatim between runs.

When to use something else

Five places to use
a different tool.

Evidara is built for evidence synthesis from published literature. That's a specific job, not every job — five places where a different tool is the right call:

01 — Patient-Level Data
Literature and safety signals, not patient records
Evidara draws on published literature and FAERS adverse-event signals — not patient charts, claims data, or cohort databases. Patient-level outcomes and real-world treatment-pattern analysis are a different tool's job.
02 — Screening Model
AI-assisted with human override, not blinded dual review
Screening is AI-driven, with every decision open to human override. A regulatory submission that requires two independent blinded human screeners with a formal inter-rater reliability check needs that specific workflow — Evidara's model is built for speed with a human in the loop, not blinded duplication.
03 — Absent Evidence
When the evidence isn't there, Evidara says so
GRADE certainty can land at Insufficient. Extraction can return Data Absent rather than reach into an abstract for something it doesn't say. Evidara reports what the literature does and doesn't establish — it doesn't generate evidence that doesn't exist.
04 — Publication-Ready Review
A defensible starting point, not a finished methodologist's review
Every output — including a Full Systematic Review — is marked Provisional and flagged for human review before external or decision use. That's the right starting point for a submission-track review, not a substitute for the methodologist team that finishes one.
05 — Full-Text Depth
Abstracts and open access, not every paywall
Full-text hydration runs through Unpaywall's open-access index. What's open, Evidara reads in full. What sits behind a publisher paywall with no open-access copy, it doesn't reach — and says so rather than reasoning from the abstract alone.
Bounded by open-access availability
See it in context

Run it on your
own question.

The clearest way to understand a grading system is to watch it grade something you actually care about.