How we certify the accuracy of clinical AI, and why the result holds up
MedAttest tests clinical AI against controlled synthetic encounters with fully known ground truth. Every assertion is graded one claim at a time, every omission is weighted by severity, contested calls are decided by clinicians, and tiers are awarded on the conservative bound of statistical confidence. The complete methodology is published below as a versioned document.
Current methodology document
MedAttest Methodology & Defensibility
v1.0.7 · published Jul 23, 2026 · 352 KB
Regenerated PDF without browser print headers/footers (file path, timestamp, page counter). Content unchanged.
How accuracy is measured
Every certification is scored on the same categories. They stay constant across versions of the methodology; each published document sets out the exact thresholds and procedures in force for that version.
Fabrication rate
Assertions the AI states that the source never supports, or that contradict it. The most heavily weighted signal, because invented facts are the most dangerous failure mode.
Severity-weighted omission rate
Required facts the AI dropped, weighted by severity from 1 (minor) to 4 (safety-critical). A missed high-stakes fact about a patient counts far more than a missed administrative detail.
Unsupported inference rate
Claims that reach beyond the evidence without being outright fabrications, such as turning a reported input into a stated conclusion.
Cross-patient containment
A scan for any detail from one patient appearing in another patient’s output. Reported as pass or fail and treated as a hard gate, never blended into any average.
Statistical confidence
Every rate is reported with a confidence interval, and tiers are awarded on the conservative bound, so a result is never rated stronger than the evidence supports.
Clinician adjudication
Clinicians review and can override every consequential judgment before a score is final, and the attestation is human signed.
Why this document matters
A certification is only worth as much as the process behind it. Clinical AI is being deployed on trust that buyers cannot independently verify: vendors grade themselves on private data, pilots have no ground truth, and generic benchmarks do not speak the language of operational risk. Publishing our methodology in full is how we earn that trust rather than asking for it. It lets a buyer, a clinician, a regulator, or an auditor check our work, compare vendors on identical terms, and understand precisely what a tier does and does not claim.
Independent and transparent
The full criteria are public, with no private scoring adjustments. Anyone can read exactly how a result was produced before they trust it.
Scientifically grounded
Outputs are tested against known ground truth, graded as individual claims, and reported with confidence intervals. Tiers gate on the conservative bound, so a result is never stronger than its evidence.
Human authority
A clinician has final say on every consequential judgment, and the attestation is human signed. The machine accelerates the work; it never has the last word.
Immutable and versioned
The standard in force at test time is recorded with every result, and scoring rules are frozen into each run. Re-tuning the standard can never silently rewrite a past certification.
Designed to be audited, not taken on faith
Because the standard is versioned and every result records the exact version it was scored under, a reader can always reconstruct the rules that applied to any certification, even years later. Nothing is hidden in a black box and nothing changes retroactively. This page exists so the reasoning behind every MedAttestattestation is open to scrutiny.
Version history
Earlier published versions remain available. The version in force when a certification was issued is recorded with that attestation, so a result keeps its original meaning for the life of the record.
- Download ↗
v1.0.7MedAttest Methodology & Defensibility
Modified version number to be included on the documentation
published Jun 14, 2026 · 374 KB
- Download ↗
v1.0.0MedAttest Methodology & Defensibility
Our initial offering
published Jun 13, 2026 · 375 KB
This methodology document is an independent description of the MedAttest evaluation process. It is not a government or regulatory approval. All evaluations use synthetic cases and no real patient data is ever used.