PHI INPUTS: PROHIBITEDPHASE: 0 — NOT FOR PHIBAA TIER: IN BUILD
Status

STATUS: DRAFT PROTOCOL · IN DEVELOPMENT · NO RESULTS PUBLISHED

Evaluation protocol. Draft.

This page describes the evaluation protocol we are developing. It is a draft. We have not completed evaluations under it. No model is scored, ranked, compared, or recommended anywhere on this site. None will be until the protocol is final, the results survive it, and per-provider publication rights are cleared.

Context

Multiple-choice medicine is a solved test, not a solved problem.

Static multiple-choice medical benchmarks are saturated: an independent leaderboard archived MedQA after nearly all recently released models scored above 95 percent. vals.ai/benchmarks/medqa · 2026-07-16

The field has moved to physician-rubric, open-ended evaluation and to task taxonomies validated with clinicians. HealthBench grades multi-turn health conversations against rubrics written by 262 physicians. MedHELM evaluates against a clinician-validated taxonomy of 121 tasks developed with 29 clinicians. arXiv 2505.08775 / 2505.23802 · 2026-07-16

One caveat travels with all of it: the dominant open health benchmarks are graded by the vendor's own models and their headline results are self-reported. Published meta-evaluation against physician panels mitigates that; it does not eliminate it.

A benchmark score is a property of a model snapshot under a specific evaluation configuration. It is not a property of the model, and it is never evidence of clinical safety or regulatory fitness.

Protocol

What we are building instead.

Each element below is a design commitment of the draft. Nothing here describes completed work.

Reviewer roles
Evaluations are designed and adjudicated by licensed physicians with current or recent clinical practice. Reviewer identity, specialty, and practice context are recorded per evaluation. Engineering staff prepare harnesses; they do not grade clinical content.
Conflicts policy
Reviewers disclose financial and professional relationships with any model provider under evaluation. A conflicted reviewer does not grade that provider's outputs. Disclosures are versioned with the protocol.
Task provenance
Tasks are authored from real clinical workflows, documented as to origin and construction, and reviewed for construct validity before use. No task is drawn from a public benchmark's test set, to avoid contamination.
Adjudication
No single-reviewer verdicts on contested items. Disagreements are resolved by a documented multi-physician adjudication step, and the disagreement rate itself is recorded and will be reported.
Grader validation
If any automated grader is used, it is meta-evaluated against physician panels before use, and its agreement statistics are published alongside anything it grades. A grader that has not been validated against physicians grades nothing we publish.
Uncertainty
Published results will carry uncertainty, not just point estimates: sample sizes, agreement rates, and the protocol version they were produced under.
Versioning
The protocol is versioned. Results are bound to the protocol version and model snapshot that produced them, with dates. Numbers produced under different versions are not comparable and will not be presented as comparable.
Correction, retraction
A published result that fails re-verification is corrected or retracted in place, with the change dated and the original preserved. The correction channel is customerservice@healthit.com.
Publication rights
Before any result naming a provider is published, that provider's terms are reviewed for benchmark-publication restrictions and a rights dossier is completed. No dossier, no publication.
Score context
Any score this site ever publishes will state, inline: benchmark or task-set version, grader model and settings, scoring adjustments, evaluated model snapshot and date, and who ran the evaluation.FORMAT (ILLUSTRATIVE, NO REAL VALUES SHIP UNTIL RESULTS EXIST): PROTOCOL vX · GRADER: MODEL+SETTINGS · SNAPSHOT: MODEL_VERSION · RUN: DATE · RUN BY: OPERATOR

This draft will change. When it does, the version number changes with it, and this page keeps its history.