What we are building instead.
Each element below is a design commitment of the draft. Nothing here describes completed work.
- Reviewer roles
- Evaluations are designed and adjudicated by licensed physicians with current or recent clinical practice. Reviewer identity, specialty, and practice context are recorded per evaluation. Engineering staff prepare harnesses; they do not grade clinical content.
- Conflicts policy
- Reviewers disclose financial and professional relationships with any model provider under evaluation. A conflicted reviewer does not grade that provider's outputs. Disclosures are versioned with the protocol.
- Task provenance
- Tasks are authored from real clinical workflows, documented as to origin and construction, and reviewed for construct validity before use. No task is drawn from a public benchmark's test set, to avoid contamination.
- Adjudication
- No single-reviewer verdicts on contested items. Disagreements are resolved by a documented multi-physician adjudication step, and the disagreement rate itself is recorded and will be reported.
- Grader validation
- If any automated grader is used, it is meta-evaluated against physician panels before use, and its agreement statistics are published alongside anything it grades. A grader that has not been validated against physicians grades nothing we publish.
- Uncertainty
- Published results will carry uncertainty, not just point estimates: sample sizes, agreement rates, and the protocol version they were produced under.
- Versioning
- The protocol is versioned. Results are bound to the protocol version and model snapshot that produced them, with dates. Numbers produced under different versions are not comparable and will not be presented as comparable.
- Correction, retraction
- A published result that fails re-verification is corrected or retracted in place, with the change dated and the original preserved. The correction channel is customerservice@healthit.com.
- Publication rights
- Before any result naming a provider is published, that provider's terms are reviewed for benchmark-publication restrictions and a rights dossier is completed. No dossier, no publication.
- Score context
- Any score this site ever publishes will state, inline: benchmark or task-set version, grader model and settings, scoring adjustments, evaluated model snapshot and date, and who ran the evaluation.FORMAT (ILLUSTRATIVE, NO REAL VALUES SHIP UNTIL RESULTS EXIST): PROTOCOL vX · GRADER: MODEL+SETTINGS · SNAPSHOT: MODEL_VERSION · RUN: DATE · RUN BY: OPERATOR
This draft will change. When it does, the version number changes with it, and this page keeps its history.