Science and methods · reviewed 19 Aug 2026

How medical AI agents should be evaluated

We separate a product hypothesis from demonstrated medical benefit. Each agent needs a defined purpose, user, data boundary, failure set, stopping criteria and professional oversight.

Scientific principles

Evaluate before deployment

Technical accuracy alone does not demonstrate clinical benefit. Evaluation covers the data, model, interface, workflow and the human–AI team.

01

Intended use and boundaries

Define one task, target user, permitted actions, and the conditions that require the agent to stop or hand off to a person.

02

Pre-specified metrics

Set completeness, false positives, omissions, time, source quality and critical stopping errors before testing starts.

03

The human–AI team

Measure how professionals interpret, check, correct and use the output—not only the model response.

04

Versioning and generalizability

Record the model, prompt, knowledge base and dataset version; do not assume results transfer across populations or sites.

Evidence ladder

Five separate stages

Each move requires a new protocol, approvals and an explicit decision by accountable specialists.

  1. 01

    Product hypothesis

    Describe the user, problem, expected action and unacceptable harm.

  2. 02

    Synthetic evaluation

    Test fictional cases, sources, refusals, escalation and interface resilience.

    Current stage
  3. 03

    Offline validation

    Use authorized data to assess pre-specified metrics, subgroups, bias and external applicability.

  4. 04

    Early clinical evaluation

    Under a protocol, study safety, human factors and small-scale operation in a live workflow.

  5. 05

    Comparative evaluation

    Evaluate process or outcome effects against a comparator, including unintended consequences.

Research programme

What should be measured

01

Clinical-trial navigation

Search recall, criterion-matching errors, unknown information, coordinator review time and the quality of handoff to a physician.

02

OMS, VMP and self-pay oncology pathways

Currency of official sources, clarity of the next step, document completeness and absence of covert treatment or provider selection.

03

Clinical copilot

Claims with sources, unsupported conclusions, critical omissions, review time and the quality of professional checking.

04

Operational agent

Missed actions, handoff quality, manual corrections, cycle time and newly introduced risk points.

Primary and official sources

Methodological foundation

Evidence boundary

What this page does not demonstrate

A methodology and source list do not mean that the project has completed a clinical study, is a registered medical device or is ready to process real health data. Those claims require separate evidence and approvals.