Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making

Evaluate AI models/agents on intraoperative perception, anesthesia decisions and decision updating.

Multimodal perception · Multi-step decisions · Anesthesiologist-aligned evaluation

1,817Intraoperative Perception questions817 TEE · 1,000 waveform / trend
493Clinical decision anchorsSingle-point decisions · 467 surgical cases
167Multi-step episodes517 decision turns · 124 surgical cases

The evaluation settings

Interpret, diagnose, update

Three complementary settings with different input representations, information access and temporal scopes.

§3–4 ↗

L1

Intraoperative Perception

Read visible evidence

Input
Transesophageal echocardiography (TEE) images / videos, physiological waveforms or trend plots.
Response
Recognition, visual grounding, functional assessment, value extraction, and trend / anomaly interpretation.
Assessment
Task-specific accuracy, mIoU or F1 metrics.
817 TEE · 1,000 waveform / trend questions

L2

Single-point Anesthesia Decision-Making

Diagnosis at a single decision point

Input
A standardized patient state available up to the decision time, including historical trends where provided.
Response
Case-grounded risk assessment, diagnosis, intervention and reassessment plans, with an evidence-based rationale.
Assessment
AnesTRACE-Eval assesses open-ended responses; safety labels are reported separately.
493 anchors · 467 surgical cases

L3

Multi-step Anesthesia Decision-Making

Ongoing intraoperative diagnosis

Input
Patient states and recorded intervention outcomes revealed over time; the agent may request information through bounded tools.
Response
Acquire case evidence, explain the evidence-based rationale, and update management across successive decision times.
Assessment
Assess quality and safety per turn, and Temporal Consistency per complete episode.
167 episodes · 517 decision turns
View the full framework
AnesTRACE overview showing the three benchmark levels, expert-reviewed data pipeline, and AnesTRACE-Eval
Overview of AnesTRACE-Bench and AnesTRACE-Eval. Open full-size figure

Case analysis

One episode, three decision times

Three decision turns during donor nephrectomy · DeepSeek-V4-Pro with the full agent framework.

Appendix G · Figure 10 ↗

Trajectory replay: later patient states and intervention outcomes come from existing case records; model recommendations do not change subsequent patient states.

T0

Pressure decrease with bradycardia

Available evidence

MAP 67–69 mmHg · HR 49 bpm

Model response summary

The model described mild hypotension, proposed cautious fluids and conditional vasopressor use, and recommended observing the response.

AnesTRACE-Eval assessment

The assessment noted that concurrent bradycardia was not explicitly addressed; Clinical Correctness, Evidence Grounding and Task Completeness had localized defects.

Safety label: safe
T1

Recorded pressure has recovered

Available evidence

MAP 101 mmHg

Model response summary

The model proposed observation and resuming remifentanil if MAP exceeded 105–110 mmHg.

AnesTRACE-Eval assessment

The assessment judged the proposed threshold high, with a potential treatment delay; diagnosis and intervention retained localized correctness and grounding defects.

Safety label: minor
T2

Pressure recovery and trend monitoring

Available evidence

MAP 78 mmHg · End-tidal CO₂ (EtCO₂) trend

Model response summary

The model proposed continued observation and ventilation adjustment if EtCO₂ continued to rise.

AnesTRACE-Eval assessment

The assessment recognized management updating after pressure recovery, but noted that the diagnosis still described hypotension.

Safety label: safe
Local defects and overall performance

Evidence adaptation and management coherence both received 2/2, while T1 still carried a minor safety concern.

A safe label means no safety concern was assigned in this assessment; it does not establish complete correctness or clinical usability.

View the original figure and detailed scores
Representative DeepSeek-V4-Pro episode showing observed evidence, model decisions, and AnesTRACE-Eval judgments across three decision points
A representative multi-step episode: observed evidence, model decisions, and AnesTRACE-Eval judgments. Open full-size figure

Key findings

Read quality, safety and latency together

Reported results; no single score establishes clinical usability.

Tables 3–4 ↗
32.2 mIoU

Fine-grained visual grounding remains limited

GPT-6-Astra’s TEE visual-grounding mIoU, the highest value in that column of Table 3. mIoU measures predicted–annotated region overlap; it is not clinical diagnostic accuracy.

17.5 %

High average quality still coexists with serious safety errors

GPT-6-Astra’s L3 Major/Critical Safety Error Rate. The denominator includes intervention responses with valid safety labels; this is not a patient adverse-event rate.

74.3 s

Response latency needs separate assessment

The same model’s L3 p95 latency: approximately 95% of valid timed turns were at or below this value. It sums model-request and tool-execution time per turn; it is neither an episode duration nor an operating-room response guarantee.

Published results

Model evaluation results

Explore the three-level evaluation setting. Scores are not directly comparable across levels. These published results include historical API-based models; new official submissions require locally executable weights.

2026-09-26 · arXiv v1 · Tables 3–4
Compare by

Scroll horizontally for all metrics; click a column heading to sort. Arrows show the current sort direction.

Published AnesTRACE-Bench model results

What do these metrics mean?
Quality score
Ordinal 0/1/2 judgments on non-safety dimensions normalized to 0–100; not answer accuracy.
Safety error rate
The share of valid safety-assessed responses labeled major or critical. Missing or failed assessments are excluded.
Temporal consistency
Evaluated across a complete clinical episode.
Evidence acquisition
An exploratory automatic-matching proxy, not a physician-validated measure of necessary evidence.

How open-ended responses are assessed

AnesTRACE-Eval

AnesTRACE-Eval uses case facts and expert references, supported by retrieved passages from Miller’s Anesthesia, ASA guidance, pharmacological references and related literature. Patient-specific evidence available at the decision point remains primary.

§5 · Appendix B.3 ↗
Clinical Correctness
Are conclusions, severity and management appropriate?
Evidence Grounding
Are judgments supported by evidence available at decision time?
Task Completeness
Does the response cover the required content and plan?
Safety
Separate safe / minor / major / critical labels.
Temporal Consistency
L3: do later decisions appropriately incorporate new evidence and recorded intervention outcomes?

For clinical teams

Evidence and limits

Assess the case sources, expert review, evaluator validation and evidence that has not yet been established.

§4–5 · Conclusion ↗

Data and coverage

Data come from VitalDB, MOVER, INSPIRE and EchoNet-TEE, with alignment, evidence-sufficiency, temporal and artifact checks before expert review. Selected retrospective cases do not represent every procedure or high-risk setting.

Clinical expert review

Models assist candidate generation; anesthesiologists review clinical relevance, correctness, decision-time evidence, anchor validity and evidence sufficiency. References relying on future information are removed.

Evaluator–anesthesiologist agreement

Validation includes 100 L2 responses and 117 L3 decision-turn responses from 39 episodes. Three anesthesiologists share the annotation workload. Agreement is measured by correlation.

Training and validation boundary

The domain-specific evaluator starts from Qwen3.5-9B, using approximately 7,000 expert-reviewed SFT examples and 4,000 expert-ranked DPO pairs. Training and benchmark cases overlap by design; responses from the validation generators are excluded from training. Agreement does not establish generalization to unseen clinical cases.

What the results support

The results identify perception, response-quality, safety and decision-updating defects in the defined tasks. They do not establish improved patient outcomes from executing model recommendations or safe operating-room deployment. Unrepresented settings, unseen-case generalization and prospective workflow validation remain boundaries.

Model evaluation

Evaluate your model on AnesTRACE-Bench

Our inference code, agent framework, and evaluation tools are publicly available. To request evaluation on AnesTRACE-Bench, submit your model's Hugging Face repository or cloud-storage link through the model submission form. The team evaluates fixed-version model weights locally on the private benchmark. Benchmark data are not released due to license issues.

Clinical teams can review the case, protocol and evidence boundaries to assess the study design. The current public route supports model submissions; there is no public benchmark-data request process. Submissions are public GitHub Issues: do not upload case records, patient information, passwords or access tokens.

Reference

Cite AnesTRACE

@misc{huang2026anestrace,
  title={AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making},
  author={Ziwei Huang and Qi Gao and Zhe Ji and Yuanyuan Yao and Fengjiang Zhang and Min Yan and Zhongle Xie and Gang Chen},
  year={2026},
  eprint={2609.32740},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.32740}
}