π» Code implementation for the Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains paper.
This repository contains the ICU-REACT benchmark, the Clin-REACT model family, and the training and evaluation code accompanying the manuscript.
Critical care decision-making requires physicians to first decide what information to retrieve and then reason over it. Most medical LLM benchmarks skip the first step, presenting models with a curated multiple-choice stem. ICU-REACT evaluates both: the retrieval of clinically relevant variables and the reasoning built on top of them.
The project has two components:
- ICU-REACT β a clinician-curated benchmark for information retrieval and clinical reasoning in critical care. Cases were curated with 19 clinicians and scored against rubrics. Retrieved variables are aligned to an OMOP-based taxonomy spanning 8 domains (diagnosis, physiology, laboratory values, imaging, medications, interventions, severity scores and clinical assessments, and administrative variables).
- Clin-REACT β a family of LLMs fine-tuned on an augmented version of the ICU-REACT training data, released at 8B, 14B, 31B, and 70B parameter scales. Models are available on Hugging Face (links below).
| Variant | Backbone |
|---|---|
| Clin-REACT 8B | Llama 3.1 8B Instruct |
| Clin-REACT 14B | Baichuan M1 14B Instruct |
| Clin-REACT 31B | Gemma 4 31B |
| Clin-REACT 70B | Llama 3.3 70B Instruct |
- Clin-REACT improves performance across every clinical reasoning benchmark tested, achieving the best results among the open-source models evaluated.
- Gains are not confined to critical care: improvements on ICU-REACT categories transfer to categories in MedRBench and VivaBench, indicating cross-taxonomy transfer rather than in-domain memorization.
- Stronger multiple-choice benchmark performance does not reliably indicate stronger clinical reasoning. Several medical LLMs appear tuned toward multiple-choice formats without corresponding gains on retrieval-and-reasoning tasks.
Reasoning and retrieval benchmarks: ICU-REACT, SCT-Bench, ER-Reason, MedRBench, VivaBench.
Multiple-choice medical benchmarks: MedQA, MedMCQA, MedXpertQA, MMLU-Med.
icureact/ Python package (installed via `pip install -e .`)
create_seed_launch.py Stage 1: seed dataset creation (CLI)
dataset_augmentation_launch.py Stage 2: dataset augmentation (CLI)
redundancy_analyzer_launch.py Stage 3a: near-duplicate removal (CLI)
llm_council_curator_launch.py Stage 3b: LLM-council difficulty labeling (CLI)
filter_dataset_by_difficulty.py Stage 3c: keep only the target difficulty (CLI)
run_curation_pipeline.py Runs 3a -> 3b -> 3c as one command
pipeline_launch.py Stage 4: LoRA training + evaluation (CLI)
evaluation_launch.py Stage 5: baseline/zero-shot evaluation (CLI)
create_seed/ Core functions used by create_seed_launch.py
dataset_augmentation/ Core functions used by dataset_augmentation_launch.py
evaluation/ One evaluator class per benchmark
trainer/, models/, rewards/ SFT/GRPO training internals
prompts/ Prompt templates (by dataset_version)
configs/ model_configs.json and other small runtime configs
path_utils.py ${PROJECT_DIR}/${LLM_DIR} placeholder expansion
run_configs/ Per-experiment JSON configs, one subfolder per stage
templates/ Clean example configs to copy for your own runs
notebooks/ Analysis, figure-generation, and Gradio apps (not a package)
scripts/ Misc one-off utilities
requirements.txt Core Python dependencies
requirements-cuda.txt Optional exact CUDA 12.8 wheel pins
pyproject.toml Packaging (pip install -e .)
git clone <this-repo> icureact
cd icureact
python -m venv .venv && source .venv/bin/activate # or your preferred env manager
pip install -r requirements.txt
# GPU users who want the exact frozen CUDA 12.8 wheel pins used to train Clin-REACT:
# pip install -r requirements-cuda.txt
pip install -e .Requires Python >= 3.10. requirements.txt is a broad list covering every pipeline stage (LLM API clients, training, evaluation, analysis, Gradio apps) β trim it if you only need a subset.
Create a .env file at the repo root (loaded automatically via python-dotenv):
| Variable | Used for | Notes |
|---|---|---|
PROJECT_DIR |
Base directory for datasets/, annotations/, models/, outputs/, analyses/ and for expanding ${PROJECT_DIR} in config files |
Defaults to .; set this to your repo checkout path |
LLM_DIR |
Base directory for local HuggingFace model weights, expanded via ${LLM_DIR} in icureact/configs/model_configs.json |
Defaults to . |
MODEL_CONFIGS_PATH |
Path to the model registry | Defaults to icureact/configs/model_configs.json |
NAVIGATOR_API_KEY |
LLM calls routed through "environment": "navigator" in model_configs.json |
Navigator is a University of Florida-internal LLM gateway; external users should either request the config for their own gateway or repoint model_configs.json entries at "environment": "azure" / "hf" (see below) |
NAVIGATOR_BASE_URL |
Navigator endpoint | Defaults to https://api.ai.it.ufl.edu |
AZURE_API_KEY_REGION2 |
LLM calls routed through "environment": "azure" |
|
AZURE_OPENAI_ENDPOINT, AZURE_OPENAI_API_VERSION |
Azure endpoint/API version | |
WANDB_API_KEY, WANDB_PROJECT |
Training run logging via --monitor in pipeline_launch.py |
Optional |
EXT_BENCHMARKS_DIR |
Root directory for external benchmark data (SCT-Bench, ER-Reason, MedRBench, VivaBench) | Defaults to external_benchmarks |
If you don't have access to the UF Navigator gateway, every script that takes --environment/--env also supports azure (set the two Azure variables above) or hf (point --hf-model-path/model_configs.json at local weights) β there is no hard dependency on Navigator specifically.
- Clin-REACT checkpoints are on Hugging Face β see the model table in Overview.
- ICU-REACT seed and training datasets are on Hugging Face: https://huggingface.co/datasets/macontreras98/ICU-REACT. You can start reproduction from here without running Stage 1 (seed creation): use the seed set to start from Stage 2, or use the pre-built training sets directly to skip straight to Stage 4.
- ICU-REACT Test (the held-out evaluation set) is also on the same Hugging Face dataset: https://huggingface.co/datasets/macontreras98/ICU-REACT.
- Raw clinician annotations (the input to seed creation, i.e. Stage 1 below) are not publicly available at this time.
Each benchmark must be obtained directly from its own source below β none of them are redistributed in this repo.
| Benchmark | Source | Where it goes |
|---|---|---|
| ICU-REACT Test | https://huggingface.co/datasets/macontreras98/ICU-REACT | Same dataset as the seed/training sets above β no separate download needed |
| SCT-Bench | https://github.com/SCT-Bench/sctpublic | $EXT_BENCHMARKS_DIR/sct_bench/ |
| ER-Reason | https://physionet.org/content/er-reason/1.0.0/ (requires PhysioNet credentialed access) | $EXT_BENCHMARKS_DIR/er_reason/ |
| MedRBench | https://github.com/MAGIC-AI4Med/MedRBench | $EXT_BENCHMARKS_DIR/medrbench/ |
| VivaBench | https://huggingface.co/datasets/chychiu/VivaBench | $EXT_BENCHMARKS_DIR/vivabench/ |
$EXT_BENCHMARKS_DIR defaults to external_benchmarks/ β see Installation & environment setup.
All commands below assume PROJECT_DIR is set to your repo checkout (or omit it and run everything from the repo root, since it defaults to .). Example configs for every stage live in run_configs/templates/ β copy one and edit the paths/models rather than starting from scratch.
create_seed_launch.py turns clinician annotation exports into the ICU-REACT-Train/Test seed sets (Methods Β§Seed dataset creation: per-round filtering β topic mapping β feedback-guided refinement β train/test split β judge rubrics). Each stage is a separate --mode, since the pipeline mixes cheap local steps with expensive LLM calls you may want to re-run independently.
Note: This stage requires the raw clinician annotation exports (
annotations/annotations.json,annotations/per_feedback_classification.jsonl) as input. These are not publicly available at this time. If you don't have access to them, skip this stage and start reproduction from the already-built seed set on Hugging Face (see Dataset augmentation below).
python -m icureact.create_seed_launch --mode split_by_round \
--annotations-path annotations/annotations.json
python -m icureact.create_seed_launch --mode filter_by_score
python -m icureact.create_seed_launch --mode map_topics
python -m icureact.create_seed_launch --mode refine_from_feedback
python -m icureact.create_seed_launch --mode add_stepwise_reasoning
python -m icureact.create_seed_launch --mode consolidate_round
python -m icureact.create_seed_launch --mode merge_seed
python -m icureact.create_seed_launch --mode split_seed
python -m icureact.create_seed_launch --mode create_judge_rubrics
# ICU-REACT-Train refinement (Fig. 1C, train branch): combines the original
# reasoning + clinician feedback into a refinement explanation + improved reasoning
python -m icureact.create_seed_launch --mode refine_reasoning
# ICU-REACT-Test refinement (Fig. 1C, test branch): feeds approved
# questions/retrieval tasks + feedback directly to the refinement agent
python -m icureact.create_seed_launch --mode refine_context_question
# Or run every stage above in order:
python -m icureact.create_seed_launch --mode all --annotations-path annotations/annotations.jsonConfig: datasets/icureact_seed/{version}/config.json (see run_configs/templates/01_create_seed_config.json) sets topic_model/refinement_model/rubric_model. --exclude-annotator-ids-json and --per-feedback-jsonl-path let you supply the clinician-annotation-round-specific inputs described in the paper without hardcoding them into the script. Run python -m icureact.create_seed_launch --help for the full flag list.
Output: datasets/icureact_seed/{version}/processed/merged/{train,test}_seed_dataset.jsonl plus the refined/rubric variants β consumed next by dataset augmentation.
dataset_augmentation_launch.py expands the seed set into the larger synthetic training corpus (Methods Β§Dataset augmentation): new context-question pairs, initial reasoning, reasoning refinements, and judge rubrics, all guided by seed-derived few-shot examples.
The clinician-validated ICU-REACT seed set (the output of Stage 1) is available pre-built on Hugging Face: https://huggingface.co/datasets/macontreras98/ICU-REACT β download it into datasets/icureact_seed/{version}/processed/merged/ to start here without running Stage 1 yourself.
You can also skip this stage entirely: the same Hugging Face dataset includes the pre-built, curated training sets that are otherwise the output of Stage 2 + Stage 3 below β download those directly and go straight to Stage 4.
# Generate context-question pairs (--size-target is the target sample count;
# shard across parallel jobs with --num-batches/--batch-id for large runs)
python -m icureact.dataset_augmentation_launch \
--config run_configs/templates/02_dataset_augmentation_config.json \
--mode pair_gen_only --n-per-record 10 --size-target 5000 \
--num-batches 1 --batch-id 0
# Merge batch shards into one file
python -m icureact.dataset_augmentation_launch \
--config run_configs/templates/02_dataset_augmentation_config.json \
--mode combine_context --dataset-size 5000
# Reasoning-refinement subset (used for the LoRA training target)
python -m icureact.dataset_augmentation_launch \
--config run_configs/templates/02_dataset_augmentation_config.json \
--mode reasoning_refinement_only --reasoning-refinement \
--n-per-record 10 --size-target 5000 --num-batches 1 --batch-id 0
python -m icureact.dataset_augmentation_launch \
--config run_configs/templates/02_dataset_augmentation_config.json \
--mode combine_reasoning_refinement --reasoning-refinement --dataset-size 5000For large runs, --num-batches/--batch-id shard generation across parallel jobs (e.g. a SLURM array), then a --mode combine_* call merges the shards. See dataset_augmentation.sh / dataset_augmentation2.sh for worked SLURM examples of this batching pattern, and python -m icureact.dataset_augmentation_launch --help for every mode.
Output: datasets/icureact_augmented/{version}/dataset_{size}/reasoning_refinement/augmented_train_reasoning_refinement_subset.jsonl (one file per size bucket, e.g. 5k/10k/20k) β consumed next by curation.
Three scripts run in sequence, linked only by a shared filename convention (*_filtered.jsonl β *_filtered_difficulty.jsonl β *_filtered_hard.jsonl). run_curation_pipeline.py runs all three as one command and derives each stage's input path from the previous stage's output automatically:
python -m icureact.run_curation_pipeline \
--redundancy-config run_configs/templates/03_redundancy_analyzer_config.json \
--council-config run_configs/templates/04_llm_council_curator_config.json \
--filter-config run_configs/templates/05_filter_dataset_by_difficulty_config.jsonWhat's happening under the hood (run these individually if you need more control)
# 3a. Redundancy dedup: vectorize + TF-IDF nearest-neighbor group + ROUGE-L >= 0.7 removal
python -m icureact.redundancy_analyzer_launch \
--config run_configs/templates/03_redundancy_analyzer_config.json
# 3b. LLM council difficulty labeling: >=3 models vote easy/medium/hard per sample
python -m icureact.llm_council_curator_launch \
--config run_configs/templates/04_llm_council_curator_config.json
# 3c. Keep only "hard" samples (pure post-processing, no LLM calls)
python -m icureact.filter_dataset_by_difficulty \
--config run_configs/templates/05_filter_dataset_by_difficulty_config.jsonOutput: *_filtered_hard.jsonl β the final ICU-REACT-Train dataset, referenced as data.train_dataset.train_jsonl_path in a pipeline config.
pipeline_launch.py fine-tunes a base model with LoRA on the curated dataset (Methods Β§Model training: assistant-only loss over the refinement-explanation + improved-reasoning target) and evaluates the resulting checkpoint:
python -m icureact.pipeline_launch --config run_configs/templates/06_pipeline_config.jsonmodel.namemust match an entry inicureact/configs/model_configs.json.pipeline.phasessupports multiple chained training phases (each phase's checkpoint feeds the next); set"skip": trueon a phase to evaluate an already-trained checkpoint without retraining.--skip-trainre-runs evaluation only, against a previously trained model.--do-inference/--do-variable-metrics/--do-rubric-judgetoggle evaluation sub-stages independently, so a re-score doesn't require re-running inference.
Output: models/{dataset_bucket}/{model_bucket}/{run_tag}/{phase_tag}/... (LoRA checkpoint) and benchmarks/{benchmark}/... (evaluation results).
evaluation_launch.py (also called internally by pipeline_launch.py) scores any model β zero-shot backbone or trained checkpoint β across the benchmark suite:
python -m icureact.evaluation_launch --config run_configs/templates/07_eval_config.jsonbenchmarks accepts any of: icureact, sct_bench, er_reason, medrbench, vivabench, multiple_choice_bench. models lists entries from model_configs.json.
run_configs/{stage}/β the authors' actual experiment configs, kept for transparency (hundreds of files spanning every model/dataset-size/ablation combination reported in the paper). Not meant as a starting point.run_configs/templates/β clean, minimal example configs (one per pipeline stage) using${PROJECT_DIR}placeholders; copy these for your own runs.${PROJECT_DIR}/${LLM_DIR}placeholders are expanded automatically by every launch script's config loader (icureact/path_utils.py) β safe to use inside any JSON config value.pipeline_config.json(used bypipeline_launch.pyandevaluation_launch.py) supports two schemas: a legacy single-phase form (run/model/data/train/evaltop-level keys) and a multi-phase form (run/model/data/pipeline, wherepipeline.phasesis a list merged againstpipeline.train_defaults/eval_defaults). Seeicureact/config/pipeline_config.pyfor the validator andrun_configs/templates/06_pipeline_config.json/07_eval_config.jsonfor examples of each.icureact/configs/model_configs.jsonis the model registry: each entry sets"environment"(navigator|azure|hf) and, forhfentries, anhf_model_path(optionally with apeft_model_pathfor a LoRA adapter on top of a base model). Add your own models here to point the pipeline at different backends.
These scripts consume benchmarks//dataset output and reproduce specific figures/tables β they are not part of the reproduction chain above, but read its outputs.
| Script | Produces |
|---|---|
icureact/benchmark_statistical_analysis_focused.py |
Paired bootstrap + Wilcoxon signed-rank comparisons (Methods Β§Statistical Analyses). Canonical script β .py/_selected.py are earlier/alternate variants, not the ones behind the reported numbers. |
icureact/category_stratify.py |
Supplemental "Performance stratified by clinical content" |
icureact/summarize_icureact_variable_coverage.py, summarize_icureact_rubric_dimensions.py |
Supplemental dataset/rubric statistics (variable coverage, rubric item counts) |
notebooks/task_ablation.py, plot_task_ablation.py, plot_size_ablation.py |
Supplemental Ablations (training-task composition and dataset-size scaling) |
notebooks/frontier_comparison.py |
Supplemental "Comparison with Frontier Models" |
notebooks/generate_seed_plots.py, generate_train_dataset_figures.py, generate_rubrics_figure.py, umap_visualize.py, questions_umap.py |
t-SNE/UMAP dataset and rubric figures (seed creation, train dataset, rubric stats) |
icureact/benchmark_visualizer.py, model_comparison_visualizer.py, multiple_choice_benchmark_visualizer.py, sct_model_comparison_app.py |
Gradio leaderboard / sample-comparison apps for interactive exploration |
icureact/data_contamination_analysis.py (embedding/exact-overlap check), icureact/data_contamination_completion.py (completion-based leakage probe) |
Data contamination analysis reported in the paper |
Remaining one-off scripts under notebooks/ not listed above (annotation-feedback tooling, OMOP variable/topic mapping utilities, etc.) were used during dataset/benchmark development but do not correspond to a specific reported figure or table.
The commands above (python -m icureact.<script> --config ...) work anywhere and are the recommended way to run each stage. If you have access to a SLURM cluster, .slurm templates (run_eval.slurm, run_train.slurm, run_pipeline_launch.slurm) and submit_*.sh wrappers are also provided; *_local.sh gives non-SLURM nohup equivalents. All of these resolve the repo root via $PROJECT_DIR (or the script's own location) β you'll still need to adjust the conda activate <env> and module load cuda/... lines to match your own cluster's environment, and create the outputs/ directory before submitting (#SBATCH --output= paths are resolved relative to your working directory at submission time and are not auto-created).
Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena, Jiaqing Zhang, Heng Sun, Hruday Tej Akkaladevi, Peiyu Lu, Jordan Rosen, Sumit Kapoor, Sasank Desaraju, Grace R. Thompson, Jacob Purcell, Michael Petrauskis, Philip KW Hong, Meghan Brennan, Sarah Chrabaszcz, Tierra Smith, Ronnie Ren, Michel S. Kabbash, Ceyhun Haziroglu, Rushi Patel, Gabriel Gomez, Charlotte Chaiklin, Randy Leung, Kenneth N. John, Whitman Wiggins, Philip Kayser, Vincent Bird, Maria Bruzzone, Tyler J. Loftus, Azra Bihorac, and Parisa Rashidi. "Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains." arXiv:2608.22622 (2026). https://arxiv.org/abs/2608.22622
Corresponding author: Parisa Rashidi (parisa.rashidi@ufl.edu)
@article{contreras2026icureact,
title = {Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains},
author = {Contreras, Miguel and Siegel, Scott and Nerella, Subhash and Sena, Jessica and Zhang, Jiaqing and
Sun, Heng and Akkaladevi, Hruday Tej and Lu, Peiyu and Rosen, Jordan and Kapoor, Sumit and
Desaraju, Sasank and Thompson, Grace R. and Purcell, Jacob and Petrauskis, Michael and
Hong, Philip KW and Brennan, Meghan and Chrabaszcz, Sarah and Smith, Tierra and Ren, Ronnie and
Kabbash, Michel S. and Haziroglu, Ceyhun and Patel, Rushi and Gomez, Gabriel and Chaiklin, Charlotte and
Leung, Randy and John, Kenneth N. and Wiggins, Whitman and Kayser, Philip and Bird, Vincent and
Bruzzone, Maria and Loftus, Tyler J. and Bihorac, Azra and Rashidi, Parisa},
journal = {arXiv preprint arXiv:2608.22622},
year = {2026},
doi = {10.48550/arXiv.2608.22622},
url = {https://arxiv.org/abs/2608.22622}
}GNU General Public License v3.0. See LICENSE / https://www.gnu.org/licenses/gpl-3.0.en.html for the full text.
