Skip to content

Latest commit

Β 

History

86 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Clin-REACT logo

ICU-REACT & Clin-REACT

πŸ’» Code implementation for the Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains paper.

This repository contains the ICU-REACT benchmark, the Clin-REACT model family, and the training and evaluation code accompanying the manuscript.


πŸ“‹ Overview

Critical care decision-making requires physicians to first decide what information to retrieve and then reason over it. Most medical LLM benchmarks skip the first step, presenting models with a curated multiple-choice stem. ICU-REACT evaluates both: the retrieval of clinically relevant variables and the reasoning built on top of them.

The project has two components:

  • ICU-REACT β€” a clinician-curated benchmark for information retrieval and clinical reasoning in critical care. Cases were curated with 19 clinicians and scored against rubrics. Retrieved variables are aligned to an OMOP-based taxonomy spanning 8 domains (diagnosis, physiology, laboratory values, imaging, medications, interventions, severity scores and clinical assessments, and administrative variables).
  • Clin-REACT β€” a family of LLMs fine-tuned on an augmented version of the ICU-REACT training data, released at 8B, 14B, 31B, and 70B parameter scales. Models are available on Hugging Face (links below).
Variant Backbone
Clin-REACT 8B Llama 3.1 8B Instruct
Clin-REACT 14B Baichuan M1 14B Instruct
Clin-REACT 31B Gemma 4 31B
Clin-REACT 70B Llama 3.3 70B Instruct

πŸ”‘ Key findings

  • Clin-REACT improves performance across every clinical reasoning benchmark tested, achieving the best results among the open-source models evaluated.
  • Gains are not confined to critical care: improvements on ICU-REACT categories transfer to categories in MedRBench and VivaBench, indicating cross-taxonomy transfer rather than in-domain memorization.
  • Stronger multiple-choice benchmark performance does not reliably indicate stronger clinical reasoning. Several medical LLMs appear tuned toward multiple-choice formats without corresponding gains on retrieval-and-reasoning tasks.

πŸ“Š Evaluation suite

Reasoning and retrieval benchmarks: ICU-REACT, SCT-Bench, ER-Reason, MedRBench, VivaBench.

Multiple-choice medical benchmarks: MedQA, MedMCQA, MedXpertQA, MMLU-Med.


πŸ—‚οΈ Repository structure

icureact/                    Python package (installed via `pip install -e .`)
  create_seed_launch.py        Stage 1: seed dataset creation (CLI)
  dataset_augmentation_launch.py  Stage 2: dataset augmentation (CLI)
  redundancy_analyzer_launch.py   Stage 3a: near-duplicate removal (CLI)
  llm_council_curator_launch.py   Stage 3b: LLM-council difficulty labeling (CLI)
  filter_dataset_by_difficulty.py Stage 3c: keep only the target difficulty (CLI)
  run_curation_pipeline.py        Runs 3a -> 3b -> 3c as one command
  pipeline_launch.py              Stage 4: LoRA training + evaluation (CLI)
  evaluation_launch.py            Stage 5: baseline/zero-shot evaluation (CLI)
  create_seed/                    Core functions used by create_seed_launch.py
  dataset_augmentation/           Core functions used by dataset_augmentation_launch.py
  evaluation/                     One evaluator class per benchmark
  trainer/, models/, rewards/     SFT/GRPO training internals
  prompts/                        Prompt templates (by dataset_version)
  configs/                        model_configs.json and other small runtime configs
  path_utils.py                   ${PROJECT_DIR}/${LLM_DIR} placeholder expansion
run_configs/                  Per-experiment JSON configs, one subfolder per stage
  templates/                    Clean example configs to copy for your own runs
notebooks/                    Analysis, figure-generation, and Gradio apps (not a package)
scripts/                      Misc one-off utilities
requirements.txt              Core Python dependencies
requirements-cuda.txt         Optional exact CUDA 12.8 wheel pins
pyproject.toml                Packaging (pip install -e .)

βš™οΈ Installation & environment setup

git clone <this-repo> icureact
cd icureact
python -m venv .venv && source .venv/bin/activate   # or your preferred env manager
pip install -r requirements.txt
# GPU users who want the exact frozen CUDA 12.8 wheel pins used to train Clin-REACT:
# pip install -r requirements-cuda.txt
pip install -e .

Requires Python >= 3.10. requirements.txt is a broad list covering every pipeline stage (LLM API clients, training, evaluation, analysis, Gradio apps) β€” trim it if you only need a subset.

Create a .env file at the repo root (loaded automatically via python-dotenv):

Variable Used for Notes
PROJECT_DIR Base directory for datasets/, annotations/, models/, outputs/, analyses/ and for expanding ${PROJECT_DIR} in config files Defaults to .; set this to your repo checkout path
LLM_DIR Base directory for local HuggingFace model weights, expanded via ${LLM_DIR} in icureact/configs/model_configs.json Defaults to .
MODEL_CONFIGS_PATH Path to the model registry Defaults to icureact/configs/model_configs.json
NAVIGATOR_API_KEY LLM calls routed through "environment": "navigator" in model_configs.json Navigator is a University of Florida-internal LLM gateway; external users should either request the config for their own gateway or repoint model_configs.json entries at "environment": "azure" / "hf" (see below)
NAVIGATOR_BASE_URL Navigator endpoint Defaults to https://api.ai.it.ufl.edu
AZURE_API_KEY_REGION2 LLM calls routed through "environment": "azure"
AZURE_OPENAI_ENDPOINT, AZURE_OPENAI_API_VERSION Azure endpoint/API version
WANDB_API_KEY, WANDB_PROJECT Training run logging via --monitor in pipeline_launch.py Optional
EXT_BENCHMARKS_DIR Root directory for external benchmark data (SCT-Bench, ER-Reason, MedRBench, VivaBench) Defaults to external_benchmarks

If you don't have access to the UF Navigator gateway, every script that takes --environment/--env also supports azure (set the two Azure variables above) or hf (point --hf-model-path/model_configs.json at local weights) β€” there is no hard dependency on Navigator specifically.


πŸ“¦ Data & model availability

πŸ§ͺ Benchmarks

Each benchmark must be obtained directly from its own source below β€” none of them are redistributed in this repo.

Benchmark Source Where it goes
ICU-REACT Test https://huggingface.co/datasets/macontreras98/ICU-REACT Same dataset as the seed/training sets above β€” no separate download needed
SCT-Bench https://github.com/SCT-Bench/sctpublic $EXT_BENCHMARKS_DIR/sct_bench/
ER-Reason https://physionet.org/content/er-reason/1.0.0/ (requires PhysioNet credentialed access) $EXT_BENCHMARKS_DIR/er_reason/
MedRBench https://github.com/MAGIC-AI4Med/MedRBench $EXT_BENCHMARKS_DIR/medrbench/
VivaBench https://huggingface.co/datasets/chychiu/VivaBench $EXT_BENCHMARKS_DIR/vivabench/

$EXT_BENCHMARKS_DIR defaults to external_benchmarks/ β€” see Installation & environment setup.


πŸš€ Reproducing results β€” step by step

All commands below assume PROJECT_DIR is set to your repo checkout (or omit it and run everything from the repo root, since it defaults to .). Example configs for every stage live in run_configs/templates/ β€” copy one and edit the paths/models rather than starting from scratch.

🌱 1. Seed dataset creation

create_seed_launch.py turns clinician annotation exports into the ICU-REACT-Train/Test seed sets (Methods Β§Seed dataset creation: per-round filtering β†’ topic mapping β†’ feedback-guided refinement β†’ train/test split β†’ judge rubrics). Each stage is a separate --mode, since the pipeline mixes cheap local steps with expensive LLM calls you may want to re-run independently.

Note: This stage requires the raw clinician annotation exports (annotations/annotations.json, annotations/per_feedback_classification.jsonl) as input. These are not publicly available at this time. If you don't have access to them, skip this stage and start reproduction from the already-built seed set on Hugging Face (see Dataset augmentation below).

python -m icureact.create_seed_launch --mode split_by_round \
    --annotations-path annotations/annotations.json

python -m icureact.create_seed_launch --mode filter_by_score
python -m icureact.create_seed_launch --mode map_topics
python -m icureact.create_seed_launch --mode refine_from_feedback
python -m icureact.create_seed_launch --mode add_stepwise_reasoning
python -m icureact.create_seed_launch --mode consolidate_round
python -m icureact.create_seed_launch --mode merge_seed
python -m icureact.create_seed_launch --mode split_seed
python -m icureact.create_seed_launch --mode create_judge_rubrics

# ICU-REACT-Train refinement (Fig. 1C, train branch): combines the original
# reasoning + clinician feedback into a refinement explanation + improved reasoning
python -m icureact.create_seed_launch --mode refine_reasoning

# ICU-REACT-Test refinement (Fig. 1C, test branch): feeds approved
# questions/retrieval tasks + feedback directly to the refinement agent
python -m icureact.create_seed_launch --mode refine_context_question

# Or run every stage above in order:
python -m icureact.create_seed_launch --mode all --annotations-path annotations/annotations.json

Config: datasets/icureact_seed/{version}/config.json (see run_configs/templates/01_create_seed_config.json) sets topic_model/refinement_model/rubric_model. --exclude-annotator-ids-json and --per-feedback-jsonl-path let you supply the clinician-annotation-round-specific inputs described in the paper without hardcoding them into the script. Run python -m icureact.create_seed_launch --help for the full flag list.

Output: datasets/icureact_seed/{version}/processed/merged/{train,test}_seed_dataset.jsonl plus the refined/rubric variants β€” consumed next by dataset augmentation.

🧬 2. Dataset augmentation

dataset_augmentation_launch.py expands the seed set into the larger synthetic training corpus (Methods Β§Dataset augmentation): new context-question pairs, initial reasoning, reasoning refinements, and judge rubrics, all guided by seed-derived few-shot examples.

The clinician-validated ICU-REACT seed set (the output of Stage 1) is available pre-built on Hugging Face: https://huggingface.co/datasets/macontreras98/ICU-REACT β€” download it into datasets/icureact_seed/{version}/processed/merged/ to start here without running Stage 1 yourself.

You can also skip this stage entirely: the same Hugging Face dataset includes the pre-built, curated training sets that are otherwise the output of Stage 2 + Stage 3 below β€” download those directly and go straight to Stage 4.

# Generate context-question pairs (--size-target is the target sample count;
# shard across parallel jobs with --num-batches/--batch-id for large runs)
python -m icureact.dataset_augmentation_launch \
    --config run_configs/templates/02_dataset_augmentation_config.json \
    --mode pair_gen_only --n-per-record 10 --size-target 5000 \
    --num-batches 1 --batch-id 0

# Merge batch shards into one file
python -m icureact.dataset_augmentation_launch \
    --config run_configs/templates/02_dataset_augmentation_config.json \
    --mode combine_context --dataset-size 5000

# Reasoning-refinement subset (used for the LoRA training target)
python -m icureact.dataset_augmentation_launch \
    --config run_configs/templates/02_dataset_augmentation_config.json \
    --mode reasoning_refinement_only --reasoning-refinement \
    --n-per-record 10 --size-target 5000 --num-batches 1 --batch-id 0

python -m icureact.dataset_augmentation_launch \
    --config run_configs/templates/02_dataset_augmentation_config.json \
    --mode combine_reasoning_refinement --reasoning-refinement --dataset-size 5000

For large runs, --num-batches/--batch-id shard generation across parallel jobs (e.g. a SLURM array), then a --mode combine_* call merges the shards. See dataset_augmentation.sh / dataset_augmentation2.sh for worked SLURM examples of this batching pattern, and python -m icureact.dataset_augmentation_launch --help for every mode.

Output: datasets/icureact_augmented/{version}/dataset_{size}/reasoning_refinement/augmented_train_reasoning_refinement_subset.jsonl (one file per size bucket, e.g. 5k/10k/20k) β€” consumed next by curation.

🧹 3. Curation: dedup β†’ difficulty labeling β†’ filtering

Three scripts run in sequence, linked only by a shared filename convention (*_filtered.jsonl β†’ *_filtered_difficulty.jsonl β†’ *_filtered_hard.jsonl). run_curation_pipeline.py runs all three as one command and derives each stage's input path from the previous stage's output automatically:

python -m icureact.run_curation_pipeline \
    --redundancy-config run_configs/templates/03_redundancy_analyzer_config.json \
    --council-config run_configs/templates/04_llm_council_curator_config.json \
    --filter-config run_configs/templates/05_filter_dataset_by_difficulty_config.json
What's happening under the hood (run these individually if you need more control)
# 3a. Redundancy dedup: vectorize + TF-IDF nearest-neighbor group + ROUGE-L >= 0.7 removal
python -m icureact.redundancy_analyzer_launch \
    --config run_configs/templates/03_redundancy_analyzer_config.json

# 3b. LLM council difficulty labeling: >=3 models vote easy/medium/hard per sample
python -m icureact.llm_council_curator_launch \
    --config run_configs/templates/04_llm_council_curator_config.json

# 3c. Keep only "hard" samples (pure post-processing, no LLM calls)
python -m icureact.filter_dataset_by_difficulty \
    --config run_configs/templates/05_filter_dataset_by_difficulty_config.json

Output: *_filtered_hard.jsonl β€” the final ICU-REACT-Train dataset, referenced as data.train_dataset.train_jsonl_path in a pipeline config.

πŸ‹οΈ 4. Model training + evaluation

pipeline_launch.py fine-tunes a base model with LoRA on the curated dataset (Methods Β§Model training: assistant-only loss over the refinement-explanation + improved-reasoning target) and evaluates the resulting checkpoint:

python -m icureact.pipeline_launch --config run_configs/templates/06_pipeline_config.json
  • model.name must match an entry in icureact/configs/model_configs.json.
  • pipeline.phases supports multiple chained training phases (each phase's checkpoint feeds the next); set "skip": true on a phase to evaluate an already-trained checkpoint without retraining.
  • --skip-train re-runs evaluation only, against a previously trained model.
  • --do-inference / --do-variable-metrics / --do-rubric-judge toggle evaluation sub-stages independently, so a re-score doesn't require re-running inference.

Output: models/{dataset_bucket}/{model_bucket}/{run_tag}/{phase_tag}/... (LoRA checkpoint) and benchmarks/{benchmark}/... (evaluation results).

πŸ“ 5. Baseline / zero-shot evaluation

evaluation_launch.py (also called internally by pipeline_launch.py) scores any model β€” zero-shot backbone or trained checkpoint β€” across the benchmark suite:

python -m icureact.evaluation_launch --config run_configs/templates/07_eval_config.json

benchmarks accepts any of: icureact, sct_bench, er_reason, medrbench, vivabench, multiple_choice_bench. models lists entries from model_configs.json.


πŸ› οΈ Configuration reference

  • run_configs/{stage}/ β€” the authors' actual experiment configs, kept for transparency (hundreds of files spanning every model/dataset-size/ablation combination reported in the paper). Not meant as a starting point.
  • run_configs/templates/ β€” clean, minimal example configs (one per pipeline stage) using ${PROJECT_DIR} placeholders; copy these for your own runs.
  • ${PROJECT_DIR} / ${LLM_DIR} placeholders are expanded automatically by every launch script's config loader (icureact/path_utils.py) β€” safe to use inside any JSON config value.
  • pipeline_config.json (used by pipeline_launch.py and evaluation_launch.py) supports two schemas: a legacy single-phase form (run/model/data/train/eval top-level keys) and a multi-phase form (run/model/data/pipeline, where pipeline.phases is a list merged against pipeline.train_defaults/eval_defaults). See icureact/config/pipeline_config.py for the validator and run_configs/templates/06_pipeline_config.json / 07_eval_config.json for examples of each.
  • icureact/configs/model_configs.json is the model registry: each entry sets "environment" (navigator | azure | hf) and, for hf entries, an hf_model_path (optionally with a peft_model_path for a LoRA adapter on top of a base model). Add your own models here to point the pipeline at different backends.

πŸ”¬ Complementary analyses shown in the paper

These scripts consume benchmarks//dataset output and reproduce specific figures/tables β€” they are not part of the reproduction chain above, but read its outputs.

Script Produces
icureact/benchmark_statistical_analysis_focused.py Paired bootstrap + Wilcoxon signed-rank comparisons (Methods Β§Statistical Analyses). Canonical script β€” .py/_selected.py are earlier/alternate variants, not the ones behind the reported numbers.
icureact/category_stratify.py Supplemental "Performance stratified by clinical content"
icureact/summarize_icureact_variable_coverage.py, summarize_icureact_rubric_dimensions.py Supplemental dataset/rubric statistics (variable coverage, rubric item counts)
notebooks/task_ablation.py, plot_task_ablation.py, plot_size_ablation.py Supplemental Ablations (training-task composition and dataset-size scaling)
notebooks/frontier_comparison.py Supplemental "Comparison with Frontier Models"
notebooks/generate_seed_plots.py, generate_train_dataset_figures.py, generate_rubrics_figure.py, umap_visualize.py, questions_umap.py t-SNE/UMAP dataset and rubric figures (seed creation, train dataset, rubric stats)
icureact/benchmark_visualizer.py, model_comparison_visualizer.py, multiple_choice_benchmark_visualizer.py, sct_model_comparison_app.py Gradio leaderboard / sample-comparison apps for interactive exploration
icureact/data_contamination_analysis.py (embedding/exact-overlap check), icureact/data_contamination_completion.py (completion-based leakage probe) Data contamination analysis reported in the paper

🧭 Exploratory analyses not reported in the manuscript

Remaining one-off scripts under notebooks/ not listed above (annotation-feedback tooling, OMOP variable/topic mapping utilities, etc.) were used during dataset/benchmark development but do not correspond to a specific reported figure or table.


πŸ–₯️ Cluster/SLURM usage

The commands above (python -m icureact.<script> --config ...) work anywhere and are the recommended way to run each stage. If you have access to a SLURM cluster, .slurm templates (run_eval.slurm, run_train.slurm, run_pipeline_launch.slurm) and submit_*.sh wrappers are also provided; *_local.sh gives non-SLURM nohup equivalents. All of these resolve the repo root via $PROJECT_DIR (or the script's own location) β€” you'll still need to adjust the conda activate <env> and module load cuda/... lines to match your own cluster's environment, and create the outputs/ directory before submitting (#SBATCH --output= paths are resolved relative to your working directory at submission time and are not auto-created).


πŸ“– Citation

Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena, Jiaqing Zhang, Heng Sun, Hruday Tej Akkaladevi, Peiyu Lu, Jordan Rosen, Sumit Kapoor, Sasank Desaraju, Grace R. Thompson, Jacob Purcell, Michael Petrauskis, Philip KW Hong, Meghan Brennan, Sarah Chrabaszcz, Tierra Smith, Ronnie Ren, Michel S. Kabbash, Ceyhun Haziroglu, Rushi Patel, Gabriel Gomez, Charlotte Chaiklin, Randy Leung, Kenneth N. John, Whitman Wiggins, Philip Kayser, Vincent Bird, Maria Bruzzone, Tyler J. Loftus, Azra Bihorac, and Parisa Rashidi. "Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains." arXiv:2608.22622 (2026). https://arxiv.org/abs/2608.22622

Corresponding author: Parisa Rashidi (parisa.rashidi@ufl.edu)

@article{contreras2026icureact,
  title   = {Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains},
  author  = {Contreras, Miguel and Siegel, Scott and Nerella, Subhash and Sena, Jessica and Zhang, Jiaqing and
             Sun, Heng and Akkaladevi, Hruday Tej and Lu, Peiyu and Rosen, Jordan and Kapoor, Sumit and
             Desaraju, Sasank and Thompson, Grace R. and Purcell, Jacob and Petrauskis, Michael and
             Hong, Philip KW and Brennan, Meghan and Chrabaszcz, Sarah and Smith, Tierra and Ren, Ronnie and
             Kabbash, Michel S. and Haziroglu, Ceyhun and Patel, Rushi and Gomez, Gabriel and Chaiklin, Charlotte and
             Leung, Randy and John, Kenneth N. and Wiggins, Whitman and Kayser, Philip and Bird, Vincent and
             Bruzzone, Maria and Loftus, Tyler J. and Bihorac, Azra and Rashidi, Parisa},
  journal = {arXiv preprint arXiv:2608.22622},
  year    = {2026},
  doi     = {10.48550/arXiv.2608.22622},
  url     = {https://arxiv.org/abs/2608.22622}
}

βš–οΈ License

GNU General Public License v3.0. See LICENSE / https://www.gnu.org/licenses/gpl-3.0.en.html for the full text.

About

ICU-REACT is a curated dataset designed to train and evaluate large language models (LLMs) for clinical reasoning in intensive care settings. It features LLM-generated ICU-relevant questions, expert-reviewed annotations, structured EHR retrieval tasks, and reasoning traces grounded in real-world workflows.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages