Repository navigation
Add minimal eval CI gate for contextual precision - #6
Conversation
There was a problem hiding this comment.
Pull request overview
Adds an evaluation-based CI gate to prevent retrieval regressions by running the project’s eval suite in GitHub Actions and enforcing a minimum ContextualPrecisionMetric threshold.
Changes:
- Introduces a new GitHub Actions workflow (
eval-gate.yml) that runs the FastAPI app +eval/run_eval.pyand fails CI if ContextualPrecisionMetric< 0.64. - Adds a small CLI utility (
eval/compare_runs.py) to compare baseline vs degraded eval outputs. - Updates
.gitignoreto ignore Deepeval artifacts and eval run output text files.
Reviewed changes
Copilot reviewed 2 out of 3 changed files in this pull request and generated 4 comments.
| File | Description |
|---|---|
eval/compare_runs.py |
New helper script to parse and tabulate eval output comparisons. |
.gitignore |
Ignores Deepeval folder and eval run output artifacts. |
.github/workflows/eval-gate.yml |
New CI gate workflow that runs evals and enforces a threshold. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| - name: Install dependencies | ||
| run: pip install -r requirements.txt |
There was a problem hiding this comment.
The workflow only installs requirements.txt, but eval/run_eval.py imports deepeval and requirements.txt currently does not include it. This job will fail at runtime unless deepeval is added to requirements or installed explicitly in this step.
| - name: Optionally override retrieval_k | ||
| if: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.retrieval_k != '' }} | ||
| run: | | ||
| python - <<'PY' | ||
| from pathlib import Path | ||
| import yaml | ||
|
|
||
| config_path = Path("config.yaml") | ||
| data = yaml.safe_load(config_path.read_text(encoding="utf-8")) or {} | ||
| data["retrieval_k"] = int("${{ github.event.inputs.retrieval_k }}") | ||
| config_path.write_text(yaml.safe_dump(data, sort_keys=False), encoding="utf-8") | ||
| print(f"Using retrieval_k={data['retrieval_k']}") | ||
| PY |
There was a problem hiding this comment.
config.yaml in this repo sets corpus_path to a Windows-local path (e.g. D:/...), and RAGEngine.ingest() raises if the corpus folder is missing. On ubuntu-latest this will cause Uvicorn to crash on startup and the gate to fail. Consider adding a CI-specific corpus directory committed to the repo (even a tiny fixture) and overriding corpus_path here (or via env/config) before starting Uvicorn.
| for i in {1..30}; do | ||
| if curl -sf http://127.0.0.1:8000/docs > /dev/null; then | ||
| break | ||
| fi | ||
| sleep 2 | ||
| done | ||
|
|
||
| python eval/run_eval.py | tee eval_output.txt |
There was a problem hiding this comment.
The readiness loop doesn't fail the job if the service never becomes available (it always continues after 30 tries). If Uvicorn crashes or startup ingestion fails, the workflow will proceed to run_eval.py and error later with a less clear failure. After the loop, add an explicit check that the endpoint is reachable and exit 1 if not.
| threshold = 0.64 | ||
| text = Path("eval_output.txt").read_text(encoding="utf-8", errors="ignore") | ||
| match = re.search(r"ContextualPrecisionMetric: avg_score=([0-9.]+)", text) | ||
| if not match: | ||
| raise SystemExit("Could not find aggregate ContextualPrecisionMetric in eval output") | ||
|
|
There was a problem hiding this comment.
The gate regex only matches numeric avg_score values ([0-9.]+). If eval/run_eval.py prints avg_score=N/A (e.g., when a metric errors and yields no numeric scores), this will fail with the misleading message "Could not find aggregate ...". Consider handling the N/A case explicitly and failing with a clearer error (or parsing the sectioned output similarly to eval/compare_runs.py).
Adds a minimal GitHub Actions eval gate for the Evalens tracer bullet.
Current gate:
Purpose: