Skip to content

Add minimal eval CI gate for contextual precision - #6

Merged
NaveenBuidl merged 5 commits into
mainfrom
ci-eval-gate
Apr 9, 2026
Merged

NaveenBuidl merged 5 commits into
mainfrom
ci-eval-gate

Conversation

@NaveenBuidl

Copy link
Copy Markdown
Owner

Adds a minimal GitHub Actions eval gate for the Evalens tracer bullet.

Current gate:

  • ContextualPrecisionMetric only
  • fails if < 0.64

Purpose:

  • prove a degraded retrieval config can be blocked by CI

Copilot AI review requested due to automatic review settings April 8, 2026 20:35

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds an evaluation-based CI gate to prevent retrieval regressions by running the project’s eval suite in GitHub Actions and enforcing a minimum ContextualPrecisionMetric threshold.

Changes:

  • Introduces a new GitHub Actions workflow (eval-gate.yml) that runs the FastAPI app + eval/run_eval.py and fails CI if ContextualPrecisionMetric < 0.64.
  • Adds a small CLI utility (eval/compare_runs.py) to compare baseline vs degraded eval outputs.
  • Updates .gitignore to ignore Deepeval artifacts and eval run output text files.

Reviewed changes

Copilot reviewed 2 out of 3 changed files in this pull request and generated 4 comments.

File Description
eval/compare_runs.py New helper script to parse and tabulate eval output comparisons.
.gitignore Ignores Deepeval folder and eval run output artifacts.
.github/workflows/eval-gate.yml New CI gate workflow that runs evals and enforces a threshold.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +23 to +24
- name: Install dependencies
run: pip install -r requirements.txt

Copilot AI Apr 8, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The workflow only installs requirements.txt, but eval/run_eval.py imports deepeval and requirements.txt currently does not include it. This job will fail at runtime unless deepeval is added to requirements or installed explicitly in this step.

Copilot uses AI. Check for mistakes.
Comment on lines +26 to +38
- name: Optionally override retrieval_k
if: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.retrieval_k != '' }}
run: |
python - <<'PY'
from pathlib import Path
import yaml

config_path = Path("config.yaml")
data = yaml.safe_load(config_path.read_text(encoding="utf-8")) or {}
data["retrieval_k"] = int("${{ github.event.inputs.retrieval_k }}")
config_path.write_text(yaml.safe_dump(data, sort_keys=False), encoding="utf-8")
print(f"Using retrieval_k={data['retrieval_k']}")
PY

Copilot AI Apr 8, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

config.yaml in this repo sets corpus_path to a Windows-local path (e.g. D:/...), and RAGEngine.ingest() raises if the corpus folder is missing. On ubuntu-latest this will cause Uvicorn to crash on startup and the gate to fail. Consider adding a CI-specific corpus directory committed to the repo (even a tiny fixture) and overriding corpus_path here (or via env/config) before starting Uvicorn.

Copilot uses AI. Check for mistakes.
Comment on lines +50 to +57
for i in {1..30}; do
if curl -sf http://127.0.0.1:8000/docs > /dev/null; then
break
fi
sleep 2
done

python eval/run_eval.py | tee eval_output.txt

Copilot AI Apr 8, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The readiness loop doesn't fail the job if the service never becomes available (it always continues after 30 tries). If Uvicorn crashes or startup ingestion fails, the workflow will proceed to run_eval.py and error later with a less clear failure. After the loop, add an explicit check that the endpoint is reachable and exit 1 if not.

Copilot uses AI. Check for mistakes.
Comment on lines +63 to +68
threshold = 0.64
text = Path("eval_output.txt").read_text(encoding="utf-8", errors="ignore")
match = re.search(r"ContextualPrecisionMetric: avg_score=([0-9.]+)", text)
if not match:
raise SystemExit("Could not find aggregate ContextualPrecisionMetric in eval output")

Copilot AI Apr 8, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The gate regex only matches numeric avg_score values ([0-9.]+). If eval/run_eval.py prints avg_score=N/A (e.g., when a metric errors and yields no numeric scores), this will fail with the misleading message "Could not find aggregate ...". Consider handling the N/A case explicitly and failing with a clearer error (or parsing the sectioned output similarly to eval/compare_runs.py).

Copilot uses AI. Check for mistakes.
@NaveenBuidl
NaveenBuidl merged commit 15c2c9b into main Apr 9, 2026
1 check passed
@NaveenBuidl
NaveenBuidl deleted the ci-eval-gate branch April 9, 2026 08:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants