Repository navigation
Add minimal eval CI gate for contextual precision #6
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
668ad6a
00c894b
815435a
5ab4ade
25d2d71
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,83 @@ | ||
| name: Eval Gate | ||
|
|
||
| on: | ||
| pull_request: | ||
| workflow_dispatch: | ||
| inputs: | ||
| retrieval_k: | ||
| description: "Optional retrieval_k override (use 1 to run degraded config)" | ||
| required: false | ||
| default: "" | ||
|
|
||
| jobs: | ||
| contextual-precision-gate: | ||
| runs-on: ubuntu-latest | ||
| timeout-minutes: 30 | ||
| steps: | ||
| - uses: actions/checkout@v4 | ||
|
|
||
| - uses: actions/setup-python@v5 | ||
| with: | ||
| python-version: "3.11" | ||
|
|
||
| - name: Install dependencies | ||
| run: pip install -r requirements.txt | ||
|
|
||
| - name: Optionally override retrieval_k | ||
| if: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.retrieval_k != '' }} | ||
| run: | | ||
| python - <<'PY' | ||
| from pathlib import Path | ||
| import yaml | ||
|
|
||
| config_path = Path("config.yaml") | ||
| data = yaml.safe_load(config_path.read_text(encoding="utf-8")) or {} | ||
| data["retrieval_k"] = int("${{ github.event.inputs.retrieval_k }}") | ||
| config_path.write_text(yaml.safe_dump(data, sort_keys=False), encoding="utf-8") | ||
| print(f"Using retrieval_k={data['retrieval_k']}") | ||
| PY | ||
|
Comment on lines
+26
to
+38
|
||
|
|
||
| - name: Run eval and enforce ContextualPrecisionMetric gate | ||
| env: | ||
| OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} | ||
| GROQ_API_KEY: ${{ secrets.GROQ_API_KEY }} | ||
| EVALENS_CORPUS_PATH: corpus/ci_smoke | ||
| run: | | ||
| set -euo pipefail | ||
|
|
||
| uvicorn app.main:app --host 127.0.0.1 --port 8000 > uvicorn.log 2>&1 & | ||
| UVICORN_PID=$! | ||
| trap 'kill $UVICORN_PID || true' EXIT | ||
|
|
||
| for i in {1..30}; do | ||
| if curl -sf http://127.0.0.1:8000/docs > /dev/null; then | ||
| echo "App is up" | ||
| READY=1 | ||
| break | ||
| fi | ||
| sleep 2 | ||
| done | ||
|
|
||
| if [ "${READY:-0}" != "1" ]; then | ||
| echo "Uvicorn failed to start. Dumping log:" | ||
| cat uvicorn.log || true | ||
| exit 1 | ||
| fi | ||
|
|
||
| python eval/run_eval.py | tee eval_output.txt | ||
|
Comment on lines
+52
to
+67
|
||
|
|
||
| python - <<'PY' | ||
| import re | ||
| from pathlib import Path | ||
|
|
||
| threshold = 0.64 | ||
| text = Path("eval_output.txt").read_text(encoding="utf-8", errors="ignore") | ||
| match = re.search(r"ContextualPrecisionMetric: avg_score=([0-9.]+)", text) | ||
| if not match: | ||
| raise SystemExit("Could not find aggregate ContextualPrecisionMetric in eval output") | ||
|
|
||
|
Comment on lines
+73
to
+78
|
||
| score = float(match.group(1)) | ||
| print(f"ContextualPrecisionMetric avg_score={score:.3f} (threshold={threshold:.2f})") | ||
| if score < threshold: | ||
| raise SystemExit(f"CI gate failed: ContextualPrecisionMetric {score:.3f} < {threshold:.2f}") | ||
| PY | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -6,4 +6,5 @@ chroma_db/ | |
| data/ | ||
| *.log | ||
| .chroma/ | ||
| .deepeval/ | ||
| .deepeval/ | ||
| eval/run_eval_output_*.txt | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,11 @@ | ||
| # Pricing and plans quick reference | ||
|
|
||
| - There is **no Pro plan**. There is a **Pro add-on** for reporting. | ||
| - Pro add-on includes: CX Score, Topics Explorer, Trends, Recommendations, Monitors, and Custom Scorecards. | ||
| - Pro add-on pricing: $99/month base for up to 1,000 conversations; conversation-volume based (not seat-based). | ||
| - Fin AI Agent with Zendesk appears in two pricing representations in docs: | ||
| - $0.99 per outcome with commitments. | ||
| - $49/month includes 50 outcomes, then $0.99 per additional outcome. | ||
| - Copilot: all plans include limited usage for full-seat teammates (10 free conversations/month). | ||
| - Lite seats cannot use Copilot. | ||
| - Unlimited Copilot requires add-on pricing references of $29/agent/month (annual) or $35/seat/month (monthly). |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,13 @@ | ||
| # Early Stage program quick reference | ||
|
|
||
| Eligibility requirements: | ||
| 1. Startup has raised up to $10M in funding. | ||
| 2. Startup has fewer than 15 employees. | ||
| 3. Startup is not currently an Intercom customer. | ||
|
|
||
| Clarifications: | ||
| - Existing paid customers (including annual plans) are not eligible. | ||
| - If someone only started a trial, they can still apply. | ||
| - Early Stage discount is only available on the Advanced plan. | ||
| - Early Stage discount does not apply to the Expert plan. | ||
| - If Expert features are needed, Expert is purchased at standard pricing. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,9 @@ | ||
| # Scope boundaries and abstention guidance | ||
|
|
||
| This corpus covers product and plan information only. | ||
|
|
||
| Out of scope examples: | ||
| - Intercom stock price is not included in this corpus. | ||
| - Platform uptime guarantee percentage is not included in this corpus. | ||
|
|
||
| If asked these questions, answer that the indexed corpus does not contain that information. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,92 @@ | ||
| from __future__ import annotations | ||
|
|
||
| import argparse | ||
| import re | ||
| from pathlib import Path | ||
|
|
||
|
|
||
| def read_text(path: Path) -> str: | ||
| raw = path.read_bytes() | ||
| for enc in ("utf-8", "utf-16", "utf-16-le", "utf-16-be"): | ||
| try: | ||
| return raw.decode(enc).replace("\x00", "") | ||
| except UnicodeDecodeError: | ||
| continue | ||
| return raw.decode("utf-8", errors="replace").replace("\x00", "") | ||
|
|
||
|
|
||
| def parse_sections(path: Path) -> tuple[dict[str, float], dict[str, float]]: | ||
| text = read_text(path) | ||
| lines = [line.strip() for line in text.splitlines()] | ||
|
|
||
| aggregate: dict[str, float] = {} | ||
| categories: dict[str, float] = {} | ||
|
|
||
| in_aggregate = False | ||
| in_categories = False | ||
| avg_re = re.compile(r"^-\s+(\w+):\s+avg_score=([^\s]+)") | ||
|
|
||
| for line in lines: | ||
| if line == "Aggregate metric averages:": | ||
| in_aggregate = True | ||
| in_categories = False | ||
| continue | ||
| if line == "Simple average by category:": | ||
| in_categories = True | ||
| in_aggregate = False | ||
| continue | ||
| if not line: | ||
| continue | ||
|
|
||
| m = avg_re.match(line) | ||
| if not m: | ||
| continue | ||
|
|
||
| name, value = m.groups() | ||
| if value == "N/A": | ||
| continue | ||
| try: | ||
| score = float(value) | ||
| except ValueError: | ||
| continue | ||
|
|
||
| if in_aggregate: | ||
| aggregate[name] = score | ||
| elif in_categories: | ||
| categories[name] = score | ||
|
|
||
| return aggregate, categories | ||
|
|
||
|
|
||
| def print_table(title: str, baseline: dict[str, float], degraded: dict[str, float]) -> None: | ||
| keys = sorted(set(baseline) | set(degraded)) | ||
| print(f"\n{title}") | ||
| print("-" * len(title)) | ||
| print(f"{'name':35} {'baseline':>10} {'degraded':>10} {'delta':>10}") | ||
| for key in keys: | ||
| b = baseline.get(key) | ||
| d = degraded.get(key) | ||
| if b is None or d is None: | ||
| print(f"{key:35} {str(b):>10} {str(d):>10} {'N/A':>10}") | ||
| else: | ||
| print(f"{key:35} {b:10.3f} {d:10.3f} {d - b:+10.3f}") | ||
|
|
||
|
|
||
| def main() -> None: | ||
| parser = argparse.ArgumentParser(description="Compare baseline vs degraded eval outputs.") | ||
| parser.add_argument("baseline", type=Path, help="Path to baseline eval output file") | ||
| parser.add_argument("degraded", type=Path, help="Path to degraded eval output file") | ||
| args = parser.parse_args() | ||
|
|
||
| base_agg, base_cat = parse_sections(args.baseline) | ||
| deg_agg, deg_cat = parse_sections(args.degraded) | ||
|
|
||
| print(f"Baseline: {args.baseline}") | ||
| print(f"Degraded: {args.degraded}") | ||
|
|
||
| print_table("Aggregate metric averages", base_agg, deg_agg) | ||
| print_table("Category averages", base_cat, deg_cat) | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| main() |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -6,3 +6,4 @@ sentence-transformers | |
| groq | ||
| python-dotenv | ||
| pyyaml | ||
| deepeval | ||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The workflow only installs
requirements.txt, buteval/run_eval.pyimportsdeepevalandrequirements.txtcurrently does not include it. This job will fail at runtime unlessdeepevalis added to requirements or installed explicitly in this step.