Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 83 additions & 0 deletions .github/workflows/eval-gate.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
name: Eval Gate

on:
pull_request:
workflow_dispatch:
inputs:
retrieval_k:
description: "Optional retrieval_k override (use 1 to run degraded config)"
required: false
default: ""

jobs:
contextual-precision-gate:
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v4

- uses: actions/setup-python@v5
with:
python-version: "3.11"

- name: Install dependencies
run: pip install -r requirements.txt
Comment on lines +23 to +24

Copilot AI Apr 8, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The workflow only installs requirements.txt, but eval/run_eval.py imports deepeval and requirements.txt currently does not include it. This job will fail at runtime unless deepeval is added to requirements or installed explicitly in this step.

Copilot uses AI. Check for mistakes.

- name: Optionally override retrieval_k
if: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.retrieval_k != '' }}
run: |
python - <<'PY'
from pathlib import Path
import yaml

config_path = Path("config.yaml")
data = yaml.safe_load(config_path.read_text(encoding="utf-8")) or {}
data["retrieval_k"] = int("${{ github.event.inputs.retrieval_k }}")
config_path.write_text(yaml.safe_dump(data, sort_keys=False), encoding="utf-8")
print(f"Using retrieval_k={data['retrieval_k']}")
PY
Comment on lines +26 to +38

Copilot AI Apr 8, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

config.yaml in this repo sets corpus_path to a Windows-local path (e.g. D:/...), and RAGEngine.ingest() raises if the corpus folder is missing. On ubuntu-latest this will cause Uvicorn to crash on startup and the gate to fail. Consider adding a CI-specific corpus directory committed to the repo (even a tiny fixture) and overriding corpus_path here (or via env/config) before starting Uvicorn.

Copilot uses AI. Check for mistakes.

- name: Run eval and enforce ContextualPrecisionMetric gate
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
GROQ_API_KEY: ${{ secrets.GROQ_API_KEY }}
EVALENS_CORPUS_PATH: corpus/ci_smoke
run: |
set -euo pipefail

uvicorn app.main:app --host 127.0.0.1 --port 8000 > uvicorn.log 2>&1 &
UVICORN_PID=$!
trap 'kill $UVICORN_PID || true' EXIT

for i in {1..30}; do
if curl -sf http://127.0.0.1:8000/docs > /dev/null; then
echo "App is up"
READY=1
break
fi
sleep 2
done

if [ "${READY:-0}" != "1" ]; then
echo "Uvicorn failed to start. Dumping log:"
cat uvicorn.log || true
exit 1
fi

python eval/run_eval.py | tee eval_output.txt
Comment on lines +52 to +67

Copilot AI Apr 8, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The readiness loop doesn't fail the job if the service never becomes available (it always continues after 30 tries). If Uvicorn crashes or startup ingestion fails, the workflow will proceed to run_eval.py and error later with a less clear failure. After the loop, add an explicit check that the endpoint is reachable and exit 1 if not.

Copilot uses AI. Check for mistakes.

python - <<'PY'
import re
from pathlib import Path

threshold = 0.64
text = Path("eval_output.txt").read_text(encoding="utf-8", errors="ignore")
match = re.search(r"ContextualPrecisionMetric: avg_score=([0-9.]+)", text)
if not match:
raise SystemExit("Could not find aggregate ContextualPrecisionMetric in eval output")

Comment on lines +73 to +78

Copilot AI Apr 8, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The gate regex only matches numeric avg_score values ([0-9.]+). If eval/run_eval.py prints avg_score=N/A (e.g., when a metric errors and yields no numeric scores), this will fail with the misleading message "Could not find aggregate ...". Consider handling the N/A case explicitly and failing with a clearer error (or parsing the sectioned output similarly to eval/compare_runs.py).

Copilot uses AI. Check for mistakes.
score = float(match.group(1))
print(f"ContextualPrecisionMetric avg_score={score:.3f} (threshold={threshold:.2f})")
if score < threshold:
raise SystemExit(f"CI gate failed: ContextualPrecisionMetric {score:.3f} < {threshold:.2f}")
PY
3 changes: 2 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -6,4 +6,5 @@ chroma_db/
data/
*.log
.chroma/
.deepeval/
.deepeval/
eval/run_eval_output_*.txt
15 changes: 13 additions & 2 deletions app/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,13 @@ class Settings:
groq_api_key: str


def _resolve_corpus_path(raw_path: str, config_dir: Path) -> str:
candidate = Path(raw_path).expanduser()
if not candidate.is_absolute():
candidate = (config_dir / candidate).resolve()
return str(candidate)


def load_settings(config_path: str = "config.yaml") -> Settings:
cfg_file = Path(config_path)
if not cfg_file.exists():
Expand All @@ -32,9 +39,13 @@ def load_settings(config_path: str = "config.yaml") -> Settings:
cfg = yaml.safe_load(f) or {}

groq_api_key = os.getenv("GROQ_API_KEY", "")
raw_corpus_path = os.getenv(
"EVALENS_CORPUS_PATH",
cfg.get("corpus_path", "D:/Evalens/corpus/intercom_external/raw_pdfs"),
)

return Settings(
corpus_path=cfg.get("corpus_path", "D:/Evalens/corpus/intercom_external/raw_pdfs"),
corpus_path=_resolve_corpus_path(raw_corpus_path, cfg_file.parent.resolve()),
chunk_size=int(cfg.get("chunk_size", 1000)),
chunk_overlap=int(cfg.get("chunk_overlap", 150)),
retrieval_k=int(cfg.get("retrieval_k", 4)),
Expand All @@ -43,4 +54,4 @@ def load_settings(config_path: str = "config.yaml") -> Settings:
chroma_path=cfg.get("chroma_path", ".chroma"),
collection_name=cfg.get("collection_name", "intercom_pdfs"),
groq_api_key=groq_api_key,
)
)
11 changes: 11 additions & 0 deletions corpus/ci_smoke/01_pricing_and_plans.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
# Pricing and plans quick reference

- There is **no Pro plan**. There is a **Pro add-on** for reporting.
- Pro add-on includes: CX Score, Topics Explorer, Trends, Recommendations, Monitors, and Custom Scorecards.
- Pro add-on pricing: $99/month base for up to 1,000 conversations; conversation-volume based (not seat-based).
- Fin AI Agent with Zendesk appears in two pricing representations in docs:
- $0.99 per outcome with commitments.
- $49/month includes 50 outcomes, then $0.99 per additional outcome.
- Copilot: all plans include limited usage for full-seat teammates (10 free conversations/month).
- Lite seats cannot use Copilot.
- Unlimited Copilot requires add-on pricing references of $29/agent/month (annual) or $35/seat/month (monthly).
13 changes: 13 additions & 0 deletions corpus/ci_smoke/02_early_stage_program.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Early Stage program quick reference

Eligibility requirements:
1. Startup has raised up to $10M in funding.
2. Startup has fewer than 15 employees.
3. Startup is not currently an Intercom customer.

Clarifications:
- Existing paid customers (including annual plans) are not eligible.
- If someone only started a trial, they can still apply.
- Early Stage discount is only available on the Advanced plan.
- Early Stage discount does not apply to the Expert plan.
- If Expert features are needed, Expert is purchased at standard pricing.
9 changes: 9 additions & 0 deletions corpus/ci_smoke/03_scope_boundaries.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# Scope boundaries and abstention guidance

This corpus covers product and plan information only.

Out of scope examples:
- Intercom stock price is not included in this corpus.
- Platform uptime guarantee percentage is not included in this corpus.

If asked these questions, answer that the indexed corpus does not contain that information.
92 changes: 92 additions & 0 deletions eval/compare_runs.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
from __future__ import annotations

import argparse
import re
from pathlib import Path


def read_text(path: Path) -> str:
raw = path.read_bytes()
for enc in ("utf-8", "utf-16", "utf-16-le", "utf-16-be"):
try:
return raw.decode(enc).replace("\x00", "")
except UnicodeDecodeError:
continue
return raw.decode("utf-8", errors="replace").replace("\x00", "")


def parse_sections(path: Path) -> tuple[dict[str, float], dict[str, float]]:
text = read_text(path)
lines = [line.strip() for line in text.splitlines()]

aggregate: dict[str, float] = {}
categories: dict[str, float] = {}

in_aggregate = False
in_categories = False
avg_re = re.compile(r"^-\s+(\w+):\s+avg_score=([^\s]+)")

for line in lines:
if line == "Aggregate metric averages:":
in_aggregate = True
in_categories = False
continue
if line == "Simple average by category:":
in_categories = True
in_aggregate = False
continue
if not line:
continue

m = avg_re.match(line)
if not m:
continue

name, value = m.groups()
if value == "N/A":
continue
try:
score = float(value)
except ValueError:
continue

if in_aggregate:
aggregate[name] = score
elif in_categories:
categories[name] = score

return aggregate, categories


def print_table(title: str, baseline: dict[str, float], degraded: dict[str, float]) -> None:
keys = sorted(set(baseline) | set(degraded))
print(f"\n{title}")
print("-" * len(title))
print(f"{'name':35} {'baseline':>10} {'degraded':>10} {'delta':>10}")
for key in keys:
b = baseline.get(key)
d = degraded.get(key)
if b is None or d is None:
print(f"{key:35} {str(b):>10} {str(d):>10} {'N/A':>10}")
else:
print(f"{key:35} {b:10.3f} {d:10.3f} {d - b:+10.3f}")


def main() -> None:
parser = argparse.ArgumentParser(description="Compare baseline vs degraded eval outputs.")
parser.add_argument("baseline", type=Path, help="Path to baseline eval output file")
parser.add_argument("degraded", type=Path, help="Path to degraded eval output file")
args = parser.parse_args()

base_agg, base_cat = parse_sections(args.baseline)
deg_agg, deg_cat = parse_sections(args.degraded)

print(f"Baseline: {args.baseline}")
print(f"Degraded: {args.degraded}")

print_table("Aggregate metric averages", base_agg, deg_agg)
print_table("Category averages", base_cat, deg_cat)


if __name__ == "__main__":
main()
1 change: 1 addition & 0 deletions requirements.txt
Original file line number Diff line number Diff line change
Expand Up @@ -6,3 +6,4 @@ sentence-transformers
groq
python-dotenv
pyyaml
deepeval
Loading