Automated experimentation layer for Yalo's conversational pipeline. Registers experiments, splits audiences deterministically, and runs frequentist + Bayesian statistical tests to determine winning variants.
Target GCP pipeline (from design doc):
Production mapping (from design doc):
| PoC component | Production GCP service |
|---|---|
poc.db (SQLite) |
Firestore (assignments + actions) + BigQuery (results) |
| CSV audience file | GCS bucket |
uvicorn process |
Cloud Run |
Manual /execute-experiment |
Cloud Run Jobs + Cloud Scheduler |
All data is stored in SQLite (poc.db).
| Table | Purpose |
|---|---|
experiments |
One row per registered experiment — config, metric, params, status |
experiment_{exp_id} |
Per-experiment event log. One row per event: ASSIGNED (group membership) and action events BUY, CLICK, IGNORE |
experiment_results |
One snapshot per /execute-experiment run — z-score, p-value, P(B>A) |
uv syncuv run python -m scripts.generate_datasetThis generates:
data/audience.csv— 600 synthetic users with conversation IDsdata/poc.db— SQLite database with the experiment record and a per-experiment event-log table (ASSIGNED + action events)- Group A (control): ~10 % conversion rate
- Group B (treatment): ~25 % conversion rate → expected z ≈ 4.5, p < 0.0001
uv run uvicorn src.main:app --reloadhttp://localhost:8000
Use the Dashboard tab to see the seeded experiment, then click ▶ Execute to run the statistical test and view the results.
Interactive API docs are available at http://localhost:8000/docs.
Register a new experiment and split the audience.
{
"experiment_owner": "yalo_client_poc",
"audience_file_path": "audience.csv",
"metric": "conversion_rate",
"descr": "Improved prompt vs baseline",
"n_min_samples": 100,
"alpha": 0.05
}Metrics: conversion_rate (BUY events) · click_rate (CLICK events)
Log a user action for an experiment. The user must already be assigned to the experiment.
{
"exp_id": "<experiment-uuid>",
"user_id": "user_0001",
"conv_id": "<conversation-uuid>",
"event_type": "BUY",
"payload": null
}Event types: BUY · CLICK · IGNORE
Run the statistical test pipeline for an existing experiment.
{
"id": "<experiment-uuid>",
"owner": "yalo_client_poc",
"metric": "conversion_rate"
}Returns z-score, p-value, P(B > A), and a significance verdict. Marks the experiment DONE automatically when significant.
List all experiments with group sizes and latest result snapshot.
Return all result snapshots for one experiment (newest first). Multiple runs per experiment are supported — each run is an independent statistical snapshot.
Both methods run on every /execute-experiment call. The experiment is flagged significant when p_value < alpha (frequentist criterion).
Goal: decide whether the observed difference in conversion rates between groups A and B is statistically distinguishable from noise.
Hypotheses:
Observed rates:
Pooled proportion (best estimate of the common rate under
Pooled standard error:
Test statistic:
Under
where math.erfc to avoid a scipy dependency:
The result is significant when
Goal: compute the posterior probability that group B's true conversion rate exceeds group A's,
Model: conversion counts are Binomial; we place a conjugate Beta prior on each group's unknown rate
Prior (uniform — no preference before seeing data):
Posteriors (closed-form via conjugacy):
Monte Carlo estimate of
A value close to 1 means the data strongly supports B outperforming A. Unlike the p-value, this quantity has a direct probability interpretation and does not depend on a significance threshold.
