Skip to content

About

Repository to develop PoC for Yalo AB Tool

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

Yalo A/B Testing Platform

Automated experimentation layer for Yalo's conversational pipeline. Registers experiments, splits audiences deterministically, and runs frequentist + Bayesian statistical tests to determine winning variants.


Architecture

Target GCP pipeline (from design doc):

Pipeline con capa de experimentación

Production mapping (from design doc):

PoC component Production GCP service
poc.db (SQLite) Firestore (assignments + actions) + BigQuery (results)
CSV audience file GCS bucket
uvicorn process Cloud Run
Manual /execute-experiment Cloud Run Jobs + Cloud Scheduler

Data model

All data is stored in SQLite (poc.db).

Table Purpose
experiments One row per registered experiment — config, metric, params, status
experiment_{exp_id} Per-experiment event log. One row per event: ASSIGNED (group membership) and action events BUY, CLICK, IGNORE
experiment_results One snapshot per /execute-experiment run — z-score, p-value, P(B>A)

Quick start

1. Install dependencies

uv sync

2. Seed the database with mock data

uv run python -m scripts.generate_dataset

This generates:

  • data/audience.csv — 600 synthetic users with conversation IDs
  • data/poc.db — SQLite database with the experiment record and a per-experiment event-log table (ASSIGNED + action events)
  • Group A (control): ~10 % conversion rate
  • Group B (treatment): ~25 % conversion rate → expected z ≈ 4.5, p < 0.0001

3. Start the server

uv run uvicorn src.main:app --reload

4. Open the dashboard

http://localhost:8000

Use the Dashboard tab to see the seeded experiment, then click ▶ Execute to run the statistical test and view the results.

Interactive API docs are available at http://localhost:8000/docs.


API reference

POST /new-experiment

Register a new experiment and split the audience.

{
  "experiment_owner": "yalo_client_poc",
  "audience_file_path": "audience.csv",
  "metric": "conversion_rate",
  "descr": "Improved prompt vs baseline",
  "n_min_samples": 100,
  "alpha": 0.05
}

Metrics: conversion_rate (BUY events) · click_rate (CLICK events)

POST /action

Log a user action for an experiment. The user must already be assigned to the experiment.

{
  "exp_id": "<experiment-uuid>",
  "user_id": "user_0001",
  "conv_id": "<conversation-uuid>",
  "event_type": "BUY",
  "payload": null
}

Event types: BUY · CLICK · IGNORE

POST /execute-experiment

Run the statistical test pipeline for an existing experiment.

{
  "id": "<experiment-uuid>",
  "owner": "yalo_client_poc",
  "metric": "conversion_rate"
}

Returns z-score, p-value, P(B > A), and a significance verdict. Marks the experiment DONE automatically when significant.

GET /experiments

List all experiments with group sizes and latest result snapshot.

GET /results/{exp_id}

Return all result snapshots for one experiment (newest first). Multiple runs per experiment are supported — each run is an independent statistical snapshot.


Statistical methods

Both methods run on every /execute-experiment call. The experiment is flagged significant when p_value < alpha (frequentist criterion).

Frequentist: two-proportion z-test

Goal: decide whether the observed difference in conversion rates between groups A and B is statistically distinguishable from noise.

Hypotheses:

$$H_0: p_A = p_B \qquad H_1: p_A \neq p_B$$

Observed rates:

$$\hat{p}_A = \frac{\text{conv}_A}{n_A}, \qquad \hat{p}_B = \frac{\text{conv}_B}{n_B}$$

Pooled proportion (best estimate of the common rate under $H_0$):

$$\hat{p}_{\text{pool}} = \frac{\text{conv}_A + \text{conv}_B}{n_A + n_B}$$

Pooled standard error:

$$SE = \sqrt{\hat{p}_{\text{pool}},(1 - \hat{p}_{\text{pool}})\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}$$

Test statistic:

$$z = \frac{\hat{p}_B - \hat{p}_A}{SE}$$

Under $H_0$, $z$ follows a standard normal distribution $\mathcal{N}(0, 1)$ for large samples. The two-tailed p-value is:

$$p = 2,\Bigl(1 - \Phi(|z|)\Bigr)$$

where $\Phi$ is the standard normal CDF, implemented via math.erfc to avoid a scipy dependency:

$$\Phi(x) = \frac{1}{2},\mathrm{erfc}!\left(\frac{-x}{\sqrt{2}}\right)$$

The result is significant when $p &lt; \alpha$ (default $\alpha = 0.05$).


Bayesian: Beta-Binomial model

Goal: compute the posterior probability that group B's true conversion rate exceeds group A's, $P(\theta_B &gt; \theta_A \mid \text{data})$.

Model: conversion counts are Binomial; we place a conjugate Beta prior on each group's unknown rate $\theta$.

Prior (uniform — no preference before seeing data):

$$\theta_A,, \theta_B ;\sim; \text{Beta}(1,, 1)$$

Posteriors (closed-form via conjugacy):

$$\theta_A \mid \text{data} ;\sim; \text{Beta}!\left(\text{conv}_A + 1,; n_A - \text{conv}_A + 1\right)$$

$$\theta_B \mid \text{data} ;\sim; \text{Beta}!\left(\text{conv}_B + 1,; n_B - \text{conv}_B + 1\right)$$

Monte Carlo estimate of $P(\theta_B &gt; \theta_A)$ with $N = 10,000$ samples:

$$P(\theta_B > \theta_A) ;\approx; \frac{1}{N}\sum_{i=1}^{N} \mathbf{1}!\left[\theta_B^{(i)} > \theta_A^{(i)}\right], \quad \theta_k^{(i)} \overset{\text{iid}}{\sim} \text{Beta}(\alpha_k, \beta_k)$$

A value close to 1 means the data strongly supports B outperforming A. Unlike the p-value, this quantity has a direct probability interpretation and does not depend on a significance threshold.


About

Repository to develop PoC for Yalo AB Tool

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages