A cargo plugin for finding and tracking flaky tests in Rust projects using Bayesian inference.
Flaky tests -- tests that sometimes pass and sometimes fail without code changes -- erode trust in your test suite, waste CI time, and mask real regressions. cargo ninety-nine runs each test multiple times, applies Bayesian statistical analysis to compute a flakiness probability, and tracks results over time.
- Bayesian flakiness detection -- computes posterior probability of flakiness using Beta distribution, not just pass/fail ratios
- Multi-phase diagnose -- stress full binaries, isolate candidates, classify Contention / Intrinsic / Broken, optional rr recording
- Interactive TUI -- terminal interface with scrollable tables, category filtering, sort cycling, and detail drill-down
- Filter DSL -- expressive query language to target tests by name, package, binary, kind, or flakiness status
- Pattern analysis -- detects time-of-day and environmental (CI vs local) failure patterns
- Trend tracking -- monitors whether tests are improving, stable, or degrading over time
- Duration regression detection -- flags tests whose execution time has significantly increased
- Quarantine management -- manually or automatically quarantine flaky tests that exceed thresholds
- Multiple export formats -- JUnit XML, HTML reports, CSV, and JSON
- CI workflow generation -- ready-to-use GitHub Actions and GitLab CI configurations
- Dual storage backends -- SQLite (default, zero-config) or PostgreSQL for shared environments
# Install
cargo install cargo-ninety-nine
# Initialize configuration
cargo ninety-nine init
# Run flaky test detection (10 iterations per test)
cargo ninety-nine test -n 10
# Run only flaky, non-quarantined tests
cargo ninety-nine test "flaky & !quarantined"
# Check a specific test's history
cargo ninety-nine status tests::my_flaky_test
# Browse all scores interactively
cargo ninety-nine status
# View session history
cargo ninety-nine history
# Export results as JSON
cargo ninety-nine export json results.json- Discover -- compiles test binaries and lists all test cases
- Filter -- applies filter expressions to select which tests to run
- Execute -- runs each test N times with configurable concurrency and timeouts
- Analyse -- applies Bayesian inference to compute P(flaky) with credible intervals
- Report -- displays results interactively or as text, with category labels (Stable, Occasional, Moderate, Frequent, Critical)
- Store -- persists scores and run history for trend analysis across sessions
| Command | Purpose |
|---|---|
test |
Run tests repeatedly and compute flakiness scores |
status |
Browse current flakiness scores (interactive TUI) |
history |
Browse past detection sessions (interactive TUI) |
init |
Create a default .ninety-nine.toml configuration file |
export |
Export results to JUnit XML, HTML, CSV, or JSON |
quarantine |
Manage test quarantine (list, add, remove) |
ci |
Generate CI workflow files |
Pass --non-interactive (-N) to disable the TUI for CI pipelines or scripted use.
cargo ninety-nine init creates a .ninety-nine.toml in your project root. Key settings:
[detection]
min_runs = 10 # iterations per test
confidence_threshold = 0.95 # Bayesian confidence required
parallel_runs = 3 # concurrent test executions
[quarantine]
enabled = true
auto_quarantine = false # set true to auto-quarantine flaky tests
[storage]
backend = "Sqlite" # or "Postgres"
retention_days = 90See the full documentation for all configuration options.
Full documentation is available at glottologist.github.io/ninety-nine, covering:
- Rust 1.85+
cargo testorcargo-nextest(auto-detected)
MIT
