Skip to content
maida-aiPublic

About

Pre-merge behavioral regression gate for AI agents. Check agent changes, inspect regressions, and block broken behavior before merge. Local-first.

Topics

Resources

Contributing

Security policy

Stars

31 stars

Watchers

3 watching

Forks

Repository files navigation

Maida

Don't let broken agent changes merge.

PyPI version Python versions Tests License Docs

Your coding agent returns a plausible answer and the tests pass, but it now loops, skips verification, or rewrites a test to hide a bug. Maida checks an agent change before merge. Output tests and evals may pass; Maida also checks how the agent worked.

This is the core product repository: the engine, CLI, and public contracts. Start here, use maida-tutorials for the canonical runnable experience, and add maida-assert for the GitHub PR boundary.

Try Maida in your repository

Use Python 3.12–3.14 and your own Git repository. Claude Code and Codex follow the same init → task → exit → check → printed view flow. Codex is unreleased; use a development installation containing this stack before selecting it. The published installation example below supports Claude Code.

Coding agent Launch a new session Setup distinction
Claude Code claude maida init --agent claude-code
Codex (unreleased) codex maida init --agent codex; review/trust native Maida hooks when requested

Complete one bounded task, exit the agent, then run maida check and its exact printed viewer command. See Codex capture and native acceptance for the development install, coverage and recovery details.

uv tool install "maida-ai==0.6.1"

cd my-repo
maida init

# Run one normal Claude Code task and exit the session.

maida check
# Then run the exact "View:" command printed by Maida.

Approve init's setup preview, then start a new session with claude. Complete one normal task and exit normally. Success looks like 3 active checks passed, your task's trace ID, and its viewer command. Open it to see the execution timeline. No agent-code changes or tutorial clone are needed.

For example: maida view 83aa19e3. Use the command from your own report.

If Maida gave you a useful signal on your agent, ⭐ star the repo — it helps other teams find the project.

Runs on your machine or CI runner. No Maida cloud account required. Task evidence is not uploaded to Maida; your coding agent still uses its normal provider, permissions, and costs.

If Maida is already installed in the project's uv environment, use uv run maida init, uv run maida check, and the printed viewer command (for example, uv run maida view 83aa19e3). Init connects that installation to the selected agent; launch its command from the table afterward.

Get your first report →

Investigate a regression

Run the exact View: command printed in the report. See the tool calls and named failure, repair the cause, and repeat the task.

The Maida timeline viewer showing a demo support agent with seven tool calls and search_kb repeated five times

The first maida check checks successful completion, recorded loops, and guardrail events. It does not compare a baseline or use an existing policy. Capture observes tool activity and lifecycle, not answer correctness or complete model-call, token, or latency coverage. Missing or unfinished capture gives recovery guidance instead of showing an older task. Keep ordinary correctness tests alongside Maida.

Protect the next agent change

Review the behavior you observed, keep a baseline and policy, then check the next change against them. Instructions, skills, tools, model configuration, harness code, and application code can all change behavior. Maida gates the resulting change whether a human or the agent authored it.

Follow Protect the next agent change to review a small contract, reproduce a safe failure, and repair it without replacing the baseline. The starter requirements are candidates for human review; they do not automatically protect test execution or prevent test rewriting.

Read the comparative verdict: PASS, FAIL, or INCONCLUSIVE. It applies to the observed evidence and selected requirements. Exit 0 includes INCONCLUSIVE, so process success alone is not approval.

See why green tests are not enough

In the canonical storefront demo, a coding agent simplifies shipping, then rewrites the VIP regression test to approve $15 shipping instead of $0. All four application tests pass. Maida fails the agent change for rewriting the protected test. The same project shows skipped verification and an agent weakening its own instructions.

The deterministic rehearsal uses released Maida v0.6.1 and runs offline after installation. It is optional practice, not a prerequisite for your own repository. For a quick canned report without cloning anything:

maida demo --regression

Expect FAIL and a PR-comment preview. The rehearsal exits 0 when the expected failure is reproduced; an actual failed check exits 1.

Add the PR gate

Once a local pass → safe failure → repair works, follow the Action setup and repository protection requirements. CI needs a repeatable task, pinned versions, a reviewed baseline and policy, and configured required checks. Test the boundary on an actual PR head and again after intentional acceptance; generating a workflow alone does not establish enforcement.

Integrate another agent/framework

Your coding-agent repository can use any language. Building a Python tool-calling agent? Follow the secondary Python walkthrough, including installation into the project environment.

Integration Setup Guide
Claude Code maida init --agent claude-code Capture and recovery
Codex (unreleased) Development installation: maida init --agent codex Capture and recovery
LangChain / LangGraph maida-ai[langchain]==0.6.1 Guide
OpenAI Agents SDK maida-ai[openai]==0.6.1 Guide
CrewAI Unsupported installation path; adapter retained Compatibility
Langfuse import Built in, read-only Guide
Another native emitter maida validate-trace Emitter guide

Adapters are optional; the core works without a framework installed. See the integration overview for coverage and setup.

Documentation and reference

Start at maida.ai/docs or the local documentation index.

Setup help and privacy

Init previews automatic local capture setup and preserves existing settings and other hooks. If detection is ambiguous, use maida init --agent claude-code or, with the unreleased development installation, maida init --agent codex. Upgrade an older standalone install with uv tool install --force "maida-ai==0.6.1", rerun init, and follow its recovery guidance. To stop capture, use maida detach --agent claude-code; restart the agent session after setup or detach. Saved evidence is preserved. See the init reference for detailed setup and upgrade handling.

Redaction is on by default and large fields are truncated. Configuration explains storage and redaction settings. No task evidence is uploaded to Maida by default. Capture integrations may use a local telemetry receiver; that is not telemetry to Maida. Optional usage counts require explicit consent and a configured collector; none is configured by default.

🧪 Development

git clone https://github.com/maida-ai/maida.git
cd maida
uv venv && uv sync && uv pip install -e .
uv run pytest
No uv? Use pip instead.
python -m venv .venv && source .venv/bin/activate
pip install -e .
pytest

Contributions welcome -- see CONTRIBUTING.md and SECURITY.md.

📄 License

Apache License 2.0. See LICENSE.

About

Pre-merge behavioral regression gate for AI agents. Check agent changes, inspect regressions, and block broken behavior before merge. Local-first.

Topics

Resources

Contributing

Security policy

Stars

31 stars

Watchers

3 watching

Forks

Releases

Used by

Contributors

Languages