A test runner for agentskills.io-style AI agent skills
-
Updated
Aug 5, 2026 - TypeScript
A test runner for agentskills.io-style AI agent skills
SWE-bench for your codebase — mine your merged PRs into local, contamination-free coding-agent benchmarks. Adapters: claude-code, aider (Opus 4.7 / GPT-5.5 / Sonnet 4.6 / Gemini 3.1 Pro).
AI agent evolving strategies through automated self-play overnight. Generic framework with GEPA-inspired feedback loop and Elo tracking.
An implementation of the Anthropic's paper and essay on "A statistical approach to model evaluations"
Replay real agent traces through cheaper models to prove which swaps are safe.
Create your self-hosted, open-source Operator model.
Observe an agent run and GroundEval drafts the policy and diagram for you, no hand-written policy required, then scores what it checked, what it skipped, and what it wasn't allowed to touch.
The trust layer for your AI agent on any platform. Run evaluation. Trace your agent steps. Full Observability. Run Model portability analysis. Autotune your agent.
BondLens: evidence-first Chinese bond analysis agent — deterministic tools, optional LLM narration under guardrails, Trust Layer, live/snapshot/static data, Docker/CI.
Test-driven harness engineering for Python agents: own the loop, capture failures, and gate every change.
Public Agent Anvil leaderboard submissions and generated index
A cli tool for evaluating coding agent plugins with a multi-tier approach.
LOAB: A benchmark for evaluating LLM agents on end-to-end mortgage lending operations under real regulatory constraints.
Legal Action Boundary Eval (LABE): public proxy eval for legal AI workflows at the action boundary
Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Offline prompt evolution engine for multi-agent systems, benchmark suites, and inspectable prompt DNA.
Evidence, governance, and static reports for agentic runs
Portable, inspectable knowledge bundles for LLM agents.
Build a private evaluation dataset to optimize your organization's token costs.
Reproducible evaluation harness for hidden coordination variables in multi-agent LLM systems.
To associate your repository with the agent-evals topic, visit your repo's landing page and select "manage topics."