Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
-
Updated
Sep 1, 2026 - TypeScript
Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
Meta-harness optimization loop wired onto Islo sandboxes. POC: 0/5→5/5 in four proposer steps. Built on islo.dev.
Observe an agent run and GroundEval drafts the policy and diagram for you, no hand-written policy required, then scores what it checked, what it skipped, and what it wasn't allowed to touch.
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
Scenario Testing for AI Agents
AI-operated company. Building agent-friend: universal tool adapter for AI agents. @tool → OpenAI, Claude, Gemini, MCP. Live 24/7 on Twitch.
Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.
Project page for Meta-harness on Islo (POC). https://zozo123.github.io/meta-harness-on-islo-page/
A reasoning benchmark runner for comparing LLMs as OpenClaw agents use them. 52 prompts, 3 eval sets, 11 traps, LLM-as-judge, tier-based leaderboard.
Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilot CLI, Codex CLI).
LLM Agentic Workflow Evaluation & Configuration System
Transcript-first evaluation tool for comparing coding-agent sessions across Codex, Claude Code, and Pi.
Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, trigger quality, functional quality, regression protection, baseline value, model variance, rollout safety. Never gradients.
PandaProbe harness turns agent failures into fixes
开源通用 AI Agent 真实任务评测 · 同 Prompt、客观开奖、评分细则全公开 | Open-source evaluation of general-purpose AI Agents on real-world tasks with verifiable outcomes — by PingWest / 硅星人
Documented, reproducible finding: VulcanBench declarative grader mis-scored all functional tasks as 0.0 due to a repo-root pytest-cov addopts leak. Filed upstream issue #79.
Auto-generate evaluation rubrics from agent audit-log trajectories (PhoneWorld pattern applied to action logs)
Trustra Agent Eval Harness Open-core evaluation and tracing harness for AI agents. Run test datasets against an agent, score results, capture production traces, and build a tamper-evident hash-chained record of what an agent actually did.
Detect pytest config-leakage that corrupts agent-benchmark grading. Deterministic, no-LLM. Catches the VulcanBench --cov addopts mis-scoring bug.
Deterministic agent-eval / benchmark-grading hygiene audit CLI. Detects the silent mis-scoring bug class (pytest cov-gate leakage). Free lead-magnet for the paid harness-audit service.
To associate your repository with the agent-eval topic, visit your repo's landing page and select "manage topics."