Provider-agnostic multi-turn loop, safe tool use, and 77.8% to 79.3% cost savings via context caching (real benchmark with Gemini, 5 repetitions per condition).
Agent Harness is a pure-Python implementation (no heavy frameworks or external runtime dependencies) of a harness for autonomous software engineering agents. It runs multi-turn loops with tool resolution (tool calling), supports multiple providers (Google Gemini and OpenAI-compatible / DeepSeek), and manages the session context to enable and measure the efficiency of Context Caching.
- Highlights
- Architecture
- Security
- Command Execution Policy
- Benchmark Results
- Installation and Setup
- Usage Examples
- Running the Test Suite
- Roadmap
- License
- Zero External Dependencies in Production: Strictly uses the Python standard library (
urllib,json,dataclasses,subprocess,argparse).pytestis the only development dependency. - Provider-Neutral: Unified message and tool protocol, dynamically adapted to Gemini's native schema (including Gemini 3+) or to the OpenAI / DeepSeek standard.
- Measured Context Caching: Stable deterministic prefixes (prefix invariance) let the Gemini API reuse cached tokens from the 2nd turn onward, reducing costs by more than 77% (up to 79.3% counterfactual, 5 repetitions per condition).
- Measured on a Second Provider: The same benchmark on
deepseek-flashreaches 91.7% direct savings (3 conditions, N = 5, 45 executions). These percentages are not comparable between providers: the cache discount is a property of the tariff — cached input costs 1/50 of the input price on DeepSeek and 1/10 on Gemini. - Active Context Pruning: head-body-tail algorithm that keeps the cacheable prefix intact at the top (head), discards intermediate turns when the token budget overflows (body), and preserves the most recent turns (tail).
- Safe Tools with Blocklist and Path Protection: Terminal commands, file reading, writing, and search come with strict restrictions against destructive commands and path traversal.
On each turn, the harness sends the structured history to the model. If the model replies requesting tool calls (tool calls), the harness dispatches the executions to the tool registry, appends the outputs as tool response messages, and triggers the next turn. When the model returns plain text, the cycle ends.
+-----------------------------+
| User (Instruction/CLI) |
+--------------+--------------+
|
v
+---------------------------+
| Context Assembly |
| [Head + Body + Tail] |
+-------------+-------------+
|
v
+---------------------------+
| Provider Adapter |
| (Gemini / DeepSeek/OpenAI)|
+-------------+-------------+
|
v
+---------------+
| LLM Call |
+-------+-------+
|
+-----------------+-----------------+
| |
[Tool Calls] [Final Text]
| |
v v
+----------------------+ +-------------------+
| Tool Registry | | Execution End |
| - executar_comando | | Cost Summary |
| - ler_arquivo | +-------------------+
| - escrever_arquivo |
| - buscar_no_projeto |
+----------+-----------+
|
v
+----------------------+
| Append Results |
| Prune if Over Limit |
+----------+-----------+
|
+----> Next Turn (Loop)
Modern LLM providers such as Google Gemini support automatic implicit context caching for static prefixes exceeding 4,096 tokens. For the cache to be used, the conversation's initial prefix must be identical on every turn (Prefix Invariance).
The context module splits the window into 3 regions:
+--------------------------------------------------------------------+
| 1. HEAD (Deterministic & Stable) |
| - System Prompt |
| - Repository Context (when provided) |
| * NEVER contains timestamps, volatile hashes, or random order |
| * Primary cache-hit target across turns 2..N |
+--------------------------------------------------------------------+
| 2. BODY (History Pruning) |
| - Intermediate tool-call and reasoning turns |
| - Dropped in pairs (call + response) if tokens overflow |
+--------------------------------------------------------------------+
| 3. TAIL (Local Preservation) |
| - Most recent turns (guaranteeing immediate coherence) |
| - Agent's final answer |
+--------------------------------------------------------------------+
When the --no-cache flag is provided, a dynamic header containing a timestamp (time.time(), float seconds) is deliberately injected at the beginning of the user message, breaking prefix invariance and forcing a cache miss on every turn for comparison purposes.
The harness exposes 4 native tools to the model:
executar_comando: Execution of whitelist-controlled commands without a shell (shell=False) in the project directory, with a 1 MB cap on subprocess output.ler_arquivo: Reading of text files with a maximum limit of 200 KB per file.escrever_arquivo: Creation and overwriting of files with a 1 MB limit.buscar_no_projeto: Regex pattern search over the contents of the project's code/text files (ignores the.venv,__pycache__,.git,.pytest_cache,build, anddistfolders; allows optional filtering by extension; ignores binaries and files larger than 1 MB; maximum of 50 results; 10-second timeout and extended ReDoS protection). It does not search by filename.
- Critical Command Blocklist: Regex-based blocking of the dangerous patterns mapped in
PADROES_BLOQUEADOS:format,diskpart,shutdown,rd /s(or/q),rmdir /s(or/q),rm -rf,reg delete,del /s(or/for/q),erase /s(or/for/q),cipher /w, andtaskkill /f /im. - Configurable Protected Paths (
CAMINHOS_PROTEGIDOS): Centralized inconfig.pyand checked by the authoritative functionresolver_caminho_seguro:bloqueio_total(reading and writing blocked):.env(blocking.env,.env.*, and.env_*, with a safe exception for templates such as.env.example,.env.sample, and.env.template) and.git/(all repository objects, configs, and references);somente_escrita(reading allowed if needed, writing categorically blocked):.github/(protects GitHub Actions workflows and automation files against overwriting or corruption by the model).
- Protection against Path Traversal and Symlink Escape: Strict validation via
resolver_caminho_seguroensuring that no path accesses folders above the project root (..forbidden) or escapes the project tree through symlinks or NTFS junctions. - Hard Timeouts and Memory Caps: Each command execution has a default limit of 30 seconds and chunked reading capped at no more than 1 MB of subprocess output, preventing hangs or memory exhaustion caused by noisy scripts. File search is limited to 10 seconds.
- Kwargs Validation in the Loop against the Schema: When the loop dispatches tools, any argument not expressly declared in the canonical schema (
TOOLS) is summarily rejected before execution, preventing the injection of spurious parameters (base_dir, etc.). - Cross-Platform Protection via Whitelist: Shell-less process execution (
shell=False) restricted to a whitelist of allowed binaries prevents destructive commands from running on both Windows (cmd.exe) and POSIX/Linux environments (rm -rf,mkfs, etc.). Native handlers such astype,dir,where, andfindstrhave transparent emulation in Python for full portability.
The combination of command blocklist, redirection inspection, ReDoS protection, and protected-path validation is a pragmatic mitigation layer (defense-in-depth) designed for assisted local development and testing, NOT a formal security sandbox:
- What the blocklist does NOT protect against:
- Arbitrary binaries invoked by the model: Tools such as
curl,wget, orbitsadmindownloading external scripts or executables. - Execution of arbitrary code in interpreters: Commands such as
python -c "..."orpowershellrunning arbitrary, obfuscated, or Base64-encoded dynamic payloads. - Complex scripts and substitutions: Batch scripts or complex chains with delayed expansion of environment variables (
cmd /v:on /c "%VAR%").
- Arbitrary binaries invoked by the model: Tools such as
- Untrusted Execution Requires a Formal Sandbox:
- Environments that execute arbitrary or untrusted code require formal kernel-level isolation via a container (Docker sandbox / gVisor) or an ephemeral microVM (Firecracker), as foreseen in the architectural roadmap (W7+).
Warning
Security Notice (Honest Disclaimer): String blocklist-based heuristics reduce accidents, but they do not replace a real isolation environment against adversarial agents. For production environments open to arbitrary code, isolation in containers (Docker sandbox / gVisor) or ephemeral virtual machines is the mandatory next step.
As of Milestone W7 (and the W7.7 consolidation), the harness adopts a strict whitelist execution policy without a shell (shell=False), eliminating shell grammar interpretation as an attack and evasion vector.
-
Primary Layer (Shell-less Whitelist):
- Execution without an Interpreter (
shell=False): The harness invokes external executables directly via subprocess without a shell. There is no pass throughcmd.exeorsh. Metacharacters such as|,&,;,>,<,$VAR, and%VAR%are not interpreted by the operating system, being treated strictly as literal arguments or rejected. - Own Tokenizer (
_tokenizar): Splits the command line by spaces while respecting single and double quotes (the content between quotes becomes a single token without the surrounding quotes), rejecting commands with unbalanced quotes before any execution. - Authorized Executables (
COMANDOS_PERMITIDOS):dir: Inspection of project directories (executed natively in Python to avoid depending on the shell).type: Quick reading of files (executed natively with path validation and blocking of protected files).python: Strict execution of Python scripts inside the project (python <file>.py). Inline interpreter flags (-c,-m,-i), references with.., and absolute paths are categorically blocked.git: Pure Metadata via Strict Allowlist (git status,git ls-files, andgit log --oneline).- Subcommands that display content or diffs (
diff,show,log -p, etc.) were completely removed from the whitelist. - Flag Allowlist: For
status, only metadata flags are accepted (--short,-s,--porcelain,--branch,-b,--untracked-files,-u,--ignored,--long), blocking-v,-vv,--verbose,-z,--null. Forls-files, only metadata is accepted (--cached,-c,--others,-o,--stage,-s,-t,--full-name,--exclude-standard, etc.). Forlog,--onelineis required and--stat,-n <N>,-n<N>,--max-count=<N>, and safe paths are accepted. - Pathspec Magic: Any argument with a
:prefix (e.g.,:(top).env) is categorically rejected to prevent evasion of path filters. - Fail-Safe Redactor with NUL Support: Outputs are inspected line by line and by NUL records (
\x00). Lines or records that cite protected files (including rename patterns{old => new}) are summarily omitted. - Over-Redaction Trade-off and Audit Metric: Commit messages that cite protected files (e.g.,
add .env) cause the entire corresponding line ofgit log --onelineto be omitted from the response to the model. The redactor counts the total number of omitted lines/records and emits an audit log tostderrfor the developer, without exposing metrics to the model, in order to avoid inference about the existence of protected files.
- Subcommands that display content or diffs (
findstr: Fast text search (with transparent fallback on non-Windows platforms, acting as a direct substring search in files without support for advanced flags of Windows' native findstr).where: Location of safe executables in PATH (with a cross-platform fallback viashutil.which, automatically handling the mapping ofpythontopython3when necessary).echo: Printing text to the terminal (without allowing redirection via the shell).
- Any binary outside the whitelist (e.g.,
rm,del,curl,powershell,cmd,bash,sh,nc) is blocked immediately with exit code-1and an explanatory message.
- Execution without an Interpreter (
-
Secondary Layer (Verb Blocklist and Redirection Protection — Defense in Depth):
- Exclusive Application on the Verb (Elimination of False Positives in Reading): The destructive pattern blocklist (
PADROES_BLOQUEADOS) is evaluated exclusively against the verb of the command (and the internal commands ofcmd /corpowershell -cwrappers), never over the body of the arguments. Evidence: legitimate searches such asfindstr "rm -rf" DOC.mdorfindstr shutdown DOC.mdwere unduly refused (rc=-1), whilefindstr "format C:"passed. With shell-less execution (shell=False), authorized reading binaries do not execute commands embedded in text, which makes it safe and necessary to restrict the blocklist to the executed verb. - As an additional backstop safeguard, destructive verbs (
format,diskpart,shutdown,rm,del, etc.) and redirections to protected files (.env,.git/,.github/) are categorically intercepted.
- Exclusive Application on the Verb (Elimination of False Positives in Reading): The destructive pattern blocklist (
The migration from the pure blocklist approach to the shell-less whitelist was driven by empirical evidence obtained over 3 rounds of multi-model code review (v1, v2, v3), in which 6 evasion vectors against blocklist-based execution with shell=True were analyzed (5 reproduced in the audits and 1 identified in the architectural analysis):
- Command Chaining via
&or&&: Allowed commands masking subsequent dangerous instructions (e.g.,dir & type .env). - Output Redirections to Critical Files: Use of stream operators (
echo x > .envortype a > b && echo x >> .git/config) to corrupt credentials or git history. - Interpreter Wrappers: Invocation through secondary interpreters (e.g.,
cmd /c "type .env"orpowershell -c "Get-Content .env"), bypassing simple lexical checks. - Redirection with Numeric Descriptors: Use of stream descriptors (
echo x 1> .envorecho x 2>> .env) that escaped standard redirection regexes. - Escaping via Quotes and Spaces: Variations with nested quotes and obfuscated relative paths that the shell interpreter decoded at runtime.
- Inline Code Execution via Flags (identified in the architectural analysis): Use of
python -c "import os; os.system('...')"to run arbitrary code without triggering shell keywords.
These tests demonstrated that no regular-expression-based blocklist is capable of exhaustively covering the recursive grammar of a shell (cmd.exe or sh). Disabling the shell (shell=False) and limiting execution to a rigorous whitelist with semantic validation of arguments eliminates this entire class of attacks by architectural definition.
- Execution of Local Python Scripts: The agent is allowed to run
python <file>.pyto execute its own tests. A script generated by the model with malicious behavior can still be executed if it is written into the project. - Need for a Sandbox for Unrestricted Autonomy: For environments open to arbitrary and unsupervised tasks, container-level isolation (Docker sandbox / gVisor) or an ephemeral microVM (Firecracker) remains the definitive containment standard (W8+).
The automated benchmark (python -m harness --bench) runs 3 real software engineering tasks against the repository, comparing Cache Enabled mode (invariant prefix) versus Cache Disabled mode (prefix intentionally broken on every turn).
The suite currently runs three conditions with N repetitions each (default 5, minimum 3):
- ON — cache ON + history pruning ON (baseline).
- OFF — cache OFF + history pruning ON (isolates the cache effect).
- SEM_PODA — cache OFF + history pruning OFF. Declared explicitly, not inferred: the third condition isolates the effect of the pruning policy, and its comparison base is the OFF condition.
The set of conditions actually measured is selectable (--condicoes ON,OFF) and the header declares exactly which ones ran: a 6-cell design is never reported as a 9-cell one.
Each execution is recorded individually and never aggregated before being written. Per task × condition the suite reports the median and range of cost, latency, turns and tokens, plus the success rate (n of N). Because the validator changed after the pilot was observed, the record stores both criteria — the current one and the strict historical one (test file containing the substring assert) — and the summary reports both success rates side by side, so the effect of the change is visible instead of chosen after the fact. Raw results are written to bench_<date>_<provider>.json under a header stating date, driver and version, effective model, reasoning_effort, max_tokens, tariff window (peak/off-peak), tariff used, account coverage, number of repetitions, replacement policy and repository commit.
Failures are classified instead of lumped together. A task failure (there were turns, the validation did not pass) counts in the denominator as a normal failure. An abort (0 turns, or a connection/API error) means nothing was measured — it is instrument failure, not a model attempt: it is recorded with tipo_falha, excluded from the success-rate denominator, and replaced, up to 2 replacements per cell. If that ceiling is exceeded, the collection stops and reports why in the header.
Measured on the gemini-3.8-flash model with the repository context. A single run per condition — which is exactly the fragility that the E1 phase of the measurement removed:
| Task | Mode | Turns | Prompt Tokens | Cached Tokens | Cost (USD) | Savings | Success |
|---|---|---|---|---|---|---|---|
| T1: Module Map | ON | 5 | 184,217 | 130,804 | $0.052799 | 62.6% | YES |
| (create bench_mapa.md with the modules of src/harness) | OFF | 4 | 146,818 | 0 | $0.112510 | 0.0% | YES |
| T2: Generation + Own Test | ON | 5 | 176,620 | 163,510 | $0.022876 | 82.8% | YES |
| (create bench_math.py + bench_test_math.py and run them) | OFF | 5 | 176,743 | 0 | $0.133562 | 0.0% | YES |
| T3: Spec + Own Test | ON | 6 | 211,947 | 196,185 | $0.027912 | 82.6% | YES |
| (create bench_contador.py + bench_test_contador.py and run them) | OFF | 6 | 212,945 | 0 | $0.161520 | 0.0% | YES |
| TOTAL CACHE ON | ON | 16 | 572,784 | 490,499 | $0.103586 | 76.2%* | 3/3 |
| TOTAL CACHE OFF | OFF | 15 | 536,506 | 0 | $0.407592 | 0.0% | 3/3 |
The absolute costs of the two measurements are not comparable with each other: the context head has grown since then (T1 ON went from 184,217 prompt tokens in 5 turns to 306,385 in 8). The comparable quantity is the ratio — and that is what the table below reports.
Gemini, 6 cells (T1/T2/T3 × ON/OFF), 5 executions per cell, 30 executions, thinking_level=medium, max_output_tokens=65536, off-peak tariff window (Gemini has no peak/off-peak tariff), commit 9484abe. Per cell the figures are the median of the 5 executions and the range (min–max) of cost; the TOTAL rows are sums over the 15 executions, not medians:
| Task | Mode | Executions | Turns (median) | Prompt Tokens (median) | Cached Tokens (median) | Median Cost (USD) | Cost Range (USD) | Counterfactual Savings | Success |
|---|---|---|---|---|---|---|---|---|---|
| T1: Module Map | ON | 5 | 8 | 306,385 | 269,724 | $0.049482 | $0.045228 – $0.065823 | 77.7% | 4/5 |
| OFF | 5 | 7 | 260,349 | 0 | $0.198745 | $0.171860 – $0.234203 | 0.0% | 5/5 | |
| T2: Generation + Own Test | ON | 5 | 5 | 173,654 | 155,326 | $0.026202 | $0.024325 – $0.026697 | 80.2% | 5/5 |
| OFF | 5 | 6 | 208,058 | 0 | $0.157206 | $0.131211 – $0.157969 | 0.0% | 5/5 | |
| T3: Spec + Own Test | ON | 5 | 6 | 210,751 | 192,092 | $0.032047 | $0.031169 – $0.034912 | 80.8% | 5/5 |
| OFF | 5 | 5 | 175,164 | 0 | $0.133046 | $0.132314 – $0.159830 | 0.0% | 5/5 | |
| TOTAL (sum of 15) | ON | 15 | 97 | 3,523,952 | 3,134,671 | $0.552130 | — | 79.3% | 14/15 |
| TOTAL (sum of 15) | OFF | 15 | 91 | 3,274,330 | 0 | $2.486055 | — | 0.0% | 15/15 |
The same suite was also run on deepseek-flash (3 conditions, N = 5, 45 executions), with 91.7% direct savings and 91.6% counterfactual (total USD 0.7556). The percentages are not comparable between providers: the cache discount is a property of the tariff, not of the harness — cached input costs 1/50 of the input price on DeepSeek and 1/10 on Gemini. The same prefix-stability work is worth more where the discount is larger, and that is not a merit of the harness.
- Tasks Evaluated:
- T1 (map): Create
bench_mapa.mdlisting each module ofsrc/harnesswith one sentence about its responsibility, based on the repository context and without changing existing files. - T2 (generation-with-test): Create
bench_math.pywith asoma(a, b)function andbench_test_math.pywith non-zero output validation in case of error, and runpython bench_test_math.py. - T3 (file-spec): Create
bench_contador.pywith acontar_palavras(t)function andbench_test_contador.pywith specific test cases, and runpython bench_test_contador.py.
- T1 (map): Create
- The model autonomously decides the number of turns for each task (in the current measurement, medians of 8 turns with cache and 7 without in T1, exploring files independently).
- Alternating Execution Order (Bias Mitigation): To prevent a fixed order (ON always before OFF) from introducing cache warm-up bias or latency advantages on the provider's server, the benchmark rotates the order of the conditions on every round, so no condition is always measured first. In the N = 1 run this was done per task: T1 ran ON -> OFF, T2 ran OFF -> ON, and T3 ran ON -> OFF.
- The Two Bases for Calculating Savings (Full Transparency):
- 77.8% — Direct Comparison between Distinct Runs: The real cost of the complete suite with Cache OFF was $2.486055 (91 turns), while with Cache ON it was $0.552130 (97 turns). The direct ratio
($2.486055 - $0.552130) / $2.486055results in 77.8% real savings, even with the agent running 6 extra turns in the cached round. - 79.3% — Turn-by-Turn Counterfactual Savings (*): The value reported in the benchmark table (
$2.115903 / 79.3%) is the counterfactual sum calculated by the harness over the exact Cache ON run: what those 97 turns would have cost if no token had been served from cache ($2.668033 counterfactual) versus what they actually cost with the cache discount ($0.552130 real).
- 77.8% — Direct Comparison between Distinct Runs: The real cost of the complete suite with Cache OFF was $2.486055 (91 turns), while with Cache ON it was $0.552130 (97 turns). The direct ratio
- The cache only starts acting from the 2nd turn of each task, when the conversation's initial prefix has already been ingested and recognized by the provider.
- Python 3.10 or higher (tested up to Python 3.14).
- Google Gemini API key (
GEMINI_API_KEY) and/or DeepSeek (DEEPSEEK_API_KEY).
# Clone the repository
git clone https://github.com/brmarcosbr/harness.git
cd harness
# Create and activate the virtual environment
python -m venv .venv
# Windows (PowerShell)
.\.venv\Scripts\Activate.ps1
# Linux / macOS
source .venv/bin/activate
# Install the local package in editable mode with dev dependencies
pip install -e .[dev]Create the .env file at the project root:
cp .env.example .envEdit the .env with your credentials:
GEMINI_API_KEY=sua_chave_gemini_aqui
DEEPSEEK_API_KEY=sua_chave_deepseek_aqui
HARNESS_PROVIDER=gemini
# Optional: reasoning effort and output ceiling sent to the OpenAI-compatible endpoint.
# Defaults: high and 65536. Accepted effort values: low, high, max (medium and xhigh
# are aliases that resolve to high). Invalid values are ignored with a warning on stderr.
# HARNESS_REASONING_EFFORT=high
# HARNESS_MAX_TOKENS=65536
# Optional: output ceiling and reasoning level sent to the Gemini endpoint
# (generationConfig.maxOutputTokens and generationConfig.thinkingConfig.thinkingLevel).
# Defaults: 65536 (the model's documented output limit) and medium (the model's documented
# default level, so declaring it does not change behaviour — it only makes the choice
# auditable). Accepted levels: low, medium, high. Invalid values are ignored with a warning.
# HARNESS_GEMINI_MAX_OUTPUT_TOKENS=65536
# HARNESS_GEMINI_THINKING_LEVEL=mediumpython -m harness --tarefa "liste os arquivos desta pasta"python -m harness --provider gemini --contexto-repo --tarefa "Analise a arquitetura de providers e sugira um novo adaptador"python -m harness --provider gemini --contexto-repo --no-cache --tarefa "Run the same analysis"python -m harness --provider deepseek --tarefa "Write a function that computes the Fibonacci sequence"python -m harness --benchBelow is the transcript of the real output of a simple run with the gemini-3.8-flash model:
============================================================
AGENT HARNESS
Provider: GEMINI | Modelo: gemini-3.8-flash
============================================================
Tarefa: liste os arquivos desta pasta
Cache de contexto: ON (head sha256: c7ef1665ff5cbcc2)
Diretório atual: D:\Projetos\Antigravity\Harness
------------------------------------------------------------
>>> TURNO 1 / 8
Prompt tokens: 640
Completion tokens: 19
Total tokens: 862
Tamanho do texto gerado: 0 caracteres
[Tool Call] Função: 'executar_comando'
[Tool Exec] Executando 'executar_comando' com args: {'comando': 'dir /b'}
[Tool Output] Código: 0 | Stdout: 112 chars | Stderr: 0 chars
>>> TURNO 2 / 8
Prompt tokens: 930
Completion tokens: 20
Total tokens: 1044
Tamanho do texto gerado: 0 caracteres
[Tool Call] Função: 'ler_arquivo'
[Tool Exec] Executando 'ler_arquivo' com args: {'caminho': 'README.md'}
[Tool Output] Status: OK | Chaves: ['sucesso', 'conteudo', 'tamanho_bytes']
>>> TURNO 3 / 8
Prompt tokens: 1093
Completion tokens: 87
Total tokens: 1193
Tamanho do texto gerado: 214 caracteres
[Resposta Final do Modelo]:
Arquivos e diretórios presentes na pasta atual:
- `.env`
- `.env.example`
- `.github/`
- `.gitignore`
- `.pytest_cache/`
- `.venv/`
- `LICENSE`
- `pyproject.toml`
- `README.md`
- `src/`
- `tests/`
- `__pycache__/`
============================================================
RESUMO DA EXECUÇÃO
============================================================
Turnos utilizados: 3 de 8
Total Prompt Tokens: 2663
Total Cached Tokens: 0
Total Completion Tokens: 126
Total Geral de Tokens: 2789
Tokens estimados do historico final: 666
Custo real (com cache): $0.002470 USD
Custo se sem cache: $0.002470 USD | Economia: $0.000000 USD (0.0%)
Modelo final: gemini-3.8-flash
============================================================
(Note: In this short 3-turn task, the prefix contained only the system prompt with 640 tokens, below the 4,096-token threshold for activating Gemini's implicit cache. The totalTokenCount reported by the Gemini API may include thinking/internal reasoning tokens in models that support native thinking. When providing --contexto-repo, the single largest prompt observed in the E1 measurement reaches 43,293 tokens (prompt_tokens_max in bench_20260911_gemini.json) and the counterfactual savings reach 79.3%, as demonstrated in the benchmark).
The 175 unit tests run 100% offline (they use mocks and fake providers, with no network dependency or API quota consumption):
pytest tests/ -qExpected output:
............................................................................................................................................................................... [100%]
175 passed in 4.85s
The tests cover:
- Price calculation and accuracy (Gemini and DeepSeek with cache windows, verification date, and
PRECO_*environment overrides). - Reasoning effort and output ceiling declared in the request body (
HARNESS_REASONING_EFFORT/HARNESS_MAX_TOKENSoverrides) instead of depending on server defaults. - Benchmark collection rigor: three conditions (cache ON, cache OFF and no pruning), N repetitions with median and range per task × condition, peak/off-peak tariff window as a pure function, and agent instruction files (
AGENTS.md,CLAUDE.md,.claude/) kept out of the repository context head. - Per-turn prefix telemetry (head hash and stability flag) and latency decomposed into model and tool time.
- Collection robustness: printing that cannot kill a paid run (UTF-8 reconfiguration plus a stream that degrades instead of raising), abort versus task failure classification with the success rate computed only over non-aborted executions, a replacement policy capped per cell, and results written explicitly as UTF-8 (
ensure_ascii=False), so the artifact does not depend on the machine's local encoding. - Tool security resolution and validation (shell-less whitelist, strict verb blocklist, timeouts, path traversal).
- Auditable guarantee that metacharacters are literals (
dir & rm -rf /,type a.txt > b.txtwithout overwriting). - Strict validation of git, python, findstr, and where arguments, and protection against symlink/junction traversal on all surfaces.
- Audit metric for git over-redaction with notification on
stderr(without leaking to the model). - Visibility with a warning on
stderrwhen the positional tool fallback in the loop is triggered. - Sanitization of environment credentials against leaking into subprocesses (by token segment/position).
- Inversion of execution layers with primary whitelist validation without false positives on safe commands (
echo format,findstr "rm -rf"). - Serialization and conversion of schemas in the Gemini and OpenAI formats.
- Normalization and parsing of multi-turn responses with tool calling.
- Context pruning and prefix invariance guarantee.
- Main loop execution and benchmark suite.
- Docker Sandboxing: Run tools in ephemeral, isolated containers.
- Streaming Support: Real-time text response via SSE/WebSockets.
- LLM Summarization: Summarize pruned history blocks instead of merely discarding them.
- Multi-Agent Extension: Orchestration of sub-agents with distinct specialties (research, coding, and review).
Distributed under the MIT license. See LICENSE for more information.
Author: Bruno Marcos Bonifacio