Repository navigation
feat(hermes): native stats from Hermes token accounting #12
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,42 @@ | ||
| # Hermes integration | ||
|
|
||
| Hermes loads cavemenko as a skill. Three parts: | ||
|
|
||
| ## 1. Install the skill | ||
|
|
||
| ```bash | ||
| # copy the ruleset into a Hermes skill | ||
| mkdir -p ~/.hermes/skills/productivity/cavemenko | ||
| cp skills/cavemenko/SKILL.md ~/.hermes/skills/productivity/cavemenko/ | ||
| ``` | ||
|
|
||
| Or point Hermes at the repo directly — the skill is one `SKILL.md`, no build step. | ||
|
|
||
| ## 2. Auto-load it | ||
|
|
||
| `~/.hermes/config.yaml`: | ||
|
|
||
| ```yaml | ||
| skills: | ||
| auto_load: | ||
| - cavemenko | ||
| ``` | ||
|
|
||
| Loaded every session, so no `/cavemenko` needed. Levels: `/cavemenko lite|full|ultra`, off via `звичайний режим`. | ||
|
|
||
| ## 3. Real stats | ||
|
|
||
| Hermes already bills every call into `state.db` (table `session_model_usage`), so token counts need no estimating: | ||
|
|
||
| ```bash | ||
| python3 integrations/hermes/cavemenko-stats.py # latest session | ||
| python3 integrations/hermes/cavemenko-stats.py --last 5 # per-session table | ||
| ``` | ||
|
|
||
| Prints output/input/cache-read tokens, API calls, average tokens per call — all provider-reported. | ||
|
|
||
| **It does not compute a savings percentage.** There is no baseline session to compare against, and a made-up ratio is worse than none. Compare two sessions with `--last 5` and read the ratio yourself. | ||
|
|
||
| Override the DB path with `HERMES_STATE_DB=/path/to/state.db` if it is not at `/opt/data/state.db`. | ||
|
|
||
| The Claude Code `Stop` hook (`hooks/cavemenko-stats.js`) is a separate path for that host; this script is the Hermes equivalent and reads the same underlying idea from Hermes' own accounting rather than counting transcript characters. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,80 @@ | ||
| #!/usr/bin/env python3 | ||
| """cavemenko stats for Hermes — real numbers from Hermes' own token accounting. | ||
|
|
||
| Hermes already bills every call into /opt/data/state.db (table | ||
| session_model_usage: output_tokens, input_tokens, cache_read_tokens). | ||
| There is no need to estimate anything: this reads the provider-reported | ||
| counts. Nothing here is inferred, extrapolated or invented. | ||
|
|
||
| Usage: | ||
| cavemenko-stats.py # latest session | ||
| cavemenko-stats.py --all # per-session table | ||
| cavemenko-stats.py --last N # last N sessions | ||
| """ | ||
| import os | ||
| import sqlite3 | ||
| import sys | ||
| import time | ||
|
|
||
| DB = os.environ.get("HERMES_STATE_DB", "/opt/data/state.db") | ||
|
|
||
|
|
||
| def load(db): | ||
| if not os.path.exists(db): | ||
| print(f"no Hermes state db at {db} — nothing measured yet") | ||
| return None | ||
| con = sqlite3.connect(f"file:{db}?mode=ro", uri=True) | ||
| rows = con.execute( | ||
| """ | ||
| SELECT session_id, model, | ||
| SUM(input_tokens), SUM(output_tokens), SUM(cache_read_tokens), | ||
| SUM(api_call_count), MAX(last_seen) | ||
| FROM session_model_usage | ||
| GROUP BY session_id, model | ||
| ORDER BY MAX(last_seen) DESC | ||
|
Comment on lines
+33
to
+34
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🔴 Latest session omits other models' tokens When a session uses multiple models, Learn moreThe query returns one row for each session-and-model pair, ordered by each model's latest use. The default display then takes only the first row, so it omits calls made to any other model during that session. The table slices those same rows, so a session can take multiple slots in Example: Session A uses model X for 10 output tokens and model Y for 20; session B uses only X. With A's Y call newest, the default prints 20 rather than A's total of 30. Recommended fix: Select the latest N distinct session IDs first. Aggregate token and call counts across all models for each session when printing session-level totals. If model-level detail is needed, render it beneath each selected session without counting it against N. Was this helpful? React with 👍 or 👎 to provide feedback. |
||
| """ | ||
| ).fetchall() | ||
| con.close() | ||
| return rows | ||
|
|
||
|
|
||
| def main(): | ||
| args = sys.argv[1:] | ||
| rows = load(DB) | ||
| if not rows: | ||
| return 0 | ||
|
|
||
| if "--all" in args or "--last" in args: | ||
| n = 10 | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. When the database contains more than ten session/model rows, Useful? React with 👍 / 👎. |
||
| if "--last" in args: | ||
| try: | ||
| n = int(args[args.index("--last") + 1]) | ||
| except (IndexError, ValueError): | ||
| pass | ||
| print(f"{'session':28} {'model':26} {'calls':>6} {'out':>9} {'in':>10} {'cache_read':>12}") | ||
| for sid, model, inp, out, cache, calls, last in rows[:n]: | ||
|
Comment on lines
+47
to
+55
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🟡 All-session view silently drops older sessions With more than ten usage rows, Learn moreThe table path is shared by Example: A database has 12 usage rows. Running Recommended fix: Keep the Was this helpful? React with 👍 or 👎 to provide feedback. |
||
| print(f"{sid[:28]:28} {model[:26]:26} {calls:>6} {out:>9,} {inp:>10,} {cache:>12,}") | ||
| return 0 | ||
|
|
||
| if not rows: | ||
| return 0 | ||
| sid, model, inp, out, cache, calls, last = rows[0] | ||
| when = time.strftime("%Y-%m-%d %H:%M", time.localtime(last)) if last else "unknown" | ||
| print("cavemenko — виміряно в Hermes (provider-reported, не оцінка)") | ||
| print(f" session {sid}") | ||
| print(f" model {model}") | ||
| print(f" остання {when}") | ||
| print(f" api calls {calls}") | ||
| print(f" output {out:,} токенів") | ||
| print(f" input {inp:,} токенів") | ||
| print(f" cache read {cache:,} токенів") | ||
| if out: | ||
| print(f" avg out {out // max(calls, 1):,} токенів на виклик") | ||
| print() | ||
| print(" Економія порівнюється з іншою сесією — подивись --last 5.") | ||
| print(" Без базової сесії відсоток не рахується: краще 0, ніж вигадка.") | ||
| return 0 | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| sys.exit(main()) | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,37 @@ | ||
| const fs = require('fs'); | ||
| const path = require('path'); | ||
| const { execFileSync } = require('child_process'); | ||
|
|
||
| const SCRIPT = path.join(__dirname, '..', 'integrations', 'hermes', 'cavemenko-stats.py'); | ||
|
|
||
| // Hermes integration: the stats script must exist, be executable Python, and | ||
| // never claim a savings percentage without a baseline. | ||
| describe('Hermes integration', () => { | ||
| test('stats script exists and is valid Python', () => { | ||
| expect(fs.existsSync(SCRIPT)).toBe(true); | ||
| const src = fs.readFileSync(SCRIPT, 'utf8'); | ||
| // A syntax error here is the failure mode we care about: a broken script | ||
| // that only fails at runtime on the user's machine. | ||
| execFileSync('python3', ['-c', `compile(open(${JSON.stringify(SCRIPT)}).read(), 'x', 'exec')`]); | ||
| expect(src).toContain('session_model_usage'); | ||
| }); | ||
|
|
||
| test('reads Hermes state db, not a transcript', () => { | ||
| const src = fs.readFileSync(SCRIPT, 'utf8'); | ||
| expect(src).toContain('state.db'); | ||
| expect(src).toContain('output_tokens'); | ||
| }); | ||
|
|
||
| test('declines to invent a savings percentage', () => { | ||
| const src = fs.readFileSync(SCRIPT, 'utf8'); | ||
| expect(src).toMatch(/не рахується|краще 0, ніж вигадка/); | ||
| }); | ||
|
|
||
| test('has an integration README', () => { | ||
| const readme = path.join(__dirname, '..', 'integrations', 'hermes', 'README.md'); | ||
| expect(fs.existsSync(readme)).toBe(true); | ||
| const text = fs.readFileSync(readme, 'utf8'); | ||
| expect(text).toContain('auto_load'); | ||
| expect(text).toContain('cavemenko-stats.py'); | ||
| }); | ||
| }); |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
When a Hermes session uses more than one model, grouping by both
session_idandmodelcreates multiple rows for that session. The default path then reports only the most recently used model's counters, while--last Nmay return the same session several times and omit older requested sessions. This makes the advertised per-session token totals incomplete; aggregate by session or select all model rows belonging to each chosen session.Useful? React with 👍 / 👎.