Repository navigation
feat(hermes): native stats from Hermes token accounting - #12
Conversation
Hermes already bills every call into state.db (session_model_usage), so token counts need no estimating. integrations/hermes/cavemenko-stats.py reads the provider-reported output/input/cache-read counts and refuses to compute a savings percentage without a baseline session. Documents auto_load in config.yaml so the ruleset applies every session without typing /cavemenko, and points the Hermes skill at the same script. Four new tests, including a compile check so a broken script fails in CI rather than on the user's machine.
There was a problem hiding this comment.
Devin Review found 2 potential issues.
3 flags not posted on this PR by your GitHub settings β view them in Devin Review. (Configure)
| GROUP BY session_id, model | ||
| ORDER BY MAX(last_seen) DESC |
There was a problem hiding this comment.
π΄ Latest session omits other models' tokens
When a session uses multiple models, main reports only the most recently used model's tokens as the latest session. load returns separate rows per model, so --last N also counts model rows instead of sessions.
Learn more
The query returns one row for each session-and-model pair, ordered by each model's latest use. The default display then takes only the first row, so it omits calls made to any other model during that session. The table slices those same rows, so a session can take multiple slots in --last N.
Example: Session A uses model X for 10 output tokens and model Y for 20; session B uses only X. With A's Y call newest, the default prints 20 rather than A's total of 30. --last 2 can show A twice and omit B.
Recommended fix: Select the latest N distinct session IDs first. Aggregate token and call counts across all models for each session when printing session-level totals. If model-level detail is needed, render it beneath each selected session without counting it against N.
Was this helpful? React with π or π to provide feedback.
| if "--all" in args or "--last" in args: | ||
| n = 10 | ||
| if "--last" in args: | ||
| try: | ||
| n = int(args[args.index("--last") + 1]) | ||
| except (IndexError, ValueError): | ||
| pass | ||
| print(f"{'session':28} {'model':26} {'calls':>6} {'out':>9} {'in':>10} {'cache_read':>12}") | ||
| for sid, model, inp, out, cache, calls, last in rows[:n]: |
There was a problem hiding this comment.
π‘ All-session view silently drops older sessions
With more than ten usage rows, --all prints only ten and silently drops older sessions. n stays at its default of ten unless --last is present.
Learn more
The table path is shared by --all and --last, but both pass through a slice using n, which defaults to 10. The --all option never updates that limit, so it does not show the complete history.
Example: A database has 12 usage rows. Running --all shows ten rows, while the two oldest rows disappear without a warning.
Recommended fix: Keep the --all path unlimited; apply the numeric slice only when --last is requested. Apply that slice to distinct sessions rather than model rows.
Was this helpful? React with π or π to provide feedback.
There was a problem hiding this comment.
π‘ Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 83ab53d817
βΉοΈ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with π.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| SUM(input_tokens), SUM(output_tokens), SUM(cache_read_tokens), | ||
| SUM(api_call_count), MAX(last_seen) | ||
| FROM session_model_usage | ||
| GROUP BY session_id, model |
There was a problem hiding this comment.
Aggregate all models for each requested session
When a Hermes session uses more than one model, grouping by both session_id and model creates multiple rows for that session. The default path then reports only the most recently used model's counters, while --last N may return the same session several times and omit older requested sessions. This makes the advertised per-session token totals incomplete; aggregate by session or select all model rows belonging to each chosen session.
Useful? React with πΒ / π.
| return 0 | ||
|
|
||
| if "--all" in args or "--last" in args: | ||
| n = 10 |
There was a problem hiding this comment.
When the database contains more than ten session/model rows, --all still initializes n to 10 and the loop slices to rows[:n], so the command silently prints only ten entries despite being documented as the unrestricted per-session table. Use the full result set for --all and apply a limit only for --last.
Useful? React with πΒ / π.
What
The Claude Code
Stophook counts tokens by reading the transcript. Hermes needs no such thing β it already bills every call intostate.db(tablesession_model_usage) with provider-reported counts. This PR reads those instead of estimating anything.--last Ngives a per-session table so two sessions can be compared by hand.It deliberately does not print a percentage
There is no baseline session to compare against, so any ratio would be invented β the same failure mode the Claude Code path was fixed for in v2.1.0, where a bare prompt asked the model to report savings and the model made one up. The script states that instead of guessing.
Also
integrations/hermes/README.mddocuments installing the skill,auto_loadinconfig.yamlso the ruleset applies every session without typing/cavemenko, and the stats script/opt/data/skills/productivity/cavemenkogained the same sectioncavemanskill (English-only, upstream) now defers tocavemenkofor Ukrainian, instead of the two competing on every loadTests
135 pass (4 new). One compiles the stats script, so a syntax error fails in CI rather than on the user's machine.