Skip to content

feat(hermes): native stats from Hermes token accounting - #12

Merged
ruslanlap merged 1 commit into
masterfrom
feat/hermes-native-stats
Sep 30, 2026
Merged

ruslanlap merged 1 commit into
masterfrom
feat/hermes-native-stats

Conversation

@ruslanlap

@ruslanlap ruslanlap commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

What

The Claude Code Stop hook counts tokens by reading the transcript. Hermes needs no such thing β€” it already bills every call into state.db (table session_model_usage) with provider-reported counts. This PR reads those instead of estimating anything.

$ python3 integrations/hermes/cavemenko-stats.py
cavemenko β€” виміряно Π² Hermes (provider-reported, Π½Π΅ ΠΎΡ†Ρ–Π½ΠΊΠ°)
  session    20260930_064310_159746ac
  model      stealth/space-bunny-alpha
  api calls  230
  output     65,447 Ρ‚ΠΎΠΊΠ΅Π½Ρ–Π²
  input      282,020 Ρ‚ΠΎΠΊΠ΅Π½Ρ–Π²
  cache read 24,454,973 Ρ‚ΠΎΠΊΠ΅Π½Ρ–Π²
  avg out    284 Ρ‚ΠΎΠΊΠ΅Π½Ρ–Π² Π½Π° Π²ΠΈΠΊΠ»ΠΈΠΊ

--last N gives a per-session table so two sessions can be compared by hand.

It deliberately does not print a percentage

There is no baseline session to compare against, so any ratio would be invented β€” the same failure mode the Claude Code path was fixed for in v2.1.0, where a bare prompt asked the model to report savings and the model made one up. The script states that instead of guessing.

Also

  • integrations/hermes/README.md documents installing the skill, auto_load in config.yaml so the ruleset applies every session without typing /cavemenko, and the stats script
  • main README Hermes section points at both
  • the local Hermes skill at /opt/data/skills/productivity/cavemenko gained the same section
  • the separate caveman skill (English-only, upstream) now defers to cavemenko for Ukrainian, instead of the two competing on every load

Tests

135 pass (4 new). One compiles the stats script, so a syntax error fails in CI rather than on the user's machine.


Devin Review

Hermes already bills every call into state.db (session_model_usage), so
token counts need no estimating. integrations/hermes/cavemenko-stats.py
reads the provider-reported output/input/cache-read counts and refuses to
compute a savings percentage without a baseline session.

Documents auto_load in config.yaml so the ruleset applies every session
without typing /cavemenko, and points the Hermes skill at the same
script. Four new tests, including a compile check so a broken script
fails in CI rather than on the user's machine.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

3 flags not posted on this PR by your GitHub settings β€” view them in Devin Review. (Configure)

Devin Review

Comment on lines +33 to +34
GROUP BY session_id, model
ORDER BY MAX(last_seen) DESC

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

πŸ”΄ Latest session omits other models' tokens

When a session uses multiple models, main reports only the most recently used model's tokens as the latest session. load returns separate rows per model, so --last N also counts model rows instead of sessions.

Learn more

The query returns one row for each session-and-model pair, ordered by each model's latest use. The default display then takes only the first row, so it omits calls made to any other model during that session. The table slices those same rows, so a session can take multiple slots in --last N.

Example: Session A uses model X for 10 output tokens and model Y for 20; session B uses only X. With A's Y call newest, the default prints 20 rather than A's total of 30. --last 2 can show A twice and omit B.

Recommended fix: Select the latest N distinct session IDs first. Aggregate token and call counts across all models for each session when printing session-level totals. If model-level detail is needed, render it beneath each selected session without counting it against N.

Devin Review


Was this helpful? React with πŸ‘ or πŸ‘Ž to provide feedback.

Comment on lines +47 to +55
if "--all" in args or "--last" in args:
n = 10
if "--last" in args:
try:
n = int(args[args.index("--last") + 1])
except (IndexError, ValueError):
pass
print(f"{'session':28} {'model':26} {'calls':>6} {'out':>9} {'in':>10} {'cache_read':>12}")
for sid, model, inp, out, cache, calls, last in rows[:n]:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟑 All-session view silently drops older sessions

With more than ten usage rows, --all prints only ten and silently drops older sessions. n stays at its default of ten unless --last is present.

Learn more

The table path is shared by --all and --last, but both pass through a slice using n, which defaults to 10. The --all option never updates that limit, so it does not show the complete history.

Example: A database has 12 usage rows. Running --all shows ten rows, while the two oldest rows disappear without a warning.

Recommended fix: Keep the --all path unlimited; apply the numeric slice only when --last is requested. Apply that slice to distinct sessions rather than model rows.

Devin Review


Was this helpful? React with πŸ‘ or πŸ‘Ž to provide feedback.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

πŸ’‘ Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 83ab53d817

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with πŸ‘.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

SUM(input_tokens), SUM(output_tokens), SUM(cache_read_tokens),
SUM(api_call_count), MAX(last_seen)
FROM session_model_usage
GROUP BY session_id, model

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Aggregate all models for each requested session

When a Hermes session uses more than one model, grouping by both session_id and model creates multiple rows for that session. The default path then reports only the most recently used model's counters, while --last N may return the same session several times and omit older requested sessions. This makes the advertised per-session token totals incomplete; aggregate by session or select all model rows belonging to each chosen session.

Useful? React with πŸ‘Β / πŸ‘Ž.

return 0

if "--all" in args or "--last" in args:
n = 10

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Return every row for --all

When the database contains more than ten session/model rows, --all still initializes n to 10 and the loop slices to rows[:n], so the command silently prints only ten entries despite being documented as the unrestricted per-session table. Use the full result set for --all and apply a limit only for --last.

Useful? React with πŸ‘Β / πŸ‘Ž.

@ruslanlap
ruslanlap merged commit ed0ced1 into master Sep 30, 2026
4 checks passed
@ruslanlap
ruslanlap deleted the feat/hermes-native-stats branch September 30, 2026 08:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant