Skip to content

feat(bench): honest 3-arm A/B vs JuliusBrussee/caveman, drop invented abbr - #10

Merged
ruslanlap merged 1 commit into
masterfrom
feat/honest-3arm-benchmark
Sep 30, 2026
Merged

ruslanlap merged 1 commit into
masterfrom
feat/honest-3arm-benchmark

Conversation

@ruslanlap

@ruslanlap ruslanlap commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

What

@JuliusBrussee/caveman (108k⭐) is the real English project behind this idea, and its eval harness does something my previous benchmark did not: it adds a control arm — a plain Answer concisely. prompt.

Its own README says it plainly:

The honest delta for any skill is <skill> vs __terse__. Comparing a skill to the no-system-prompt baseline conflates the skill with the generic terseness ask, which is what an earlier version of this harness did and is why its numbers were inflated.

That is exactly what v2.1.0's numbers did. This PR re-runs it correctly.

Result — 3 arms, claude-opus-4-6, real maintainer tasks

Arm chars facts vs base vs terse control
base (no instructions) 9625 5/13 — +66%
terse (control) 5793 5/13 −40% —
caveman (upstream, en) 4606 4/13 −53% −21%
cavemenko (uk) 2970 5/13 −70% −49%

Two things this shows:

  1. cavemenko is stronger on Ukrainian — −49% vs upstream's −21% on top of the same control. Upstream has no Ukrainian rules (pro-drop, short forms, зламано/треба/ok), and it lost a ground-truth fact on the Ukrainian case while cavemenko lost none.
  2. My previous 57–71% was inflated — that was a two-arm comparison against no instructions at all. Corrected number is −49% on top of generic terseness.

What I took from upstream

Its measured rule, now in skills/cavemenko/SKILL.md:

never invent new abbreviations (cfg/impl/req/res/fn) tokenizer split them same as full word: zero token saved, reader still decode. Full word cheaper AND clearer. No causal arrows (→) either own token, save nothing.

v2.1.0's ultra level recommended exactly those. It now bans them. Invented abbr and → arrows save zero tokens; Ukrainian words that are genuinely shorter (кеш, БД) stay.

Also

  • README leads with the 3-arm table and credits caveman as the English sibling, instead of claiming a number that came from a broken harness
  • repo description updated to match
  • 131 tests pass; the abbreviation test now asserts the ban rather than the old table

🤖 Generated with Claude Code


Devin Review

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

2 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)

Devin Review

Comment thread skills/cavemenko/SKILL.md
- **Підрядні:** `функція-source error` не `функція, яка викликає error`. `вчорашній код` не `код, який було написано вчора`.
- **English де коротше:** auth, bug, fix, run, check, do, cache, deploy, update, delete, config, repo, etc, eg, ASAP.
- **Українське де коротше:** кеш, БД.
- **Ukrainian wins over inventing a new one:** `кеш`, `БД`. **Custom abbr і стрілки `→` — ні:** токенайзер ріже `імпл` так само як `implementation`, 0 токенів econ, читається гірше. Якщо скорочення не коротше за повне слово — пиши повне.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Configured abbreviations stop working

With custom abbreviations configured, Custom abbr і стрілки tells Claude to avoid them. The activation hook still injects those abbreviations, so user-defined shorthand conflicts with the active rules.

Learn more

The SessionStart hook loads user and project abbreviations and appends their table to the ruleset after the main skill. The new blanket prohibition on custom abbreviations also reaches abbreviations deliberately configured by users. As a result, Claude receives conflicting instructions for a documented feature in every active mode.

Example: A user sets "КБ": "кодова база" in their abbreviation file. The activation hook supplies КБ in its custom table, but the skill says not to use custom abbreviations.

Recommended fix: Restrict the ban to invented, unconfigured abbreviations. Explicitly exempt abbreviations supplied by the user's or project's configuration in the skill rule and its ultra description.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment thread skills/cavemenko/SKILL.md
| **lite** | ~40% baseline | Cut воду/ввічливість/hedging. Повні речення. Для docs/пояснень |
| **full** | ~25-30% baseline | + pro-drop, тире, short forms, наказовий, фрагменти. Default |
| **ultra** | ~15-20% baseline | + abbr (БД/фн/імпл/конф/env/dep), arrows `X → Y`, 1 sentence де вистачає |
| **ultra** | ~15-20% baseline | + strip conjunctions, одне слово де вистачає, кожен факт один раз. **Без** вигаданих abbr і **без** стрілок `→` (токенайзер ріже `імпл` так само як `implementation` — 0 токенів, гірше читається) |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Ultra examples still promote banned arrows

In ultra mode, the new arrow ban conflicts with the retained ultra examples. The activation hook injects those arrow-filled examples, so Claude still receives instructions to produce them.

Learn more

The activation hook filters the skill to the active level, preserving the ultra examples later in the skill. Both ultra examples use →, while the new ultra rule forbids it. The runtime prompt therefore pairs the new ban with examples that demonstrate precisely the banned syntax.

Example: For a React rerender question in ultra mode, Claude sees Inline obj → new ref → rerender as its ultra example despite being told to avoid →.

Recommended fix: Rewrite the ultra examples without arrows, including the token-expiry comparison, so the examples match the new ultra-mode rule.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

… abbr

Adds the control arm the upstream caveman eval uses: a plain
'Answer concisely.' prompt. Comparing a skill to a no-instructions
baseline conflates the skill with generic terseness and inflates the
number — that is exactly what our previous 57-71% figures did.

3-arm result on real maintainer tasks, claude-opus-4-6:
  base   9625 chars, 5/13 facts
  terse  5793 chars, 5/13 facts   (control)
  caveman (upstream, en) 4606 chars, 4/13 facts  -> -21% vs control
  cavemenko (uk)         2970 chars, 5/13 facts  -> -49% vs control

cavemenko wins on Ukrainian because upstream has no Ukrainian rules
(pro-drop, short forms) and lost a fact on the Ukrainian case.

Adopts upstream's measured rule: invented abbreviations and '->' arrows
save zero tokens under the tokenizer, so ultra now bans them instead of
recommending them. README reports the 3-arm table; repo description
updated to match.

Also fixes the project etymology: 'менко' is a nod to the Ukrainian
surname Menko (like Brussee upstream), not a diminutive suffix.
@ruslanlap
ruslanlap force-pushed the feat/honest-3arm-benchmark branch from 221e86c to da6d043 Compare September 30, 2026 07:56
@ruslanlap
ruslanlap merged commit 303fbf2 into master Sep 30, 2026
2 of 3 checks passed
@ruslanlap
ruslanlap deleted the feat/honest-3arm-benchmark branch September 30, 2026 07:57

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 221e86c9fa

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/cavemenko/SKILL.md
| **lite** | ~40% baseline | Cut воду/ввічливість/hedging. Повні речення. Для docs/пояснень |
| **full** | ~25-30% baseline | + pro-drop, тире, short forms, наказовий, фрагменти. Default |
| **ultra** | ~15-20% baseline | + abbr (БД/фн/імпл/конф/env/dep), arrows `X → Y`, 1 sentence де вистачає |
| **ultra** | ~15-20% baseline | + strip conjunctions, одне слово де вистачає, кожен факт один раз. **Без** вигаданих abbr і **без** стрілок `→` (токенайзер ріже `імпл` так само як `implementation` — 0 токенів, гірше читається) |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Remove arrow examples from the ultra prompt

When ultra mode is activated, hooks/cavemenko-activate.js retains the ultra: examples, so the injected prompt bans → here but then demonstrates it at lines 57 and 62 (Inline obj → ... and < → <=). The README also still describes ultra as using “Abbr, arrows” and retains the фн/імпл/конф table, leaving both runtime instructions and user guidance contradictory; update those examples and stale documentation so the new ban is actually communicated.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant