Skip to content

Fix #2461: [Bug] Content between embedder limit and chat_window_max_tokens is silently stor - #2462

Open
Memtensor-AI wants to merge 2 commits into
MemTensor:dev-v2.0.36from
Memtensor-AI:bugfix/autodev-2461-20261008063231554
Open

Memtensor-AI wants to merge 2 commits into
MemTensor:dev-v2.0.36from
Memtensor-AI:bugfix/autodev-2461-20261008063231554

Conversation

@Memtensor-AI

Copy link
Copy Markdown
Collaborator

Description

Fixed the silent no-vector write for memory items whose token count falls in (embedder per-item limit, chat_window_max_tokens] (issue #2461). Such an item was previously neither split (the splitter only fired above chat_window_max_tokens, hard-coded to 1024 and never passed by either api/config.py mem_reader assembly site) nor embeddable (it exceeded the embedder's limit). /product/add still returned 200, but the item was persisted with metadata.embedding=None and vector_sync != "success", and search_by_embedding filters on vector_sync == "success" — so the memory was permanently invisible to semantic search. A secondary cause: _split_large_memory_item trusted chunker output and never re-checked the token budget per chunk, and the sentence chunker (chonkie) does not split CJK punctuation (it can also raise on an incompatible installed version), in which case the except branch returned the over-budget item unchanged.

The fix adds an explicit, env-configurable chat_window_max_tokens field on BaseMemReaderConfig (MEM_READER_CHAT_WINDOW_MAX_TOKENS, default 1024 so existing deployments are unchanged), wired into all three mem_reader assembly sites via APIConfig.get_chat_window_max_tokens() with fault-tolerant parsing. The effective split budget is now min(chat_window_max_tokens, embedder per-item limit). A new _chunk_within_budget helper re-checks every chunker-produced chunk and hard-splits any that still exceeds the budget — preferring a punctuation boundary (CJK and Latin terminators) and falling back to a binary-search character window — and a chunker that raises or yields nothing now hard-splits instead of passing the over-budget item through. _concat_multi_modal_memories uses the effective budget for both the split trigger and the window accumulator, plus a guard that emits over-budget items directly rather than aggregating them into a larger window. Finally, the per-item embedding retry (and the base reader's _make_memory_item path) truncates to budget before calling the embedder, so one over-limit item can no longer leave a whole batch without vectors.

Verification: a new 18-case regression suite (tests/mem_reader/test_embed_budget_split.py) passes, covering budget derivation, punctuation-free CJK splitting within budget, boundary preference, empty/failed chunker fallback, contiguous re-indexing, the 700-token in-window case, and the mixed-batch embedding fallback. Affected suites (mem_reader, configs, api, memories, mem_cube) report 341 passed with 4 pre-existing failures confirmed on the untouched baseline (markitdown extra not installed; unrelated kv-cache API). ruff check and ruff format both clean. An end-to-end reproduction using the real SentenceChunker and tiktoken counter with a 512-token embedder limit turns the issue's 780-token Chinese paragraph into chunks of 512 + 268 tokens, all embedded, while the same script's pre-fix branch reproduces the reported embedded=[False].

Out of scope: backfilling vectors for memories already stored without one.

Related Issue (Required): Fixes #2461

Type of change

Please delete options that are not relevant.

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Refactor (does not change functionality, e.g. code style improvements, linting)
  • Documentation update

How Has This Been Tested?

Not run; documentation-only change.

  • Unit Test
  • Test Script Or Test Steps (please provide)
  • Pipeline Automated API Test (please provide)

Checklist

  • I have performed a self-review of my own code
  • I have commented my code in hard-to-understand areas
  • I have added tests that prove my fix is effective or that my feature works
  • I have created related documentation issue/PR in MemOS-Docs (if applicable)
  • I have linked the issue to this PR (if applicable)
  • I have mentioned the person who will review this PR

@WeiminLee please review this PR.

Reviewer Checklist

Items whose token count fell in (embedder limit, chat_window_max_tokens]
were neither split nor embeddable, so /product/add returned 200 while the
item was stored with metadata.embedding=None and vector_sync != "success",
permanently invisible to semantic search (issue MemTensor#2461).

- expose chat_window_max_tokens as an explicit, env-configurable field on
  BaseMemReaderConfig (MEM_READER_CHAT_WINDOW_MAX_TOKENS), default 1024
- derive the split budget as min(window, embedder per-item limit)
- re-check chunker output and hard-split any over-budget chunk, preferring a
  punctuation boundary and falling back to a binary-search character window;
  a chunker that fails or yields nothing now hard-splits instead of passing
  the over-budget item through
- truncate to budget before the per-item embedding retry so one over-limit
  item cannot leave a whole batch without vectors
@Memtensor-AI Memtensor-AI added ai:generated Generated or modified by AI | 由 AI 生成或修改 area:api 云服务 / FastAPI / OpenAPI / MCP area:core MOS 编排层 / 框架底座 / 跨模块问题 area:memory 记忆存储、检索、更新、召回逻辑 status:in-progress Someone or AI is working on it | 人工或 AI 正在处理 labels Oct 8, 2026
@Memtensor-AI

Memtensor-AI commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

🤖 Open Code Review

Target: PR #2462
Task: 00c0f8f3534623df
Base: dev-v2.0.36
Head: bugfix/autodev-2461-20261008063231554
Head SHA: 243ef50068787b1d00a6a45599d3d3b60b3f9481

✅ OpenCodeReview: Review complete: 0 finding(s) across 6 selected item(s).

Generated by cloud-assistant via Open Code Review.

@Memtensor-AI

Copy link
Copy Markdown
Collaborator Author

🔧 Open Code Review requested Agent fix

Open Code Review found 4 issue(s). I have resumed the development Agent to fix them.

  • Task: 00c0f8f3534623df
  • Fix attempt: 1/2
  • Finding delta: 0 repeated / 4 new / 0 likely resolved

The Agent will push a new commit to this PR branch. OCR will recheck after the commit is pushed.

Addresses the review findings on the MemTensor#2461 fix:

- multi_modal_struct: the window loop and the overlap trim still called
  `_count_tokens` directly, so a reader built without a live tokenizer made
  its budget decision with `_count_tokens_safe` and then raised
  `AttributeError` a few lines later, dropping the whole scene. Both call
  sites now use the safe counter.
- multi_modal_struct: `_chunk_within_budget` wrapped the per-chunk token
  check and hard split in the same `try` as the chunker, so a tokenizer
  error was reported as "Chunker failed" and the already-validated chunks
  were thrown away. The guard now covers only `chunker.chunk(text)`.
- simple_struct: `_find_hard_split_index` could return `best + 1`, letting
  a terminator just past the budget be kept on the left side. That made the
  reported prefix `budget + 1` tokens; `_truncate_to_budget` used it
  directly, so the last guard before `embedder.embed([payload])` could still
  hand the provider an over-limit payload — the exact rejection this change
  set out to remove. The scan now stays inside the fitting prefix.
- simple_struct: drop the unreachable range guard in the scan loop.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai:generated Generated or modified by AI | 由 AI 生成或修改 area:api 云服务 / FastAPI / OpenAPI / MCP area:core MOS 编排层 / 框架底座 / 跨模块问题 area:memory 记忆存储、检索、更新、召回逻辑 status:in-progress Someone or AI is working on it | 人工或 AI 正在处理

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants