Skip to content

fix(chat): prevent long-history work from blocking the main thread - #1157

Open
3316891527 wants to merge 2 commits into
AAswordman:devfrom
3316891527:fix/long-history-main-thread-anr
Open

3316891527 wants to merge 2 commits into
AAswordman:devfrom
3316891527:fix/long-history-main-thread-anr

Conversation

@3316891527

@3316891527 3316891527 commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

变更说明 / Description

长历史聊天在删除、回滚或发送消息时,会把上下文重建、thinking/search 清洗、角色卡配置解析和窗口估算带入 Main.immediate 路径。长对话因此可能出现明显卡顿,严重时触发 ANR。

本 PR 将这些高成本操作移出主线程,并让历史变更后的窗口刷新异步、按聊天可取消。用户侧的聊天上下文裁剪语义不变,但删除/发送/插入总结等操作不再同步等待整段历史处理。

背景与动机 / Context and motivation

vivo 上长历史聊天复现了删除或回滚后界面无响应并被系统静默杀死的问题。代码核对确认:有 summary 时,summary 之前的历史本来就不会进入运行时上下文;问题在于剩余上下文的重建和清洗被放在了错误的线程,并且发送路径还存在同步阻塞和重复工作。

改动范围 / Changes

  • 将上下文窗口重算和 memory 构建移到 Dispatchers.IO。
  • 为窗口刷新增加按 chatId 管理的可取消异步任务;删除、回滚、编辑消息和变体操作不再等待完整重算。
  • 缓存 think/search 正则;processAiMessage 仅在跨角色桥接时清洗内容。
  • 移除发送入口及角色卡/上下文配置解析中的 runBlocking,改为挂起调用和 IO 调度。
  • 将手动插入总结的历史读取、总结生成和写入移出主线程,并复用缓存的 memory 正则。
  • 明确不包含:summary 历史裁剪规则、复制路径、数据库 schema、网络 API 和配置格式均未改变。

兼容性与风险 / Compatibility and risks

  • 无数据库迁移、公开 API、配置格式或网络协议变化。
  • 历史变更后的 token/window 估算改为异步最终更新;同一聊天的新刷新会取消旧刷新,避免过时结果覆盖最新状态。
  • 发送和总结内容处理规则保持不变;主要变化是执行线程和调度时序。
  • 未进行真机/QQ 回归,原因是当前环境没有可用的 Android 真机测试条件。

关联 Issue / Related issue

N/A — 本 PR 针对 QQ/vivo 长历史 ANR 反馈;当前没有提供可关联的 Issue 编号。

验证方式 / Verification

检查或命令:
- ./gradlew :app:compileDebugKotlin --no-daemon
- ./gradlew :app:testDebugUnitTest --no-daemon --tests com.ai.assistance.operit.util.ChatUtilsTest --tests com.ai.assistance.operit.util.ChatUtilsThinkingTest --tests com.ai.assistance.operit.util.ChatUtilsThinkingEdgeTest --tests com.ai.assistance.operit.util.ChatMarkupRegexSearchTest
- git diff --check
- GitHub Actions: Android Build / assembleDebug
- GitHub Actions: Android Tests
环境与变体:
- 本地 Ubuntu 工作区;Debug Kotlin 编译和 ChatUtils 相关 JVM 单测
- GitHub Actions Ubuntu runner;Android Debug 构建
结果:
- 本地 Kotlin 编译:通过,退出码 0
- 本地针对性 ChatUtils 单测:通过,退出码 0
- diff 检查:通过
- Android Build / assembleDebug:通过
- Android Tests:测试源码编译阶段被 upstream/dev 中既有的 DeepseekProviderMediaRoleTest 和 XaiProviderReasoningTest 阻断;本 PR 未修改这两个测试文件,也未将该上游问题作为本 PR 的修复范围
- 未运行真机/QQ 回归:当前没有可用 Android 真机环境

证据 / Evidence

检查清单 / Checklist

  • 我已记录可复现验证和未运行项原因 / Reproducible verification and reasons for unrun checks are recorded
  • 日常开发 PR 的目标分支为 dev;如目标为 main,我已说明这是维护者发布同步 / The target branch is dev for regular work; if it is main, I explained why this is a maintainer-led release sync
  • 我已确认 Candidate checks 覆盖改动范围,并已区分上游测试源码失败 / Candidate checks covers the change scope and the upstream test-source failure is identified separately
  • 最终 diff 无无关、临时、生成、二进制或敏感内容 / Final diff has no unrelated, temporary, generated, binary, or secret content
  • 已提供对应的回归、UI、文档/字符串或兼容性证据 / Relevant regression, UI, docs/strings, or compatibility evidence is provided

@CATMIAOZHI
CATMIAOZHI self-requested a review September 25, 2026 07:17

@CATMIAOZHI CATMIAOZHI left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

审计结论:✅ 批准合并

改动(5 文件,+182/−143):把长历史聊天的高成本操作(上下文重建、thinking/search 清洗、角色卡配置解析、窗口估算)移出主线程,修复删除/回滚/发送时的 ANR;历史变更后的窗口刷新改为按 chatId 管理、可取消的异步任务。

通过的检查:

  • 目标分支为 dev ✓;PR 描述与 diff 一致。
  • 三处缓存正则与原内联正则逐字等价(think/search/capture),语义不变。
  • getMemoryFromMessages 确实移入 Dispatchers.IO;发送入口与角色卡/上下文配置解析中的 runBlocking 已全部改为挂起调用 + withContext(Dispatchers.IO)。
  • processAiMessage 仅在跨角色桥接时清洗:原代码非桥接路径本来就直接返回原 content,行为等价,只是少做了无用功。
  • scheduleStableContextWindowRefresh 的 per-chatId job 管理(put + cancel 旧 job + 按条件 remove)无泄漏;CancellationException 正确 rethrow,不会被吞。
  • ChatMarkupRegex.memoryTag 与原内联正则一致。
  • StateFlow 的跨线程 .value 读取是线程安全的;toast 等 UI 回调仍回到主线程执行。

问题:

  1. [P2] sendUserMessage 改为 fire-and-forget 后,MessageProcessingDelegate.sendUserMessage 里 isLoading 的 check-then-set 从主线程搬到了 IO 线程。之前主线程串行化保证了快速双击不可能重复发送;现在两个协程可能同时通过检查、各自发送(输入框清空也延后到了 IO 线程,进一步加大了窗口)。建议:在 delegate 的 sendUserMessage 入口、调用线程上先做一次同步的发送中检查,或用 per-chat 的 Mutex/AtomicBoolean 把发送入口串行化。
  2. [P2] scheduleStableContextWindowRefresh 里"新 job 直接 cancel 旧 job"并不能保证旧 job 先停:如果旧 job 已经进入 withContext(Main) 写回估算,取消对其中非挂起的 StateFlow 写无效,过期结果仍可能覆盖新结果,与"避免过时结果覆盖最新状态"的目标相悖。建议 cancel 后 join 旧 job,或用 Mutex 把刷新串行化。
  3. [P3] FloatingChatService.onSendMessage 的 try/catch 现在捕获不到预处理阶段的异常了:异常会抛到 runtimeScope 的未捕获异常处理器(默认行为是崩溃),之前会被 catch 打日志。建议在新 launch 的协程内加 runCatching,或给 scope 加 CoroutineExceptionHandler。
  4. [P3] sendMessageInternal 里"解析角色卡对话模型绑定失败"的 catch (e: Exception) 会吞掉 CancellationException;该函数现在是 suspend,吞取消的危害更大。建议排除 CancellationException 后再 catch。
  5. [question] StandardChatManagerTool 在调 sendUserMessage 后轮询 getResponseStream 等待新流:发送前处理现在多了若干 IO 跳,流建立整体变慢,RESPONSE_STREAM_ACQUIRE_TIMEOUT 是否仍然够用?AI 代理连续发送的场景实测过吗?

后续可跟进(不阻塞):

  • 作者已披露 Android Tests 在 fork 上被上游既有的两个测试文件编译失败阻断(非本 PR 范围),合入前建议确认主仓 CI 状态。

由 水晴喵的muse 审计

chatModelIndexOverride = chatModelIndexOverride,
turnOptions = turnOptions
)
// 已有对话,异步进入发送流程;内部的配置和历史读取会切到 IO

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] 这里改成 fire-and-forget 后,快速双击发送按钮可能走到两个并发的 sendMessageInternal:MessageProcessingDelegate.sendUserMessage 里的 isLoading check-then-set 现在跑在 IO 线程,不再原子;输入框清空也延后到了 IO 线程。之前主线程串行化时这不可能发生。建议在 sendUserMessage 入口(调用线程)先做一次同步的发送中检查,或用 per-chat Mutex/AtomicBoolean 串行化。

AppLogger.w(TAG, "异步刷新上下文窗口失败: chatId=$targetChatId", e)
}
}
stableWindowRefreshJobsByChatId.put(targetChatId, refreshJob)?.cancel()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] put + cancel 旧 job 并不能保证旧 job 先停:如果旧 job 已经进入 withContext(Dispatchers.Main) 写回 setTokenCounts,取消对其中非挂起的 StateFlow 写无效,过期估算仍可能覆盖新值。建议 cancel 后 join 旧 job,或用 Mutex 把刷新串行化,才能真正保证“新结果覆盖旧结果”。

@3316891527
3316891527 force-pushed the fix/long-history-main-thread-anr branch from a59e3a2 to 628f1b5 Compare September 25, 2026 16:17

@CATMIAOZHI CATMIAOZHI left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复审(针对新提交 628f1b53):💬 仍有 P2/P3 未处理

本次 push 内容:分支 rebase 到最新 dev(29d67812),旧提交 a59e3a20 的 5 文件 diff 变为 4 文件——ChatUtils 的正则缓存改动被移除(dev 已有 ChatMarkupRegex 统一缓存正则),核心改动(长历史高成本工作移出主线程、按 chatId 管理的可取消窗口刷新)保持不变。评论区无作者回复,视为上轮意见尚未回应。

上轮问题核验:

  1. [P2] isLoading check-then-set 竞态 → 经核验不成立,撤销:check-then-set 仍在 MessageProcessingDelegate.sendUserMessage 函数体内、调用线程(Main)上同步完成;fire-and-forget 的 launch 发生在 coordinator 层且继承 Main 调度,快速双击仍被主线程串行化,不会重复发送。
  2. [P2] scheduleStableContextWindowRefresh cancel 不 join → 未修复,保留(见下)。
  3. [P3] FloatingChatService.onSendMessage try/catch 失效 → 未修复,且随 fire-and-forget 范围扩大而加重,升级为 [P2](见下)。
  4. [P3] sendMessageInternal 的 catch (e: Exception) 吞 CancellationException → 未修复,保留。
  5. [question] RESPONSE_STREAM_ACQUIRE_TIMEOUT 是否够用 → 新 head 中该常量已不存在(疑似重命名/移除),且流建立逻辑未被本 PR 改动,撤回该问题。

仍存在的问题:

  1. [P2] sendUserMessage 改为 fire-and-forget 后,sendMessageInternal 预处理阶段的异常无处可去:coroutineScope 来自 ChatRuntimeHolder.runtimeScope = CoroutineScope(SupervisorJob() + Dispatchers.Main.immediate),没有 CoroutineExceptionHandler;sendMessageInternal 只有角色卡解析那一段有 try/catch,其余部分(总结检查、附件/UI 读取、delegate 调用等)一旦抛异常,会直达进程未捕获异常处理器导致崩溃。之前它是 Main 同步调用,FloatingChatService.onSendMessage 的 try/catch 还能兜住。建议:在 coroutineScope.launch { sendMessageInternal(...) } 的协程体内加 try/catch(打日志+toast),或给 runtimeScope 加 CoroutineExceptionHandler。
  2. [P2] scheduleStableContextWindowRefresh 的 put(targetChatId, refreshJob)?.cancel() 仍无 join:refreshStableContextWindow 末尾有 withContext(Dispatchers.Main) 写回 setTokenCounts,非挂起的 StateFlow 写不受取消影响,旧 job 的过期估算仍可能覆盖新值。建议 cancel 后 join 旧 job,或用 Mutex 把刷新串行化。
  3. [P3] sendMessageInternal 中"解析角色卡对话模型绑定失败"的 catch (e: Exception) 仍会吞掉 CancellationException;该函数现在是 suspend 且跑在 fire-and-forget 协程里,吞取消的危害更大。建议先 catch (e: CancellationException) { throw e }。
  4. [P3] FloatingChatService.onSendMessage 的 try/catch 现在形同虚设(见 [P2]-1):chatCore.sendUserMessage 立即返回,预处理异常进不到这个 catch。修好 [P2]-1 后可恢复其兜底作用,或删掉这个误导性的 try/catch。

通过的检查:rebase 后 diff 与上轮审计结论一致——processAiMessage 仅跨角色桥接时清洗、getMemoryFromMessages 移入 Dispatchers.IO、ChatMarkupRegex.memoryTag 与原内联正则一致;stableWindowRefreshJobsByChatId 的 put/cancel/remove 模式无泄漏;scheduleStableContextWindowRefresh 内对 CancellationException 正确 rethrow。

由 水晴喵的muse 审计

@CATMIAOZHI CATMIAOZHI left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复审(针对新提交 d0fdd16):✅ 批准合并

本轮只审新增 diff(628f1b53..d0fdd163,1 文件 +121/−82,MessageCoordinationDelegate.kt):fix(chat): serialize window refresh and handle async send errors。

上轮问题修复情况:

  • [P2#2 已修] 窗口刷新的旧结果覆盖新结果:新增按 chatId 的 Mutex(stableWindowRefreshMutexesByChatId),refreshStableContextWindow 把估算→持久化→UI 写回整体串行,并在 token 统计落库前、切回 Main 写回前加了 ensureActive() 检查,被取消的旧任务会在写回前退出;调度侧改为 LAZY 启动 + synchronized 登记(先 put 新 job 再 cancel 旧 job),并发调度时旧 job 不会再反过来取消新 job。两者结合,过期结果无法覆盖最新状态。
  • [P3#4 已修] 角色卡对话模型绑定解析的 catch (e: Exception) 现在先 rethrow CancellationException;群组编排 runCatching 的回退路径也补了同样的处理 + 回退前 ensureActive()。
  • [P3#3 已解决(换了实现方式)] 三个发送入口统一走新增的 launchMessageSend/runMessageSend:异常被捕获后经 reportSendFailure 打日志并在主线程弹 toast(R.string.message_send_failed 多语言资源已确认存在),不再抛到 scope 的未捕获异常处理器。FloatingChatService.onSendMessage 自带的 try/catch 现在冗余但无害。
  • [P2#1 已缓解] 三个用户触发的发送入口改为 Dispatchers.Main.immediate 启动,MessageProcessingDelegate.sendUserMessage 的 isLoading 检查+置位回到主线程串行执行,快速双击重复发送的竞态窗口已关闭。

本轮新增观察(不阻塞):

  1. [P3] stableWindowRefreshMutexesByChatId 只 put 不 remove(job map 有 invokeOnCompletion 清理),已删除聊天的 Mutex 会一直堆积。建议在 job map 清理时一并移除已无任务的 mutex。
  2. [question] 上轮问的 RESPONSE_STREAM_ACQUIRE_TIMEOUT 是否仍然够用(发送前处理多了 Main.immediate + runMessageSend 跳转),暂未看到作者回复——AI 代理连续发送场景建议实测确认一下。

由 水晴喵的muse 审计

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants