Skip to content

fix(monitor): recover legacy poll batches atomically with original successor identities - #5648

Open
hhyykk wants to merge 7 commits into
loopx-project:mainfrom
hhyykk:codex/monitor-recovery-transaction
Open

hhyykk wants to merge 7 commits into
loopx-project:mainfrom
hhyykk:codex/monitor-recovery-transaction

Conversation

@hhyykk

@hhyykk hhyykk commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

A legacy quota monitor-poll can stop after updating the Monitor but before creating both successor roles. If an already-created successor is later renamed or completed, retry can create duplicate work. This change commits the observation, successors and immutable business receipt in one locked atomic Markdown replacement; exact retry preserves the original IDs and public results after rename, completion, archival or canonical promotion.

The existing TypeScript Monitor planner now serves both storage paths. Python retains legacy parsing, lifecycle admission, storage and public response projection. Shared Todo serializers preserve the add/update fields, and the existing typed Next Action rule runs within the same batch. Ordinary retry also recovers a prepared shadow outbox before promotion. Native leases and CAS retain their existing owner.

Compatibility changes: fresh legacy batches reject conflicting requested successor semantics instead of silently updating existing work; matching successors are reused. Old pending legacy operations without an immutable business receipt require reconciliation. Old canonical pending plans and completed quota receipts remain recoverable. Dry-run and shadow diagnostics describe one atomic batch; post-commit shadow diagnostics are outside the immutable receipt. Drain pending operations before downgrading.

Validation for head c5c82f526bfcaa54787a7adc119625ca27b6642c:

  • 20 real CLI recovery tests pass; fresh installed wheel and sdist each pass the same 20 outside the checkout. Package source hashes match all 12 changed product files.
  • 187 shared Todo tests pass. The native transaction qualification has 426 passing tests, including isolated real PostgreSQL, on the unchanged native path before the final legacy-only corrections; 19 final typed planner/Next Action tests pass. Read-only legacy/File/PostgreSQL rehearsal preserves source bytes, relative order and unrelated state.
  • Full public-response comparisons cover normal, User-gate, quiet, dry-run, invalid input, semantic duplicate/conflict and crowded-state cases. Packaged Chat reads back the renamed original task and User action after a real crash/retry.
  • Typecheck, Ruff, semantic validation and premerge checks pass. The pre-existing maintainability advisory remains non-blocking. Independent review found two regressions (whitespace replay and Next Action reconciliation); both were reproduced, fixed and rechecked.
  • 20 alternating uninstrumented CLI pairs pass: execute p95 1859.3 → 1817.7 ms; exact retry 1133.7 → 1119.9 ms. These bounded samples meet the RFC regression limits.

Retained receipts have a storage cost: one synthetic two-successor poll adds about 8.9 KB, versus about 1.2 KB on main. This advances the existing T2 recovery checkpoint; legacy storage, retention/soak qualification and the wider migration remain open. No new frontend, Lark or CLI operation is introduced. Maintainer merge only.

@hhyykk
hhyykk marked this pull request as ready for review October 5, 2026 13:15
hhyykk added 7 commits October 5, 2026 21:50
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
…ivery

Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
@hhyykk
hhyykk force-pushed the codex/monitor-recovery-transaction branch from c5c82f5 to 74d70c7 Compare October 5, 2026 13:54

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; reasoning_effort=xhigh

动机

维护外部依赖观察任务的用户,会遇到“已经记下变化,但后续工作只创建了一半”的中断。以前观察结果、Agent 后续任务和用户待办分别保存;重试时又按当前任务文本匹配,任务一旦改名或完成,就可能再建一份,长期积累重复工作。这个 PR 试图一次保存整个批次,并用不可变回执找回原来的后续任务。原身份恢复的改善已经实测,但把工作交给另一位已注册 agent 的原有路径发生回归。本次只核验 monitor 批次保存与重试,不把模拟当作真实 agent 接收、执行或整个长期 Goal 完成。

改动思路

接受依据为 docs/reference/protocols/quota-monitor-observation-receipt-v0.md 的 Canonical leased observations and recovery,以及 docs/architecture/rfcs/typescript-control-plane-migration-v0.md 的 T2 — close monitor writeback and its atomic follow-up.:原文要求 malformed claims 在写观察前失败、valid claim aliases 在读回被当作相同 route。判断基于固定 pre-change c46f397c0f8b6115ed6efa0b06ab9bb6ca9e73dc 的 monitor observation protocol 和 TS migration RFC 的 T2 acceptance,再与当前 exact head 74d70c7f89298d347af15a32ff286f94dceadc68 对照。独立判据是:一次观察和两种 successor 必须一起提交,重启恢复原身份,已有工作不能被回放改掉,注册/claim/lease/gate/source 边界继续有效;不是“有回执字段就成功”。

复用 canonical 的 TypeScript planner 是合理方向:规则仍归 typed owner,Python 只负责 Markdown 编解码、锁和存储效果。新 legacy adapter 在一把既有锁下规划、准备 shadow capture,再一次替换 state;回执在 Todo 块外,归档任务不会直接删除它。canonical 的 CAS、lease 和 command receipt owner 保留。当前两处 blocker 都在这个接缝:复用普通 Todo-create 的自认领规则,改变了既有 peer 路由;其抛出的协议输入错误又被新 runtime handler 当作内部故障。

具体改动

18 个文件(+1549/-476):11 个生产文件、2 个协议/RFC 文档、5 个测试文件。scheduler/monitor_batch.ts 统一原 canonical 规划;legacy_monitor_poll.py 保存观察、successor 和不可变回执的一次替换;monitor_poll_writeback.py 区分历史 legacy replay 与 promoted provider,显式 lease proof 仍不能落回 legacy writer。quota/monitor_poll_commit.ts 为新 frozen plan 添加 legacy_batch_version:1,旧 pending 没有可证明的 business receipt 时拒绝猜身份,已有 canonical/completed quota receipt 仍可恢复。

todos/mutation_response.py 提取原 add/update 公共序列化,保持嵌套响应字段;todo_create.ts 统一 omitted-empty capabilities,并在相同 role/text 但元数据冲突时先拒绝整批,替代旧 partial upsert。这项默认行为变化已有双语披露。共享 planner/serializer 是已应用的 bounded future-facing refactor,能减少以后同一规则的两份实现;不需要另建 capability、配置开关、框架或 Python 决策源。现有 CLI 入口、前端/Lark 调用 owner 未增新用户操作;本轮实际执行了 CLI 和真实 provider,未声称 installed 前端/Lark journey 被完整操作过。

[P2] 已注册 peer 的 successor 路由回归:monitor_batch.ts:119-121 将观察 actor 原样传给 planCoordinationTodoCreate,而这个 owner 要求 claimed create 的 actor 等于 owner。隔离 Goal 注册 observer 和 builder,monitor 属于 observer;已有 quota monitor-poll ... --next-agent-todo 'Validate domain dependency' --next-action-kind validate --next-claimed-by builder --execute 在固定 base 成功并创建 builder 的任务,在 head 返回 exit1、无业务写入、quota_unexpected_collection_error。自指派在两端成功,所以不是注册或环境问题。跨领域 observer→implementer 的原入口因此失效;新增测试只覆盖 self/unclaimed 路由,没有覆盖它。请在既有 typed authority/handoff owner 内保留授权 peer routing,或明确迁移/退役该公开输入并提供可用替代与兼容测试;不要简单置空 actor 来绕过 claim 权限。回归用例应分别证明 self、registered peer、unknown peer,并把 canonical 的既有 stricter actor rule 与 legacy compatibility 分开。

[P2] 输入拒绝被误报为 quota 健康故障:相同接缝上,unknown peer 使 planCoordinationTodoCreate 抛出 AuthorityStoreProtocolError,新的 scheduler.monitor_batch.plan 没把它转换为 typed request rejection;legacy_monitor_poll.py:64-68 只处理 EffectRuntimeRejected。独立真实 CLI 得到 blocked_health、waiting_on:codex 和“fix quota/status collection before spending automatic compute”,没有指出无效 successor owner。以前 CLI 会明确列出注册 owner。head 的 fail-before-write 是改善,但误导恢复指令会把普通输入修正引向 quota 自修复。请由 owning handler 转换预期协议输入错误,保留 internal failure 的真实故障语义;非法 owner 的测试需检查 actionable error、业务状态与 quota/event 不变,不能只断言退出失败。

对主干的风险

独立场景用 read-only authority/profile 关系映射到隔离 synthetic Goals,19 个不同 role-qualified 组合分别执行 dry-run、正常保存、User action/gate、非法 owner、任务改名和 exact retry。head 的19/19重试保持原 successor ID/原响应且不改当前文件;base 的0/19保住这项判据,都会重建改名后的任务。两端 dry-run 和 User actor scope、旁支任务保护均保持。此证据覆盖不同职责的控制面输入,不证明真实 receiver adoption、agent 模型表现或个人数据业务效果。

独立新进程在第一次 ACTIVE state 替换后直接 os._exit(87):base 只存了观察,0个 successor;head 已一起保存观察、2个 successor 和回执。后续 exact retry 恢复成功。作者新增的20项真实 CLI恢复测试也全部通过,覆盖 before/after replace、quota commit 中断、旧 pending、坏回执、promotion、shadow、归档/完成、writer fence 与 Next Action。共享 Python 两端各93通过;native TS File/SQLite/monitor/create 两端718/725通过;隔离真实 PostgreSQL16.2 两端各341通过、0跳过,只用合成 tenant/临时服务器,结束已关闭清理。上述没有用活跃 Goal 做故障注入。

两端 typecheck 与完整 semantic-vocabulary smoke 通过,当前 diff advisory 为0 supported carriers;新3个协议字面值及 version marker 已逐项人工核验,0不是语义等价证明。改动9个 Python路径 Ruff通过;完整 ruff check . 两端同为606错误,逐项按路径/规则/message归一比较相同、均在未改路径,保留为 baseline lint debt,没有称全树绿色。初次独立 harness 误用 update 完成任务被原政策正确拒绝,已改成公开 rename 场景;这不是产品回归,日志保留。CI 按 resolved policy 不查询、不轮询、不等待。

需要正视长程成本:独立12次“无变化”保存,原 state 从521到677 bytes,head 到44881 bytes,保留12个 inline回执,约3.70KB/次。恢复身份有价值,但把完整公共响应编码进主 state 会随观察次数增长;没有测过长时间累计、压缩/淘汰、磁盘耐久性或大型 state 的解析成本。短样本两端单次中位时间约0.16s,存在并发负载,不能据此宣称提速。协议已经明确不把 indefinite retention/whole T2 promotion 当作完成。保留历史回执有独立接受基础,迁移/回退需先 drain pending,不能删 receipt/fence 掩盖恢复问题。

我的整体评价

REQUEST_CHANGES。原身份恢复和原子保存是已证实的正向增量,能减少重复任务和人工拼回半批次;代价是逐次增长的回执体积,尚无持续效率结论。整体验收目前不能通过,因为已注册 peer 路由失效、普通输入错误又误导成 quota 健康故障,影响多领域 agent 的长期分工和恢复体验。这两个问题在当前源、真实 CLI、同输入 base/head 对照中复现,不是等待 CI 的理由。

修复应留在本 PR 的 typed batch/authority/error 接缝,补 registered-peer 与 unknown-peer 的 public CLI 回归,并保持刚验证的原身份恢复、lease/gate/source no-effect 和真实 provider parity。当前变更控制面与 persisted receipt 合同,评审不授予自合并或 Goal 完成权限;没有执行 merge、安装升级或真实 agent 消费任务。

English verdict: REQUEST_CHANGES — exact head 74d70c7f89298d347af15a32ff286f94dceadc68 independently improves atomic legacy recovery and preserves original successor identity, but breaks the existing registered-peer successor route and reports expected owner-input rejection as quota health failure. Fixed base succeeds for observer→registered builder; head fails atomically with a misleading collection error. Current real CLI recovery, File/SQLite and isolated real PostgreSQL parity pass; baseline lint debt and linear receipt retention cost remain explicit. Repair the typed authority/error seam with public CLI regressions before approval. No merge, installed UI or sustained-efficiency qualification is claimed.

if (todo.task_class === "user_gate") Object.assign(todo, {action_kind: "gate", unblocks_todo_id: monitor.todo_id});
}
const created = planCoordinationTodoCreate({goal_id: input.goal_id, operation_id: input.operation_id,
actor_agent_id: actor, registered_agents: input.registered_agents, dry_run: input.dry_run,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve registered-peer successor routing

An isolated real CLI Goal registers observer and builder, with observer owning the Monitor. The existing quota monitor-poll ... --next-agent-todo "Validate domain dependency" --next-action-kind validate --next-claimed-by builder --execute succeeds at immutable base c46f397 and creates a builder-owned successor, but this head exits1 without creating work. Passing the observer actor to ordinary Todo-create now triggers its actor==claimed-owner guard. Self/unclaimed routes pass, so existing tests miss the cross-role regression. Preserve permitted peer routing via the existing typed authority/handoff owner or explicitly migrate this public input with a usable alternative; do not null the actor to bypass authority. Add self/registered-peer/unknown-peer public CLI cases.


def _plan(request: dict[str, Any], **facts: Any) -> dict[str, Any]:
try:
result = effect_runtime_result("scheduler.monitor_batch.plan", {**request, **facts})

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Keep expected owner-input errors actionable

Unknown successor owners make the shared Todo-create planner throw AuthorityStoreProtocolError; the new batch runtime handler does not translate that expected input error to request_rejected, so this adapter receives EffectRuntimeInternalError rather than the exception caught here. Independent actual CLI returns quota_unexpected_collection_error, blocked_health and advice to repair quota/status, hiding the invalid owner. The immutable base gives the registered-owner validation error. The head correctly prevents partial business writes, but callers cannot follow truthful recovery. Translate expected protocol/input errors at the owning typed handler, preserving actual internal failures, and assert actionable error plus no business/quota/event effect in public CLI negatives.

@huangruiteng

Copy link
Copy Markdown
Collaborator

这里确实值得优化

@mergify

mergify Bot commented Oct 6, 2026

Copy link
Copy Markdown

This pull request has merge conflicts with main and cannot be merged
until they are resolved. Please rebase or merge the base branch, @hhyykk.

Choose the remote for the base repository, not an out-of-date fork.
For a fork clone, first inspect git remote -v; upstream must point
to https://github.com/loopx-project/loopx.git. If it is absent, add it
with git remote add upstream https://github.com/loopx-project/loopx.git.
Then run:

git fetch upstream
git rebase upstream/main
# Resolve each conflict, git add the resolved files, then git rebase --continue.
git push --force-with-lease origin HEAD

For a same-repository clone whose origin points to
https://github.com/loopx-project/loopx.git, use origin instead of
upstream for fetch/rebase. If you prefer merging the base, use
git merge <base-remote>/main and push normally.

Keep the DCO Signed-off-by trailer on every commit when you rebase.
https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase Mergify: the pull request has merge conflicts with its base branch label Oct 6, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs-rebase Mergify: the pull request has merge conflicts with its base branch

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants