Add live RLM operator behavior acceptance #242
Closed
lost-rob0t wants to merge 20 commits from
rage/183-live-operator-behavior into main
pull from: rage/183-live-operator-behavior
merge into: nsaspy:main
nsaspy:main
nsaspy:feat/prolog-rlm-language
nsaspy:codex/claude-api-provider
nsaspy:codex/symbolic-experts
nsaspy:codex/openai-api-provider
nsaspy:feat/zai-provider-protocols
nsaspy:hardening/project-semantic-20260911
nsaspy:hardening/ws7-nix-ci-pinning
nsaspy:chore/track-prolog
nsaspy:rage/98-semantic-project-knowledge
nsaspy:fix/339-agent-zero-skill-graph
nsaspy:ci/tree-sitter-runner-labels
nsaspy:issue-293
nsaspy:research/text-streaming
nsaspy:docs/agent-editor-handoff
nsaspy:icon-add
nsaspy:hydra/fleet-jobs-20260904
nsaspy:rage/355-d6-11-plan-native-dispatch
nsaspy:cd/nix-installer-action-url
nsaspy:issue-355
nsaspy:rage/97-query-capture-apis
nsaspy:research/expert-direct-tool-projection
nsaspy:rage/96-versioned-syntax-facts
nsaspy:fix/direct-deadline-typing
nsaspy:dogfood/auto-dig-native-tools
nsaspy:feature/rage-feature-freeze
nsaspy:fix/323-native-batch-cardinality
nsaspy:fix/328-direct-user-namespace-text-string
nsaspy:fix/transient-provider-retry-2026-09-01
nsaspy:fix/openrouter-dotted-tool-result-name
nsaspy:rage/325-native-call-isolation
nsaspy:fix/316-original-batch-effect-isolation
nsaspy:fix/313-per-call-preflight
nsaspy:fix/312-peek-schema-contract
nsaspy:docs/agents-worktree-rule
nsaspy:fix/child-native-capability-narrowing
nsaspy:test/paid-lane-glm-53-flash
nsaspy:docs/spec-seeded-symbolic-plans
nsaspy:rage/290-spec-plan-authority
nsaspy:fix/304-context-peek-selector-contract
nsaspy:fix/298-native-any-json-schema
nsaspy:rage/288-spec-plan-graph-executor
nsaspy:feat/spec-plan-flow-api
nsaspy:prolog-rlm-v1
nsaspy:gpt-5-6-sol-high/questions-for-rlm-prolog-and-lambda-rlm
nsaspy:rage/277-planner-protocol-context
nsaspy:research-approval/rlm-research-026-task-deadlines-20260827135838
nsaspy:research-approval/rlm-research-025-lem-ui-20260827135751
nsaspy:agent/127-agentprolog-config
nsaspy:recovery/132-agentprolog-monorepo
nsaspy:rage/223-durable-context-mount-recovery
nsaspy:rage/223-constraint-benchmark-recovery
nsaspy:research-approval/rlm-research-011-managed-context-tool-discovery-20260827054019
nsaspy:feat/research-approval-schema
nsaspy:agent/rrlm-control-plane-research
nsaspy:docs/219-adrrd-review
nsaspy:salvage/231-runtime-status
nsaspy:salvage/231-runtime-status-run
nsaspy:tmp
nsaspy:tmp2
nsaspy:tmp3
nsaspy:rage/245-planner-structural-retry
nsaspy:rage/168-skill-selection-eval
nsaspy:rage/250-skill-catalog-graph
nsaspy:rage/56-result-acceptance
nsaspy:rage/172-parent-resume-replan
nsaspy:rage/175-deadline-policy-recovery
nsaspy:rage/257-provider-tool-choice-normalization
nsaspy:agent/tool-result-projection-presets
nsaspy:fix/234-numeric-schema-bounds
nsaspy:feature/231-rlm-cli-reference-harness
nsaspy:feature/223-durable-context-mounts
nsaspy:feature/223-real-constraint-benchmark
nsaspy:feature/223-real-constraint-benchmark-clean
nsaspy:feature/223-real-constraint-benchmark-final
nsaspy:feature/223-real-constraint-benchmark-impl
nsaspy:feature/223-real-constraint-benchmark-now
nsaspy:feature/223-real-constraint-benchmark-tdd
nsaspy:feature/223-real-constraint-benchmark-work
nsaspy:rage/176-root-planner-tool-projection
nsaspy:rage/172-typed-delegation-policy
nsaspy:rage/175-subagent-deadline-policy
nsaspy:rage/206-prompt-command-runtime
nsaspy:rage/203-subagent-skill-role-provenance
nsaspy:rage/200-permanent-rlm-context
nsaspy:fix/190-cli-help-success
nsaspy:archive/pr-132-agentprolog-config-20260827
nsaspy:feature/117-prolog-skill-activation-linear
nsaspy:feature/117-prolog-skill-activation
nsaspy:fix/194-completion-budget-usage
nsaspy:fix/191-capability-filtered-tool-schemas
nsaspy:fix/185-reasoning-effort
nsaspy:ci/report-workflow-failures-20260825
nsaspy:cleanup/186-remove-legacy-harnesses
nsaspy:fix/185-reasoning-routing
nsaspy:agent/evolution-async-evaluator
nsaspy:rage/181-binding-replay-race
nsaspy:agent/124-deepseek-harness-prolog
nsaspy:rage/144-fallback-closure
nsaspy:rage/164-conversation-metadata
nsaspy:agent/subagent-supervised-call-conformance
nsaspy:fix/160-context-adapter-closed-data
nsaspy:codex/move-agent-zero-adaptor
nsaspy:codex/sol-high-integration
nsaspy:fix/151-plunit-gate
nsaspy:fix/165-evolution-closed-data
nsaspy:fix/162-async-control-exceptions
nsaspy:fix/158-prompt-compiler-closed-dicts
nsaspy:fix/151-plunit-main-ownership
nsaspy:fix/156-registry-destroy-hooks
nsaspy:fix/154-anonymous-dict-canonicalization
nsaspy:fix/flake-lock-reproducibility
nsaspy:agent/issue-142-evolution-kernel
nsaspy:agent/141-flake-runtime-package
nsaspy:agent/rlm-subagent-runtime
nsaspy:agent/144-subagent-fallback-a
nsaspy:agent/107-bound-adapter-metadata
nsaspy:validation/clean-pack-install
nsaspy:fix/45-authoritative-nested-model-events
nsaspy:fix/46-router-safe-live-streaming
nsaspy:agent/prompt-context-compiler
nsaspy:agent/42-canonical-recursive-fingerprints
nsaspy:agent/136-static-load-errors-fail-ci
nsaspy:agent/44-completion-error-usage
nsaspy:fix/67-loader-registry-cleanup
nsaspy:agent/opentui-solid-reference-client-current
nsaspy:agent/95-project-source-registry
nsaspy:feature/issue-117-prolog-skill-compiler
nsaspy:backlog/roadmap-86-merged
nsaspy:agent/opentui-solid-reference-client
nsaspy:79-tool-effect-boundary
nsaspy:agent/prolog-agent-ui-research
nsaspy:agent/conversation-cold-context
nsaspy:agent/94-tree-sitter-ffi
nsaspy:agent/conversation-warm-context
nsaspy:111-opentui-solid-reference-client
nsaspy:109-prolog-agent-ui-v1
nsaspy:agent/conversation-runtime
nsaspy:agent/spec-mode-language
nsaspy:agent/spec-verify-foundation
nsaspy:agent/prompt-compiler-research
nsaspy:reconcile-backlog-postmerge
nsaspy:reconcile-backlog-issues
nsaspy:agent/prolog-agent-roadmap
nsaspy:feature/issue-84-effect-store-migration
nsaspy:fix/issue-80-effect-substrate-adversarial-hardening
nsaspy:feature/issue-57-effect-identity
nsaspy:agent/mcp-declaration-security
nsaspy:agent/external-tool-category-boundary
nsaspy:feature/issue-53-authority-pending-async
nsaspy:feature/issue-54-agent-graph-canonical-async
nsaspy:feature/issue-54-tools-mcp-canonical-async
nsaspy:feature/issue-54-async-canonical-runtime
nsaspy:agent/dual-sync-async-runtime
nsaspy:agent/reconcile-todo-status
nsaspy:feature/issue-20-deep-recursion-experiments
nsaspy:feature/issue-19-cli-demo-trace
nsaspy:feature/issue-18-benchmark-conformance
nsaspy:feature/issue-17-adaptive-recursion
nsaspy:feature/issue-16-durable-artifacts
nsaspy:feature/issue-15-mcp-2026-dual-version
nsaspy:feature/issue-14-mcp-2025-11-25
nsaspy:hotfix/live-tool-native-openrouter
nsaspy:feature/issue-13-chain-runtime
nsaspy:fix/stable-live-repair-gate
nsaspy:fix/live-repair-strategy-parser
nsaspy:feature/issue-12-durable-graph
nsaspy:feature/issue-11-agent-supervision
nsaspy:feature/issue-10-structured-outcomes-repair
nsaspy:feature/issue-9-rlm-completion
nsaspy:feature/issue-8-capability-tools
nsaspy:feature/issue-7-typed-plan-runtime
nsaspy:feature/issue-6-context-store
nsaspy:fix/openrouter-reasoning-response
nsaspy:fix/live-openrouter-smoke-stability
nsaspy:feature/issue-5-openrouter-provider
nsaspy:feature/issue-4-swi-bootstrap
nsaspy:agent/agentic-harness-research
nsaspy:agent/prolog-rlm-foundation
No reviewers
Labels
Clear labels
bug
Something isn't working
documentation
Improvements or additions to documentation
duplicate
This issue or pull request already exists
enhancement
New feature or request
good first issue
Good for newcomers
help wanted
Extra attention is needed
invalid
This doesn't seem right
question
Further information is requested
wontfix
This will not be worked on
No labels
bug
documentation
duplicate
enhancement
good first issue
help wanted
invalid
question
wontfix
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set.
Reference
nsaspy/prolog-rlm!242
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "rage/183-live-operator-behavior"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Advances #183.
RAGE slice
TDD-first from exact canonical
main267697bef10a3fffff7c093e1435ece770e7444b.The remaining non-overlapping #183 gap is behavioral provider evidence. Existing
live_completion_openrouter_test.plsupplies exact plan JSON throughplanner_instruction/1, so it cannot prove the model independently learned how to operate the RLM runtime from the default permanent skills.Initial falsifiable contract
This draft adds a credential-backed OpenRouter behavior suite on the ordinary
rlm_completion/4path with no exact planner instruction:behavior_lookuptool rather than guess;rlmrecursion and use at least two model calls.Assertions use canonical runtime transition/usage/recursion evidence, not model self-report.
Adversarial boundary
rlmtransition/depth evidence;Gate
This PR is intentionally draft until the same-repository PR workflow runs the credential-backed REAL OpenRouter suite. If the behavioral contract fails, preserve the failure as evidence and refine the approved #183 implementation rather than adding an exact plan hint or weakening the test.
Live gate evidence is useful and should remain a hard failure for this candidate. The decomposable fixture exposed
rlmat the root ([rlm, context(slice), model(openrouter)]) and the branch strengthensrlm-recurseto treat independently investigable evidence as a strong recursion signal. The provider nevertheless selected two context slices + two direct model calls and synthesized successfully withrecursive_calls:0/max_depth:0. This is not a missing capability or malformed-plan failure, and the trivial/no-recursion + typed-tool cases both passed. Keep the recursion assertion; do not convert successful direct decomposition into acceptance for #183's explicit recursive-behavior case. Next design/debug target is why the permanent recursion guidance is insufficiently behavior-shaping for this model/provider, without spoon-feeding an exact plan.Second live attempt on exact head
f0e09b9ddfe59d91772ae4b98a2c3b8a5ba08ab4is still a useful hard failure.The strengthened generic
rlm-recurserule closed the previous wording loophole (multiple direct rootmodelcalls no longer count as satisfying an explicit recursion signal), butopenai/gpt-oss-120bstill selected the same basic root strategy: twocontext(slice)operations followed by two direct rootmodel(...)calls and norlmtransition (recursive_calls:0,max_depth:0). Trivial/no-recursion and typed-tool behavior still pass.More importantly, the direct fallback is semantically bad rather than merely non-recursive: the first leaf received a truncated raw evidence chunk as its entire prompt and asked what task it was supposed to perform; the second leaf received the other raw chunk and emitted another typed-plan JSON instead of evidence. The final result therefore also failed to preserve ALPHA-17/BETA-42.
This points at the next boundary to inspect: how planner-selected leaf model calls and nested
rlmplans carry task framing + retrieved evidence. Do not weaken the recursion assertion, and do not add an exact plan hint. Before changing skill prose again, inspect the model-step expression/prompt composition surface and the nested-RLM execution contract to determine whether the planner actually has a clean domain-neutral way to delegateanalyze this evidence for this subgoalrather than feeding raw context bytes as a standalone prompt.RAGE exact-head gate split for
3e7f04b4787be408158f97f92376236514bad1ca:The #183 behavior contract itself has now crossed an important line: in both the normal CI REAL OpenRouter job and the pinned Paid OpenRouter job,
Run REAL OpenRouter core suiteis green. This branch wireslive_rlm_operator_behavior_openrouterinto that exact core suite, so the repaired terms/context(peek)fixture plus currentrlm-operate/rlm-recurseskills now satisfy the three live no-spoon-feed cases on this immutable head: trivial stays non-recursive, unknown information uses the typed tool, and the decomposable two-record task uses the required bounded recursion/evidence path.Do not promote the PR yet. Both live jobs then fail at the separate pre-existing
Run REAL depth 0/1/2 recursion experimentstep; deterministic CI is green. That is now the bug-first blocker. Treat it as a potential regression until isolated—do not waive it just because the new #183 acceptance finally passes, and do not weaken/remove the depth gate.Next decision gate: compare the failing depth experiment on this head against a clean baseline containing the same current-main runtime, then inspect whether the new permanent skill text changes fixed-plan child model behavior or whether this is an already-red baseline/provider behavior. If branch-caused, preserve the #183 live success while repairing the smallest generic skill/runtime incompatibility. If baseline-red, keep that failure owned by its canonical benchmark transaction rather than falsely attributing it to #183. Exact-head full-gate promotion remains HOLD.
RAGE exact-head failure isolation on
3e7f04b4787be408158f97f92376236514bad1ca:The new #183 behavioral acceptance itself is green in the REAL OpenRouter core suite: trivial/no-recursion, typed unknown-info tool use, and decomposable bounded recursion all passed. Deterministic CI, Nix, clean-pack, and Tree-sitter are also green.
The remaining CI failure is the older injected-plan
deep_openrouter_experiment, and the logs make the cause deterministic enough to classify: depth 0 and depth 1 pass; depth 2 reaches the requested recursion depth but is rejected by the benchmark's fixedmax_total_tokens:3000after accumulating3795tokens (prompt_tokens:3636,completion_tokens:159,model_calls:4in the structured error). The benchmark source still gives every depth the same 3000-token ceiling even though the fixed depth-2 plan necessarily performs three nested provider calls, and the now-permanent RLM operating context materially increases each provider-visible prompt (~1212 prompt tokens/call in this run).Adversarial decision: this is not evidence that depth-2 recursion execution regressed, and it is not a reason to weaken/remove the depth gate. It is a stale benchmark-budget contract exposed by the larger required permanent context. The repair should be TDD-first and benchmark-local: make the live depth fixture's token ceiling explicitly sufficient for its expected fixed number of provider calls (prefer a depth-derived/tested ceiling or another deterministic contract), while retaining a finite hard token budget and all depth/provider-call assertions. Do not change production runtime budgeting, skip the live lane, or borrow the successful #183 behavior result as proof for this separate gate.
PR remains HOLD/draft until a changed exact head passes the full gate.
RAGE reconciliation on exact head
cff9f518c609286509c49e6be2877ce6936b75cd:mainat87189f7128bacfe97cd09874e958dd07e03322d3; the old #242 branch had diverged by 34 main commits and still carried benchmark/runtime patches whose ownership/implementation has since moved on main.rlm-operate/rlm-recurseguidance.rlmrecursion and preserve requested evidence/provenance identifiers.query; ordinary model steps are now explicitly told not to emit another planner object unless asked to plan.Fresh-head evidence so far: deterministic unit/load gate SUCCESS; Nix flake SUCCESS; Clean SWI pack SUCCESS; Tree-sitter FFI SUCCESS. Credential-backed REAL OpenRouter core and Paid OpenRouter lanes are still running. PR remains draft/HOLD until those exact-head live gates finish; no old-head success is being reused.
WIP: Add live RLM operator behavior acceptanceto Add live RLM operator behavior acceptancePull request closed