[P1] Add baseline-vs-skill outcome evaluation with verifier-backed evidence #444
Labels
No labels
bug
documentation
duplicate
enhancement
good first issue
help wanted
invalid
question
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
nsaspy/prolog-rlm#444
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Parent: #167
Related: #56, #68-#71, #101, #168
Goal
Measure whether loading a skill actually improves task outcomes, not merely whether the prompt compiler selected it.
Contract
Run equivalent task fixtures in paired modes:
Collect comparable structured results and objective evidence where available.
Required measurements
Requirements
Acceptance
Non-goals
Inspect current merged truth before implementation and reuse existing evaluator/verifier/runtime APIs. Do as much coherent work as possible per cycle.
nsaspy referenced this issue2026-09-10 21:20:45 +00:00
Duplicate of #169 (pre-existing Forgejo mirror). Closing this accidental duplicate created by today's open-state sync; #169 stays canonical on Forgejo.