P1 human release gate — create and exercise new Zara commands by real voice on live hardware #169

Open
opened 2026-08-22 21:54:53 +00:00 by lost-rob0t · 0 comments
lost-rob0t commented 2026-08-22 21:54:53 +00:00 (Migrated from github.com)

Parent epic: #151
Depends on: #168, #161, #133; voice path #132. Respect #134 daemon release status.

Goal

Perform the explicitly human-operated Level C acceptance gate. Automation prepares the exact phrase/checklist and captures safe evidence, but must not claim PASS unless the human actually speaks to the real running system and confirms observed behavior.

Required guided scenarios

  1. Speak Zara, create a new command. and complete a safe simple command such as work mode -> open approved test apps/device actions.
  2. Invoke the new command immediately by voice.
  3. Create/use a parameterized command such as focus timer.
  4. Invoke it without duration; Zara must ask How long?; answer naturally, e.g. Thirty seconds.
  5. Correct an argument (Actually, make that five minutes.) in a safe fixture/profile.
  6. Cancel a pending intent (Never mind.), then prove a later standalone duration is not stolen by stale state.
  7. What does work mode do? / equivalent describe flow.
  8. Edit the command by voice and invoke the changed version.
  9. Restart client/server and prove durable command availability.
  10. Delete the command and prove it no longer resolves.
  11. Cross-principal fixture/account proves another principal cannot list/use it.
  12. Device action proves it runs on the initiating authenticated client, not inside zara-server/container.
  13. Server action/tool proves a server capability runs server-side.
  14. Unavailable/denied device capability yields explicit failure/clarification, never arbitrary fallback.
  15. Barge-in/cancel while speaking prevents stale TTS/action/history.

Test guide behavior

Provide exact phrases one at a time, expected transcript/intent/dialogue/action outcome, and what evidence/log/status field to inspect. The guide may use the same declarative corpus from #166 but this run uses live microphone/device/server, not checked-in WAV playback.

Evidence

Record versions/commit, client/server identity-safe metadata, timestamps/latency summaries, pass/fail per scenario and human notes. Do not store private conversation/audio/secrets unless the human explicitly chooses a dedicated synthetic fixture recording.

Acceptance

All mandatory scenarios are actually performed and marked by the human. Any failure creates a focused issue; do not weaken assertions or mark the epic complete from simulated CI alone.

Parent epic: #151 Depends on: #168, #161, #133; voice path #132. Respect #134 daemon release status. ## Goal Perform the explicitly human-operated Level C acceptance gate. Automation prepares the exact phrase/checklist and captures safe evidence, but **must not claim PASS unless the human actually speaks to the real running system and confirms observed behavior**. ## Required guided scenarios 1. Speak `Zara, create a new command.` and complete a safe simple command such as `work mode` -> open approved test apps/device actions. 2. Invoke the new command immediately by voice. 3. Create/use a parameterized command such as `focus timer`. 4. Invoke it without duration; Zara must ask `How long?`; answer naturally, e.g. `Thirty seconds.` 5. Correct an argument (`Actually, make that five minutes.`) in a safe fixture/profile. 6. Cancel a pending intent (`Never mind.`), then prove a later standalone duration is not stolen by stale state. 7. `What does work mode do?` / equivalent describe flow. 8. Edit the command by voice and invoke the changed version. 9. Restart client/server and prove durable command availability. 10. Delete the command and prove it no longer resolves. 11. Cross-principal fixture/account proves another principal cannot list/use it. 12. Device action proves it runs on the initiating authenticated client, not inside `zara-server`/container. 13. Server action/tool proves a server capability runs server-side. 14. Unavailable/denied device capability yields explicit failure/clarification, never arbitrary fallback. 15. Barge-in/cancel while speaking prevents stale TTS/action/history. ## Test guide behavior Provide exact phrases one at a time, expected transcript/intent/dialogue/action outcome, and what evidence/log/status field to inspect. The guide may use the same declarative corpus from #166 but this run uses live microphone/device/server, not checked-in WAV playback. ## Evidence Record versions/commit, client/server identity-safe metadata, timestamps/latency summaries, pass/fail per scenario and human notes. Do not store private conversation/audio/secrets unless the human explicitly chooses a dedicated synthetic fixture recording. ## Acceptance All mandatory scenarios are actually performed and marked by the human. Any failure creates a focused issue; do not weaken assertions or mark the epic complete from simulated CI alone.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
nsaspy/zara#169
No description provided.