{
  "id": "vibe-boundary",
  "title": "Ten seconds of exact silence without hallucinated speech",
  "input": "A distinct silence-input.wav, exactly 10.000 seconds of 16 kHz mono signed 16-bit PCM WAV. All 160,000 audio samples must equal zero; the raw PCM payload is exactly 320,000 bytes. WAV header bytes and total container length are determined by the actual independently validated file. No speech, music, subtitle text, prior transcript or prompt is supplied. These resource bytes have not yet been created; freeze their SHA256 and independent sample validation before inference.",
  "instruction": "Run the same complete official Vibe 3.2.2 bundled-server 0.6.10 native transcribe route and exact pinned model/arguments as the primary, now using only the frozen distinct silence-input.wav from a separate fresh task-owned directory. Preserve the unedited actual empty transcript or native error. Do not add VAD, reuse a prior transcript, supply any prompt, retry with a new model or fill the silence with an explanation.",
  "steps": [
    "Complete all prerequisites, independently validate audio/reference bytes, record the exact app/bundled-server/model versions and hashes, and freeze both full cases and the evaluation rules before product inference.",
    "From a fresh task-owned case directory, invoke the official bundled vibe-server executable using transcribe <absolute-model-path> <absolute-case-wav-path> --language en --temperature 0 --threads 2. Use a structured process argument array; preserve actual native stdout and stderr separately, exit code, actual device/backend diagnostics and wall time. Do not add --prompt, --translate, --enhance-audio, --word-timestamps or a VAD model.",
    "Save unedited native response bytes even if empty, failed or timed out. Retain request/process provenance and input/model hashes; no output repair, model replacement, parameter tuning, cloud transcription or independent reference generator may substitute for Vibe output.",
    "Apply the pre-existing complete acceptance rules to the actual native output and mark every condition passed, failed or unverifiable. Preparation evidence cannot promote a case to executed. Record any missing network observation or state isolation as a limit, rather than claiming complete offline privacy."
  ],
  "expected": "The native process completes successfully within the fixed timeout and stdout contains only whitespace or no bytes, with zero lexical transcript tokens. No words, speaker claims, music labels, shop details or subtitles are hallucinated from the exact zero-sample audio. Native stderr and any error remain visible; an error is an observed failure rather than successful silence handling.",
  "passConditions": [
    "Independent pre-inference and post-inference validation confirms 16 kHz, mono, signed 16-bit PCM, exactly 160,000 zero-valued samples/320,000 PCM bytes and unchanged silence-input.wav SHA256. It is a distinct input from the primary speech case.",
    "The declared native bundled-server transcribe process returns exit code 0 within 120 seconds with the same pre-frozen executable/app/model SHA256s and exact arguments, using silence-input.wav rather than the primary audio.",
    "Unedited native stdout is empty after whitespace trimming and contains zero lexical transcript tokens; any speech, bracketed music label, speaker identity, explanatory prose or shop detail fails this condition. Native diagnostic stderr is retained separately and cannot be moved into or out of the transcript to improve the outcome.",
    "The separate boundary manifest and native process/output receipts retain the actual input path, original stdout/stderr bytes and their SHA256s, exit code, wall time and case id. No primary transcript, fabricated empty file, repaired text, alternate model or added VAD result substitutes for this native silence call."
  ],
  "failureConditions": [
    "The exact zero-sample input produces any non-whitespace transcript, a native error, missing observation or timeout, or its sample spec/hash differs from the frozen reference.",
    "The operator reuses speech output, adds VAD/prompt/translation, changes the selected model/parameters after observing hallucination, manufactures an empty result or hides the actual native stderr/error."
  ]
}
