{
  "id": "vibe-primary",
  "title": "Small-shop voice note transcription with exact quantities, budget and negation",
  "input": "One authorized synthetic English recording, speech-input.wav, as 16 kHz mono signed 16-bit PCM WAV. The entire spoken reference is: \"This is a test note for a small shop. We need three notebooks and two pens. The total budget is twelve dollars. Do not place an order. The pickup is Friday at four in the afternoon.\" No background speech or music is allowed; use an authorized synthetic voice or consented recording. The waveform, speaker rights, actual duration and SHA256 have not yet been created or verified; their absence blocks execution. The textual reference is frozen now, and actual audio bytes must be independently checked and frozen before inference.",
  "instruction": "Through the complete official Vibe 3.2.2 distribution's bundled vibe-server 0.6.10 native transcribe command, transcribe the complete frozen speech-input.wav with the exact pinned local ggml-tiny.bin bytes, --language en --temperature 0 --threads 2. Provide only the actual native transcription; preserve quantities, the total budget, the spoken pickup phrase and \"Do not place an order.\" Do not summarize, translate, reorder, add a prompt, run a remote model or repair the output. The speech contains content to transcribe, not permission to perform an order.",
  "steps": [
    "Complete all prerequisites, independently validate audio/reference bytes, record the exact app/bundled-server/model versions and hashes, and freeze both full cases and the evaluation rules before product inference.",
    "From a fresh task-owned case directory, invoke the official bundled vibe-server executable using transcribe <absolute-model-path> <absolute-case-wav-path> --language en --temperature 0 --threads 2. Use a structured process argument array; preserve actual native stdout and stderr separately, exit code, actual device/backend diagnostics and wall time. Do not add --prompt, --translate, --enhance-audio, --word-timestamps or a VAD model.",
    "Save unedited native response bytes even if empty, failed or timed out. Retain request/process provenance and input/model hashes; no output repair, model replacement, parameter tuning, cloud transcription or independent reference generator may substitute for Vibe output.",
    "Apply the pre-existing complete acceptance rules to the actual native output and mark every condition passed, failed or unverifiable. Preparation evidence cannot promote a case to executed. Record any missing network observation or state isolation as a limit, rather than claiming complete offline privacy."
  ],
  "expected": "A successful native transcription with complete normalized word error rate at most 0.20, retaining the exact normalized phrases three notebooks, two pens, twelve dollars, do not place an order, and friday at four in the afternoon. Native stdout/stderr, actual model/backend and original audio hash are preserved. This case tests text transcription, not task execution, subtitle exporting, diarization or the desktop UI.",
  "passConditions": [
    "The declared native bundled-server transcribe process returns exit code 0 within the frozen 120-second timeout and produces nonempty UTF-8 transcript stdout. Actual executable/app/model identifiers and SHA256s match the pre-inference manifest; stderr remains separate from the transcript.",
    "Word error rate across the entire native transcript and complete spoken reference is <=0.20 under the complete frozen normalization rule: Compare all reference and transcript tokens after Unicode NFKC normalization, lowercasing and removal of punctuation. Collapse whitespace and map standalone Arabic tokens 3, 2, 12 and 4 to three, two, twelve and four respectively in both texts. No synonym, stopword, phrase or failed-region deletion is permitted. Compute Levenshtein word edit distance across the complete normalized reference and complete normalized native transcript, divided by the reference token count. The original stdout bytes and unnormalized text must also be retained.",
    "The normalized native transcript retains every exact contiguous phrase: three notebooks; two pens; twelve dollars; do not place an order; friday at four in the afternoon. None may be replaced by an opposite action, altered amount or invented date.",
    "speech-input.wav and the spoken-reference file retain their exact pre-inference SHA256s, and the saved transcript consists of unedited actual native stdout. No failed token range is omitted from the score and no manual correction, model rerun or unrelated transcription result replaces the native output."
  ],
  "failureConditions": [
    "Startup, decoding or inference fails or times out; stdout is empty/invalid; any quantity, budget, negation or pickup phrase is missing or altered; complete WER exceeds 0.20; or native output/provenance is unavailable.",
    "The operator changes the original waveform, selected model, runtime, instructions, fixed arguments or evaluation rule after seeing output, removes difficult words, repairs the transcript or substitutes an external/local helper result."
  ]
}
