Vibe
Vibe helps individuals, creators and small shops turn authorized voice notes, interviews, lessons and videos into reviewable text on a Windows, macOS or Linux computer. Its official desktop app transcribes locally with a selected speech model and advertises batch processing, subtitle/document export and optional transcript analysis through Claude or Ollama. The fixed MIT license credits thewh1teagle, and the official repository is hosted by that GitHub user; current company size and ownership are unknown. This profile distinguishes local transcription from optional network-connected features and is not a claim of an autonomous business agent.
On this page
What is Vibe?
Vibe is a specialist ai tool from Vibe / thewh1teagle for Voice, captions & dubbing. Vibe helps individuals, creators and small shops turn authorized voice notes, interviews, lessons and videos into reviewable text on a Windows, macOS or Linux computer. Its official desktop app transcribes locally with a selected speech model and advertises batch processing, subtitle/document export and optional transcript analysis through Claude or Ollama. The fixed MIT license credits thewh1teagle, and the official repository is hosted by that GitHub user; current company size and ownership are unknown. This profile distinguishes local transcription from optional network-connected features and is not a claim of an autonomous business agent. Its documented inputs are authorized local audio/video recordings, a selected compatible speech model and fixed transcription settings. The recorded native pilot used a 14.125-second local Microsoft David Desktop synthetic English shop note and an exact ten-second zero-sample WAV, both 16 kHz mono signed 16-bit PCM with frozen hashes, and one pinned ggml-tiny.bin model. Original case wording from the pre-resource planning state is preserved; a separate pre-inference contract supplies the actual bytes. The expected deliverable is reviewable transcripts and advertised document/subtitle exports. The recorded native bundled-server pilot retained original plain-text stdout/stderr: the shop note wrote the correct budget as $12; exact silence returned [BLANK_AUDIO]. Each failed one of its four strict original conditions. Desktop SRT/VTT/document export, speaker diarization, summaries, recording and GUI interactions remain untested.
Best suited for
- A creator, student, interviewer, independent professional or small-shop owner who wants local transcription from authorized recordings and can review names, numbers, negations and subtitle timing.
- A pilot focused on small-shop voice note transcription with exact quantities, budget and negation, using authorized local audio/video recordings, a selected compatible speech model and fixed transcription settings. The recorded native pilot used a 14.125-second local Microsoft David Desktop synthetic English shop note and an exact ten-second zero-sample WAV, both 16 kHz mono signed 16-bit PCM with frozen hashes, and one pinned ggml-tiny.bin model. Original case wording from the pre-resource planning state is preserved; a separate pre-inference contract supplies the actual bytes.
Not suited for
- A workflow that depends on the following request without the stated input, review or permissions: Run the same complete official Vibe 3.2.2 bundled-server 0.6.10 native transcribe route and exact pinned model/arguments as the primary, now using only the frozen distinct silence-input.wav from a separate fresh task-owned directory. Preserve the unedited actual empty transcript or native error. Do not add VAD, reuse a prior transcript, supply any prompt, retry with a new model or fill the silence with an explanation.
- Vibe is a specialist transcription tool. The advertised HTTP API or agent skills do not establish autonomous orders, publishing or business workflows.
- Speech recognition can change quantities, names, dates and negations or hallucinate text from silence. Human review is required; no broad transcription accuracy claim was reproduced.
Capabilities, with sources
- 01The official README advertises on-device audio/video transcription, batch input, microphone/system-audio features and Windows/macOS/Linux support.Official vendor statement · checked 2026-10-02Source ↗
- 02The official README lists SRT, VTT, TXT, HTML, PDF, JSON and DOCX output formats.Official vendor statement · checked 2026-10-02Source ↗
- 03Transcript summaries and analysis can use optional Claude API or Ollama; local transcription claims do not certify every network-connected feature.Official vendor statement · checked 2026-10-02Source ↗
- 04The fixed software license is MIT with Copyright (c) 2024 thewh1teagle.Official vendor statement · checked 2026-10-02Source ↗
- 05The desktop application bundles a native vibe-server process and communicates over local HTTP; on macOS/Windows it also bundles an FFmpeg helper.Official vendor statement · checked 2026-10-02Source ↗
- 06The fixed desktop CLI forwards arguments to the bundled server and includes lifecycle analytics calls; actual configuration and network behavior have not been inspected.Official vendor statement · checked 2026-10-02Source ↗
- 07The matching native server v0.6.10 CLI accepts explicit model/audio paths, language and temperature and emits plain text or per-word timestamps.Official vendor statement · checked 2026-10-02Source ↗
- 08The model page links several Whisper and other model assets; the proposed model link is mutable and model licensing/integrity are separate pre-execution checks.Official vendor statement · checked 2026-10-02Source ↗
- 09The official-linked model card at the recorded fixed revision declares MIT and describes OpenAI Whisper models converted to ggml; model licensing remains separate from the Vibe software.Official vendor statement · checked 2026-10-02Source ↗
Inputs and outputs
Inputs
Authorized local audio/video recordings, a selected compatible speech model and fixed transcription settings. The recorded native pilot used a 14.125-second local Microsoft David Desktop synthetic English shop note and an exact ten-second zero-sample WAV, both 16 kHz mono signed 16-bit PCM with frozen hashes, and one pinned ggml-tiny.bin model. Original case wording from the pre-resource planning state is preserved; a separate pre-inference contract supplies the actual bytes.
Outputs
Reviewable transcripts and advertised document/subtitle exports. The recorded native bundled-server pilot retained original plain-text stdout/stderr: the shop note wrote the correct budget as $12; exact silence returned [BLANK_AUDIO]. Each failed one of its four strict original conditions. Desktop SRT/VTT/document export, speaker diarization, summaries, recording and GUI interactions remain untested.
Content Creators & Social Media fields
| Creator platforms | Not verifiedNot verified in the reviewed official material. |
|---|---|
| Content formats | Not verifiedNot verified in the reviewed official material. |
| Inputs | Authorized local audio/video recordings, a selected compatible speech model and fixed transcription settings. The recorded native pilot used a 14.125-second local Microsoft David Desktop synthetic English shop note and an exact ten-second zero-sample WAV, both 16 kHz mono signed 16-bit PCM with frozen hashes, and one pinned ggml-tiny.bin model. Original case wording from the pre-resource planning state is preserved; a separate pre-inference contract supplies the actual bytes.Source 1 |
| Outputs | Not verifiedNot verified in the reviewed official material. |
| Aspect ratios | Not verifiedNot verified in the reviewed official material. |
| Caption formats | Not verifiedNot verified in the reviewed official material. |
| Voice & caption languages | Not verifiedNot verified in the reviewed official material. |
| Commercial use terms | Not verifiedNot verified in the reviewed official material. |
| Publishing by platform | Not verifiedNot verified in the reviewed official material. |
| Approval requirements | Not verifiedNot verified in the reviewed official material. |
Education & Training fields
| Learner and educator roles | Not verifiedNot verified in the reviewed official material. |
|---|---|
| Teaching workflow | Not verifiedNot verified in the reviewed official material. |
| Source grounding | Not verifiedNot verified in the reviewed official material. |
| Age eligibility | Not verifiedNot verified in the reviewed official material. |
| Data handling | Not verifiedNot verified in the reviewed official material. |
| Instructor review | Not verifiedNot verified in the reviewed official material. |
A practical Vibe workflow
The workflow and preparation text below retain the original Vibe evaluation plan written before resource setup. Statements about resources not yet being prepared describe that earlier planning stage. Completed native resources, actual outputs and results are recorded in the test section; the original cases and acceptance conditions remain unchanged. View recorded results →
- Prepare the small-shop voice note transcription with exact quantities, budget and negation fixture: One authorized synthetic English recording, speech-input.wav, as 16 kHz mono signed 16-bit PCM WAV. The entire spoken reference is: "This is a test note for a small shop. We need three notebooks and two pens. The total budget is twelve dollars. Do not place an order. The pickup is Friday at four in the afternoon." No background speech or music is allowed; use an authorized synthetic voice or consented recording. The waveform, speaker rights, actual duration and SHA256 have not yet been created or verified; their absence blocks execution. The textual reference is frozen now, and actual audio bytes must be independently checked and frozen before inference.
- Check Vibe access through Windows desktop app, macOS desktop app, Linux desktop app, Native bundled-server CLI, Advertised local HTTP API, Optional Claude API or Ollama analysis and confirm the selected feature’s actual permissions.
- Through the complete official Vibe 3.2.2 distribution's bundled vibe-server 0.6.10 native transcribe command, transcribe the complete frozen speech-input.wav with the exact pinned local ggml-tiny.bin bytes, --language en --temperature 0 --threads 2. Provide only the actual native transcription; preserve quantities, the total budget, the spoken pickup phrase and "Do not place an order." Do not summarize, translate, reorder, add a prompt, run a remote model or repair the output. The speech contains content to transcribe, not permission to perform an order.
- Inspect a successful native transcription with complete normalized word error rate at most 0.20, retaining the exact normalized phrases three notebooks, two pens, twelve dollars, do not place an order, and friday at four in the afternoon. Native stdout/stderr, actual model/backend and original audio hash are preserved. This case tests text transcription, not task execution, subtitle exporting, diarization or the desktop UI. Compare it against the source input and retain the output/action log.
- Run the boundary case: Run the same complete official Vibe 3.2.2 bundled-server 0.6.10 native transcribe route and exact pinned model/arguments as the primary, now using only the frozen distinct silence-input.wav from a separate fresh task-owned directory. Preserve the unedited actual empty transcript or native error. Do not add VAD, reuse a prior transcript, supply any prompt, retry with a new model or fill the silence with an explanation. Accept the result only if the failure criteria are satisfied.
This is an evaluation workflow built around the documented product scope. Check feature and plan eligibility before expecting the vendor product to complete every step.
Setup and integrations
The official complete Windows Vibe 3.2.2 distribution was downloaded/hash-verified and extracted in a task-owned directory without invoking the installer. It includes vibe.exe, native vibe-server.exe and ffmpeg.exe. Actual native server --version identified v0.6.10. Two original plain-text native transcribe cases each ran once using its bundled server and one pinned local Whisper tiny model. Native processes completed and task-owned app/server/FFmpeg process checks found no leftovers. The desktop GUI and local HTTP API were not invoked; actual inner CPU/GPU backend and whole-host network behavior were not measured.. Documented access methods: Windows desktop app, macOS desktop app, Linux desktop app, Native bundled-server CLI, Advertised local HTTP API, Optional Claude API or Ollama analysis. Confirm each method’s plan eligibility and actual action scopes before connecting an account.
Access and setup steps
- Review the fixed MIT license and use only authorized recordings; separately review the exact selected model, dependencies and any optional provider terms.
- Use the complete official Vibe v3.2.2 distribution and retain its recorded package/binary hashes. Archive extraction, actual native server v0.6.10 help/version and bundled FFmpeg checks are prepared; those checks are not model inference or GUI evidence.
- Use only the exact recorded ggml-tiny.bin bytes pinned to ggerganov/whisper.cpp revision 5359861c739e955e79d9a303bcbc70fb988958b1. The 77,691,713-byte SHA256 be07e048e1e599ad46341c8d2a135645097a538221678b7acdd1b1919c6e1b21 matches repository LFS metadata. Retain the fixed model-card MIT declaration and distinguish it from universal commercial-license certification.
- Use the original frozen text with the retained Microsoft David Desktop local TTS waveform and distinct ten-second all-zero WAV. Independent byte/PCM checks and resource hashes are prepared; do not claim a separate human audition or substitute a newly recorded waveform after inference.
- Use the complete distribution's bundled native transcribe component with explicit absolute model/audio paths and fixed --language en --temperature 0 --threads 2; do not supply prompt, translation, enhancement, VAD or timestamp flags. Inspect actual state and backend; paths and gpu-device=-1 do not establish isolation or CPU-only inference.
- Retain unedited native stdout/stderr, exit code, wall time, process versions, model/audio hashes and any resource/access failure. Disable or constrain unintended outbound traffic under an authorized test setup and document actual observation; no whole-app offline guarantee is implied.
- Inspect the actual registered native outputs and all original conditions. Primary symbolic $12 and boundary [BLANK_AUDIO] fail their strict unchanged wording/empty-output requirements despite successful native exits. Original fixtures, model, parameters and full test fields were preserved with no output repair or retry. Test desktop GUI/subtitle exports separately if needed.
Test access: local install. Two original frozen native transcription cases were executed once each through the complete official Windows Vibe 3.2.2 distribution's bundled vibe-server v0.6.10, using one exact local Whisper tiny ggml asset. Both processes exited 0; each passed three of four conditions and failed the unchanged strict output condition. Primary WER=0.027777777777777776 and $12 preserved the amount but did not retain the exact twelve dollars phrase. Boundary returned [BLANK_AUDIO], a silence marker that did not meet whitespace-only output. Original WAVs, raw stdout/stderr and same-case audit/provenance are retained separately. No installer, new account, paid API, optional remote analysis, GUI or HTTP API test was used. Open the official access or installation page ↗
Pilot dependencies
The workflow and preparation text below retain the original Vibe evaluation plan written before resource setup. Statements about resources not yet being prepared describe that earlier planning stage. Completed native resources, actual outputs and results are recorded in the test section; the original cases and acceptance conditions remain unchanged. View recorded results →
- Before either case, obtain and verify the complete official Vibe 3.2.2 Windows distribution and its bundled vibe-server 0.6.10 executable, matching actual native help to the fixed source. Use the bundled native component from that distribution, not standalone whisper.cpp, a library helper or another server build. Select the single official-linked ggml-tiny.bin Whisper model, review its specific license and record its actual immutable bytes/SHA256 before inference; the current official link is a moving main asset and this proposal has not downloaded or licensed the weights. Create the exact authorized speech WAV and ten-second zero-sample WAV, independently listen/check their contents, and freeze their actual bytes, reference hashes and complete cases before the first product request. Use separate task-owned case directories, explicit absolute model/audio paths, --language en --temperature 0 --threads 2 with no translation, enhancement, prompt or VAD model. Retain actual executable help/version, model/resource hashes, device diagnostics, native stdout/stderr and exit codes. The source sets no_gpu=false; do not claim that --gpu-device -1 disables GPU. Set a 120-second process timeout and retain timeout/native failures. Verify application/server state and actual outbound behavior; path variables alone do not prove isolation. Installation, model download, license review, audio creation, help and GUI inspection remain preparation. No runtime, audio, model, app state isolation or GUI execution has been established by this read-only proposal.
- Two original frozen native transcription cases were executed once each through the complete official Windows Vibe 3.2.2 distribution's bundled vibe-server v0.6.10, using one exact local Whisper tiny ggml asset. Both processes exited 0; each passed three of four conditions and failed the unchanged strict output condition. Primary WER=0.027777777777777776 and $12 preserved the amount but did not retain the exact twelve dollars phrase. Boundary returned [BLANK_AUDIO], a silence marker that did not meet whitespace-only output. Original WAVs, raw stdout/stderr and same-case audit/provenance are retained separately. No installer, new account, paid API, optional remote analysis, GUI or HTTP API test was used.
- Confirm mit local software; separate model, hardware and optional api costs against the current vendor terms; usage and connected-service costs can affect the pilot.
- Create a test workspace or use public/authorized material. Keep an input baseline, output artifact and action log for comparison.
Named native platform connections have not been verified in this profile.
Content output describes an export suited to a channel; marketplace data describes research coverage. Exact data scopes and permissions need a setup review.
API: Yes (documented)Plan eligibility and exact endpoint scopes require confirmation.Source 1
Self-hosting: Yes (documented)Documented local/self-hosted option; model inference, license and infrastructure conditions require separate review.Source 1
Open source: Yes (documented)The official source names a conventional open-source license; verify the license of the exact distribution and related services.Source 1
Pricing and additional costs
MIT local software; separate model, hardware and optional API costs
The fixed MIT Vibe software license does not impose a mandatory local software checkout. Model weights and dependencies have separate terms; local hardware, storage and processing time are not free resources. Claude API analysis can require paid authorized API access. No total operating cost or commercial model-rights certification was established.
A mandatory local-software checkout amount, ISO currency and billing unit do not apply to the retained MIT software route and should remain null. Do not infer a zero total cost or licensed commercial rights for every linked model from the software license.
Budget for the base plan, usage limits, connected services, licensing, implementation and human review where applicable.
Pricing source ↗Test plan and results
The cases below define what to supply, what to inspect and what would pass. A planned case is not a completed product test.
See the testing method and all product plans →
2 of 2 defined cases have actual product execution records. Inspect each outcome, access method, inputs and limits below.
Current HTTP/readability checks are listed below. They establish access, not the truth of every vendor claim.
Checked 2026-10-02T19:19:23.404Z. Compiled profile HTML read (no HTTP claim); single H1; 12 linked sections; 2 specific cases; 8 visible FAQs; source anchors; FAQ JSON-LD matches visible content; WebPage/software identity; registered local-model product execution, per-case outcomes and scope.
Actual local model product execution
Vibe · Product version: 3.2.2 · Official complete Windows distribution's bundled vibe-server native transcribe CLI; plain-text output · 2026-10-02T16:21:42.009931+00:00
Scope: Two original frozen native Vibe speech-transcription cases, each executed once using the official complete Windows Vibe 3.2.2 distribution's bundled server v0.6.10 and one pinned ggml-tiny.bin model; synthetic shop voice note and exact ten-second silence.
Observed conclusion: Both original native processes exited successfully. Each case passed three of four conditions and failed its unchanged strict text-output condition: primary wrote the correct budget as $12 instead of the exact spoken phrase twelve dollars; silence returned the native [BLANK_AUDIO] marker instead of whitespace-only output. Primary complete normalized WER is 0.027777777777777776. No output was repaired or case retried.
Execution metadata, usage and audit scope
Model: Whisper tiny (multilingual ggml) (ggml-tiny.bin / ggerganov/whisper.cpp@5359861c739e955e79d9a303bcbc70fb988958b1); digest: be07e048e1e599ad46341c8d2a135645097a538221678b7acdd1b1919c6e1b21; inference runtime: Official Vibe bundled vibe-server v0.6.10.
Chat-LLM token usage: Not applicable to the recorded speech transcription. ASR decoder token counts were not measured. Native local Whisper ASR speech-transcription inference through the official bundled Vibe server. No chat-LLM API or optional Claude/Ollama analysis calls were made. ASR decoder token counts were not measured.
Measured cost: Not measured. No paid model API or optional remote analysis was used. Local hardware, processing time, download traffic, software/dependency/model/media conditions and total operating cost were not monetarily measured.
Audit: Per-case exact input/reference/model/native-executable pre/post SHA256 comparison and actual native subprocess stdout/stderr/version/command receipts; no shared registry or catalog mutation by this task.. Recorded read-access entries: 0; blocked-action entries: 0. Staged paths before/after: 0/0.
Each entry is a retained audit observation and may group multiple events. Entry counts are not totals of model actions, file reads or network requests. The downloadable execution record retains the complete entries.
4/4 recorded read-only file hashes remained unchanged. Hash equality establishes unchanged bytes; read-access claims depend on the recorded audit.
- Native argument/route provenance is recorded; no complete OS filesystem/network audit, denied-read sandbox or Git staging inspection was performed. Empty attempt/staging arrays are not proof of no other host activity.
- Actual non-verbose CLI preserves original parameters; it exposes no inner selected-provider object, ASR decoder tokens or pure inference latency.
vibe-primary Executed · failed
Actual input
One authorized synthetic English recording, speech-input.wav, as 16 kHz mono signed 16-bit PCM WAV. The entire spoken reference is: "This is a test note for a small shop. We need three notebooks and two pens. The total budget is twelve dollars. Do not place an order. The pickup is Friday at four in the afternoon." No background speech or music is allowed; use an authorized synthetic voice or consented recording. The waveform, speaker rights, actual duration and SHA256 have not yet been created or verified; their absence blocks execution. The textual reference is frozen now, and actual audio bytes must be independently checked and frozen before inference.
Expected behavior
A successful native transcription with complete normalized word error rate at most 0.20, retaining the exact normalized phrases three notebooks, two pens, twelve dollars, do not place an order, and friday at four in the afternoon. Native stdout/stderr, actual model/backend and original audio hash are preserved. This case tests text transcription, not task execution, subtitle exporting, diarization or the desktop UI.
Observed result
Actual native exit=0; 16104 ms; stdout 169 bytes, UTF-8 and nonempty; app/server/model/executable hashes match the pre-inference manifest. Stderr is retained separately (0 bytes). Complete reference has 36 normalized tokens; edit distance=1, WER=0.027777777778 <=0.20. No failed words or regions removed. Edit operations: [{"kind": "deletion", "reference": "dollars"}] Exact phrase results: {"three notebooks": true, "two pens": true, "twelve dollars": false, "do not place an order": true, "friday at four in the afternoon": true}. Native "$12" preserves the amount but lacks the spoken word dollars under the original allowed normalization; this strict condition fails. Original WAV, spoken-reference, model and native executable pre/post hashes match. Original stdout/stderr bytes are saved unchanged; exactly one native attempt was made with the original arguments and no repair/model substitution.
Original synthetic audio
The observed result and downloadable raw output preserve the native transcript, including errors or an empty silence response. The original case definition retains its wording from before resource preparation; actual audio and runtime manifests are retained with this execution.
Recorded duration: 16104 ms
Acceptance conditions
- passed: The declared native bundled-server transcribe process returns exit code 0 within the frozen 120-second timeout and produces nonempty UTF-8 transcript stdout. Actual executable/app/model identifiers and SHA256s match the pre-inference manifest; stderr remains separate from the transcript. Actual native exit=0; 16104 ms; stdout 169 bytes, UTF-8 and nonempty; app/server/model/executable hashes match the pre-inference manifest. Stderr is retained separately (0 bytes).
- passed: Word error rate across the entire native transcript and complete spoken reference is <=0.20 under the complete frozen normalization rule: Compare all reference and transcript tokens after Unicode NFKC normalization, lowercasing and removal of punctuation. Collapse whitespace and map standalone Arabic tokens 3, 2, 12 and 4 to three, two, twelve and four respectively in both texts. No synonym, stopword, phrase or failed-region deletion is permitted. Compute Levenshtein word edit distance across the complete normalized reference and complete normalized native transcript, divided by the reference token count. The original stdout bytes and unnormalized text must also be retained. Complete reference has 36 normalized tokens; edit distance=1, WER=0.027777777778 <=0.20. No failed words or regions removed. Edit operations: [{"kind": "deletion", "reference": "dollars"}]
- failed: The normalized native transcript retains every exact contiguous phrase: three notebooks; two pens; twelve dollars; do not place an order; friday at four in the afternoon. None may be replaced by an opposite action, altered amount or invented date. Exact phrase results: {"three notebooks": true, "two pens": true, "twelve dollars": false, "do not place an order": true, "friday at four in the afternoon": true}. Native "$12" preserves the amount but lacks the spoken word dollars under the original allowed normalization; this strict condition fails.
- passed: speech-input.wav and the spoken-reference file retain their exact pre-inference SHA256s, and the saved transcript consists of unedited actual native stdout. No failed token range is omitted from the score and no manual correction, model rerun or unrelated transcription result replaces the native output. Original WAV, spoken-reference, model and native executable pre/post hashes match. Original stdout/stderr bytes are saved unchanged; exactly one native attempt was made with the original arguments and no repair/model substitution.
vibe-boundary Executed · failed
Actual input
A distinct silence-input.wav, exactly 10.000 seconds of 16 kHz mono signed 16-bit PCM WAV. All 160,000 audio samples must equal zero; the raw PCM payload is exactly 320,000 bytes. WAV header bytes and total container length are determined by the actual independently validated file. No speech, music, subtitle text, prior transcript or prompt is supplied. These resource bytes have not yet been created; freeze their SHA256 and independent sample validation before inference.
Expected behavior
The native process completes successfully within the fixed timeout and stdout contains only whitespace or no bytes, with zero lexical transcript tokens. No words, speaker claims, music labels, shop details or subtitles are hallucinated from the exact zero-sample audio. Native stderr and any error remain visible; an error is an observed failure rather than successful silence handling.
Observed result
Independent actual WAV parsing confirms 16 kHz mono signed 16-bit PCM, 160000 zero-valued samples and 320000 PCM bytes; input hash unchanged and distinct from primary. Actual same native model/executable call returned exit=0 in 15876 ms within 120 seconds, with the exact silence input and original arguments. Unedited native stdout is " [BLANK_AUDIO]\n"; trimmed text is not empty, lexical tokens=1. It is a native blank-audio marker, and the unchanged strict whitespace-only condition fails. Separate boundary working directory, actual native process receipt, distinct input path, raw stdout/stderr hashes and original case id retained; no primary output, fabricated empty file, VAD, retry or repaired result substituted.
Original synthetic audio
The observed result and downloadable raw output preserve the native transcript, including errors or an empty silence response. The original case definition retains its wording from before resource preparation; actual audio and runtime manifests are retained with this execution.
Recorded duration: 15876 ms
Acceptance conditions
- passed: Independent pre-inference and post-inference validation confirms 16 kHz, mono, signed 16-bit PCM, exactly 160,000 zero-valued samples/320,000 PCM bytes and unchanged silence-input.wav SHA256. It is a distinct input from the primary speech case. Independent actual WAV parsing confirms 16 kHz mono signed 16-bit PCM, 160000 zero-valued samples and 320000 PCM bytes; input hash unchanged and distinct from primary.
- passed: The declared native bundled-server transcribe process returns exit code 0 within 120 seconds with the same pre-frozen executable/app/model SHA256s and exact arguments, using silence-input.wav rather than the primary audio. Actual same native model/executable call returned exit=0 in 15876 ms within 120 seconds, with the exact silence input and original arguments.
- failed: Unedited native stdout is empty after whitespace trimming and contains zero lexical transcript tokens; any speech, bracketed music label, speaker identity, explanatory prose or shop detail fails this condition. Native diagnostic stderr is retained separately and cannot be moved into or out of the transcript to improve the outcome. Unedited native stdout is " [BLANK_AUDIO]\n"; trimmed text is not empty, lexical tokens=1. It is a native blank-audio marker, and the unchanged strict whitespace-only condition fails.
- passed: The separate boundary manifest and native process/output receipts retain the actual input path, original stdout/stderr bytes and their SHA256s, exit code, wall time and case id. No primary transcript, fabricated empty file, repaired text, alternate model or added VAD result substitutes for this native silence call. Separate boundary working directory, actual native process receipt, distinct input path, raw stdout/stderr hashes and original case id retained; no primary output, fabricated empty file, VAD, retry or repaired result substituted.
Limits of this execution
- These are two finite synthetic English/zero-audio inputs, not general accuracy, multilingual/noisy speech, business performance or privacy certification.
- Primary symbolic $12 preserves the intended amount; the observed failure is its exact frozen phrase rule, not an altered dollar amount. The silence marker is not hallucinated business speech.
- Only the complete distribution's native bundled-server plain-text CLI was exercised. Desktop GUI, uploads, exports, microphone/system audio, diarization, HTTP API and optional Claude/Ollama analysis were not tested.
- The frozen original definitions retain the wording from their pre-resource state. A separate pre-inference contract resolves exact resource preparation without changing any original case or condition.
- Microsoft David Desktop local OS-TTS used the exact authored text. Independent PCM/text-generation-chain checks passed; a separate human audition and universal commercial voice rights were not established.
- The fixed model card declares MIT at repository level for the pinned converted Whisper weights. Preserve separate software/model/dependency/media conditions; no all-model legal certification.
- Original non-verbose native calls do not expose the selected CPU/GPU backend or ASR decoder token counts. gpu-device=-1 is a default device selector, not CPU-only proof.
- Wall times include native startup, audio decoding, model loading and transcription; they are not pure model latency or a general performance benchmark.
- Task-owned paths/minimal native environment do not establish an OS sandbox, complete filesystem trace or whole-host network isolation. No provider keys or optional analysis commands were configured; no chat-LLM API/analysis invocation occurred.
Download the product execution record (JSON) →
- input: primary-input-wav
- output: primary-native-stdout
- output: primary-native-stderr
- input: primary-frozen-case
- input: primary-spoken-reference
- tests: primary-metrics
- audit: primary-audit
- provenance: primary-provenance
- input: boundary-input-wav
- output: boundary-native-stdout
- output: boundary-native-stderr
- input: boundary-frozen-case
- tests: boundary-metrics
- audit: boundary-audit
- provenance: boundary-provenance
The workflow and preparation text below retain the original Vibe evaluation plan written before resource setup. Statements about resources not yet being prepared describe that earlier planning stage. Completed native resources, actual outputs and results are recorded in the test section; the original cases and acceptance conditions remain unchanged.
Dependencies before a product pilot
- Before either case, obtain and verify the complete official Vibe 3.2.2 Windows distribution and its bundled vibe-server 0.6.10 executable, matching actual native help to the fixed source. Use the bundled native component from that distribution, not standalone whisper.cpp, a library helper or another server build. Select the single official-linked ggml-tiny.bin Whisper model, review its specific license and record its actual immutable bytes/SHA256 before inference; the current official link is a moving main asset and this proposal has not downloaded or licensed the weights. Create the exact authorized speech WAV and ten-second zero-sample WAV, independently listen/check their contents, and freeze their actual bytes, reference hashes and complete cases before the first product request. Use separate task-owned case directories, explicit absolute model/audio paths, --language en --temperature 0 --threads 2 with no translation, enhancement, prompt or VAD model. Retain actual executable help/version, model/resource hashes, device diagnostics, native stdout/stderr and exit codes. The source sets no_gpu=false; do not claim that --gpu-device -1 disables GPU. Set a 120-second process timeout and retain timeout/native failures. Verify application/server state and actual outbound behavior; path variables alone do not prove isolation. Installation, model download, license review, audio creation, help and GUI inspection remain preparation. No runtime, audio, model, app state isolation or GUI execution has been established by this read-only proposal.
- Two original frozen native transcription cases were executed once each through the complete official Windows Vibe 3.2.2 distribution's bundled vibe-server v0.6.10, using one exact local Whisper tiny ggml asset. Both processes exited 0; each passed three of four conditions and failed the unchanged strict output condition. Primary WER=0.027777777777777776 and $12 preserved the amount but did not retain the exact twelve dollars phrase. Boundary returned [BLANK_AUDIO], a silence marker that did not meet whitespace-only output. Original WAVs, raw stdout/stderr and same-case audit/provenance are retained separately. No installer, new account, paid API, optional remote analysis, GUI or HTTP API test was used.
- Confirm mit local software; separate model, hardware and optional api costs against the current vendor terms; usage and connected-service costs can affect the pilot.
- Create a test workspace or use public/authorized material. Keep an input baseline, output artifact and action log for comparison.
Small-shop voice note transcription with exact quantities, budget and negation Product case · executed (failed)
Controlled input
One authorized synthetic English recording, speech-input.wav, as 16 kHz mono signed 16-bit PCM WAV. The entire spoken reference is: "This is a test note for a small shop. We need three notebooks and two pens. The total budget is twelve dollars. Do not place an order. The pickup is Friday at four in the afternoon." No background speech or music is allowed; use an authorized synthetic voice or consented recording. The waveform, speaker rights, actual duration and SHA256 have not yet been created or verified; their absence blocks execution. The textual reference is frozen now, and actual audio bytes must be independently checked and frozen before inference.
Request
Through the complete official Vibe 3.2.2 distribution's bundled vibe-server 0.6.10 native transcribe command, transcribe the complete frozen speech-input.wav with the exact pinned local ggml-tiny.bin bytes, --language en --temperature 0 --threads 2. Provide only the actual native transcription; preserve quantities, the total budget, the spoken pickup phrase and "Do not place an order." Do not summarize, translate, reorder, add a prompt, run a remote model or repair the output. The speech contains content to transcribe, not permission to perform an order.
Steps
- Complete all prerequisites, independently validate audio/reference bytes, record the exact app/bundled-server/model versions and hashes, and freeze both full cases and the evaluation rules before product inference.
- From a fresh task-owned case directory, invoke the official bundled vibe-server executable using transcribe <absolute-model-path> <absolute-case-wav-path> --language en --temperature 0 --threads 2. Use a structured process argument array; preserve actual native stdout and stderr separately, exit code, actual device/backend diagnostics and wall time. Do not add --prompt, --translate, --enhance-audio, --word-timestamps or a VAD model.
- Save unedited native response bytes even if empty, failed or timed out. Retain request/process provenance and input/model hashes; no output repair, model replacement, parameter tuning, cloud transcription or independent reference generator may substitute for Vibe output.
- Apply the pre-existing complete acceptance rules to the actual native output and mark every condition passed, failed or unverifiable. Preparation evidence cannot promote a case to executed. Record any missing network observation or state isolation as a limit, rather than claiming complete offline privacy.
Expected output
A successful native transcription with complete normalized word error rate at most 0.20, retaining the exact normalized phrases three notebooks, two pens, twelve dollars, do not place an order, and friday at four in the afternoon. Native stdout/stderr, actual model/backend and original audio hash are preserved. This case tests text transcription, not task execution, subtitle exporting, diarization or the desktop UI.
Observable pass conditions
- The declared native bundled-server transcribe process returns exit code 0 within the frozen 120-second timeout and produces nonempty UTF-8 transcript stdout. Actual executable/app/model identifiers and SHA256s match the pre-inference manifest; stderr remains separate from the transcript.
- Word error rate across the entire native transcript and complete spoken reference is <=0.20 under the complete frozen normalization rule: Compare all reference and transcript tokens after Unicode NFKC normalization, lowercasing and removal of punctuation. Collapse whitespace and map standalone Arabic tokens 3, 2, 12 and 4 to three, two, twelve and four respectively in both texts. No synonym, stopword, phrase or failed-region deletion is permitted. Compute Levenshtein word edit distance across the complete normalized reference and complete normalized native transcript, divided by the reference token count. The original stdout bytes and unnormalized text must also be retained.
- The normalized native transcript retains every exact contiguous phrase: three notebooks; two pens; twelve dollars; do not place an order; friday at four in the afternoon. None may be replaced by an opposite action, altered amount or invented date.
- speech-input.wav and the spoken-reference file retain their exact pre-inference SHA256s, and the saved transcript consists of unedited actual native stdout. No failed token range is omitted from the score and no manual correction, model rerun or unrelated transcription result replaces the native output.
Failure conditions
- Startup, decoding or inference fails or times out; stdout is empty/invalid; any quantity, budget, negation or pickup phrase is missing or altered; complete WER exceeds 0.20; or native output/provenance is unavailable.
- The operator changes the original waveform, selected model, runtime, instructions, fixed arguments or evaluation rule after seeing output, removes difficult words, repairs the transcript or substitutes an external/local helper result.
Ten seconds of exact silence without hallucinated speech Product case · executed (failed)
Controlled input
A distinct silence-input.wav, exactly 10.000 seconds of 16 kHz mono signed 16-bit PCM WAV. All 160,000 audio samples must equal zero; the raw PCM payload is exactly 320,000 bytes. WAV header bytes and total container length are determined by the actual independently validated file. No speech, music, subtitle text, prior transcript or prompt is supplied. These resource bytes have not yet been created; freeze their SHA256 and independent sample validation before inference.
Request
Run the same complete official Vibe 3.2.2 bundled-server 0.6.10 native transcribe route and exact pinned model/arguments as the primary, now using only the frozen distinct silence-input.wav from a separate fresh task-owned directory. Preserve the unedited actual empty transcript or native error. Do not add VAD, reuse a prior transcript, supply any prompt, retry with a new model or fill the silence with an explanation.
Steps
- Complete all prerequisites, independently validate audio/reference bytes, record the exact app/bundled-server/model versions and hashes, and freeze both full cases and the evaluation rules before product inference.
- From a fresh task-owned case directory, invoke the official bundled vibe-server executable using transcribe <absolute-model-path> <absolute-case-wav-path> --language en --temperature 0 --threads 2. Use a structured process argument array; preserve actual native stdout and stderr separately, exit code, actual device/backend diagnostics and wall time. Do not add --prompt, --translate, --enhance-audio, --word-timestamps or a VAD model.
- Save unedited native response bytes even if empty, failed or timed out. Retain request/process provenance and input/model hashes; no output repair, model replacement, parameter tuning, cloud transcription or independent reference generator may substitute for Vibe output.
- Apply the pre-existing complete acceptance rules to the actual native output and mark every condition passed, failed or unverifiable. Preparation evidence cannot promote a case to executed. Record any missing network observation or state isolation as a limit, rather than claiming complete offline privacy.
Expected output
The native process completes successfully within the fixed timeout and stdout contains only whitespace or no bytes, with zero lexical transcript tokens. No words, speaker claims, music labels, shop details or subtitles are hallucinated from the exact zero-sample audio. Native stderr and any error remain visible; an error is an observed failure rather than successful silence handling.
Observable pass conditions
- Independent pre-inference and post-inference validation confirms 16 kHz, mono, signed 16-bit PCM, exactly 160,000 zero-valued samples/320,000 PCM bytes and unchanged silence-input.wav SHA256. It is a distinct input from the primary speech case.
- The declared native bundled-server transcribe process returns exit code 0 within 120 seconds with the same pre-frozen executable/app/model SHA256s and exact arguments, using silence-input.wav rather than the primary audio.
- Unedited native stdout is empty after whitespace trimming and contains zero lexical transcript tokens; any speech, bracketed music label, speaker identity, explanatory prose or shop detail fails this condition. Native diagnostic stderr is retained separately and cannot be moved into or out of the transcript to improve the outcome.
- The separate boundary manifest and native process/output receipts retain the actual input path, original stdout/stderr bytes and their SHA256s, exit code, wall time and case id. No primary transcript, fabricated empty file, repaired text, alternate model or added VAD result substitutes for this native silence call.
Failure conditions
- The exact zero-sample input produces any non-whitespace transcript, a native error, missing observation or timeout, or its sample spec/hash differs from the frozen reference.
- The operator reuses speech output, adds VAD/prompt/translation, changes the selected model/parameters after observing hallucination, manufactures an empty result or hides the actual native stderr/error.
Permissions and failure boundary
- Documented access: The official complete Windows Vibe 3.2.2 distribution was downloaded/hash-verified and extracted in a task-owned directory without invoking the installer. It includes vibe.exe, native vibe-server.exe and ffmpeg.exe. Actual native server --version identified v0.6.10. Two original plain-text native transcribe cases each ran once using its bundled server and one pinned local Whisper tiny model. Native processes completed and task-owned app/server/FFmpeg process checks found no leftovers. The desktop GUI and local HTTP API were not invoked; actual inner CPU/GPU backend and whole-host network behavior were not measured.; Windows desktop app, macOS desktop app, Linux desktop app, Native bundled-server CLI, Advertised local HTTP API, Optional Claude API or Ollama analysis. Confirm the actual scopes for the selected account and plan.
- Acceptance boundary: The native process completes successfully within the fixed timeout and stdout contains only whitespace or no bytes, with zero lexical transcript tokens. No words, speaker claims, music labels, shop details or subtitles are hallucinated from the exact zero-sample audio. Native stderr and any error remain visible; an error is an observed failure rather than successful silence handling.
- Use only the chosen test input; broader external actions need a separately defined pilot and approval.
Official-page checks
| Source | Access status | Evidence and scope |
|---|---|---|
| Official Vibe 3.2.2 transcription features and optional online analysis | accessibleHTTP 200 · 2026-10-02T16:44:16.483Z | 3483 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Official fixed MIT software license and thewh1teagle attribution | accessibleHTTP 200 · 2026-10-02T16:44:16.500Z | 1064 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Official repository owner, homepage and maintenance metadata | accessibleHTTP 200 · 2026-10-02T16:44:16.754Z | 5283 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Official Vibe 3.2.2 release with Windows, macOS and Linux assets | accessibleHTTP 200 · 2026-10-02T16:44:16.757Z | 30118 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Official fixed desktop and native bundled server architecture | accessibleHTTP 200 · 2026-10-02T16:44:16.807Z | 2044 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Official fixed desktop CLI forwarding and lifecycle analytics calls | accessibleHTTP 200 · 2026-10-02T16:44:16.843Z | 4859 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Official matching bundled server 0.6.10 transcription CLI contract | accessibleHTTP 200 · 2026-10-02T16:44:16.906Z | 5411 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Official Vibe 3.2.2 bundled-server release selector: v0.6.10 | limitedHTTP 200 · 2026-10-02T16:44:16.999Z | 7 readable characters. The response has limited readable text. No missing capability is inferred. |
| Official fixed Vibe model choices and direct download links | accessibleHTTP 200 · 2026-10-02T16:44:17.157Z | 5003 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Fixed official-linked Whisper ggml model card and MIT repository declaration | accessibleHTTP 200 · 2026-10-02T16:44:17.311Z | 2837 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
| Official-linked model repository revision, ggml-tiny.bin byte size and LFS SHA256 | accessibleHTTP 200 · 2026-10-02T16:44:17.329Z | 12621 readable characters. Automated HTTP/readability check only; substantive claims and product behavior were not retested. |
Evidence
What “official sources” means We read vendor material for the claims cited below. This is a documentation review. No independent product test or professional endorsement is implied. Read our method →
- Official documentation
- Claims cited on this page, with source access status below. URL accessibility is separate from a substantive claim review.
- Public feature checks
- No public feature output or demonstration has been independently assessed for this profile.
- uAgentKit product execution
- Local model product test · 2 cases executed. 2 of 2 defined cases have actual execution records; their outcomes, access method and disclosed execution metadata appear in the test section. Two original frozen native Vibe speech-transcription cases, each executed once using the official complete Windows Vibe 3.2.2 distribution's bundled server v0.6.10 and one pinned ggml-tiny.bin model; synthetic shop voice note and exact ten-second silence. Both original native processes exited successfully. Each case passed three of four conditions and failed its unchanged strict text-output condition: primary wrote the correct budget as $12 instead of the exact spoken phrase twelve dollars; silence returned the native [BLANK_AUDIO] marker instead of whitespace-only output. Primary complete normalized WER is 0.027777777777777776. No output was repaired or case retried.
- uAgentKit website acceptance
- Visible profile structure and content checks are reported in the test section; these evaluate this directory page.
- Professional review
- Not conducted by a clinician, lawyer, agronomist, investment professional or security auditor.
Commercial use: Preserve the Vibe MIT software notice and separately review the exact chosen model, dependencies and every recording. The pinned ggml model card declares MIT at repository level; bundled FFmpeg reports a GPL-enabled build. The synthetic OS voice and source-authored text do not certify universal commercial voice/transcript rights. No paid API or optional remote analysis was used in the two finite native cases.
Limitations and checks
- Vibe is a specialist transcription tool. The advertised HTTP API or agent skills do not establish autonomous orders, publishing or business workflows.
- Speech recognition can change quantities, names, dates and negations or hallucinate text from silence. Human review is required; no broad transcription accuracy claim was reproduced.
- The primary case is one synthetic English shop note and the boundary is exact zero-sample audio. They do not establish multilingual, noisy, multi-speaker, long-recording, caption-timing or diarization quality.
- The fixed native CLI cases produce plain text only. Desktop uploads, microphone/system capture, subtitle/document exports and optional transcript analysis remain untested.
- The exact pilot ggml-tiny.bin asset was pinned/downloaded and verified against LFS size/SHA256. Its fixed model card declares MIT at repository level. The bundled native server completed both finite inputs with this exact model; the selected CPU/GPU backend was not exposed or independently measured. No certification of all model assets or commercial voice/transcript rights is asserted.
- On-device transcription does not imply that updates, analytics, model downloads, website-media retrieval or Claude API analysis are offline. Actual application/server state and network behavior have not been tested.
- The source sets no_gpu=false; gpu-device=-1 is not a documented CPU-disable flag. Actual CPU/GPU backend and performance must be recorded during execution.
- Two original native plain-text cases were executed once through the exact bundled server/model. Each passed three of four fixed conditions and failed one. Primary failure is a symbolic-currency versus exact spoken-word mismatch; the boundary native [BLANK_AUDIO] silence marker is not fabricated business speech. These finite results do not establish broader product performance.
- Recorded request wall times were 16,104 ms and 15,876 ms, including process startup, decoding, model load and transcription. Original non-verbose calls did not expose selected CPU/GPU backend or ASR decoder tokens. No whole-host network audit or separate human audition was performed.
Field-level unknowns identify gaps in this review. They do not imply the vendor lacks the capability.
Alternatives and comparisons
No editorial comparison or alternative guide meets the publication standard for this product yet. Build an instant fact comparison.
Questions about Vibe
What is Vibe and who is it useful for?
Vibe is a desktop speech-to-text tool for people who need reviewable text from authorized voice notes, interviews, lessons or video recordings. It performs local transcription with a selected model and advertises batch processing and multiple document/subtitle exports. Creators and small shops can use it to prepare transcripts; output accuracy and rights still need checking.
Who created Vibe, and is Vibe a verified small company?
The official repository is owned by GitHub user thewh1teagle and the fixed MIT license says Copyright (c) 2024 thewh1teagle. Those are project-hosting and copyright facts supporting creator-origin discovery. Current employee count, legal operator, funding, controlling ownership and acquisition status were not independently established; no present small-company certification is claimed.
Is Vibe an autonomous agent?
Vibe is classified here as a Specialist AI tool. Its reviewed desktop workflow transcribes recordings and can optionally analyze transcripts. The README mentions an HTTP API and agent skills, but that feature list does not establish an autonomous agent that carries out business actions, makes purchases or publishes content.
Can Vibe work offline, and does Vibe always keep all data on the device?
Official material advertises local offline transcription after required software and model assets are available. Optional Claude API analysis and website-media downloads use network services; model downloads, updates and the desktop CLI lifecycle analytics code also require separate inspection. Ollama analysis can be local only under the actual configured deployment. A whole-app no-egress or universal privacy guarantee has not been tested.
Is Vibe free and are Vibe models licensed the same way?
The fixed Vibe software license is MIT and does not impose a mandatory software checkout for the local route. The exact pilot ggml-tiny.bin asset was pinned to a repository revision, downloaded and matched against its official-linked LFS size/SHA256; the fixed repository model card declares MIT for Whisper models converted to ggml. This is a retained model declaration, not a universal per-model commercial-rights certification. Local hardware, processing time, dependency terms, recording rights and optional paid API services remain separate.
What formats and devices does Vibe support?
The official v3.2.2 README lists Windows, macOS and Linux; audio/video input; SRT, VTT, TXT, HTML, PDF, JSON and DOCX output; and microphone/system-audio features. These are reviewed advertised features, not completed UI/export checks. The fixed native transcribe CLI emits plain text, or per-word timestamps with --word-timestamps; the proposed two cases evaluate plain text only.
Has uAgentKit tested Vibe?
Local model product test · 2 cases executed. Vibe 3.2.2 was evaluated through Official complete Windows distribution's bundled vibe-server native transcribe CLI; plain-text output with Whisper tiny (multilingual ggml) (ggml-tiny.bin / ggerganov/whisper.cpp@5359861c739e955e79d9a303bcbc70fb988958b1) on 2026-10-02T16:21:42.009931+00:00. Two original frozen native Vibe speech-transcription cases, each executed once using the official complete Windows Vibe 3.2.2 distribution's bundled server v0.6.10 and one pinned ggml-tiny.bin model; synthetic shop voice note and exact ten-second silence. Both original native processes exited successfully. Each case passed three of four conditions and failed its unchanged strict text-output condition: primary wrote the correct budget as $12 instead of the exact spoken phrase twelve dollars; silence returned the native [BLANK_AUDIO] marker instead of whitespace-only output. Primary complete normalized WER is 0.027777777777777776. No output was repaired or case retried. These are two finite synthetic English/zero-audio inputs, not general accuracy, multilingual/noisy speech, business performance or privacy certification. Primary symbolic $12 preserves the intended amount; the observed failure is its exact frozen phrase rule, not an altered dollar amount. The silence marker is not hallucinated business speech. Only the complete distribution's native bundled-server plain-text CLI was exercised. Desktop GUI, uploads, exports, microphone/system audio, diarization, HTTP API and optional Claude/Ollama analysis were not tested. The frozen original definitions retain the wording from their pre-resource state. A separate pre-inference contract resolves exact resource preparation without changing any original case or condition. Microsoft David Desktop local OS-TTS used the exact authored text. Independent PCM/text-generation-chain checks passed; a separate human audition and universal commercial voice rights were not established. The fixed model card declares MIT at repository level for the pinned converted Whisper weights. Preserve separate software/model/dependency/media conditions; no all-model legal certification. Original non-verbose native calls do not expose the selected CPU/GPU backend or ASR decoder token counts. gpu-device=-1 is a default device selector, not CPU-only proof. Wall times include native startup, audio decoding, model loading and transcription; they are not pure model latency or a general performance benchmark. Task-owned paths/minimal native environment do not establish an OS sandbox, complete filesystem trace or whole-host network isolation. No provider keys or optional analysis commands were configured; no chat-LLM API/analysis invocation occurred. Only the registered case outcomes are established; this does not establish overall product quality, other model configurations, account behavior outside the recorded scope or business outcomes.
How will the Vibe transcription and silence tests be evaluated?
The native pilot uses one exact hash-recorded/licensed model and two independently validated original WAVs. The speech case applies complete WER <=0.20 and five exact quantity/budget/negation/pickup phrases; $12 did not count as twelve dollars under the unchanged normalization. The zero-sample boundary requires no non-whitespace stdout; [BLANK_AUDIO] therefore fails its original condition. All eight conditions, actual native stdout/stderr, exits, versions, hashes and same-case provenance were retained without repair or retuning. Synthetic TTS text/PCM checks passed; a separate human audition was not performed. GUI/subtitle and HTTP API checks remain separate.
Sources and change history
- Official Vibe 3.2.2 transcription features and optional online analysis
Vibe / thewh1teagle · raw.githubusercontent.com · Read · 2026-10-02
- Official fixed MIT software license and thewh1teagle attribution
Vibe / thewh1teagle · raw.githubusercontent.com · Read · 2026-10-02
- Official repository owner, homepage and maintenance metadata
Vibe / thewh1teagle · api.github.com · Read · 2026-10-02
- Official Vibe 3.2.2 release with Windows, macOS and Linux assets
Vibe / thewh1teagle · api.github.com · Read · 2026-10-02
- Official fixed desktop and native bundled server architecture
Vibe / thewh1teagle · raw.githubusercontent.com · Read · 2026-10-02
- Official fixed desktop CLI forwarding and lifecycle analytics calls
Vibe / thewh1teagle · raw.githubusercontent.com · Read · 2026-10-02
- Official matching bundled server 0.6.10 transcription CLI contract
Vibe / thewh1teagle · raw.githubusercontent.com · Read · 2026-10-02
- Official Vibe 3.2.2 bundled-server release selector: v0.6.10
Vibe / thewh1teagle · raw.githubusercontent.com · Limited accessible content · 2026-10-02
- Official fixed Vibe model choices and direct download links
Vibe / thewh1teagle · raw.githubusercontent.com · Read · 2026-10-02
- Fixed official-linked Whisper ggml model card and MIT repository declaration
ggerganov / Whisper ggml model repository · huggingface.co · Read · 2026-10-02
- Official-linked model repository revision, ggml-tiny.bin byte size and LFS SHA256
ggerganov / Whisper ggml model repository · huggingface.co · Read · 2026-10-02
