Benchmark policy
A number without its evidence package is advertising, not a reproducible benchmark. VocalCode does not currently publish quantitative accuracy, latency or competitor scorecards.
Why there are no numbers here
An earlier draft used internal measurements without a complete, checked-in substantiation bundle, and part of its evaluation data was not licensed for commercial promotional use. We withdrew those claims. They will not return unless both the evidence and the data permissions withstand review.
Models selected for the release
- European-language path: an on-device NVIDIA Parakeet TDT model distributed under CC BY 4.0.
- Mandarin Chinese and English path: an on-device Paraformer model distributed under Apache 2.0.
- Korean and Japanese path: an on-device SenseVoice model derived from FunAudioLLM SenseVoice.
- Punctuation: a local CT-Transformer punctuation model distributed under Apache 2.0.
These are model-selection notes, not accuracy rankings. Results vary with the speaker, microphone, acoustic environment, language, vocabulary, utterance length, hardware and thread configuration. The 25 European language entries currently reflect upstream Parakeet model support and remain preview labels until each release binary has completed the corresponding product-level audio tests.
Evidence required before publishing a result
- Evaluation rights: a dataset licence that permits the intended commercial measurement and marketing use, with attribution and other conditions recorded.
- Exact inputs: corpus version, sample identifiers, preparation commands and input file hashes.
- Exact software: application revision, model source revisions and hashes, runtime versions, decoding settings and scoring normalisation.
- Raw evidence: references, unedited model outputs, failures, per-sample timings and the script that produces every aggregate.
- Test environment: operating system, CPU, memory, power mode, warm-up policy, thread counts and whether model-load, audio capture and text insertion are included.
- Limitations: uncertainty, excluded samples, adverse cases and a clear distinction between a model test and an end-to-end product test.
What a future comparison must not imply
Running two model files through one harness does not establish that one finished dictation product is better than another. Applications may use different model sizes, vocabulary, audio preprocessing, post-processing and cloud services. A product claim requires testing the named product itself under the same disclosed protocol.
Report a reproducibility issue
If we publish a future evidence bundle and you cannot reproduce it, please open an issue in the public documentation tracker. We will correct or withdraw a claim that its evidence does not support.