How the matcher scores models
The goal is not to invent a universal leaderboard. It is to reduce bad matches between a user’s constraints and a model’s documented capabilities.
1. Hard gate: language support
If the reviewed primary source does not list the selected language, the model receives a heavy penalty. For Slovenian we also distinguish explicit support from NVIDIA Nemotron’s current “adaptation-ready” tier, which is not the same thing as ready-to-use transcription.
2. Hardware fit
Hardware fit is a conservative qualitative score based on model size, official memory guidance where available, and the practical local runtimes exposed by the model family. We avoid pretending that one VRAM number applies to every runtime or quantization.
3. Workload fit
Long recordings, live dictation, subtitles, voice agents and edge use put different pressure on latency, memory and streaming architecture. Native streaming receives a larger boost when the user explicitly asks for it.
4. License visibility
License is displayed but does not silently decide technical quality. MIT and Apache-2.0 are labeled permissive; CC BY 4.0 is shown with attribution obligations; custom licenses are flagged for manual review. This website is not legal advice.
5. No fake precision
The fit score is an internal decision score, not a WER claim. A 92/100 match does not mean 92% transcription accuracy. Before production, test the top candidates on representative audio, language, accent, noise and hardware.
Update policy
Each database entry has a review date. We prioritize primary model cards and official repositories. If a source changes supported languages, licensing or deployment guidance, the entry should be updated before the site claims it.