
Research
Introducing Phonon-2
The most accurate open speech recognition model under 900 MB, in a 164 MB download that turns an hour of audio into text in about 20 seconds on a MacBook Air.

The most accurate open speech recognition model under 900 MB. 164 MB, 5.21 % word error on seven public test sets, 174× realtime on a MacBook Air.
Overview
Phonon-2 is the most accurate open speech recognition model under 900 MB. In a 164 MB download it averages 5.21 % word error on the Open ASR Leaderboard’s seven English sets, and every open model that scores better is at least 5.8 times its size. Set for set it holds the accuracy of its 2.5 GB full-precision teacher and beats it on meetings and parliamentary speech. Its encoder stores every weight as one of five learned levels in about 2.1 bits.
It transcribes at 174 times realtime on an M5 MacBook Air, where Parakeet TDT 0.6B v3 in FluidAudio’s Core ML runtime reaches 104.9 on the same audio; at 143 times on eight Zen 5 cores (16 vCPU); and at 6,680 times on one H100 in batches of 128. A Core ML runtime for Apple devices is coming soon. The model writes punctuated, capitalized text, and the weights are released under CC-BY-4.0. Phonon-2 is the model behind Detta, the Fermion Research dictation app for the Mac.
Models
Phonon-2 is a single file. Phonon-1 remains available beside it.
Evaluation
Phonon-2 averages 5.21 % word error on the seven public test sets of the Open ASR Leaderboard, scored with the board’s own code on the full test sets. The table sets it beside its full-precision teacher, Parakeet TDT 0.6B v3, which averages 4.96 in a 2,508 MB download, and six other open models from 178 MB to about 8 GB.
Seven-set comparison
| Model | Params | Download | LS clean | LS other | AMI | Earnings-22 | GigaSpeech | SPGISpeech | VoxPopuli | Average |
|---|---|---|---|---|---|---|---|---|---|---|
| Phonon-2Fermion Research | 0.60 B | 164 MB | 1.72 | 3.92 | 9.37 | 6.96 | 8.35 | 3.70 | 2.46 | 5.21 |
| parakeet-tdt-0.6b-v3nvidia | 0.60 B | 2,508 MB | 1.52 | 3.13 | 9.42 | 5.85 | 7.99 | 3.63 | 3.19 | 4.96 |
| Parakeet ReduxMoondream | 0.60 B | 178 MB | 1.94 | 4.35 | 9.16 | 7.90 | 8.62 | 4.01 | 3.87 | 5.69 |
| Phonon-1Fermion Research | 0.78 B | 415 MB | 2.11 | 5.03 | 10.31 | 12.34 | 8.73 | 3.67 | 3.73 | 6.56 |
| canary-180m-flashnvidia | 0.18 B | 737 MB | 1.52 | 3.42 | 12.09 | 8.33 | 8.87 | 2.04 | 3.57 | 5.69 |
| Voxtral-Mini-4B-Realtime-2602mistralai | 4.00 B | 8,000 MB* | 1.62 | 4.94 | 13.34 | 9.31 | 8.80 | 2.23 | 2.60 | 6.12 |
| whisper-large-v3-turboopenai | 0.80 B | 1,618 MB | 2.13 | 3.71 | 13.88 | 8.09 | 8.47 | 2.79 | 7.02 | 6.58 |
| nemotron-3.5-asr-streaming-0.6bnvidia | 0.64 B | 2,368 MB | 2.83 | 6.79 | 13.43 | 15.30 | 9.86 | 3.27 | 4.24 | 7.96 |
Phonon-2Fermion Research · 0.60 B · 164 MB
parakeet-tdt-0.6b-v3nvidia · 0.60 B · 2,508 MB
Parakeet ReduxMoondream · 0.60 B · 178 MB
Phonon-1Fermion Research · 0.78 B · 415 MB
canary-180m-flashnvidia · 0.18 B · 737 MB
Voxtral-Mini-4B-Realtime-2602mistralai · 4.00 B · 8,000 MB*
whisper-large-v3-turboopenai · 0.80 B · 1,618 MB
nemotron-3.5-asr-streaming-0.6bnvidia · 0.64 B · 2,368 MB
Throughput
One 164 MB file runs on every surface, and the fast path on each keeps the accuracy of the exact one.
| Surface | Times realtime | Word error, fast path against exact path |
|---|---|---|
| Apple M5 MacBook Air, GPU (MLX) | 174× | 2.94 % against 2.94 % (400 LibriSpeech utterances) |
| Apple M5 MacBook Air, CPU only | 40× | 2.33 % on a 40-clip check |
| Linux x86-64, eight Zen 5 cores (16 vCPU) | 142.8× | 3.94 % against 3.91 % (LibriSpeech test-other) |
| Linux Arm, eight Google Axion cores | 52.2× | 3.90 % against 3.91 % (LibriSpeech test-other) |
| Windows x64, 8 vCPU | 21.0× | 2.21 % on a 40-clip check |
| NVIDIA A100 80 GB | 267× one stream · 3,614× batch 128 | 5.20–5.22 % against 5.20 % (seven sets) |
| NVIDIA H100 80 GB | 465× one stream · 6,680× batch 128 | 5.20–5.22 % against 5.20 % (seven sets) |
| Runtime on the same M5 MacBook Air | Times realtime |
|---|---|
| Phonon-2 (MLX) | 174.0× |
| FluidAudio, Parakeet TDT 0.6B v3 (Core ML) | 104.9× |
| FluidAudio, Parakeet Redux (Core ML) | 27.8× |
| Moonshine tiny | 26.7× |
| whisper.cpp, large-v3-turbo (Metal) | 17.0× |
| sherpa-onnx, Parakeet TDT 0.6B v3 (int8) | 16.5× |
Run it
Detta, the dictation app for the Mac, runs Phonon-2 in any text field.
$ pip install fermion-research$ phonon transcribe meeting.wav$ phonon serve --port 8010$ phonon listen$ fermion transcribe phonon-2 meeting.wav
The fermion command line installs from PyPI. It transcribes files, serves an OpenAI-compatible endpoint and, on a Mac, transcribes the microphone live.
$ pip install mlx mlx-audio mlx-lm soundfile scipy zstandard
On Apple silicon it runs Phonon-2 on the GPU through MLX.
$ pip install --no-deps torch --index-url https://download.pytorch.org/whl/cpu$ pip install fermion-research torch safetensors soundfile scipy zstandard
On Linux (x86-64 and Arm) and Windows the same package runs the CPU engine.
$ docker run --rm -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cpu:2.0.3 transcribe phonon-2 /audio/recording.wav
The CPU container runs on amd64 and arm64.
$ docker run --rm --gpus all -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cuda:1.0.4 transcribe phonon-2 /audio/recording.wav
The CUDA container runs Phonon-2 on NVIDIA GPUs.
Weights on Hugging Facefermion-research on PyPIEngines on GitHub
Availability
The weights are released under the Creative Commons Attribution 4.0 licence, which Phonon-2 inherits from NVIDIA’s Parakeet TDT 0.6B v3. The licence permits commercial use, modification and redistribution with attribution. The command line is released under Apache 2.0.
@misc{fermionresearch2026phonon2,
title = {Phonon-2},
author = {{Fermion Research}},
year = {2026},
url = {https://fermionresearch.com/models/phonon-2/}
}