Fermion Research
Speech recognition · Available now

Phonon-2

The most accurate open speech recognition model under 900 MB. 164 MB, 5.21 % word error on seven public test sets, 174× realtime on a MacBook Air.

164 MB
Download, 177 MB on disk
5.21 %
Seven-set word error rate
174×
Realtime, one stream, M5 MacBook Air
6,680×
Realtime, batch 128, H100

Overview

Phonon-2 is the most accurate open speech recognition model under 900 MB. In a 164 MB download it averages 5.21 % word error on the Open ASR Leaderboard’s seven English sets, and every open model that scores better is at least 5.8 times its size. Set for set it holds the accuracy of its 2.5 GB full-precision teacher and beats it on meetings and parliamentary speech. Its encoder stores every weight as one of five learned levels in about 2.1 bits.

It transcribes at 174 times realtime on an M5 MacBook Air, where Parakeet TDT 0.6B v3 in FluidAudio’s Core ML runtime reaches 104.9 on the same audio; at 143 times on eight Zen 5 cores (16 vCPU); and at 6,680 times on one H100 in batches of 128. A Core ML runtime for Apple devices is coming soon. The model writes punctuated, capitalized text, and the weights are released under CC-BY-4.0. Phonon-2 is the model behind Detta, the Fermion Research dictation app for the Mac.

Models

Phonon-2 is a single file. Phonon-1 remains available beside it.

Phonon-2Encoder about 2.1 bits per stored weight, 6-bit tables elsewhere
164 MB download · 177 MB on disk
Phonon-1The earlier model
415 MB download · 455 MB on disk
Language
English
Audio inputmicrophone or file
16 kHz
Output
Punctuated, capitalized text

The Phonon-1 spec

Evaluation

Phonon-2 averages 5.21 % word error on the seven public test sets of the Open ASR Leaderboard, scored with the board’s own code on the full test sets. The table sets it beside its full-precision teacher, Parakeet TDT 0.6B v3, which averages 4.96 in a 2,508 MB download, and six other open models from 178 MB to about 8 GB.

Seven-set comparison

  • Phonon-2Fermion Research · 0.60 B · 164 MB

    LS clean
    1.72
    LS other
    3.92
    AMI
    9.37
    Earnings-22
    6.96
    GigaSpeech
    8.35
    SPGISpeech
    3.70
    VoxPopuli
    2.46
    Average
    5.21
  • parakeet-tdt-0.6b-v3nvidia · 0.60 B · 2,508 MB

    LS clean
    1.52
    LS other
    3.13
    AMI
    9.42
    Earnings-22
    5.85
    GigaSpeech
    7.99
    SPGISpeech
    3.63
    VoxPopuli
    3.19
    Average
    4.96
  • Parakeet ReduxMoondream · 0.60 B · 178 MB

    LS clean
    1.94
    LS other
    4.35
    AMI
    9.16
    Earnings-22
    7.90
    GigaSpeech
    8.62
    SPGISpeech
    4.01
    VoxPopuli
    3.87
    Average
    5.69
  • Phonon-1Fermion Research · 0.78 B · 415 MB

    LS clean
    2.11
    LS other
    5.03
    AMI
    10.31
    Earnings-22
    12.34
    GigaSpeech
    8.73
    SPGISpeech
    3.67
    VoxPopuli
    3.73
    Average
    6.56
  • canary-180m-flashnvidia · 0.18 B · 737 MB

    LS clean
    1.52
    LS other
    3.42
    AMI
    12.09
    Earnings-22
    8.33
    GigaSpeech
    8.87
    SPGISpeech
    2.04
    VoxPopuli
    3.57
    Average
    5.69
  • Voxtral-Mini-4B-Realtime-2602mistralai · 4.00 B · 8,000 MB*

    LS clean
    1.62
    LS other
    4.94
    AMI
    13.34
    Earnings-22
    9.31
    GigaSpeech
    8.80
    SPGISpeech
    2.23
    VoxPopuli
    2.60
    Average
    6.12
  • whisper-large-v3-turboopenai · 0.80 B · 1,618 MB

    LS clean
    2.13
    LS other
    3.71
    AMI
    13.88
    Earnings-22
    8.09
    GigaSpeech
    8.47
    SPGISpeech
    2.79
    VoxPopuli
    7.02
    Average
    6.58
  • nemotron-3.5-asr-streaming-0.6bnvidia · 0.64 B · 2,368 MB

    LS clean
    2.83
    LS other
    6.79
    AMI
    13.43
    Earnings-22
    15.30
    GigaSpeech
    9.86
    SPGISpeech
    3.27
    VoxPopuli
    4.24
    Average
    7.96
Table 1Word error rate, %, on the seven sets, lower is better; bold marks the best value in each column. Leaderboard rows are its published results of 25 September 2026, and the other rows were scored with its code on the same full test sets.

Throughput

One 164 MB file runs on every surface, and the fast path on each keeps the accuracy of the exact one.

SurfaceTimes realtimeWord error, fast path against exact path
Apple M5 MacBook Air, GPU (MLX)174×2.94 % against 2.94 % (400 LibriSpeech utterances)
Apple M5 MacBook Air, CPU only40×2.33 % on a 40-clip check
Linux x86-64, eight Zen 5 cores (16 vCPU)142.8×3.94 % against 3.91 % (LibriSpeech test-other)
Linux Arm, eight Google Axion cores52.2×3.90 % against 3.91 % (LibriSpeech test-other)
Windows x64, 8 vCPU21.0×2.21 % on a 40-clip check
NVIDIA A100 80 GB267× one stream · 3,614× batch 1285.20–5.22 % against 5.20 % (seven sets)
NVIDIA H100 80 GB465× one stream · 6,680× batch 1285.20–5.22 % against 5.20 % (seven sets)
Table 2One stream at a time unless a row says batch.
Runtime on the same M5 MacBook AirTimes realtime
Phonon-2 (MLX)174.0×
FluidAudio, Parakeet TDT 0.6B v3 (Core ML)104.9×
FluidAudio, Parakeet Redux (Core ML)27.8×
Moonshine tiny26.7×
whisper.cpp, large-v3-turbo (Metal)17.0×
sherpa-onnx, Parakeet TDT 0.6B v3 (int8)16.5×
Table 3The same 20 dictations, 797 seconds of speech, on the same MacBook Air, one stream at a time with load time excluded, each runtime at its defaults.
1×3×10×30×100×300×realtime factor, one stream, Apple M5 MacBook Air 16 GB, log scaledownload WERPhonon-2 (MLX, default): 174× realtime, 5.7 ms per second of audio, WER 5.21 (our read, board scorer (accuracy within noise; 10/400 hypotheses differ))Phonon-2 (MLX, default)174× 164 MB 5.21Phonon-2 (MLX, exact decode): 109× realtime, 9.2 ms per second of audio, WER 5.21 (our read, board scorer)Phonon-2 (MLX, exact decode)109× 164 MB 5.21FluidAudio Parakeet TDT 0.6B v3 (Core ML): 104.9× realtime, 9.5 ms per second of audio, WER 4.96 (board's run)FluidAudio Parakeet TDT 0.6B v3 (Core ML)105× 483 MB 4.96Phonon-1 Big (app default): 52× realtime, 19.2 ms per second of audio, WER 6.65 (our read, board scorer)Phonon-1 Big (app default)52× 581 MB 6.65FluidAudio Parakeet Redux (Core ML): 27.8× realtime, 35.9 ms per second of audioFluidAudio Parakeet Redux (Core ML)28× 220 MB —Moonshine tiny (moonshine-voice): 26.7× realtime, 37.4 ms per second of audioMoonshine tiny (moonshine-voice)27× 44 MB —Moonshine base (moonshine-voice): 21.6× realtime, 46.3 ms per second of audioMoonshine base (moonshine-voice)22× 141 MB —whisper.cpp large-v3-turbo (ggml f16): 17× realtime, 58.7 ms per second of audio, WER 6.58 (board's run)whisper.cpp large-v3-turbo (ggml f16)17× 1,625 MB 6.58sherpa-onnx Parakeet TDT 0.6B v3 (int8): 16.5× realtime, 60.5 ms per second of audio, WER 4.96 (board's run)sherpa-onnx Parakeet TDT 0.6B v3 (int8)17× 670 MB 4.96whisper.cpp large-v3-turbo (q5_0): 12.2× realtime, 81.8 ms per second of audio, WER 6.58 (board's run)whisper.cpp large-v3-turbo (q5_0)12× 574 MB 6.58Moonshine medium-streaming (moonshine-voice): 6.1× realtime, 163.5 ms per second of audioMoonshine medium-streaming (moonshine-voice)6× 269 MB —whisper.cpp large-v3 (ggml f16): 3.2× realtime, 309.6 ms per second of audio, WER 6.04 (board's run)whisper.cpp large-v3 (ggml f16)3× 3,095 MB 6.04whisper.cpp large-v3 (q5_0): 2.7× realtime, 363.8 ms per second of audio, WER 6.04 (board's run)whisper.cpp large-v3 (q5_0)3× 1,081 MB 6.04
Figure 1Speed on a MacBook Air. Single-stream realtime factor on an Apple M5 MacBook Air, 16 GB.

Run it

Detta, the dictation app for the Mac, runs Phonon-2 in any text field.

$ pip install fermion-research
$ phonon transcribe meeting.wav
$ phonon serve --port 8010
$ phonon listen
$ fermion transcribe phonon-2 meeting.wav

The fermion command line installs from PyPI. It transcribes files, serves an OpenAI-compatible endpoint and, on a Mac, transcribes the microphone live.

$ pip install mlx mlx-audio mlx-lm soundfile scipy zstandard

On Apple silicon it runs Phonon-2 on the GPU through MLX.

$ pip install --no-deps torch --index-url https://download.pytorch.org/whl/cpu
$ pip install fermion-research torch safetensors soundfile scipy zstandard

On Linux (x86-64 and Arm) and Windows the same package runs the CPU engine.

$ docker run --rm -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cpu:2.0.3 transcribe phonon-2 /audio/recording.wav

The CPU container runs on amd64 and arm64.

$ docker run --rm --gpus all -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cuda:1.0.4 transcribe phonon-2 /audio/recording.wav

The CUDA container runs Phonon-2 on NVIDIA GPUs.

Weights on Hugging Facefermion-research on PyPIEngines on GitHub

Availability

Open weightsCC-BY-4.0Apple siliconNVIDIA GPUsCPUs

License

The weights are released under the Creative Commons Attribution 4.0 licence, which Phonon-2 inherits from NVIDIA’s Parakeet TDT 0.6B v3. The licence permits commercial use, modification and redistribution with attribution. The command line is released under Apache 2.0.

Citation

@misc{fermionresearch2026phonon2,
  title  = {Phonon-2},
  author = {{Fermion Research}},
  year   = {2026},
  url    = {https://fermionresearch.com/models/phonon-2/}
}

Related research

All research