
Introducing Neutrino.
Three foundation models in a proprietary ternary-family format eight times smaller than fp16, served by one engine built alongside them. All available today.
- 2.56 GB
- download, 8B
- 1/8
- the bits
- 763
- tok/s, drafted
The family
Three models share one container format and one set of binaries. The 8B carries 36 decoder layers at a hidden width of 4,096; both 0.6B models carry 28 at 1,024, and tie their embedding table so a single tensor serves the input lookup and the output head. Because the runtime reads all three without a rebuild, the 0.6B loads into the 8B's own process as its drafting model rather than standing up a second deployment.
Neutrino-1 8B
flagship
- Base model
- Qwen3-8B
- Parameters
- 8,190,735,360
- Decoder layers
- 36
- Hidden width
- 4,096
- Feed-forward width
- 12,288
- Attention
- 32 query, 8 key-value
- Embedding table
- untied, two tensors
- KV cache
- 144 KiB
- Context length
- 40,960 tokens
- Download
- 2.56 GB
- On disk
- 3.88 GB
- Apple M5, CPU cores
- 24.9 tok/s
- Runtimes
- pip engine, GGUF pack, MLX pack
- Role
- General assistant and tool caller
- Status
- Available 2026-07-27
Neutrino-1 0.6B
certified draft
- Base model
- Qwen3-0.6B
- Parameters
- 596,049,920
- Decoder layers
- 28
- Hidden width
- 1,024
- Feed-forward width
- 3,072
- Attention
- 16 query, 8 key-value
- Embedding table
- tied, one tensor
- KV cache
- 112 KiB
- Context length
- 40,960 tokens
- Download
- 238 MB
- On disk
- 328 MB
- Apple M5, CPU cores
- 225 to 236 tok/s
- Runtimes
- pip engine, GGUF pack, MLX pack
- Role
- Draft model for the 8B's decode
- Status
- Available 2026-07-27
Neutrino-1 0.6B Chat
conversational
- Base model
- Qwen3-0.6B
- Parameters
- 596,049,920
- Decoder layers
- 28
- Hidden width
- 1,024
- Feed-forward width
- 3,072
- Attention
- 16 query, 8 key-value
- Embedding table
- tied, one tensor
- KV cache
- 112 KiB
- Context length
- 40,960 tokens
- Download
- 238 MB
- On disk
- 328 MB
- Apple M5, CPU cores
- 170 to 218 tok/s
- Runtimes
- pip engine, GGUF pack, MLX pack
- Role
- Small talk and short helpful replies
- Status
- Available 2026-07-27
Where a parameter stops being cheap
The format only touches the transformer linears. Token embeddings stay int8 because their rows are read one token at a time instead of being multiplied against the whole activation stream, and that split moves with scale. At 8B the coded lane holds 84.8% of the parameters in 67.2% of the bytes. At 0.6B it holds 73.9% of the parameters in 50.4% of the bytes, and the untouched vocabulary swells from a third of the file to nearly half. Both 0.6B models share that split exactly: the lane sizes follow from the tensor shapes, so the chat SKU is the same 328 MB down to the byte.
Neutrino-1 8B
36 layers
Neutrino-1 0.6B
28 layers
Neutrino-1 0.6B Chat
28 layers
- Download
- 2.56 GB, lossless
- Size on disk
- 3.88 GB
- Runs on
- a 16 GB MacBook (Apple M5)
- General knowledge (MMLU)
- 72.1
- Knowledge, re-annotated set (MMLU-Redux)
- 67.8
- Instruction following (IFEval)
- 77.2
- Tool use (BFCL)
- 68.9
- Math (GSM8K, flexible)
- 53.4
- Math (GSM8K, stated format)
- 51.73
- License
- Apache 2.0, open weights
Neutrino-1 0.6B
The draft model.
Speculative decoding turns one large forward pass into several emitted tokens, and it only pays when the small model proposing those tokens is cheap enough to run six times in the gaps. Neutrino-1 0.6B is that model: 596 million parameters in a 328 MB container, 238 MB on the wire, decoding above 200 tokens a second on the CPU cores of a 16 GB Apple M5. The 8B accepts a proposal only when it equals its own argmax, which is why the drafted stream is the plain greedy stream and the certificate reads 27,648 tokens with zero divergences.
The same property makes it a model in its own right. It is the identical container format, read by the identical binaries, so a laptop that runs the 8B runs this one at nine times the rate from a file that fits on a phone.
- 596,049,92028 layers, hidden 1,024
- Geometry
16 query heads share 8 key-value heads at a head width of 128, and the embedding table is tied: one int8 tensor serves both the input lookup and the output head.
- 47.5%of a 328 MB container
- Composition
The share of the file held by that tied vocabulary. It is 26.1% of the parameters and nearly half the bytes, because the format touches the 196 transformer linears and leaves the token table alone.
- 225 to 236 tok/sApple M5, CPU cores, 9 threads
- Native rate
Single-stream decode from the shipped binary with no GPU involved. The container is memory-mapped and executed as it arrives: there is no dequantization pass at load.
- 201.2 tok/sApple M5, 16 GB, MLX pack
- Metal rate
Median of five 512-token greedy runs through the custom Metal kernels, peaking at 0.53 GiB of memory. The whole model plus its working set fits in half a gigabyte.
- 1,177 tok/sH100, inside the 8B's process
- Draft walker
Standalone decode as the C-served draft, 0.85 ms per token. That headroom is what lets six proposals be generated in the time the 8B spends on one verify pass.
- 27,648 tokensH100, dynamic draft length
- Certified drafting
Consecutive tokens drafted into the 8B and verified against plain greedy decode with zero divergences. On counting prompts every drafted token is accepted and the 8B delivers 762.6 tok/s, against 395.9 undrafted.
Against Qwen3-0.6B-Base
The control for a compressed model is its own base at full precision, graded on the same harness on the same day. The reference here is Qwen3-0.6B-Base in bf16. Both sides below carry 596,049,920 parameters, and both were run through eleven zero-shot tasks rather than one, because a single task is a cherry and a grid is a measurement. Winogrande retains 91.5% of the reference and PIQA 90.4%. Accuracy does not fall anything like as fast as the bytes do.
81.6 percent of Qwen3-0.6B-Base’s ability across seven common tasks, at 15.8 percent of the bytes.
15.8% of Qwen3-0.6B-Base, which ships 1,503 MB
78.0% averaged across all eleven tasks measured
Neutrino-1 0.6B on Hugging FaceThe drafted decode path on the 8B
Neutrino-1 0.6B Chat
A pocket conversationalist.
Conversation at 596 million parameters is a behavior install, not a capability curve. The chat SKU runs the same 28 layers at the same 1,024 width through the same tied embedding table, and it exports to a container of exactly 327,719,836 bytes, the same figure as the draft model. The behaviour is trained in; the geometry and the footprint do not move.
What it buys is measured on both shipping runtimes at one documented decode config each. On 20 held-out sentence fragments the model answers naturally 20 times out of 20, and on 30 held-out factual asks it is correct or honestly uncertain 30 times out of 30, with no template leaks across the five probes designed to draw one. Identity holds 8 of 8 on the native binary and 8 of 8 on the Apple-silicon pack. Across 200 held-out chat prompts the native binary terminates on its own 88.5% of the time at a 4.5% loop rate, and the Apple-silicon pack reads 85.0% and 5.0%.
sup
Hey there! What's up?
who you
I'm Neutrino, a tiny language model made by Fermion Research. What's up? How can I help?
what is the moon
The Moon is a natural satellite that orbits Earth. It reflects sunlight, which makes it visible in the night sky.
Are you ChatGPT?
No, I'm Neutrino. I'm a small language model made by Fermion Research. I chat and help with phrasing, but I'm not ChatGPT.
Tuned settings
Already tuned when it lands.
Every runtime ships carrying the decode settings its behaviour was measured at, so the first command you type is the tuned one. Swept across conversation, factual, tool-shaped and long-form prompts on the shipping containers. Here is what each setting is tuned for, if you want to reach for a different one.
- Neutrino-1 8B, conversationThe settings the assistant behaviour was measured at.
- temperature 0 · penalty 1.05 · 512 tokens
- Neutrino-1 8B, tool calls and JSONReturns valid, fence-free JSON on tool-shaped prompts at every temperature we swept, up to 0.7. Plain decoding keeps the calls reproducible.
- temperature 0 · no penalty · 256 tokens
- Neutrino-1 8B, long-form proseSampling buys variety here, and the budget matches how much the model actually writes.
- temperature 0.7 · top-p 0.95 · 1024 tokens
- Neutrino-1 8B, scripted and reproducibleWhat the pip engine does by default, so a script gives you the same tokens every run.
- temperature 0 · no penalty
- Neutrino-1 0.6B Chat, conversationThe settings the conversational record on this page was measured at, on all three runtimes.
- temperature 0 · penalty 1.05 · 256 tokens
- Neutrino-1 0.6B, drafting for the 8BNothing to tune. The draft sets the pace; the 8B decides every token that reaches you.
- inherits the 8B
Run it
Two commands to a streaming chat.
$ pip install fermion-research$ fermion chat
The first run pulls the container and the native binary for the host, then the prompt opens. Later runs load from cache.
- A local OpenAI-compatible server
fermion serveServes at http://127.0.0.1:8000/v1. Point any OpenAI client at that base URL; the key can be any string.
- Any of the three models
fermion chat --model fermionresearch/Neutrino-0.6B-Chat--model takes a repository id or a path on disk. --backend native pins the compiled runtime instead of the torch path.
- GGUF, through our llama.cpp fork
llama-completion -m neutrino-8b-fv5.gguf -ngl 99Build the fork, then run the pack from the model repository. The CUDA build offloads all 36 layers.
- MLX, on Apple silicon
python -m fermion_mlx --model neutrino-8b_v4.bin --mode chat --tokenizer .Run from the mlx/ folder of the model repository, after installing its requirements file.
Get the models
Weights, native binaries, GGUF pack, MLX pack.
The draft model, and a fast model on its own.
The conversational small model.
The engine, the loader, and the chat front end.
Adds the format to the llama.cpp toolchain.

