Run Qwen Flash Next Q4 on a Mac Mini (360 tks read in / 17.5 tks decode)
Forward
This article is a technical write up of how to run Qwen Flash Next (QFN) Q4, a 125b parameter leading open weight model, on a Mac Mini M5 64GB.
This has become a viable option only recent due to a) quality of open weight models b) major advances in the model itself such as hybrid linear attention and the n-gram lookup table increasing efficiency and c) use of SSD streaming as a means to read experts from disk.
Why use my approach? I made a large number of small efficiency gains and one large one - a form of carousel buffering that increased prompt processing by +30%.
Of further note, the components of QFN that make this possible are currently a preview (hence the next) of the Qwen 4th generation models where we can expect further advances.
To skip the writeup here is the GitHub https://github.com/skeggsguy/Flash-next-ssd
Contents
- Summary
- Testing methodology
- The star - carousel buffering
- The runner up - multiple ssd concurrency
- The stack and rig
- Major learnings
- Detailed Results
1. Summary
At time of writing we are in an interesting time for local AI.
Cloud AI has roughly 40-60 tokens per seconds (tk/s) output with high intelligence. In comparison, local AI users are running astronomical speeds at 200-800 tk/s on small models, however are hard blocked in running larger more intelligent models due to memory.
There is a clear gap between the two, and this project along with others are searching for a middle ground between the two extremes.
This project shows that streaming experts directly from SSD on an MOE model (Qwen Flash Next aka QFN), is a viable model within speeds that make sense for many workloads.
The following diagram shows this architecture.
Headline results
| First run | Final setup | Change | |
|---|---|---|---|
| Writing (tokens/s) | 11.88 | 17.46 | +47% |
| Time to first word (median) | 11.4 s | 5.5 s | −52% |
| Reading in (tokens/s) | 189 | 364 | +93% |
Writing (decode) - agentic flows
| Chat length so far | 13–16K | 16–32K | 32–64K |
|---|---|---|---|
| Writing (tokens/s) | 16.8 | 17.4 | 17.2 |
Writing (decode) - one off tasks outside normal use
| Kind of task | Writing (tokens/s) |
|---|---|
| Short code execution | 17.9 |
| Prose | 16.9 |
| Plans and explanations | 17.6 |
| Other (maths, data, trivia) | 14.9 |
Reading in (prompt processing)
| New tokens read in | Cold (tokens/s) | Real (tokens/s) | Real: already in context |
|---|---|---|---|
| ~0.7K | 252 | 140 | 12.8K |
| ~1.4K | 321 | 328 | 12.8K |
| 4K | 523 | – | – |
| ~7K | 503 | 383 | 26K |
| ~18K | 462 | 401 | 12.8K |
| 32K | 430 | – | – |
| 100K | 391 | – | – |
| All 120 real messages | ~440 (estimate) | 364 | 12.8K |
Here we see reading in represented by the real life test dataset (based on my personal usage of AI before the project) vs an artificial fresh read in of long prose.
This is to demonstrate real life usage with multiple prior turns stored as context weighing down the performance and also to provide a frame of reference for external benchmarks.
The cold read in does align with one use case which is a single read in of a long document or web search result result which does occur frequently in chat.
Key findings and implications
The project showed a successful implementation of SSD streaming for a large model on Mac mini hardware.
In particular this project tested new features the author is not aware of elsewhere which are carousel buffering to optimise reading in by +30% speed and dual ssd cards for +15% speed on both read and write. Refer to sections 3 and 4 for the detailed writeups.
In practice this enables (and has enabled for the author):
- Powering a private personal AI assistant, by replacing Claude Sonnet as the AI, which contains sensitive data and also leverages a large number of tools including browser navigation, scratchpad and Google MCP. I note the performance of QFN vibes higher than Claude Sonnet which reflects several benchmarks showing QFN scores at Opus4.6 level.
- Enabling long session AI coding runs (8+hrs) during the workday. Following a quick specification write up on the train on the way to work (connecting via Moshi), I will let QFN run all day on a coding task. This has changed my coding practice as this isn’t possible on a normal Anthropic or OpenAI subscription limits. At date of writing I’m considering migrating to API pricing only for Anthropic to do quick planning followed by letting QFN perform the long implementation at 0 incremental cost - this would decrease my cloud AI bill substantially.
Acknowledgements
This project is a fork from the npanj and Marian, and I lifted ideas from Slotstream.
- Marian Mihailescu (mihailescu2m/llama.cpp): wrote the SSD expert streaming (
--moe-stream), the Flash-Next support, the MTP head and union sparse attention. - npanj (npanj/llama.cpp): packaged it; this engine forks from npanj's at
f507d75f8. Also the prebuilt MTP head (nitinpanj/qwen38-flash-next-v3) and the M5 Pro MacBook Pro reference numbers. - Slotstream: runs off a different engine (MLX) but I used the cold-read disk benchmark method slotstream came up with.
- Ofcourse: Unsloth for the UD-Q4_K_XL quant, Strata for the PC comparison, and the llama.cpp project.
Sign up for the next novel AI experiment
Once a month you'll join us and get deep into the weeds of the latest AI engineering setups.
Get the next article2. Testing Methodology
The following outlines the test approach, learning heavily on a test dataset generated from on my own personal AI use to represent real results, not benchmaxxing.
The agentic workflows include agentic runs involving coding, google mcp calls, web searches, file movements, file editing, browser navigation and calendar organisation.
The following outline of the test process is AI generated:
- Real work, replayed exactly. 120 real assistant turns from my own agent (with tool results stubbed in so a multi-step turn gets the data it originally got), 80 unseen prompts (coding, data, maths, trivia, creative, non-English), and 12 long-writing prompts (4 prose, 4 code, 4 plans).
- One change at a time, A/B/A. Each test ran the old setting, the new one, then the old one again, each on a fresh server. The new one is compared against both old runs, because the box warms up and a single A-then-B flatters whichever ran second. If the two A runs disagree, the test is void.
- Cold start. The OS file cache is purged before each test, experts are read past it (
F_NOCACHE), and the first request of each run is a discarded warm-up. - Zero swap is absolute. A watchdog samples swap, memory pressure and heat every second. Any swap stops the run; any heat level above Nominal voids the turn.
- Does the output change? Changes meant to be exact must write byte-identical answers on three fixed prompts at temperature 0. Changes that move words are judged pair by pair by a Sonnet judge, with perplexity alongside.
- Quality: a fixed 20-task exam (10 from the assistant's own work, 10 standard coding tasks with unit tests). It needed 17/20 to pass; every sitting scored 18/20.
- One model process at a time. Two servers would void each other's timings.
- Writing is word-weighted. Averaging each reply's own speed favours short replies; on one test that reads +34% where the word-weighted gain is +18.7%.
3. The Star - Carousel Buffering
The main contribution of my approach, carousel buffering, is very simple and speeds up prompt processing (reading-in).
During SSD read in the GPU normally waits 50% of the time, waiting for experts to be loaded into memory.
Carousels buffering pivots the traditional approach and instead creates an entirely seperate loading desk for reading in, and continues to load them whilst GPU is working. Very simply this removes any limitation of the SSD load approach on reading in, the bottleneck is now the GPU itself.
The approach achieves an overall +32% readin speed, with blockers to further speed being the first wave of read in, last wave being an inconsistent size, and the fact short read ins less the 1k tokens don’t benefit from the carousel. Short read ins less than 1k tokens continue to the use the old batch system.
Ideal case long cold read ins have more than double speed vs the batch process as per figure 3.
Where the carousel shines in day to day use is double read in for large context such as web search results, code exploration and code dif reads.
The below figure demonstrates the prior and new approaches:
4. The runner up - 2 SSD drives for concurrent streaming
Multiple other projects reference experimenting with multiple SSD drives as an outstanding research question. This has likely not been investigated as the cost may not be within appetite for many users, however if proven as a large uplift it may earn its spot in the rig.
I found it did have an uplift of 15% for both read and write. Earlier results showed 40% faster read in however a large chunk of this was subsumed by the more efficient buffering system from section 4. This is outlined in the figure below.
5. The stack and rig
The stack
- Model: Qwen3.8-Flash-Next
- Quant used: Unsloth UD-Q4_K_XL, 111.3 GB on disk. Each expert is about 2.7 MB, 24,576 of them in all. Q6 was tested too: it writes 37% slower and reads in 30% slower, so Q4 won.
- Engine: a fork of llama.cpp (Marian Mihailescu's SSD expert streaming, via npanj's fork), with about 15 study patches.
- Machine: Mac mini M5 Pro, 64 GB. The model is nearly twice the size of RAM, so most experts are read from SSD as they are needed.
The rig - Mac Mini
There‘s not a lot of optionality in Mac Mini however there are three points of note:
- Running an approach like this only practical with a machine with 64Gib of memory. You need to leave about 16Gib of space for the system and other processes leaving you with 48Gib for the model working space.
- I purchased the Mac Mini with 512GB internal SSD. This was a mistake as smaller SSD is also lower bandwidth. I recommend purchasing 2TB of internal SSD. This wasn’t tested here however Slotstream measured 17gb+ of streaming on an internal 2TB ssd on MacBook, which if true on Mac Mini would offset some benefits of a second SSD.
Hardware specs and costs
| Part | Spec | Measured read speed |
|---|---|---|
| Mac mini | M5 Pro, 18-core CPU, 20-core GPU, 64 GB | – |
| Internal SSD | Apple 512 GB | 6.2–6.5 GB/s |
| External SSD | Samsung 990 PRO 2 TB (PCIe 4.0) | 5.7 GB/s |
| Enclosure | OWC Express 1M2 80G, Thunderbolt 5 | – |
| Both SSDs at once | – | 12.3 GB/s |
Total cost: $3,950 USD / 5,690 AUD (AUD - $4,499 Mac Mini + $709 SSD + $482 enclosure)
6. Major learnings
This sections includes both practical learnings discovered via dogfooding QFN in real life, as well as unexpected findings unearthed from the data generated from experiments attempted.
Practical learning #1: One lane for coding, seperate lane for chat
Having multiple cache lanes is not realistic within the memory available. To enable quality of life where chat can be used as well as coding I configured a two lane system.
When a chat is received by the server it kicks out the coding run and keeps a five minute clock before the coding session can continue. This enables chat to function during a long coding run. The coding run is kept alive via a keep awake and when it continues it needs to reread all history, which is ok because we are not waiting for it and it’s free.
I also recommend setting in the system prompt for your code tool to only run one subagent at a time, this saves cache where multiple requests from multiple agents need to constantly read in from scratch.
This is setup within the GitHub repo and is configurable.
Practical learning #2: Increase thinking cap in your coding harness
At time of writing the default cap in opencode is ~8k for thinking blocks. At the cap the harness will just stop silently unless you have a hook. Either way QFN commonly thinks for 30k tokens and at times higher. It is very different than other models. I set my default cap to 32k thinking tokens, and maximum of 200 total tokens for the session before compaction and haven’t had any issues since on long (8+ hr) coding runs.
Technical learning #1 MTP drafting works on tool calls and code but not on prose
Using MTP, tool calls and code on my tested dataset saw +68% tk/s however prose saw no change and in many tests got worse due to extra processing time. The end setup used a progressive mtp draft which started at guessing forward once and then increased up to 5 depending on measured confidence of sucessful guesses. Essentially have your cake and eat it too.
Technical learning #2 Open issue with Llama.cpp which changes the output of QFN based on batch sizes
Each test confirms whether the change resulted in a change in output which is not preferred. For a few reasons this is inevitable to slightly occur on deep context runs 100k+ in size however is quite minor.
However when batching is changed on speculative decoding, even at temperature 0, the model's probability for the next token moves 5–7 points when the same text is processed as a 6-token check instead of one token at a time, and up to 12 points when it is processed all at once. Drafting therefore changes output, instead of only speed.
This is a change in output where it shouldn’t occur. It’s untested what the changes to quality are.
On further investigation multiple other projects have also discovered this same issue, and it’s logged with llama cpp.
- llama.cpp issue #23302: a Qwen MTP model, Q4_K_M, Metal on an M4 Pro; raising
--spec-draft-n-maxfrom 2 to 3 changed word 38. Open, no cause yet (as of 2026-10-03). - Qwen speculative decoding on an RTX 3090: 76–80% of answers diverged within 400 tokens.
- Strata (same model, on PCs), issue #152: the same effect. Edit: strata has fixed this in their fork.
- vLLM batch invariance ships a fix as an option; background in Simon Willison's note on "Defeating nondeterminism".
Technical learning #3 There are multiple indications that local inference engines have significant room still for improvement
The following facts point to expected further benefits associated with running open weight models locally in the future. This is not criticising current work, rather the opposite, the massive speed of improvements has opened up a large number of further research opportunities to improve local inference.
- The writing part of the engine still involves the GPU waiting 40% of the time. This can definitely be improved in a similar but different manner to the carousel approach for reading in.
- BaseRT achieved 1.9x reading in and 1.7x writing based on a new kernel for the Qwen 35b a3b model, optimised for Mac metal tensors. This is weeks of expensive GPU work however indicates large performance uplifts are available through kernels for other qwen models when tailored for metal.
- The pace of repo submissions both on Llama.cp and SSD streaming forks is fast. The project is now deviated quite fairly significantly from Marian’s however that project is very active.
- This project has a backlog of 10+ ideas to test.
7. Detailed results
The following is a breakdown of key metrics across various slices to show what did and did not work.
Of note many ideas represented a juggling act where some for example made significant gains on coding and tool use but required additional memory which reduced overall hot experts and resulted in a decline in prose writing.
Similarly union attention allowed reduction in amount of GiB used with minor reduction in speed, allowing more hot experts to be held in memory.
What worked - writing tk/s
| Change | Before | After | Gain |
|---|---|---|---|
| Prefetch next layer's experts (width 10) | 9.75 | 10.46 | +7.3% |
| Second SSD, experts split 53/47 | 10.76 | 12.35 | +14.8% |
| MTP drafting, 5 tokens ahead | 12.37 | 14.68 | +18.7% |
| Two attention switches | 12.09 | 12.93 | +7.0% |
| Carousel buffering | 13.36 | 12.59 | −5.7%¹ |
| Lend the carousel back while writing | 12.79 | 13.49 | +5.5% |
| KV checkpoints every 8,192 tokens | 13.43 | 13.25 | −1.3%² |
| Drafting depth tunes itself | 14.38 | 14.43 | +0.4% |
| Sparse-attention index in 512-row slices | 12.66 | 12.93 | +2.2% |
| MTP head only records during prefill | 11.83 | 12.20 | +3.1%³ |
| N-gram table on its own 128 MiB shelf | 12.14 | 12.50 | +3.0% |
| Union sparse attention | 12.65 | 12.38 | −2.1%⁴ |
| 96K-token draft vocabulary | 15.34 | 15.74 | +2.6% |
| Cheaper per-layer stops (4 switches) | 15.52 | 16.66 | +7.3% |
| Whole climb, 120 real turns | 11.88 | 17.46 | +47% |
¹ Kept for reading in (+32%, −28% time to first word). ² Kept for time to first word (−17%) and reading in (+20%). ³ Also reading in +15%. ⁴ Kept for reading in (+5%) and memory; its perplexity at 32K is level. Also kept, with too small a measured effect to list: block-summary reuse (+0.8%).
What didn’t work
| Idea | Result |
|---|---|
| Q6 quant instead of Q4 | writing −37%, reading in −30% |
| Bigger expert cache (36 GiB) | memory brake, then a swap: void |
| Pinning the most-used experts | at most +0.3 hit-rate points: dropped |
| Fusing small GPU jobs | 542 fewer jobs a token, writing +0.3% |
| Prefetching drafted tokens' experts | +1.8%, top-3 variant −0.5% |
| Prefetching two layers ahead | the run hit swap: void |
| Deeper drafting rule | code +9%, prose −9% |
| A full draft cache for the MTP head | assistant turns +2.1% only |
| Skipping fetches for doubtful guesses | prose 40–65% slower |
| Drafting 30 tokens ahead | 4.0 vs 14.2 tokens/s |
| Prefill batch 8192 | reading in +27%, but 4.5 GiB more memory |
| Reading each expert once per check | at most ~2 ms of a 64 ms token |
Quality measure
Quality was measured pre-post the changes based on an exam of long running agentic flows and a pass mark of 17/20 which was the pass rate of Sonnet5 at time of exam.
| Setup | Exam score |
|---|---|
| Q4, two SSDs, no drafting | 18/20 |
| Final settings before the draft vocabulary | 18/20 |
| Final setup | 18/20 |
Sign up for the next novel AI experiment
Once a month you'll join us and get deep into the weeds of the latest AI engineering setups.
Get the next article
Comments ()