Run Qwen Flash Next Q4 on a Mac Mini (360 tks read in / 17.5 tks decode)

Run Qwen Flash Next Q4 on a Mac Mini (360 tks read in / 17.5 tks decode)

Forward

This article is a technical write up of how to run Qwen Flash Next (QFN) Q4, a 125b parameter leading open weight model, on a Mac Mini M5 64GB.

This has become a viable option only recent due to a) quality of open weight models b) major advances in the model itself such as hybrid linear attention and the n-gram lookup table increasing efficiency and c) use of SSD streaming as a means to read experts from disk.

Why use my approach? I made a large number of small efficiency gains and one large one - a form of carousel buffering that increased prompt processing by +30%.

Of further note, the components of QFN that make this possible are currently a preview (hence the next) of the Qwen 4th generation models where we can expect further advances.

To skip the writeup here is the GitHub https://github.com/skeggsguy/Flash-next-ssd

Contents

  1. Summary
  2. Testing methodology
  3. The star - carousel buffering
  4. The runner up - multiple ssd concurrency
  5. The stack and rig
  6. Major learnings
  7. Detailed Results

1. Summary

At time of writing we are in an interesting time for local AI.

Cloud AI has roughly 40-60 tokens per seconds (tk/s) output with high intelligence. In comparison, local AI users are running astronomical speeds at 200-800 tk/s on small models, however are hard blocked in running larger more intelligent models due to memory.

There is a clear gap between the two, and this project along with others are searching for a middle ground between the two extremes.

This project shows that streaming experts directly from SSD on an MOE model (Qwen Flash Next aka QFN), is a viable model within speeds that make sense for many workloads.

The following diagram shows this architecture.

How a 111 GB model runs on a 64 GB Mac The model is 24,576 small specialists ("experts"). They live on the SSDs; only the ones needed right now come into memory. 1 · READING YOUR MESSAGE (prefill: the whole prompt at once) Your messageplus the chat so far,e.g. 30,000 tokens Needs nearly every expertso many tokens that eachlayer uses almost all 512 The carousel (1.4 GiB)a conveyor belt: the SSDs loadthe next layer while GPU works GPU does the mathsbarely waits:the belt is always full Speed limit: whichever is slower, the SSDs or the GPU → 390–520 tokens read per second The model's home: two SSDs the same file on both, experts split 53 / 47, 12.3 GB/s together Internal SSD6.5 GB/s · 53% External SSD (TB5)5.7 GB/s · 47% streams every expert, layer by layer fetches the 1 in 4 that's missing: the GPU waits 2 · WRITING THE ANSWER (decode: a few tokens at a time) A helper guessesa small draft modelguesses up to 5 tokens Pick the expertson each of 48 layers,10 of 512 per token Expert cache (28 GiB RAM)keeps the ~10,000 used lately3 in 4 are already there GPU checks guessesall in one pass,keeps the right ones repeat until the answer is done Speed limit: waiting for missing experts to arrive from the SSDs → about 17 tokens written per second Where the time goes for one written token Milliseconds per token. The orange and grey parts are the GPU standing idle between layers. Tool calls 35 ms · guesses usually right Prose 66 ms helper guessing GPU working waiting for the SSDs handover between layers Prose is slower because more guesses are wrong, and every guess, right or wrong, needs its own experts fetched. Source: study C1 profile, measured before the final fix to the layer handover (+7% writing).

Headline results

First runFinal setupChange
Writing (tokens/s)11.8817.46+47%
Time to first word (median)11.4 s5.5 s−52%
Reading in (tokens/s)189364+93%

Writing (decode) - agentic flows

Chat length so far13–16K16–32K32–64K
Writing (tokens/s)16.817.417.2

Writing (decode) - one off tasks outside normal use

Kind of taskWriting (tokens/s)
Short code execution17.9
Prose16.9
Plans and explanations17.6
Other (maths, data, trivia)14.9

Reading in (prompt processing)

New tokens read inCold (tokens/s)Real (tokens/s)Real: already in context
~0.7K25214012.8K
~1.4K32132812.8K
4K523––
~7K50338326K
~18K46240112.8K
32K430––
100K391––
All 120 real messages~440 (estimate)36412.8K

Here we see reading in represented by the real life test dataset (based on my personal usage of AI before the project) vs an artificial fresh read in of long prose.

This is to demonstrate real life usage with multiple prior turns stored as context weighing down the performance and also to provide a frame of reference for external benchmarks.

The cold read in does align with one use case which is a single read in of a long document or web search result result which does occur frequently in chat.

Key findings and implications

The project showed a successful implementation of SSD streaming for a large model on Mac mini hardware.

In particular this project tested new features the author is not aware of elsewhere which are carousel buffering to optimise reading in by +30% speed and dual ssd cards for +15% speed on both read and write. Refer to sections 3 and 4 for the detailed writeups.

In practice this enables (and has enabled for the author):

  1. Powering a private personal AI assistant, by replacing Claude Sonnet as the AI, which contains sensitive data and also leverages a large number of tools including browser navigation, scratchpad and Google MCP. I note the performance of QFN vibes higher than Claude Sonnet which reflects several benchmarks showing QFN scores at Opus4.6 level.
  2. Enabling long session AI coding runs (8+hrs) during the workday. Following a quick specification write up on the train on the way to work (connecting via Moshi), I will let QFN run all day on a coding task. This has changed my coding practice as this isn’t possible on a normal Anthropic or OpenAI subscription limits. At date of writing I’m considering migrating to API pricing only for Anthropic to do quick planning followed by letting QFN perform the long implementation at 0 incremental cost - this would decrease my cloud AI bill substantially.

Acknowledgements

This project is a fork from the npanj and Marian, and I lifted ideas from Slotstream.

  • Marian Mihailescu (mihailescu2m/llama.cpp): wrote the SSD expert streaming (--moe-stream), the Flash-Next support, the MTP head and union sparse attention.
  • npanj (npanj/llama.cpp): packaged it; this engine forks from npanj's at f507d75f8. Also the prebuilt MTP head (nitinpanj/qwen38-flash-next-v3) and the M5 Pro MacBook Pro reference numbers.
  • Slotstream: runs off a different engine (MLX) but I used the cold-read disk benchmark method slotstream came up with.
  • Ofcourse: Unsloth for the UD-Q4_K_XL quant, Strata for the PC comparison, and the llama.cpp project.
Fresh Worktree

Sign up for the next novel AI experiment

Once a month you'll join us and get deep into the weeds of the latest AI engineering setups.

Get the next article Free · roughly monthly · no spam

2. Testing Methodology


The following outlines the test approach, learning heavily on a test dataset generated from on my own personal AI use to represent real results, not benchmaxxing.

The agentic workflows include agentic runs involving coding, google mcp calls, web searches, file movements, file editing, browser navigation and calendar organisation.

The following outline of the test process is AI generated:

  • Real work, replayed exactly. 120 real assistant turns from my own agent (with tool results stubbed in so a multi-step turn gets the data it originally got), 80 unseen prompts (coding, data, maths, trivia, creative, non-English), and 12 long-writing prompts (4 prose, 4 code, 4 plans).
  • One change at a time, A/B/A. Each test ran the old setting, the new one, then the old one again, each on a fresh server. The new one is compared against both old runs, because the box warms up and a single A-then-B flatters whichever ran second. If the two A runs disagree, the test is void.
  • Cold start. The OS file cache is purged before each test, experts are read past it (F_NOCACHE), and the first request of each run is a discarded warm-up.
  • Zero swap is absolute. A watchdog samples swap, memory pressure and heat every second. Any swap stops the run; any heat level above Nominal voids the turn.
  • Does the output change? Changes meant to be exact must write byte-identical answers on three fixed prompts at temperature 0. Changes that move words are judged pair by pair by a Sonnet judge, with perplexity alongside.
  • Quality: a fixed 20-task exam (10 from the assistant's own work, 10 standard coding tasks with unit tests). It needed 17/20 to pass; every sitting scored 18/20.
  • One model process at a time. Two servers would void each other's timings.
  • Writing is word-weighted. Averaging each reply's own speed favours short replies; on one test that reads +34% where the word-weighted gain is +18.7%.

3. The Star - Carousel Buffering

The main contribution of my approach, carousel buffering, is very simple and speeds up prompt processing (reading-in).

During SSD read in the GPU normally waits 50% of the time, waiting for experts to be loaded into memory.

Carousels buffering pivots the traditional approach and instead creates an entirely seperate loading desk for reading in, and continues to load them whilst GPU is working. Very simply this removes any limitation of the SSD load approach on reading in, the bottleneck is now the GPU itself.

The approach achieves an overall +32% readin speed, with blockers to further speed being the first wave of read in, last wave being an inconsistent size, and the fact short read ins less the 1k tokens don’t benefit from the carousel. Short read ins less than 1k tokens continue to the use the old batch system.

Ideal case long cold read ins have more than double speed vs the batch process as per figure 3.

Where the carousel shines in day to day use is double read in for large context such as web search results, code exploration and code dif reads.

The below figure demonstrates the prior and new approaches:

Before: waves. The SSD and GPU take turns The expert cache doubles as the loading area; each layer arrives in ~5 waves; time = SSD + GPU SSD GPU GPU idle done After: carousel. A separate belt keeps both busy A 1.44 GiB ring buffer (1.25 layers, 4 parts a layer) filled ahead by the SSDs; time = the slower of SSD and GPU SSD GPU done Measured: reading in +32% time to first word −28% output bit-identical SSD loading experts GPU working on them GPU waiting Schematic timing. Measured on 12 real assistant turns, 30 GiB cache, two SSDs (study test RR-room).

4. The runner up - 2 SSD drives for concurrent streaming

Multiple other projects reference experimenting with multiple SSD drives as an outstanding research question. This has likely not been investigated as the cost may not be within appetite for many users, however if proven as a large uplift it may earn its spot in the rig.

I found it did have an uplift of 15% for both read and write. Earlier results showed 40% faster read in however a large chunk of this was subsumed by the more efficient buffering system from section 4. This is outlined in the figure below.

Gain from adding a second SSD before carousel final setup Reading in +40% +15% Writing +14.8% +15.4% Time to first word (shorter) −33% −31% Before: test E-paper, 60 prompts, 32 GiB cache, no drafting. Final: test R-drives, 24 prompts, 28 GiB cache, drafting on. Both A/B/A.

5. The stack and rig

The stack

  • Model: Qwen3.8-Flash-Next
  • Quant used: Unsloth UD-Q4_K_XL, 111.3 GB on disk. Each expert is about 2.7 MB, 24,576 of them in all. Q6 was tested too: it writes 37% slower and reads in 30% slower, so Q4 won.
  • Engine: a fork of llama.cpp (Marian Mihailescu's SSD expert streaming, via npanj's fork), with about 15 study patches.
  • Machine: Mac mini M5 Pro, 64 GB. The model is nearly twice the size of RAM, so most experts are read from SSD as they are needed.

The rig - Mac Mini

There‘s not a lot of optionality in Mac Mini however there are three points of note:

  1. Running an approach like this only practical with a machine with 64Gib of memory. You need to leave about 16Gib of space for the system and other processes leaving you with 48Gib for the model working space.
  2. I purchased the Mac Mini with 512GB internal SSD. This was a mistake as smaller SSD is also lower bandwidth. I recommend purchasing 2TB of internal SSD. This wasn’t tested here however Slotstream measured 17gb+ of streaming on an internal 2TB ssd on MacBook, which if true on Mac Mini would offset some benefits of a second SSD.

Hardware specs and costs

PartSpecMeasured read speed
Mac miniM5 Pro, 18-core CPU, 20-core GPU, 64 GB–
Internal SSDApple 512 GB6.2–6.5 GB/s
External SSDSamsung 990 PRO 2 TB (PCIe 4.0)5.7 GB/s
EnclosureOWC Express 1M2 80G, Thunderbolt 5–
Both SSDs at once–12.3 GB/s

Total cost: $3,950 USD / 5,690 AUD (AUD - $4,499 Mac Mini + $709 SSD + $482 enclosure)

6. Major learnings

This sections includes both practical learnings discovered via dogfooding QFN in real life, as well as unexpected findings unearthed from the data generated from experiments attempted.

Practical learning #1: One lane for coding, seperate lane for chat

Having multiple cache lanes is not realistic within the memory available. To enable quality of life where chat can be used as well as coding I configured a two lane system.

When a chat is received by the server it kicks out the coding run and keeps a five minute clock before the coding session can continue. This enables chat to function during a long coding run. The coding run is kept alive via a keep awake and when it continues it needs to reread all history, which is ok because we are not waiting for it and it’s free.

I also recommend setting in the system prompt for your code tool to only run one subagent at a time, this saves cache where multiple requests from multiple agents need to constantly read in from scratch.

This is setup within the GitHub repo and is configurable.

Practical learning #2: Increase thinking cap in your coding harness

At time of writing the default cap in opencode is ~8k for thinking blocks. At the cap the harness will just stop silently unless you have a hook. Either way QFN commonly thinks for 30k tokens and at times higher. It is very different than other models. I set my default cap to 32k thinking tokens, and maximum of 200 total tokens for the session before compaction and haven’t had any issues since on long (8+ hr) coding runs.

Technical learning #1 MTP drafting works on tool calls and code but not on prose

Using MTP, tool calls and code on my tested dataset saw +68% tk/s however prose saw no change and in many tests got worse due to extra processing time. The end setup used a progressive mtp draft which started at guessing forward once and then increased up to 5 depending on measured confidence of sucessful guesses. Essentially have your cake and eat it too.

Technical learning #2 Open issue with Llama.cpp which changes the output of QFN based on batch sizes

Each test confirms whether the change resulted in a change in output which is not preferred. For a few reasons this is inevitable to slightly occur on deep context runs 100k+ in size however is quite minor.

However when batching is changed on speculative decoding, even at temperature 0, the model's probability for the next token moves 5–7 points when the same text is processed as a 6-token check instead of one token at a time, and up to 12 points when it is processed all at once. Drafting therefore changes output, instead of only speed.

This is a change in output where it shouldn’t occur. It’s untested what the changes to quality are.

On further investigation multiple other projects have also discovered this same issue, and it’s logged with llama cpp.

Technical learning #3 There are multiple indications that local inference engines have significant room still for improvement

The following facts point to expected further benefits associated with running open weight models locally in the future. This is not criticising current work, rather the opposite, the massive speed of improvements has opened up a large number of further research opportunities to improve local inference.

  1. The writing part of the engine still involves the GPU waiting 40% of the time. This can definitely be improved in a similar but different manner to the carousel approach for reading in.
  2. BaseRT achieved 1.9x reading in and 1.7x writing based on a new kernel for the Qwen 35b a3b model, optimised for Mac metal tensors. This is weeks of expensive GPU work however indicates large performance uplifts are available through kernels for other qwen models when tailored for metal.
  3. The pace of repo submissions both on Llama.cp and SSD streaming forks is fast. The project is now deviated quite fairly significantly from Marian’s however that project is very active.
  4. This project has a backlog of 10+ ideas to test.

7. Detailed results

The following is a breakdown of key metrics across various slices to show what did and did not work.

Of note many ideas represented a juggling act where some for example made significant gains on coding and tool use but required additional memory which reduced overall hot experts and resulted in a decline in prose writing.

Similarly union attention allowed reduction in amount of GiB used with minor reduction in speed, allowing more hot experts to be held in memory.

What worked - writing tk/s

ChangeBeforeAfterGain
Prefetch next layer's experts (width 10)9.7510.46+7.3%
Second SSD, experts split 53/4710.7612.35+14.8%
MTP drafting, 5 tokens ahead12.3714.68+18.7%
Two attention switches12.0912.93+7.0%
Carousel buffering13.3612.59−5.7%¹
Lend the carousel back while writing12.7913.49+5.5%
KV checkpoints every 8,192 tokens13.4313.25−1.3%²
Drafting depth tunes itself14.3814.43+0.4%
Sparse-attention index in 512-row slices12.6612.93+2.2%
MTP head only records during prefill11.8312.20+3.1%³
N-gram table on its own 128 MiB shelf12.1412.50+3.0%
Union sparse attention12.6512.38−2.1%⁴
96K-token draft vocabulary15.3415.74+2.6%
Cheaper per-layer stops (4 switches)15.5216.66+7.3%
Whole climb, 120 real turns11.8817.46+47%

¹ Kept for reading in (+32%, −28% time to first word). ² Kept for time to first word (−17%) and reading in (+20%). ³ Also reading in +15%. ⁴ Kept for reading in (+5%) and memory; its perplexity at 32K is level. Also kept, with too small a measured effect to list: block-summary reuse (+0.8%). 

What didn’t work

IdeaResult
Q6 quant instead of Q4writing −37%, reading in −30%
Bigger expert cache (36 GiB)memory brake, then a swap: void
Pinning the most-used expertsat most +0.3 hit-rate points: dropped
Fusing small GPU jobs542 fewer jobs a token, writing +0.3%
Prefetching drafted tokens' experts+1.8%, top-3 variant −0.5%
Prefetching two layers aheadthe run hit swap: void
Deeper drafting rulecode +9%, prose −9%
A full draft cache for the MTP headassistant turns +2.1% only
Skipping fetches for doubtful guessesprose 40–65% slower
Drafting 30 tokens ahead4.0 vs 14.2 tokens/s
Prefill batch 8192reading in +27%, but 4.5 GiB more memory
Reading each expert once per checkat most ~2 ms of a 64 ms token

Quality measure

Quality was measured pre-post the changes based on an exam of long running agentic flows and a pass mark of 17/20 which was the pass rate of Sonnet5 at time of exam.

SetupExam score
Q4, two SSDs, no drafting18/20
Final settings before the draft vocabulary18/20
Final setup18/20
Fresh Worktree

Sign up for the next novel AI experiment

Once a month you'll join us and get deep into the weeds of the latest AI engineering setups.

Get the next article Free · roughly monthly · no spam