Qwen3.8 27B delivers capable local AI work on Apple silicon, but at a speed cost

Qwen3.8 27B can handle everyday local AI tasks such as document classification, feed summaries, and long discussion recaps, according to benchmarks from a Mac Studio. However, its new architecture currently produces text much more slowly than its predecessor on Apple hardware.

Hemant Kumar writes for TerminalBytes that the 27.3-billion-parameter model generated an average of 14 tokens per second on a Mac Studio with an M3 Ultra chip and 256GB of unified memory. Under the same conditions, Qwen3.6 27B reached 28.6 tokens per second.

The lower speed did not necessarily mean slower completed work. In the tests, Qwen3.8 used about 890 to 1,090 tokens for responses that required Qwen3.6 to generate 1,950 to 3,340 tokens. The reported example put total response time at 67 seconds for Qwen3.8, compared with 72 seconds for the older model.

Smaller versions expand access

The report also tested a 1-bit compressed version of Qwen3.8. It occupied 6.7GB and generated 27.2 tokens per second, nearly twice the speed of the standard 17GB Q4 model. It retained basic factual knowledge, the report says, but struggled to settle on an answer in practical tasks such as writing a Bash command.

That limitation matters for people using AI in workflows. Unsloth, which publishes the compressed model files, advises against 1-bit versions for tool use or autonomous tasks. Its 9.8GB Q2 version is presented as the minimum suitable option for those applications.

The hardware threshold is relatively modest for compressed models. A system with 16GB of RAM can run the 1-bit or 2-bit editions, while the more capable Q4 version needs 32GB. The full BF16 model requires at least 96GB.

The article also warns that users need recent software. Older llama.cpp builds may reject Qwen3.8 because they do not recognize its “qwen35” architecture. For local AI users, the results illustrate a broader tradeoff: model compression can make advanced systems accessible, but quality losses can affect reliability more than factual recall.

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×