Experiments · 6 min

Local inference on a Mac: DeepSeek and Qwen side by side

A reproducible comparison of two quantized local models on an Apple M1 Pro with 16 GB of unified memory.

This is a small, concrete measurement from the Zyrabit work: the same question, two local models, one Apple M1 Pro with 16 GB of unified memory.

The question was: “What are the three primary colors?”

The comparison is useful because the answers are not identical. DeepSeek returned the RGB/physics interpretation — red, blue, and green — while Qwen returned the RYB/pigment interpretation — red, blue, and yellow. The difference is a reminder that a benchmark is also about interpretation, not only speed.

Models

DeepSeek-R1-Distill-Llama-8B

  • File: DeepSeek-R1-Distill-Llama-8B-Q4_K_M.gguf
  • Model file size: 4.92 GB
  • Average speed: approximately 20 tokens/s
  • End-to-end response: 3.96 seconds / 3962 ms

Qwen2.5 7B Instruct

  • File: qwen2.5-7b-instruct.gguf
  • Model file size: 4.40 GB
  • SHA256: 2bada8a7450677000f678be90653b85d364de7db25eb5ea54136ada5f3933730
  • Average speed: approximately 40 tokens/s
  • End-to-end response: 2.17 seconds / 2172 ms

The SHA256 was validated at startup. The comparison uses the model files loaded by the local runtime, not symlink aliases.

Response comparison

MetricDeepSeek-R1-Distill-Llama-8BQwen2.5 7B Instruct
End-to-end response3.96 s2.17 s
Average speed~20 tokens/s~40 tokens/s
Model file4.92 GB4.40 GB
AnswerRed, blue, green — RGB/physicsRed, blue, yellow — RYB/pigment

Qwen was almost twice as fast in this test. The practical explanation observed here is that it goes directly to the final answer instead of streaming internal reasoning tokens such as <think> before the visible response.

Memory and Apple Silicon

Apple Silicon uses unified memory shared by the CPU and GPU. For that reason, this report uses unified memory and Metal allocation rather than treating it as discrete VRAM.

Memory metricDeepSeek-R1-Distill-Llama-8BQwen2.5 7B Instruct
Process RSS~25 MB~2.81 GB / 2,814,464 KB
Metal allocation~4.92 GB~4.40 GB
Free system memory after load~14.49 GB~9.73 GB

The RSS number and the Metal allocation are different measurements. They should not be added together as if they were the same pool, and they should not be generalized to every runtime or quantization format.

Conditions and limits

The results correspond to this Apple M1 Pro, 16 GB unified memory, the recorded model builds, local Metal acceleration, the same question, and the same runtime setup. They are not a universal ranking of the models. A longer prompt, cold start, different context length, runtime version, or different quantization can change the result.

The next useful step is to record those variables beside every measurement: operating system, runtime version, model hash, prompt length, generated tokens, cold-start status, Metal usage, and memory before and after loading.

Why this matters for Zyrabit

The point is not that every local model is fast enough for every task. The point is that the boundary becomes visible on accessible hardware: what can run locally, what trade-offs the interface should communicate, and where a person should remain in control.