AI Briefing
KO

Analysis of LLM Inference Performance on Apple Silicon

·2026.05.05 22:43

Key point

Benchmark analysis results on LLM inference performance, power efficiency, and MoE model running performance across Apple Silicon chipsets have been released.

Details

Memory Bandwidth and Inference Speed: The M1 Ultra (800 GB/s) recorded 36.5 tok/s on large models, while the M4 Pro (273 GB/s) recorded 45.9 tok/s. The M3 Max (400 GB/s) showed a middle value of 155 tok/s, demonstrating that the combination of bandwidth and memory capacity is key to performance.

Power Efficiency: The M4 proved to be the most efficient chip at 0.34 W/tok. The A18 Pro had the highest efficiency at 0.19 W/tok, but showed low throughput due to its 8GB RAM limitation.

Accessibility of MoE (Mixture-of-Experts) Models: In a 128GB memory environment, large-scale MoE models such as GPT-OSS-120B (74.3 tok/s) and Qwen3.5-122B (59.5 tok/s) were confirmed to run smoothly. This suggests a high potential for utilizing large models on Mac lineups equipped with high-capacity RAM.

Notable Chip Performance Highlights:

  • M5 Max: Demonstrated strong performance by recording the highest average speed in the dataset at 173 tok/s.
  • A18 Pro: Showed its strength as an edge device by recording an overwhelming Prefill (input processing) speed of 18,600 tok/s on the Llama-3.2-3B model.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.