Apple Silicon LLM Inference Performance Analysis Results
Key point
Analysis results of LLM inference performance and power efficiency for Apple M4/M5 and A18 Pro chips have been released.
Details
Memory bandwidth has been confirmed as the key driver of LLM inference performance. The M1 Ultra (800 GB/s) and M3 Max (400 GB/s) showed excellent performance based on their high bandwidth.
The M4 series recorded outstanding power efficiency. In particular, the base M4 chip recorded 0.34W per token, showing much higher efficiency compared to the M4 Pro or Max.
Accessibility to large-scale MoE models has improved. On M4 Max and M5 Max devices equipped with 128GB RAM, 120B+ parameter models such as GPT-OSS-120B and Qwen3.5-122B can be run at stable speeds.
Key highlights by chip are as follows:
- M5 Max: Demonstrated strong performance, recording the highest average token generation speed within the dataset.
- A18 Pro: Recorded an overwhelming prefill speed of 18,600 tok/s when testing the Llama-3.2-3B model, showing its strength as an edge device.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.