AI Briefing
KO

Artificial Analysis Releases Small LLM Inference Benchmark for Mobile Devices

·2026.08.28 04:12

Key point

Artificial Analysis, in collaboration with Liquid AI, released a benchmark measuring the intelligence and inference speed of small LLMs under 8GB on the iPhone 17 Pro.

Details

Artificial Analysis, in collaboration with Liquid AI, released benchmark results measuring the actual inference performance and intelligence of small LLMs running on mobile devices. This evaluation was conducted based on the iPhone 17 Pro, targeting models that fit within 8GB of memory (including KV cache, 8K context) after quantization.

Evaluation Metrics and Methodology

The evaluation used 5 benchmarks reflecting real-world mobile usage environments (BFCL subset, IFBench, AA-Omniscience, GPQA Diamond, MATH-500), with the context limited to 16K tokens. The key metrics are as follows:

  • Model Intelligence: Average score of the 5 evaluations above
  • Inference Performance: End-to-end (E2E) time taken to process a 1024-token prompt and generate a 256-token response
  • Token Efficiency: Percentage of generations stopped due to the 16K token limit

Key Models Evaluated

Major small models were included, such as Gemma 4 31B, LFM2.5-2.6B, Ling 3.0 Tiny, the Ministral 3 series, and the Qwen3.5/3.6 series. Some models were excluded from performance measurement due to device memory limits or timeouts. Artificial Analysis stated that it independently verified Liquid AI's inference measurement process.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.