AI Briefing
KO

Ling-3.0-tiny Released

·2026.08.11 02:11

Key point

inclusionAI has released an 8B MoE model with 1.3B active parameters.

Details

inclusionAI has released Ling-3.0-tiny on Hugging Face. It is a small MoE model with 8B total parameters and 1.3B active parameters during inference.

According to the model card, performance under FP8 is as follows:

  • DGX Spark: approximately 100~105 tokens/s
  • M4 Pro MacBook: approximately 86~90 tokens/s
  • Peak memory usage at 8K context: approximately 8.34 GiB

Ling-3.0-tiny targets high inference throughput with fewer active parameters than large-scale models, demonstrating its potential for local and edge environments.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.