Free LLM API Updates
·2026.04.19 11:37
Key point
A roundup of free LLM APIs and limits from Cohere, Gemini, Mistral, and Z.AI.
Details
Updated the Free LLM API list, collecting only permanent free tiers.
-
Cohere: Command A 111B, Command R+, Command R, Command R7B.
- 256K/128K context, 4K output, 20 RPM.
- Embed 4 offers text+image embedding, 2,000 inputs/min.
-
Google Gemini: Gemini 2.5 Flash, Flash-Lite.
- 1M context, 65K output, supports text, image, audio, and video.
- 10 RPM / 250 RPD and 15 RPM / 1,000 RPD respectively.
-
Mistral AI: Small 4, Medium 3, Large 3, Nemo 12B, Codestral.
- Up to 256K context and output, generally ~1 RPS, 500K TPM.
-
Z.AI: GLM-4.7-Flash, GLM-4.5-Flash.
- 200K context, 128K output, limited to 1 concurrent request.
Focusing on always-free plans rather than trial credit, the list makes it easy to compare model names and limits at a glance.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.