AI Briefing
KOSign in

Project Maya enables local execution of GLM-5.3-Flash on single GPUs with up to 118 tokens/s

·2026.10.12 03:42

Key point

The open-source Project Maya allows running the 321-billion-parameter GLM-5.3-Flash model on consumer hardware like a single RTX 5090 or four RTX 4090s.

Details

Project Maya is an open-source tool that enables the local execution of GLM-5.3-Flash, a 321-billion-parameter model developed by zai-org under an MIT license. The tool is designed to run this large model on consumer-grade hardware, supporting configurations from a single NVIDIA GPU up to 16 GPUs, with experimental support for AMD and Windows.

Performance and Hardware Requirements

The engine optimizes memory usage by keeping frequently used experts on the GPU, less frequent ones in RAM, and the remainder on NVMe SSDs, dynamically moving them during conversation. Performance benchmarks reported include:

  • Up to 118 tokens/s on four RTX 4090s.
  • 30+ tokens/s on a single RTX 5090.
  • The model has approximately 18 billion active parameters per token and supports a context window of up to 1 million tokens.

Accuracy and Integration

Maya uses custom quantization techniques that retain 97.7-99.2% of the FP8 model's zero-shot accuracy. It provides an OpenAI- and Anthropic-compatible API, allowing integration with existing coding agents and tools. The package includes a browser-based chat interface, image support, and thinking level controls, with a self-tuning mechanism that adjusts to the user's specific GPU, CPU, RAM, and SSD configuration.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.