Convai Unveils 'Laya', a Decision Model Without Text Generation
Key point
A Jev-compatible open-weight model that returns probability distributions in a single forward pass to address LLM hallucination and latency issues.
Details
Convai Innovations has released Laya, an open-weight decision model that returns probability distributions in a single forward pass without a text generation process. This model is designed to resolve token generation latency, parsing errors, and hallucination issues in LLM-based agents, and is compatible with the API format of Jev (TypeSafe AI).
Architecture and Performance
Laya consists of a ModernBERT-large (English) or mmBERT-base (multilingual) encoder and a 2-layer transformer decision head. Depending on the question type, it returns choice, score, or true/false (noul) probabilities. On a Tesla T4, it processes a single question in 32.8ms and a batch of 10 questions in 72.3ms. Training applied the RLCD (Reinforcement Learning for Calibrated Decisions) method to enhance confidence calibration.
Key Checkpoints and Limitations
- laya: English-only, 421M parameters, context length 512
- laya-multilingual: Supports over 100 languages, 322M parameters, approximately 2x faster than the English version
- laya-typed-decisions: Checkpoints fine-tuned for specific business workflows
The base checkpoint's zero-shot performance is close to random guessing, making fine-tuning essential. Additionally, there is a limitation where accuracy drops sharply when the number of choices exceeds 20. On the Banking77 benchmark, it showed lower performance compared to Jev. Conversely, it recorded higher accuracy than Jev on specific tasks such as AG News and DAIR Emotion.
Local Execution Ecosystem
Laya is released under the Apache-2.0 license, allowing for commercial use, and comes with various local execution tools.
- Laya-MLX: MLX port for Apple Silicon GPUs. P50 7.39ms for multilingual questions on M3 Max
- Laya-CoreML: Optimized for Apple Neural Engine (ANE). P50 4.98ms for multilingual questions on M3 Max
- Ollaya: Rust-based single binary. Uses ONNX Runtime and takes 8~10ms to process 5 questions on RTX 4090. Supports other decision models besides Laya, such as winnow and decider.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.