LLM-as-Jev Paper Released, Demonstrating Decision Model Usage Without Structural Changes
Key point
Qwen3.5-4B achieves 81.4% on JevBench without training, while fine-tuning proves effective for small models
Details
The LLM-as-Jev paper (arXiv:2610.02076) proposes a method for using general-purpose LLMs as Jev-style decision models without changing the model structure, tokenizer, or vocabulary. It assigns numbers to choices and prefills the assistant response to use the log probability of candidate suffixes as decision probabilities.
Core Mechanism and Efficiency
- No Structural Changes: Forms a Trie structure without an upper limit on the number of choices to prevent collisions. Verified up to 1,000 choices.
- Parallel Evaluation: Through KV cache reuse and Teacher Forcing, even 151 choices can be processed with a maximum of 3 forward passes.
- Performance: The Qwen3.5-4B model records 81.4% accuracy on JevBench without training, showing no statistical difference from the existing SemIf model. It achieves 69.0% on Banking77 (77 choices).
Conditions and Effects of Fine-Tuning
Fine-tuning is effective only when the base model is weak or for specific tasks (such as multi-choice classification).
- Small Model (Qwen3-0.6B): Performance is poor without training, but fine-tuning significantly improves Banking77 accuracy from 22.2% to 60.5%.
- Large Model (Qwen3.5-4B): Already shows high performance without training, but LoRA fine-tuning raises JevBench accuracy to 84.0%. Full fine-tuning may actually lower the JevBench score.
- KL Anchor: To preserve general capabilities, a KL anchor (λ=1) must be applied to prevent distribution collapse and format deviation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.