AI Briefing
KO

Intel Releases OpenVINO 2026.4

·2026.09.17 18:51

Key point

Intel has released OpenVINO 2026.4, adding support for the latest models such as Gemma 4 and Qwen 3.5/3.6, along with Multi-Token Prediction inference acceleration.

Details

Intel has officially released OpenVINO 2026.4, significantly expanding the range of generative AI models supported and improving inference performance. This update focuses on strengthening integration with the latest Gen AI frameworks while minimizing code changes.

Key Model Support and Inference Optimization

  • New Model Support: Supports Gemma-3n on CPU, and Kokoro-82M, Qwen3-VL-4B (with EAGLE3), Muse Glimmer 30B, Qwen 3.8 27B, and Gemma4 12B on CPU/GPU. On NPU, FLUX.2-Klein 4B and Kokoro 82M have been added.
  • Multi-Token Prediction (MTP): OpenVINO GenAI supports MTP speculative decoding for Gemma 4, Qwen 3.5, and Qwen 3.6 on CPU and GPU, increasing throughput and reducing latency without accuracy loss.
  • Tree Drafting Support: Supports EAGLE3's Tree Drafting (Top-K) to improve throughput in VLM pipelines compared to Chain Drafting (Top-1).
  • Xe3 Graphics Optimization: Improves performance for processing long context inputs for Gemma 4 models on Intel Core Ultra Series 3 processors.

Enhanced Developer Tools and Edge Deployment Features

  • NPU Profiling Support: Extends Instrumentation and Tracing Technology (ITT) profiling to NPU, allowing analysis of CPU, GPU, and NPU execution through a consistent toolchain with Intel VTune Profiler.
  • Node.js ASR Pipeline: OpenVINO GenAI supports ASRPipeline in Node.js, enabling JavaScript developers to run speech recognition models like Whisper or Qwen3-ASR with streaming and performance metrics via an API similar to C++/Python.
  • Model Server Improvements: OpenVINO Model Server supports agentic models such as Muse Glimmer 30B and Qwen3.8 27B, and offers a preview feature for idle model management that unloads models when unused to save memory.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.