AI Briefing
KO

GPT-5.3-Codex-Spark Unveiled

·2026.02.12 19:00

Key point

OpenAI, in partnership with Cerebras, has unveiled GPT-5.3-Codex-Spark, a real-time coding model that generates over 1,000 tokens per second.

Details

OpenAI has released a research preview of GPT-5.3-Codex-Spark, a small model optimized for real-time coding. Developed through a partnership with Cerebras, this model generates over 1,000 tokens per second, delivering near-real-time response speed.

Unlike existing models that perform large-scale tasks, Codex-Spark specializes in immediate interactions such as real-time code editing and logic restructuring. It supports a 128k context window and has demonstrated strong performance on the SWE-Bench Pro and Terminal-Bench 2.0 benchmarks.

In addition, the following technical improvements were introduced to reduce latency across the entire request-response pipeline:

  • Introduction of WebSocket connections and optimization of the Responses API
  • 80% reduction in client/server round-trip overhead
  • 30% reduction in per-token overhead and 50% reduction in time to first token (TTFT)

This model runs on Cerebras's Wafer Scale Engine 3 accelerator. While GPUs are cost-effective for broad usage, Cerebras complements this by handling workflows that require extremely low latency, achieving optimal performance.

It is currently available as a research preview for ChatGPT Pro users in the Codex app, CLI, and VS Code extension.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.