AI Briefing
KOSign in

Jane Street Intern Project Explores Autoregressive Diffusion for Generating Market Data

·2026.10.09 23:56

Key point

A Jane Street intern project demonstrated that flow matching provides stable generation of synthetic order book events where DDPM diverged, though fully continuous diffusion struggles with the discrete nature of market microstructure.

1 / 8

Details

Jane Street published findings from a 2026 summer internship project exploring the use of autoregressive diffusion models to synthesize market data events. The research aimed to move beyond point-estimate price predictions to generate full order book streams, including event timing, prices, and types (trades vs. BBO updates).

Model Architecture and Flow Matching

The project utilized an encoder–diffuser architecture inspired by Autoregressive Image Generation without Vector Quantization. A causally masked transformer encoder produced latent embeddings, which fed into a small event-kind head (categorical) and a diffusion head (continuous targets like price and time).

A key technical finding was the instability of DDPM (Denoising Diffusion Probabilistic Models) in this context. DDPM trajectories exploded, with 88-95% of values deviating more than 8 standard deviations from the mean. The researchers switched to flow matching, which interpolates linearly between noise and data, resulting in stable, nearly straight trajectories that could be integrated accurately in fewer steps.

Handling Discrete Market Microstructure

The core challenge was that market data is neither fully continuous nor fully discrete. Features like order arrival times exhibit point masses (spikes) at zero seconds or whole numbers, and prices cluster around bid/ask midpoints. Standard diffusion models struggled with these jagged distributions.

Two approaches were tested:

  • Categorical Fracturing: Breaking BBO changes into 20 classes (e.g., "size only," "ask up") to handle discontinuities explicitly. This improved marginal distributions but required extensive hand-engineering and did not scale well to high-cardinality features.
  • Atom Smoothing: A procedure to smooth sharp spikes in the true distribution and then re-sculpt the probability mass. This method, combined with flow matching, produced more realistic samples, as verified by a discriminator classifier.

Results and Limitations

The model was evaluated on four years of US equities data. While the atom smoothing technique improved the realism of single-step generations, the autoregressive rollouts degraded over time, with synthetic spreads widening compared to real data. The research concluded that a viable generative model for market data must carefully balance continuous diffusion with discrete categorical modeling to capture the blend of continuous values and discrete actions inherent in market microstructure.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.