Pinterest Shifts to 4B LLM for User Journey Inference, Retaining Clustering as Fallback
Key point
Pinterest has transitioned production inference to a fine-tuned 4B LLM for user journeys, retaining the legacy clustering system as a fallback, with initial tests showing a 1.1% lift in email click-through rates.
Details
Pinterest has moved from a multi-stage clustering pipeline to a generative Large Language Model (LLM) system for inferring user journeys. Unlike traditional interest-based clustering, this approach identifies coherent goals spanning weeks or months, such as "Figuring out what’s in style for a summer wedding," rather than fragmented topics like "Summer dresses." The system uses a fine-tuned 4B parameter LLM (specifically a Qwen3 student model) to read time-series activity logs and output a ranked list of journey names in strict JSON format. The legacy clustering system remains in place as a fallback.
Architecture and Distillation
The core of the system is a Qwen3 student model distilled from a frontier teacher model. To balance quality and serving costs for hundreds of millions of users, Pinterest used Supervised Fine-Tuning (SFT) and Low-Rank Adaptation (LoRA). The 4B model was selected as the optimal balance between output quality and throughput, outperforming smaller 0.6B and 1.7B variants in consistency for complex cases. The prompt design emphasizes behavioral vocabulary (search, save, like, click, dislike) and includes negative signals to suppress irrelevant journeys.
Serving and Performance
The system is served using NVIDIA Dynamo and vLLM on L40S GPUs, achieving a median latency of 1.2 seconds and a throughput of approximately 775 requests per second across a cluster of about 100 GPUs. Initial online experiments in the US and Canada showed a 1.1% increase in email click-through rates and a 1.3% increase in push notification open rates compared to the previous clustering system. The new approach also reduces fragmentation, merging related activities into single coherent journeys and filtering out incidental noise like memes.
Future Developments
Pinterest is currently exploring Semantic IDs (SIDs) to compress Pin data into 64-bit integers for more efficient input processing. The team is also integrating offsite conversion signals (such as page visits and checkouts) to support users with sparse on-site activity, though these are treated as secondary to on-site actions. Additionally, they are implementing a user feedback loop using explicit "Yes/No" buttons and implicit engagement signals to directly optimize the model beyond the limits of teacher distillation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.