AI Briefing
KO

TaoLive Post-Training Small Models to Follow Evolving Harnesses

·2026.08.21 09:00

Key point

TaoLive unveiled the HAT technique to train small LLMs for real-time live commerce to adapt to evolving harnesses.

Details

Real-time live commerce digital avatars must respond immediately to product Q&A and changes in business strategy. To achieve this, TaoLive developed the Evolvable Harness, which separates model weights, Skills, Hooks, and system prompts to allow runtime behavior changes without retraining.

However, as harnesses continue to change, small models tend to memorize specific configurations, while large models fail to meet the low latency requirements for real-time processing. To address this, Harness-Aware Training (HAT) was introduced. HAT incorporates harness states into the training distribution, ensuring the model follows the currently provided harness.

HAT consists of three stages.

  • HSA-based SFT: Supervised fine-tuning via harness state augmentation
  • General on-policy distillation: Recovery of general capabilities
  • HSA-based agentic RL: Reinforcement learning in a production-based live room simulator

The 35B small model achieved a score of 94.8 in real live stream QA, surpassing both the base model (80.3) and the strongest general LLM (93.0). It also scored 94.6 in harness-variant QA and maintained an IFEval score of 83.5, demonstrating that general instruction-following capabilities were not compromised. In a single NVIDIA H20 GPU environment, it achieved a P50 latency of 3.407 seconds and a P95 latency of 8.114 seconds, confirming its viability for real-time operations.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.