ByteDance preparing real-time spatial video model led by Zhang Yiming (4-minute clip)
Key point
ByteDance is set to launch a cloud-rendering-based real-time spatial video model next month under the supervision of Zhang Yiming.
Details
ByteDance is preparing a real-time spatial video generation AI model under the direct supervision of founder Zhang Yiming. According to Bloomberg, the model may be released next month and aims to generate interactive virtual worlds for live streams, short-form dramas, and games.
Technical Specifications and Cloud Rendering Strategy
The model adopts cloud rendering rather than on-device processing to reduce computational load on headsets and lower hardware costs. Technically, it implements on-demand rendering at approximately 20 frames per second with low latency of around 0.05 seconds (50ms). This strategy shifts the fidelity limitations of existing XR hardware to the model and cloud capacity.
'World Models' Priority and Investment Scale
ByteDance has designated 'World Models' as its top priority among its 2026 AI priorities. The data budget allocated to this field is the largest within the company, amounting to tens of millions to hundreds of millions of yuan, which is 3 to 4 times the spending of competitors. Additionally, the company is proceeding with large-scale investments, having secured a $30 billion loan last week and reviewing capital expenditures (capex) of up to $70 billion.
Market Impact and Competitive Landscape
Based on the technical capabilities of its existing video generation system Seedance, ByteDance is expected to be benchmarked against Google's Genie. Current internal test results show it is approximately 10% behind the latest global technologies, but this launch is an attempt to close that gap. While the cloud-based model does not surpass the fidelity of the Apple Vision Pro, it is evaluated as an alternative solution to the mainstreaming failures faced by Meta and Apple, with the potential to shift the XR competitive landscape from hardware to software and infrastructure.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.