AI Briefing
Sign in

MVNT's Journey to Building AWS-Based Dance Generation AI Infrastructure: GPU Inference Separation and Serving Optimization

·2026.09.29 14:04

Key point

Switching to g7e.2xlarge instances improved generation speed by 3-5x and reduced new user wait times by 57.5%

1 / 6

Details

MVNT develops dance generation AI models and built AWS-based infrastructure to secure choreography-specific data that is difficult to handle with general-purpose models and to serve global users reliably. The core lies in a pipeline that converts large volumes of unstructured dance videos into training data and an architecture that decouples GPU inference asynchronously.

Data Pipeline: Converting Unstructured Videos into Training Data

MVNT collects data using a hybrid approach combining motion capture (accuracy) and video analysis (volume). When original videos are uploaded to Amazon S3, an event is triggered, and requests are buffered through Amazon SQS queues. Human Mesh Recovery (HMR) inference, which requires GPUs, is processed asynchronously by Amazon EC2 workers, with errors logged and reprocessed via Amazon SNS and DynamoDB upon failure. Through this process, videos are converted into structured training datasets including poses, motion, and meshes.

Inference Serving Architecture: Multi-Tier Structure and Optimization

To overcome the limitations of the initial single-server structure, web, app, and database tiers were separated within Amazon VPC. The web tier receives requests via CPU-based EC2, while the app tier handles model inference using GPU-based EC2 Auto Scaling groups. Amazon Aurora is used in the database tier to distribute read traffic and enhance availability.

Performance Improvement and Scalability Validation

Switching from g5 to g7e.2xlarge GPU instances for in-house diffusion model inference significantly improved generation speed. Generation time for 20-second audio was reduced from 3-5 minutes to approximately 35 seconds, making it 3-5x faster, and the wait time for new users' first results decreased from about 2 minutes to 51 seconds, a 57.5% reduction. Thanks to these optimizations, the infrastructure stably handled traffic from 50 on-site and 2,000 online attendees connecting simultaneously at Unreal Fest Seoul in August 2026, supporting a 3x increase in new sign-ups and a 2x increase in generation requests over one month.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.