AI Briefing
KO

Webinar Recap: How Cars24 Automated 3 Million Minutes of Sales Calls with Voice AI

·2026.04.10 14:58

Key point

Cars24 automated 25% of sales calls with multilingual multi-agent Voice AI.

Details

Cars24 operates production voice agents in 13 languages across India, UAE, and Australia, automating over 3 million minutes of sales calls. At a scale of selling 4,000–4,500 cars per month and handling 100,000+ inspections and 22,000+ test drives, phone support became a more critical customer touchpoint than the app.

75% of buyers are first-time car buyers, making decision-making burdensome, and the sales funnel spans 30–45 days. Before adopting AI, customers called the wrong teams, causing longer wait times, missed follow-ups, and mixed operational costs from blending support and sales.

Now, 25% of all calls are automated, nearly half of sales are assisted by voice AI, and call costs have dropped by 50%. They started with simple reminders, phased them into real customer traffic, then expanded to longer conversations and complex handoffs.

In Demo 1, a Hindi-speaking negotiation agent works with a seller, going beyond comparisons with competing platforms to perform real-time rebuttals and persuasion, driving a conversion decision during the call. In Demo 2, Sneha handles rebooking for a customer whose schedule slipped, and upon detecting upgrade intent, hands off context mid-call to a sales agent — even when the customer requests a callback an hour later, the conversation context is preserved to closing.

Technically, they chose multi-agent orchestration to avoid the limits of a single prompt and single agent. A small qualification agent captures initial intent, a larger model handles discovery, and once pricing or financing comes up, a dedicated loan agent takes over. Splitting flows this way means edits to one flow don't break another.

The stack consists of ElevenLabs Scribe v2 real-time STT, Flash 2.5 TTS with V3 expressive mode, GPT-4.1 mini for most use cases, the GPT-4o family for longer discovery, 24kHz PCM WebSockets, Twilio/Exotel/Plivo, HubSpot integration, and Qdrant's binary quantization to reduce latency and cost on large datasets. Compared to manually stitching together separate STT, LLM, and TTS, ElevenLabs Agents' integrated orchestration cut latency by 30–40%, and that reduction translated into a 30–40% increase in conversion.

Their operating principles are clear.

  • Start with missed appointment reminders, the lowest-cost-of-failure use case.
  • Validate with 10% of real customer traffic before relying on simulations.
  • Keep prompts within 4,000–5,000 tokens for calls under 2 minutes; if it grows longer, create a new agent.
  • Put frequently asked answers directly in the prompt rather than a KB, handling only the remaining 20% via tool calls and RAG.
  • Pre-generate the first message and make it uninterruptible to reduce drop-off.
  • Run 100% evals on every call, with manual review on failures.
  • Roll out deployments as 5% → 10% → 20% → 50% → 100%, with a 2-day hold at each stage.
  • No outbound calls without consent; if customers ask whether it's AI, disclose it's a virtual assistant, and immediately hand off to a human if requested.
  • Treat AI costs not as an expense line but as an investment measured by performance.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.