AI Briefing
KO

Launch HN: Tokenless (YC S26) – Automatic Model Switching to Cut Costs

·2026.07.30 00:55

Key point

A router service has launched that reduces inference costs by running multiple models simultaneously and quickly selecting the optimal one.

Details

Tokenless is a drop-in replacement model router designed to cut costs on API calls.

Here's how it works:

  • When a request comes in, it fans out the request to multiple models simultaneously (Fan-out).
  • It monitors each model's reasoning process, and as soon as it determines that a particular model is close to the correct answer, it immediately selects that model and stops execution of the remaining models.
  • This allows users to pay only for the tokens they actually need, maintaining quality without being forced to always use a frontier model.

Key features:

  • Provides OpenAI- and Anthropic-compatible endpoints, minimizing changes to existing code.
  • Rather than relying on simple marketing claims, it measures and provides solve rate and cost per task based on a public Agentic Benchmark as its performance metric.
  • Users can simulate expected savings based on their monthly LLM spend.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.