Claude Code Pro Max 5x plan quota exhausted in just 1.5 hours even with moderate usage
Key point
A cache_read calculation error in the Pro Max 5x plan causes quota to be consumed rapidly.
Details
On the Pro Max 5x (1M context) plan, even ordinary levels of Q&A and development work trigger a token limit exceeded in just 1.5 hours.
The cause is pointed to a bug where cache_read tokens are counted at the full rate (1.0x). Because tokens that should be saved through caching are counted at a much higher value than they actually are, the caching benefit effectively disappears in practice and usage is consumed rapidly.
Key points
- The consumption rate is abnormally fast even on the 1M context plan
- The limit is reached with only moderate levels of Q&A and development work
- The cache_read calculation error is cited as the cause
- The core of the problem is a user experience where the caching benefit seems to be nullified
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.