Cloudflare introduces on-demand CPU and memory profiling with flamegraphs for Workers and Durable Objects
Key point
Developers can now request interactive flamegraphs for active Workers and Durable Objects via the CLI or Cloudflare Dashboard to identify resource bottlenecks in production.
Details
Cloudflare has announced support for on-demand CPU and memory profiling for Workers and Durable Objects. This feature allows developers to request profiles of active production instances directly from the Workers Observability page or via the CLI, visualizing the results as interactive flamegraphs to pinpoint exact functions consuming resources.
How It Works
Profiling is initiated through the Cloudflare Dashboard or the cf CLI. Users select a specific Worker version and duration, then click Capture Profile. The resulting flamegraph displays function calls as rectangles, where width corresponds to CPU time or memory usage. For TypeScript projects, enabling source maps is recommended to avoid obfuscated function names.
The underlying architecture addresses the distributed nature of Workers. Since requests are routed to the nearest data center and Workers are replicated across multiple metals, the system must first identify the correct isolate. For Durable Objects, which are stateful and named, the runtime routes the request to the specific metal owning the actor. The profiler holds the isolate lock only during lifecycle operations, allowing normal request execution to continue during sampling.
Real-World Impact
Cloudflare teams have used this tool to optimize performance and resolve memory issues. In one case, CPU profiling of the R2 binding Worker revealed that genericR2JsonReplacer was recursively walking the JSON tree unnecessarily. Fixing this duplicate work made the function 2.7x faster. Another optimization involved caching a heavy metrics call that accounted for 1% of CPU time.
For memory issues, a Worker experiencing frequent Exceeded Memory errors at P999 (133 MB against a 128 MB limit) was profiled using pprof. The analysis showed that partially disabled Prometheus code accounted for 66.7% of allocations. Removing this code path reduced P999 memory usage to 118 MB, providing necessary headroom.
Limitations and Future Work
Current limitations include the requirement for explicit session initiation, which may miss transient issues, and the inability to capture allocations occurring during Worker start-up. Cloudflare is developing continuous profiling to automatically capture samples, addressing these gaps and enabling easier exploration of rare events.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.