New Features for Managed Service for Apache Spark Clusters
Key point
Google Cloud unveiled new features for Managed Spark clusters, including Lightning Engine, which maximizes performance, and Flexible VM, which increases resource availability.
Details
Google Cloud has reorganized its Dataproc service into Managed Service for Apache Spark to efficiently handle large-scale analytics and data science workloads, while strengthening integration with Agentic Data Cloud. This service is offered in a serverless mode that requires no infrastructure management and a managed clusters mode that allows for fine-grained customization.
Managed Spark clusters have been redesigned around three core pillars: faster, which increases execution speed; easier, which reduces operational overhead; and smarter, which integrates AI.
The most notable update is the introduction of Lightning Engine. Built on a C++ vectorized execution engine based on Velox and Gluten, it eliminates JVM bottlenecks and optimizes SIMD vectorization. This delivers up to 4.9x faster performance compared to existing open-source Spark and up to 2x the cost-performance compared to competitors, and can be applied immediately without any changes to existing code.
Additionally, the Flexible VMs feature has been made generally available to make resource acquisition easier. Users can define up to 10 prioritized machine types, and the system automatically scans available resources within the region to deploy the optimal hardware layout. This prevents cluster creation failures due to insufficient resources and maximizes Spot VM utilization.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.