AI Briefing
Sign in

Benchmark reveals unauthenticated control-plane ports in Ray and vLLM multi-node LLM deployments

·2026.09.29 18:47

Key point

A security benchmark found that default Ray and vLLM configurations on Kubernetes expose critical control-plane ports to cross-namespace attacks, which standard scanners fail to detect.

Details

A security benchmark of multi-node distributed LLM clusters on AWS EKS revealed significant vulnerabilities in the default configurations of Ray, KubeRay, and vLLM. The study found that 15 of 17 active listening sockets were not declared in container manifests, allowing a container in an unrelated namespace to connect to Ray GCS (port 6379) and raylet worker RPCs (ports 10002–10006) over unauthenticated cleartext gRPC. Additionally, a single-GPU vLLM deployment exposed 26 API routes without authentication.

Scanner Blind Spots

Standard security tools, including Trivy, Checkov, Kubescape, and kube-linter, failed to detect this runtime exposure. The report attributes this failure to the scanners' inability to inspect the RayCluster custom resource, leaving a critical gap in visibility for distributed AI infrastructure.

Mitigation Performance Costs

The benchmark evaluated mitigation strategies and their impact on inference performance:

  • Network Policies: Implementing an Ingress default-deny NetworkPolicy successfully blocked all neighbor-pod access to the Ray control plane without adding measurable latency.
  • Traffic Encryption: Encrypting inter-node GPU traffic using WireGuard (via Cilium chained with AWS VPC CNI) secured the workload but introduced performance overhead. This resulted in a 3.4% to 6.5% drop in throughput and a 3.2% to 8.4% increase in latency under load.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.