AI Briefing
KOSign in

Google Applies Externally Verifiable Differential Privacy to Gboard Model Training via TEE-Based Federated Learning

·2026.10.11 13:30

Key point

Shortens Gboard training period from 2 months to 3 weeks and improves privacy budget by 3x

1 / 5

Details

Google Research announced a Trusted Execution Environment (TEE)-based Federated Learning (FL) system, revealing a case of applying Differential Privacy (DP) to the training of Gboard's English and Japanese next-word prediction models. This system decrypts device data only within server TEEs for training and provides 'externally verifiable central DP guarantees,' allowing external third parties to verify whether DP noise has been applied via code.

Core Mechanism and Architecture

  • Chain of Trust: Establishes a chain of trust from devices to the Key Management System (KMS), Root TEE, and Worker TEE. Devices encrypt and upload data only if the access policy is registered in the public transparency log (Rekor).
  • Workload Execution: The Root TEE executes the Python program (trusted_program), distributing subtasks to the Worker TEE cluster. An end-to-end encrypted channel is established between the Root and Worker TEEs using the Noise protocol.
  • Verification: The KMS and data processing binaries can be built reproducibly (Reproducible Build) from open-source code, ensuring transparency.

Gboard Experimental Results and Improvements

  • Training Efficiency: The existing system relied on participating device availability, taking 1-2 months for training. The new system collects uploaded data first and then optimizes on the server, shortening the training period to 3 weeks.
  • Privacy Budget: A/B test results showed that the zCDP (differential privacy budget metric) for the TEE experimental group was 0.215, approximately 3x smaller than the 0.641 of the existing production model, providing stronger privacy protection.
  • Usability: Core usability metrics such as Words Per Minute and Words Modified Ratio showed neutral results, with no difference from the existing model.

Limitations and Future Challenges

  • TEE Side Channels: Current-generation TEEs may be vulnerable to side-channel analysis, which is currently considered outside the threat model and left to developers to address.
  • Model Scale: Models trained to date are limited to 10M parameters or fewer; training larger models requires addressing GPU usage and communication bottlenecks.
  • Verification Scope: The Python program and binaries are published, but side-loaded private logic or parameters are excluded from verification.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.