Google Research has introduced a new federated learning system that moves computation from devices to servers running trusted execution environments, or TEEs. The change enables external verification and auditing of data anonymization. Gboard has already deployed the system for English and Japanese next-word prediction models, which Google says are more accurate.
Gboard training depends less on device availability
The change affects how the keyboard’s models are trained. Google says training each of these models previously could take one to two months. Progress was constrained by the availability of participating devices, their computing power and competition between multiple training workloads for the same resources.
The new system significantly speeds up training by distributing computation across many servers, according to Google. Availability of TEE resources is now the limiting factor. The company also collects all device uploads before starting server-side training, so daily fluctuations in device availability no longer affect training progress.
An approved policy governs data access
Google describes a process in which devices first encrypt training examples locally, then upload them to a server. Before processing, the devices authorize an access policy specifying which TEE computations may use the data. Those computations may release only anonymized results, and devices require the policies to be published in a public transparency log.
A Key Management System, or KMS, controls access to decryption keys. Google describes it as a cluster of TEEs using the RAFT consensus protocol. Keys are issued only to server-side workloads running in TEEs whose computations match those permitted by the access policy.
Python coordinates parallel computation and recovery
According to Google, a coordinating TEE runs a training loop written in Python. It assigns tasks that can run in parallel to a group of worker TEEs. Distributed logic is expressed in the open-source Federated Language, derived from TensorFlow Federated; the training loop periodically supplies anonymized model weights to the analyst.
To handle failures in the coordinating or worker environments, the program saves recovery state at the end of a training round. That state is encrypted through the KMS. Google says this process is designed to let training resume after a failure without exposing additional sensitive information.
Auditors can inspect policies and reproduce software builds
Google says workload operators can see only metrics and differentially private model weights. Uploaded training examples can be decrypted and processed only inside TEEs running Python programs authorized by the access policy. Processing is also limited to a restricted period after upload.
For external scrutiny, the company publishes access policies in the public Rekor transparency log. Google says auditors can track the complete set of server-side workloads in which devices could participate. It also says the KMS and data-processing binaries can be reproducibly built from open-source code in the Confidential Federated Compute repository on GitHub.
Google says TEEs allow remote verification of the logic being executed, protect its internal state from observation and preserve computational integrity. The company qualifies these properties as subject to the limitations of current-generation TEEs.
Proprietary model components can be loaded at runtime
Google says auditability can be maintained even when it keeps model architectures or data-preprocessing methods proprietary. The data-processing environments therefore support loading serialized information into the Python program at runtime. Google makes externally verifiable privacy protection conditional on all privacy-relevant logic remaining hardcoded in the program.
Better participation schedules and a path to larger models
During server-side processing, Google says the program can dynamically calculate an optimal device participation schedule and use it to adjust other differential privacy parameters. The company compared the new and previous systems by training an English next-word prediction model for 5,000 rounds with cohorts of 6,500 devices in both systems. According to its research comparison, optimizing participation enabled stronger privacy guarantees, smaller multipliers for added noise, or both.
Client gradient computation has also moved to the server. Google says this removes constraints imposed by device computing power and opens a path to training larger models through federated learning. It considers integration between TEEs and accelerators important for those future uses.



