Google moves federated learning computation from devices to servers

Google moves federated learning computation from devices to servers

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
4. 10. 2026
3 minutes reading · 2 views
Listen to the article
Audio version of the article
Google moves federated learning computation from devices to servers

Google Research has introduced a new federated learning system that moves computation from devices to servers running trusted execution environments, or TEEs. The change enables external verification and auditing of data anonymization. Gboard has already deployed the system for English and Japanese next-word prediction models, which Google says are more accurate.

Gboard training depends less on device availability

The change affects how the keyboard’s models are trained. Google says training each of these models previously could take one to two months. Progress was constrained by the availability of participating devices, their computing power and competition between multiple training workloads for the same resources.

The new system significantly speeds up training by distributing computation across many servers, according to Google. Availability of TEE resources is now the limiting factor. The company also collects all device uploads before starting server-side training, so daily fluctuations in device availability no longer affect training progress.

An approved policy governs data access

Google describes a process in which devices first encrypt training examples locally, then upload them to a server. Before processing, the devices authorize an access policy specifying which TEE computations may use the data. Those computations may release only anonymized results, and devices require the policies to be published in a public transparency log.

A Key Management System, or KMS, controls access to decryption keys. Google describes it as a cluster of TEEs using the RAFT consensus protocol. Keys are issued only to server-side workloads running in TEEs whose computations match those permitted by the access policy.

Python coordinates parallel computation and recovery

According to Google, a coordinating TEE runs a training loop written in Python. It assigns tasks that can run in parallel to a group of worker TEEs. Distributed logic is expressed in the open-source Federated Language, derived from TensorFlow Federated; the training loop periodically supplies anonymized model weights to the analyst.

To handle failures in the coordinating or worker environments, the program saves recovery state at the end of a training round. That state is encrypted through the KMS. Google says this process is designed to let training resume after a failure without exposing additional sensitive information.

Auditors can inspect policies and reproduce software builds

Google says workload operators can see only metrics and differentially private model weights. Uploaded training examples can be decrypted and processed only inside TEEs running Python programs authorized by the access policy. Processing is also limited to a restricted period after upload.

For external scrutiny, the company publishes access policies in the public Rekor transparency log. Google says auditors can track the complete set of server-side workloads in which devices could participate. It also says the KMS and data-processing binaries can be reproducibly built from open-source code in the Confidential Federated Compute repository on GitHub.

Google says TEEs allow remote verification of the logic being executed, protect its internal state from observation and preserve computational integrity. The company qualifies these properties as subject to the limitations of current-generation TEEs.

Proprietary model components can be loaded at runtime

Google says auditability can be maintained even when it keeps model architectures or data-preprocessing methods proprietary. The data-processing environments therefore support loading serialized information into the Python program at runtime. Google makes externally verifiable privacy protection conditional on all privacy-relevant logic remaining hardcoded in the program.

Better participation schedules and a path to larger models

During server-side processing, Google says the program can dynamically calculate an optimal device participation schedule and use it to adjust other differential privacy parameters. The company compared the new and previous systems by training an English next-word prediction model for 5,000 rounds with cohorts of 6,500 devices in both systems. According to its research comparison, optimizing participation enabled stronger privacy guarantees, smaller multipliers for added noise, or both.

Client gradient computation has also moved to the server. Google says this removes constraints imposed by device computing power and opens a path to training larger models through federated learning. It considers integration between TEEs and accelerators important for those future uses.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Suno Speech combines spoken narration and music in a single audio trackSuno Speech combines spoken narration and music in a single audio track
Suno has opened the Speech beta, which turns text, a voice description and musical direction into a single recording of narration and an original score. It is available to all app users, though results can vary.
2 min read
4. 10. 2026
Shopify introduces Canvas for building online stores through AI chatShopify introduces Canvas for building online stores through AI chat
Canvas connects online store creation with the AI agent Sidekick. It displays changes in real time and lets merchants test interactive elements, animations and layouts across different screen sizes.
1 min read
4. 10. 2026
Hugging Face and Liquid AI bring model training to coding agents without changing their codeHugging Face and Liquid AI bring model training to coding agents without changing their code
The open stack enables reinforcement learning inside Claude Code, Codex and OpenCode. In Hugging Face and Liquid AI’s experiment, LFM2.5-2.6B’s success rate across four environments rose from 42.2% to 54.2%.
4 min read
3. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok