AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Toward provably private learning from federated data

Collected Oct 2, 2026

Google Research announced a new Federated Learning (FL) system that uses Trusted Execution Environments (TEEs) to provide externally verifiable and auditable data anonymization guarantees, with computation shifted to the server to improve training speed, accuracy, and device coverage.

The announcement was made by Katharine Daly, Software Engineer, and Daniel Ramage, Research Director. Google introduced Federated Learning in 2017, and it has been used for features including next-word prediction and Smart Compose on Gboard, reply suggestions in Google Messages, and Smart Text Selection in Android.

The company said its FL development is guided by four privacy principles: data minimization, data anonymization, transparency and control, and verifiability and auditability. It cited differential privacy algorithms including matrix factorization DP-FTRL (MF-DP-FTRL) and distributed differential privacy with Secure Aggregation.

The new system coordinates four operational concepts: client devices locally encrypt training examples and upload them under pre-authorized access policies; a Key Management System of TEEs implementing the RAFT consensus protocol releases decryption keys only to matching server-side TEE workloads; a root data processing TEE runs a Python training loop and delegates subtasks to worker TEEs using Federated Language; and the program saves a KMS-encrypted recovery state for fault tolerance. Access policies are published to Rekor, a public transparency log, and the KMS and data processing binaries can be reproducibly built from open source code in the Confidential Federated Compute Github repository.

Google said only metrics and differentially private model weights are visible to workload operators, and that data processing TEEs support sideloading serialized information into the Python program at runtime to protect proprietary model architectures and preprocessing logic.

Gboard has deployed the system to launch English and Japanese next-word prediction models. Google said training these FL models previously could take one to two months each, limited by device availability and on-device compute, and that bottlenecks have now moved to the server, where parallelization across many machines allows significant speedups currently limited by TEE resource availability. It also said collecting all device uploads before server-side training removes diurnal device-availability constraints, and that the company is experimenting with other workloads such as synthetic data generation.

Read at Google Research

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Mobile Systems