Google Moves Gboard AI Training to Protected Servers With Auditable Privacy Rules
Encrypted training examples now leave devices, but approved server workloads control their use. Google reports faster training and better accuracy; hardware limitations remain.
In an October 2 announcement, Google described a server-based approach to federated training now used for Gboard’s English and Japanese next-word prediction models. Instead of relying on phones to compute updates, the system processes locally encrypted examples inside attested Trusted Execution Environments, under access rules recorded in a public log. Google says this speeds training and improves accuracy, but reports no new end-to-end training time; progress now depends on available TEE capacity. The design makes privacy-relevant training logic more inspectable, while leaving side-channel risks and full implementation proofs unresolved.
01
Devices authorize specific server workloads before upload; a TEE-based key-management system releases decryption keys only to workloads that match those policies.
02
Google says operators see metrics and differentially private model weights, not training examples; uploaded data is processed only for a limited period.
03
The Python training logic is published, but proprietary model architecture and preprocessing can be loaded at runtime if privacy-relevant logic remains fixed.
Google’s new approach to private AI training sends encrypted examples off users’ devices—not just computation results. In an October 2 announcement, Google Research said protected server environments make that processing externally verifiable. The system already trains Gboard’s English and Japanese next-word prediction models, with Google reporting faster computation, better accuracy and stronger privacy guarantees.
The change puts a different question at the center of federated learning, a way for devices with private data to collaborate on a shared model. Rather than relying on phones to calculate training updates, Google now uses server hardware designed to restrict who can inspect or alter the work.
The keys follow the approved program
Devices encrypt training examples locally before uploading them. Each device pre-authorizes an access policy: the set of server computations allowed to process its data. Those policies must appear in Rekor, a public transparency log that outside auditors can inspect.
The server protection comes from Trusted Execution Environments, or TEEs. These provide remote attestation, letting another party verify the code being executed. They also protect the program’s internal state and execution from observation or interference, subject to the limitations of current TEE hardware.
A key-management system running inside a cluster of TEEs releases decryption keys only to workloads that match the access policy. Google says uploaded examples can be decrypted and processed only inside these approved environments, and only for a limited period after upload.
Privacy rules outsiders can inspect
The intended advance is not encryption alone. Google says earlier systems required trusting it to add random noise correctly to combined training updates. That noise supplies differential privacy, an anonymization protection for released models. The new public policies directly describe the Python program responsible for the training logic.
Google also says the key-management and data-processing binaries can be reproducibly built from open-source code in Confidential Federated Compute. Its Federated Language software, which organizes distributed work across the protected machines, is open source too. Together, the published programs and policies give auditors visibility into what may process uploaded data.
Training no longer waits on phone computation
Previous models could take one to two months to train, Google says. Progress depended on device availability, on-device computing power and competing training jobs seeking the same resources. Collecting uploads before server execution separates the training schedule from daily swings in phone availability.
A coordinating TEE runs the Python training loop and delegates parallel work to worker TEEs. Google reports substantial training speedups, but gives no new end-to-end training time. The bottleneck has moved: progress is now limited by available TEE resources.
Scheduling also affects privacy and accuracy. The underlying research paper says integrating collected data on a schedule optimized for differential privacy improves device coverage and the privacy-quality tradeoff. Its authors report better Gboard accuracy under smaller privacy budgets than the previous system. These are Google’s research results, not independent measurements.
Auditable does not mean every component is public
Google allows proprietary model architecture and data-preprocessing information to be loaded into the program at runtime. It says auditability is preserved as long as all privacy-relevant logic remains fixed in the published Python program. That makes the privacy controls inspectable without requiring every model detail to be open.
The protection still has boundaries. Google acknowledges side-channel observations—information exposed indirectly by execution—and expects future hardware and mitigation research to strengthen protection against malicious server-side attacks. Full proofs that the privacy algorithms and system software are implemented correctly remain an aspiration, not a completed feature.
Moving computation off devices also opens a path to larger federated models, Google says, with protected accelerators expected to play an important role. The company is experimenting with other Python workloads, including synthetic-data generation. Those are directions for the infrastructure; the concrete deployment described now is Gboard’s next-word prediction.
The diagram shows encrypted messages flowing from devices to storage and TEE-hosted processing, with access policies and a transparency log used in the verification flow.Source: research.google.
Sources
arxiv.orgToward provably private learning from federated data
research.googleToward provably private learning from federated data
Reader comments
Newest comments first. Replies stay oldest first.