VaultGemma: The world's most capable differentially private LLM
Google DeepMind
Google DeepMind released VaultGemma, a 1-billion-parameter language model trained with differential privacy, accompanied by new scaling laws that describe how privacy, compute, and data budgets interact during training. The model was trained with a sequence-level privacy guarantee of (ε ≤ 2.0, δ ≤ 1.1e-10) and showed no detectable memorization of training data in empirical tests. VaultGemma performs comparably to non-private models from approximately five years ago, establishing a baseline for measuring progress as privacy-preserving training methods improve.
Why it matters
We introduce VaultGemma, the most capable model trained from scratch with differential privacy.