Nvidia’s new DNA model learns what token prediction misses
The New Stack 1 month ago 48
Nvidia released JEPA-DNA, a genomic foundation model that combines traditional masked language modeling with latent-space prediction to learn DNA sequence representations beyond simple token reconstruction. The model builds on DNABERT-2's 117 million parameters by adding a second learning objective that predicts functional representations of masked segments rather than their literal character sequences. This hybrid approach enables models to capture broader biological patterns and functional meaning that token-only prediction misses, pointing toward architectures combining multiple learning objectives for complex domains.