NVIDIA releases Nemotron 3 Nano Omni multimodal model with support for text, images, video, and audio
Model release Provisional 95% confidence first seen
NVIDIA released Nemotron 3 Nano Omni, a 30-billion parameter multimodal model capable of processing text, images, video, and audio in a single inference pass with up to 256K tokens of context. The model is now available through multiple platforms including Together AI, enabling developers to build agentic applications without fragmented multi-model pipelines while achieving higher throughput than competing alternatives.