TLDRocket
Sign in

Fine-Tuning Platform Upgrades: Larger Models, Longer Contexts, Enhanced Hugging Face Integrations

Together AI

Together AI just supercharged its fine-tuning platform. Now you can train massive 100B+ models like DeepSeek-V3 and Llama 4 Maverick with way longer context windows.

Based on reporting by Together AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Together AI dropped a hefty update to its Fine-Tuning Platform this week, and the headline is simple: bigger models, longer memory, easier plumbing. The company now supports fine-tuning on giants like DeepSeek-V3.1, DeepSeek-R1-0528, Qwen3-235B-A22B, Qwen3-Coder-480B-A35B, and Meta's Llama 4 Maverick and Scout — models that until recently required serious infrastructure gymnastics just to train reliably across multiple nodes.

Context length got a quiet but meaningful overhaul too. Together says most models on the platform now support 2x to 4x longer training contexts than before, at no extra cost. Some, like Llama 3.1-8B and Gemma 3-4B, can now stretch all the way to their theoretical max of 131,000 tokens. This matters more than it sounds — if your training data doesn't match the context length you'll actually use at inference time, results suffer. Slingshot AI, which built the therapy app Ash, leaned on this directly: co-founder Daniel Cahn said the platform's long-context handling eliminated the job failures his team kept hitting elsewhere while fine-tuning on full clinical conversations.

The other big shift is Hugging Face integration, done properly this time. You can now pull any model from the Hub as a starting point for fine-tuning, and push your finished checkpoints straight back to a repository, provided you hand over an API key with the right permissions. Together frames this as removing a step that used to require manual downloading, uploading, and babysitting — useful given how many task-specific fine-tunes get published daily by the community rather than by the big labs themselves. The company flags it as experimental for now, so expect some friction with unusual model architectures.

Rounding things out, Together added more preference-optimization flavors — length-normalized DPO, DPO+NLL, and SimPO — exposed as simple training flags, plus a batch_size="max" option that auto-picks the largest batch size the platform supports for a given model and mode. Small detail, but it's the kind of thing that quietly saves people from wasted experimentation cycles.

None of this is flashy in isolation. But stacked together, it's Together AI positioning itself less as a place to rent GPUs and more as the default operating layer for anyone customizing open models at serious scale.

My take — AI-written commentary, not fact-checked reporting

What strikes me is how much of this update is really about removing friction rather than adding new capability — and that's usually where the real adoption happens, not in flashy benchmark wins. The Hugging Face round-trip integration in particular is a smart bet: it quietly cements Together as infrastructure sitting underneath an ecosystem it doesn't own, which is exactly the position you want if you believe open models keep winning share from closed ones.

Read more about this at: Together AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.