TLDRocket
Sign in

Migrating from closed to open source models, Together

Together AI

Together AI lays out a five-step move from closed models to open ones. The pitch: weeks to months, not years, if the evals and tooling are solid.

Based on reporting by Together AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Together AI is arguing that swapping closed models for open source models does not have to feel like a ruinous migration project. The company’s pitch is a five-stage playbook: discover, evaluate, adapt, decide, and take it to production. The basic claim is simple enough. With the right managed service, the blast radius shrinks, the tooling gets easier to swap, and the move can happen in weeks to months instead of months to years.

The discovery stage starts with discipline, not benchmarks. Together says teams should define the workload in plain terms: what the model needs to do, what a good answer looks like, and what the traffic actually contains. Only then do public benchmarks become useful as a filter. The article points readers to sources such as The Open Frontier, Artificial Analysis, Epoch AI, Vals AI, Scale’s leaderboards, Intelligence and Arena, and it breaks models into three rough tiers: frontier open models like GLM-5.3 and Kimi K3, value-tier options like DeepSeek-V4 Flash and GLM-5.3 Flash, and smaller edge-deployable choices such as Qwen 3.8 27B and Muse Glimmer.

But raw score is not the whole story. Together wants teams to look at cost per task, token use, number of steps, end-to-end runtime, and how quickly a correct answer can be verified. A model that scores a little lower but finishes with half the tokens and half the time may be the better fit. After that comes a reality check in a playground, using representative tasks to see whether the problem is the model itself or the harness wrapped around it.

Evaluation, in the company’s framing, comes down to two things: accuracy and performance. Accuracy covers what the model can do, including instruction following, summarization, function calling, and vision. Performance covers how fast and cheaply it does it. Together says the most useful tests come from replaying real traffic, because generic benchmarks miss the shape of a workload. If there is no traffic data, teams can still size by input and output size, plus expected cache hit rate. Gateways such as LiteLLM can help gather some of that data.

If a model does not pass, the answer is not always to give up. Together lays out a ladder of adjustments: system prompts, sampling settings, context engineering, and the harness itself, then fine-tuning or distillation if needed. Once the model clears the goal line, the company says the team can package the case for stakeholders by laying out effort, risks, and ROI. Together claims some customers have seen up to 70% cost reduction when moving to open source, and says production can even begin with canary deployments at 10% of traffic.

My take — AI-written commentary, not fact-checked reporting

This is the rare AI migration advice that sounds like it was written by people who have actually wrestled a model into production. The real tell is the insistence on replaying your own traffic instead of worshipping leaderboards like they’re carved on a mountain. That’s the right bias: less model fan fiction, more boring verification, which is usually where the money is hiding.

Read more about this at: Together AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.