TLDRocket
Sign in

Configuring Dedicated Model Inference

Together AI

Together AI's Dedicated Model Inference platform uses endpoints (stable identities), deployments (model + hardware combinations), and configs (runtime recipes) to route traffic based on capacity rather than fixed percentages. Traffic splits assign weights to deployments, and the router distributes requests proportionally to effective capacity (weight × ready replicas), so autoscaling automatically adjusts traffic without manual intervention. This architecture enables A/B tests, rollouts, shadow experiments, and zero-downtime changes, with a live experiment showing that when deployment A scaled from 1 to 2 replicas, its traffic share increased from 50% to 69.4% as the router followed the capacity change.

Why it matters

The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.