Together AI launches Dedicated Model Inference platform with autoscaling and traffic routing capabilities
Product launch Provisional 85% confidence first seen
Together AI introduced its Dedicated Model Inference platform, which enables users to autoscale LLM deployments based on inference-specific metrics such as in-flight requests, time-to-first-token, and GPU utilization. The platform features a traffic routing system that distributes requests proportionally to endpoint capacity and supports A/B testing and gradual rollouts without manual intervention.