Together AI
Together AI is an AI inference platform that provides API access to open-weight and proprietary language models, including recent partnerships to host Kimi K3 and Inkling models. The company operates a Dedicated Model Inference service with autoscaling capabilities, GPU cluster infrastructure, and Provisioned Throughput pricing options designed for production workloads, while recently expanding partnerships with model providers like Moonshot AI and Y Combinator to increase developer access to compute resources.
Updated 3 August 2026
Signals
10 stories (+233%)
Media momentum
As of 3 Aug 2026 · Visible stories in the last 30 days, compared with the 30 days before.
4 sources
Source diversity
As of 3 Aug 2026 · Distinct publications behind this entity's visible coverage.
4 events
Release activity
As of 3 Aug 2026 · Model, product and open-source release events in the last 90 days whose coverage involves this entity.
—
Funding signals
As of 3 Aug 2026 · Funding and acquisition events in the last 90 days whose coverage involves this entity.
Jan 2025 → Aug 2026
Coverage span
As of 3 Aug 2026 · First to most recent month of TLDRocket coverage of this entity.
Latest developments
Kimi K3: The Complete Developer Guide
Together AI · 2 days ago ·
47
Autoscaling endpoints for LLM inference
Together AI · 3 days ago ·
3
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
Together AI · 5 days ago ·
5
Configuring Dedicated Model Inference
Together AI · 5 days ago ·
20
China’s Low-Priced Z.ai Model Is Exposing Costly Coder Habits
IEEE Spectrum AI · 1 week ago ·
10
Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
Together AI · 2 weeks ago ·
41
What does 99.9% uptime mean for inference?
Together AI · 2 weeks ago ·
19
New in Together GPU Clusters: Reliability and control for production GPU clusters
Together AI · 2 weeks ago ·
15
Q3 2026
Together AI launches Dedicated Model Inference platform with autoscaling and traffic routing capabilities Product launch
Moonshot AI releases Kimi K3, a 2.8 trillion parameter open-weight model achieving frontier-class performance Open source release
Moonshot AI releases Kimi K3, a 2.8 trillion parameter open-weight model with API pricing and delayed weights availability Model release
Thinking Machines releases Inkling, an open-source multimodal language model with 975 billion parameters Open source release
- Kimi K3: The Complete Developer Guide
- Autoscaling endpoints for LLM inference
- Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
- Configuring Dedicated Model Inference
- China’s Low-Priced Z.ai Model Is Exposing Costly Coder Habits
- Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
- What does 99.9% uptime mean for inference?
- New in Together GPU Clusters: Reliability and control for production GPU clusters
- Together AI brings Thinking Machines Lab’s new model Inkling on day 0
- Open, convenient and predictable: Introducing Provisioned Throughput
- Announcing our $800M Series C to accelerate the shift to open-source AI
Q2 2026
NVIDIA releases Nemotron 3 Nano Omni multimodal model with support for text, images, video, and audio Model release
- Together AI at ICML 2026: frontier research across the full stack
- Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification
- Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets
- How Together AI built the world’s fastest speech-to-text stack
- Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference
- Violin: An open-source video translation skill that breaks language barriers
- Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices
- Serving DeepSeek-V4: why million-token context is an inference systems problem
- Deploy and inference any model from HuggingFace
- Foundational research powering efficient inference at scale
- Announcing Together AI and Adaption Partnership
- DeepSeek-V4 Pro now available on Together AI
- Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0
- Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams
- What is an AI Native Cloud?
- Wan 2.7 video model suite now available on Together AI
- Deepgram speech-to-text and voice models now available natively on Together AI
- Inside the Together AI kernels team
Q1 2026
Together AI Expands Fine-Tuning Service with Tool Calling, Reasoning, and Vision Support Feature update
- Deep Learning Weekly: Issue 447
- Together AI expands fine-tuning service with tool calling, reasoning, and vision support
- Together AI at NVIDIA GTC 2026: Explore our latest innovations across research and products
- Build real-time voice agents on Together AI
- Together AI Brings NVIDIA Nemotron 3 to Developers on Day 0
- New in Together GPU Clusters: Autoscaling, observability, and self-healing
- Key research and product announcements at the AI Native Conf
- Cache-aware prefill–decode disaggregation (CPD) for up to 40% faster long-context LLM serving
- Introducing Together AI’s new look
- How speech models fail where it matters the most and what to do about it
- Introducing Dedicated Container Inference: Delivering 2.6x faster inference for custom AI models
- Rime Arcana V3 Turbo and Rime Arcana V3 now available on Together AI
- Together AI welcomes Alon Gavrielov as VP of Infrastructure Strategy
- Optimizing inference speed and costs: Lessons learned from large-scale deployments
- Learn how Cursor partnered with Together AI to deliver real-time, low-latency inference at scale
Q4 2025
Together AI Integrates PyTorch Reinforcement Learning Capabilities into AI Cloud Platform Partnership
- MiniMax Speech 2.6 Turbo now available natively on Together AI
- Rime voice models now available on Together AI
- Announcing native availability of NVIDIA Nemotron 3 Nano, NVIDIA’s latest reasoning model
- How to run TorchForge reinforcement learning pipelines in the Together AI Native Cloud
- Together AI and Meta partner to bring PyTorch Reinforcement Learning to the AI Native Cloud
- Together AI delivers fastest inference for the top open-source models
Relationships
Products & technology
- Deploys Kimi K3 · 2 sources
- Supplies DeepSeek · 1 source
- Supplies Llama · 1 source
- Supplies Mistral · 1 source
- Deploys NVIDIA Blackwell · 1 source
- Supplies NVIDIA · 1 source
- Integrated with OpenAI Sora 2 · 1 source
- Integrated with Google Veo 3.0 · 1 source
- Integrated with ByteDance SeeDream · 1 source
Partnerships
- Partnered with Moonshot AI · 1 source
- Partnered with Hypertec · 1 source
- Partnered with 5C Group · 1 source