Together AI
Company www.together.ai ● Covered in 79 stories + Follow ✴ AI Graph
Together AI is a platform provider offering managed services for model fine-tuning, dedicated model inference, and production deployment workflows. Recent coverage highlights updates to Together Fine-Tuning (more open-weight model support, live experiment tracking, and added training controls), the addition of endpoint-level A/B testing, and autoscaling of inference endpoints based on metrics such as in-flight requests and GPU utilization. Together AI has also partnered with Moonshot AI to natively serve Kimi K3 on its API and worked with Y Combinator to make a dedicated GPU cluster available to YC portfolio startups.
Updated 13 September 2026
Signals
2 stories (-71%)
Media momentum
As of 17 Sep 2026 · Visible stories in the last 30 days, compared with the 30 days before.
5 sources
Source diversity
As of 17 Sep 2026 · Distinct publications behind this entity's visible coverage.
4 events
Release activity
As of 17 Sep 2026 · Model, product and open-source release events in the last 90 days whose coverage involves this entity.
—
Funding signals
As of 17 Sep 2026 · Funding and acquisition events in the last 90 days whose coverage involves this entity.
Jan 2025 → Sep 2026
Coverage span
As of 17 Sep 2026 · First to most recent month of TLDRocket coverage of this entity.
Management
updated 2 Sep 2026-
Vipul Ved Prakash
Co-founder & CEO
-
Kai Mak
Chief Revenue Officer
-
Ce Zhang
Co-founder, CTO
- Charles Zedlewski Chief Product Officer
- Timothy Yen Accounting & Financial Operations Lead
- John C. Lee Director of Financial Operations & Accounting
-
James Barker
VP EMEA
-
Nicolette L.
Director Of HR & Operations
-
Vanessa Hsu
Technical Advisor To The CEO
- Kae Lim Executive Assistant To Co-founder And CEO
-
Kae Ike Lim
Executive Assistant To Co-founder And CEO
-
Cyrus L.
Founding GTM
+ 6 more tracked
Recent changes
- Meicheng Shi left (was SVP Finance)
- Tri Dao Founding Chief Scientist → Chief Scientist
Latest developments
Saudi fires up Gulf AI race with $15bn tech deals at LEAP
Fortune ·
46
A/B test models in production
Together AI · 1 month ago ·
53
Kimi K3: The Complete Developer Guide
Together AI · 1 month ago ·
50
Autoscaling endpoints for LLM inference
Together AI · 1 month ago ·
7
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
Together AI · 1 month ago ·
12
Configuring Dedicated Model Inference
Together AI · 1 month ago ·
26
China’s Low-Priced Z.ai Model Is Exposing Costly Coder Habits
IEEE Spectrum · 1 month ago ·
13
September 2026
Humain and Together AI announce a partnership Partnership
- Together AI expands fine-tuning service with more models, live metrics, and finer controls
- Saudi fires up Gulf AI race with $15bn tech deals at LEAP
August 2026
July 2026
Together AI launches Dedicated Model Inference platform with autoscaling and traffic routing capabilities Product launch
Moonshot AI releases Kimi K3, a 2.8 trillion parameter open-weight model achieving frontier-class performance Open source release
Moonshot AI releases Kimi K3, a 2.8 trillion parameter open-weight model with API pricing and delayed weights availability Model release
Thinking Machines releases Inkling, an open-source multimodal language model with 975 billion parameters Open source release
- Autoscaling endpoints for LLM inference
- Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
- Configuring Dedicated Model Inference
- China’s Low-Priced Z.ai Model Is Exposing Costly Coder Habits
- Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
- What does 99.9% uptime mean for inference?
- New in Together GPU Clusters: Reliability and control for production GPU clusters
- Together AI brings Thinking Machines Lab’s new model Inkling on day 0
- Open, convenient and predictable: Introducing Provisioned Throughput
- Announcing our $800M Series C to accelerate the shift to open-source AI
June 2026
- Together AI at ICML 2026: frontier research across the full stack
- Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification
- Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets
May 2026
- How Together AI built the world’s fastest speech-to-text stack
- Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference
- Violin: An open-source video translation skill that breaks language barriers
- Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices
- Serving DeepSeek-V4: why million-token context is an inference systems problem
- Deploy and inference any model from HuggingFace
- Foundational research powering efficient inference at scale
April 2026
NVIDIA releases Nemotron 3 Nano Omni multimodal model with support for text, images, video, and audio Model release
- Announcing Together AI and Adaption Partnership
- DeepSeek-V4 Pro now available on Together AI
- Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0
- Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams
- What is an AI Native Cloud?
- Wan 2.7 video model suite now available on Together AI
- Deepgram speech-to-text and voice models now available natively on Together AI
- Inside the Together AI kernels team
March 2026
Together AI Expands Fine-Tuning Service with Tool Calling, Reasoning, and Vision Support Feature update
- Deep Learning Weekly: Issue 447
- Together AI expands fine-tuning service with tool calling, reasoning, and vision support
- Together AI at NVIDIA GTC 2026: Explore our latest innovations across research and products
- Build real-time voice agents on Together AI
- Together AI Brings NVIDIA Nemotron 3 to Developers on Day 0
- New in Together GPU Clusters: Autoscaling, observability, and self-healing
- Key research and product announcements at the AI Native Conf
- Cache-aware prefill–decode disaggregation (CPD) for up to 40% faster long-context LLM serving
- Introducing Together AI’s new look
February 2026
- How speech models fail where it matters the most and what to do about it
- Introducing Dedicated Container Inference: Delivering 2.6x faster inference for custom AI models
- Rime Arcana V3 Turbo and Rime Arcana V3 now available on Together AI
- Together AI welcomes Alon Gavrielov as VP of Infrastructure Strategy
January 2026
- Optimizing inference speed and costs: Lessons learned from large-scale deployments
- Learn how Cursor partnered with Together AI to deliver real-time, low-latency inference at scale
December 2025
Relationships
Products & technology
- Deploys Kimi K3 · 2 sources
- Supplies DeepSeek · 1 source
- Supplies Llama · 1 source
- Supplies Mistral · 1 source
- Deploys NVIDIA Blackwell · 1 source
- Supplies NVIDIA · 1 source
- Integrated with OpenAI Sora 2 · 1 source
- Integrated with Google Veo 3.0 · 1 source
- Integrated with ByteDance SeeDream · 1 source
Partnerships
- Partnered with Moonshot AI · 1 source
- Partnered with Hypertec · 1 source
- Partnered with 5C Group · 1 source