Menlo’s Investment in Fireworks: The Runtime for Specialized Intelligence
Menlo Ventures Menlo Ventures
Menlo Ventures is backing Fireworks AI's $1.5 billion Series D. The company runs the infrastructure that lets custom and open-source AI models actually ship in production, and it's now doing 43 trillion tokens a day.
Based on reporting by Menlo Ventures, Menlo Ventures — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Money is pouring into the unglamorous plumbing of AI right now, and Fireworks just landed a big chunk of it. Menlo Ventures announced it's joining Fireworks' $1.5 billion Series D, betting that the real growth story in AI infrastructure isn't just the handful of labs chasing raw capability, but the layer underneath that actually gets specialized models into production cheaply and fast.
The numbers behind that bet are hard to ignore. Fireworks' daily token volume has nearly tripled since late last year, climbing from 15 trillion to 43 trillion, and annualized revenue has reached $1 billion. Menlo points to a broader shift too: open-source usage on OpenRouter has grown more than 10x since the start of the year, as models became capable of sustaining long-running tasks without someone babysitting them the whole way through.
What Fireworks actually sells is the messy engineering work that sits between a trained model and a usable product — continued pretraining, fine-tuning, and reinforcement learning living on the same stack as inference, so teams can take an open model, adapt it to their own data, and keep improving it from live usage. The company's FireAttention stack automates the search across hardware, quantization, sharding, speculative decoding, batching, and kernel choices that determines whether a model runs fast and cheap or slow and expensive. Customers are already leaning on it for serious work: Cursor trained its Composer 2 coding model on the platform, Vercel says it got a 40x latency improvement for v0 users through reinforcement fine-tuning and speculative decoding with Fireworks, and Factory is using it to squeeze up to 15x more work out of the same budget with open-source options.
The team behind it has a pedigree that matches the ambition. CEO Lin Qiao previously ran PyTorch at Meta, and CTO Dmytro Dzhulgakov was one of its core maintainers, and the rest of the seven-person founding group came out of Meta's ads infrastructure, News Feed machine learning, and compiler work, plus a former AI lead from Google Vertex. President George Hu, who joined more recently, previously helped scale Salesforce 50x to $5 billion and then led a 10x growth stretch at Twilio.
Menlo frames this as one piece of a larger thesis about what it calls the token path — the compute, data, and orchestration turning raw intelligence into shippable product. Its portfolio already includes Anthropic at the model layer, OpenRouter in routing, Gimlet and Modal in compute, Neon and Pinecone in databases, and Unstructured in data infrastructure. Fireworks, in Menlo's telling, sits right at the center of that stack.
My take — AI-written commentary, not fact-checked reporting
Everyone keeps arguing about which frontier lab has the smartest model, but the money quietly piling into companies like Fireworks suggests the more durable business is in making models cheap and fast enough to actually run. A $1.5 billion round for infrastructure that most end users will never hear about says a lot about where investors think the durable margins actually sit — not with whoever ships the flashiest benchmark, but with whoever owns the plumbing everyone else has to rent.
Read more about this at: Menlo Ventures