TLDRocket
Sign in

Kog is going deeper to squeeze more inference out of GPUs

TechCrunch Anna Heim

French startup Kog says it can make standard data center GPUs run AI much faster. That could cut wait times and costs without waiting for new chips.

Based on reporting by TechCrunch, Anna Heim — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

The AI inference race is pushing some companies toward new hardware, but Kog is betting on the opposite path: wringing more speed out of the GPUs data centers already have. The French startup first got attention in May with a preview meant to show that “extremely fast single-request decoding” could run on standard enterprise GPUs, including AMD MI300X and NVIDIA H200 chips.

That pitch landed because inference has become a real bottleneck. Kog says it got 200 tangible business leads after the demo, and CEO Gaël Delalleau expects software engineering to be the first place customers use it. The startup is aiming at people who already feel the pain of long waits, including Claude Code users and teams that pay extra for Claude’s Fast Mode. It also has design partners building games and apps from prompts, where faster output can mean more revenue.

But there’s a gap between the demo and the claim on the tin. Kog says its proof point hit 3,000 per-request tokens per second, but that was on Laneformer 2B, a small open-source model with about 2 billion parameters. The company is now focused on getting larger models to run well enough to back its promise of “30x faster LLM inference.” Delalleau says the team learned its prospective customers are not eager to fine-tune small models, which pushed Kog toward larger ones instead.

Delalleau argues that GPUs are misunderstood, not obsolete. He says newer chips have more memory bandwidth waiting to be used, and he compares Kog’s approach to deep GPU research more than broad software portability. That focus comes from his background in solid-state physics, offensive cybersecurity and low-level reverse engineering. It also comes with a cost: for each new GPU, Kog says it can spend weeks or even months digging into the hardware, which limits how many chips a team of 11 can tackle.

The company is backed by Scaleway, Bpifrance and French Tech 2030’s program, with Varsity VC co-leading the seed round. Kog’s bigger test is still ahead. Delalleau says the first major model at 10x speed should arrive in September, and only then will the startup be able to show customer traction and go after a Series A.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of contrarian bet: not more expensive silicon, but less laziness in software. The AI market keeps rewarding whoever makes the same hardware feel less ordinary, and GPU vendors should be sweating that a small French team thinks it can still find room in the cracks. Hardly a comforting thought for the chip stack.

Read more about this at: TechCrunch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.