Cerebras got a warm IPO reception in May, and the message from markets seemed clear: purpose-built silicon is where the inference story goes next. But Kog, a French startup founded by a former white-hat hacker with a physics degree, is making a different argument entirely. According to TechCrunch, the company believes there is dramatically more performance sitting untapped inside the GPUs enterprises already own, and that the right software can get to it.
Kog hit the front page of Hacker News in May with a tech preview demonstrating 3,000 tokens per second on a single request, running on AMD MI300X and Nvidia H200 GPUs. That is the kind of speed that tends to make developers stop scrolling. The catch: it was achieved using Laneformer 2B, a purpose-built 2-billion parameter model that Kog has since open-sourced. Scaling that result to the large language models customers actually want to run is the harder problem, and the one Kog is now working to solve.
The startup claims to target 30x faster LLM inference through its Kog Inference Engine. CEO Gaël Delalleau says the demo generated 200 concrete business leads, with software engineering emerging as the clearest early use case. Anyone who has run extended Claude Code sessions knows the pain of waiting hours for results. Anthropic has already signaled that speed carries real monetary value, charging a premium for Claude’s Fast Mode. Kog is positioning itself for the customers who hit that ceiling regularly.
Early conversations also revealed something the startup had not fully anticipated. Prospective customers are not ready to fine-tune smaller models. So Kog shifted focus toward accelerating larger models to match actual demand. That is a significant pivot, and it raises the bar for what the company needs to demonstrate before a Series A becomes realistic. Delalleau says a first major model running at 10x speed is expected in September, after which he plans to show customer traction and begin fundraising.
Kog is not alone in this space. ZML, also French, released hardware-agnostic inference software that bypasses CUDA to support multiple chip vendors. But Delalleau draws a sharper comparison to Hazy Research at Stanford, a lab known for going deep into GPU-level optimization rather than building abstraction layers on top. Kog’s approach is closer to that end of the spectrum.
That philosophy comes directly from Delalleau’s background. He studied solid-state physics at École Polytechnique, then spent years in offensive cybersecurity, reaching finalist status at DEFCON’s CTF competition four times. The hacker mindset, he says, means reverse-engineering systems at the assembly and binary level to make them do things they were not originally designed for. Applied to GPUs, the argument is that newer hardware has far more memory bandwidth available than most inference stacks actually use.
The tradeoff is speed of coverage. For each new GPU architecture, Kog spends weeks or months on low-level engineering research before it can support that chip well. With a team of 11, that limits how many chips the company can work with at once. The longer-term plan involves agent-based pipelines that could help automate more of that work and expand coverage over time.
Backing includes Varsity VC, co-led by Kamel Zeroual, Delalleau’s former co-founder from his first startup Stribe, plus Bpifrance and France’s French Tech 2030 program. Scaleway is also a supporting partner. The European angle matters here. As EU institutions push to build sovereign AI infrastructure, a French startup with deep GPU expertise and regional backing is well-positioned to benefit from that momentum. But first, Kog needs to prove the approach works at real model scale. September is the near-term test.




