logo-darklogo-darklogo-darklogo-dark
  • Tool Categories
    • 🎨Art & Creative Design505
    • 🏢Business Management644
    • 💻Coding & Development514
    • 👮Detection83
    • 🧠General Use728
    • 🏥Health & Wellness55
    • 📷Image & Photo Analysis100
    • 🖼️Image Generation & Editing618
    • 📐Interior & Architectural Design37
    • 🎓Learning & Education483
    • ⚖️Legal & Finance90
    • 🎭Lifestyle & Entertainment236
    • 📢Marketing & Advertising627
    • 🎧Music & Audio138
    • 👔Office & Workplace1,014
    • 🔬Research & Data Analysis373
    • 👥Social Media245
    • 🎥Video Generation & Editing426
    • 👧🏻Virtual Companion135
    • 🎤Voice Generation & Editing381
    • ✍️Writing & Editing808
    • All Categories
    • AI Use Cases
  • News
  • Events
    • Academic Conferences
    • Developer Conferences
    • Expos / Trade Shows
    • Industry Summits
    • Workshops / Training
    • All Events
    • Past Events
  • Saved Tools
  • Suggest a Tool
✕
Home › News › Kog thinks software, not silicon, is the answer to faster AI inference

Kog thinks software, not silicon, is the answer to faster AI inference

August 14, 2026
Kog thinks software, not silicon, is the answer to faster AI inference

Cerebras got a warm IPO reception in May, and the message from markets seemed clear: purpose-built silicon is where the inference story goes next. But Kog, a French startup founded by a former white-hat hacker with a physics degree, is making a different argument entirely. According to TechCrunch, the company believes there is dramatically more performance sitting untapped inside the GPUs enterprises already own, and that the right software can get to it.

Kog hit the front page of Hacker News in May with a tech preview demonstrating 3,000 tokens per second on a single request, running on AMD MI300X and Nvidia H200 GPUs. That is the kind of speed that tends to make developers stop scrolling. The catch: it was achieved using Laneformer 2B, a purpose-built 2-billion parameter model that Kog has since open-sourced. Scaling that result to the large language models customers actually want to run is the harder problem, and the one Kog is now working to solve.

The startup claims to target 30x faster LLM inference through its Kog Inference Engine. CEO Gaël Delalleau says the demo generated 200 concrete business leads, with software engineering emerging as the clearest early use case. Anyone who has run extended Claude Code sessions knows the pain of waiting hours for results. Anthropic has already signaled that speed carries real monetary value, charging a premium for Claude’s Fast Mode. Kog is positioning itself for the customers who hit that ceiling regularly.

Early conversations also revealed something the startup had not fully anticipated. Prospective customers are not ready to fine-tune smaller models. So Kog shifted focus toward accelerating larger models to match actual demand. That is a significant pivot, and it raises the bar for what the company needs to demonstrate before a Series A becomes realistic. Delalleau says a first major model running at 10x speed is expected in September, after which he plans to show customer traction and begin fundraising.

Kog is not alone in this space. ZML, also French, released hardware-agnostic inference software that bypasses CUDA to support multiple chip vendors. But Delalleau draws a sharper comparison to Hazy Research at Stanford, a lab known for going deep into GPU-level optimization rather than building abstraction layers on top. Kog’s approach is closer to that end of the spectrum.

That philosophy comes directly from Delalleau’s background. He studied solid-state physics at École Polytechnique, then spent years in offensive cybersecurity, reaching finalist status at DEFCON’s CTF competition four times. The hacker mindset, he says, means reverse-engineering systems at the assembly and binary level to make them do things they were not originally designed for. Applied to GPUs, the argument is that newer hardware has far more memory bandwidth available than most inference stacks actually use.

The tradeoff is speed of coverage. For each new GPU architecture, Kog spends weeks or months on low-level engineering research before it can support that chip well. With a team of 11, that limits how many chips the company can work with at once. The longer-term plan involves agent-based pipelines that could help automate more of that work and expand coverage over time.

Backing includes Varsity VC, co-led by Kamel Zeroual, Delalleau’s former co-founder from his first startup Stribe, plus Bpifrance and France’s French Tech 2030 program. Scaleway is also a supporting partner. The European angle matters here. As EU institutions push to build sovereign AI infrastructure, a French startup with deep GPU expertise and regional backing is well-positioned to benefit from that momentum. But first, Kog needs to prove the approach works at real model scale. September is the near-term test.

Share

Related news

Perplexity’s India giveaway is paying off, but the real test is just starting
August 18, 2026

Perplexity’s India giveaway is paying off, but the real test is just starting


Read more
Nous Research ships Bot Mode for Hermes Agent, enabling multi-agent collaboration out of the box
August 18, 2026

Nous Research ships Bot Mode for Hermes Agent, enabling multi-agent collaboration out of the box


Read more
Michael Burry calls out Big Tech’s $3 trillion AI spending problem hiding in plain sight
August 18, 2026

Michael Burry calls out Big Tech’s $3 trillion AI spending problem hiding in plain sight


Read more

Recent Posts

  • Perplexity’s India giveaway is paying off, but the real test is just starting
  • Nous Research ships Bot Mode for Hermes Agent, enabling multi-agent collaboration out of the box
  • Michael Burry calls out Big Tech’s $3 trillion AI spending problem hiding in plain sight
  • OpenAI launches ChatGPT for Teens with automatic age detection and stricter safety limits
  • Honor’s next wide foldable may run on a 2nm Snapdragon chip with a 7,000mAh battery
Best AI Tools

Discover the best AI tools for any use case

Explore
  • Tool Categories
  • AI Use Cases
  • AI Events
  • AI News
  • Saved Tools
Company
  • About Us
  • Contact Us
  • Media & Partnerships
  • Suggest a Tool
Legal
  • Privacy Policy
  • Terms of Service
Copyright © 2026 Best AI Tools 415 Mission Street, 37th Floor, San Francisco, CA 94105