/ Sep 20, 2026
Trending
French startup Kog is challenging the assumption that specialized AI chips are necessary for fast inference, betting instead that conventional GPUs can be pushed much further through software optimization. The company’s recent demo, which gained attention on Hacker News, showcased extremely fast single-request decoding on standard data center GPUs, including the AMD MI300X and NVIDIA H200.
The demo achieved an impressive 3,000 tokens per second but used a small, purpose-built model with roughly 2 billion parameters. Now, Kog is aiming to deliver on its promise of 30x faster LLM inference with much larger models, a goal its CEO Gaël Delalleau believes is achievable. He argues that newer GPUs have significant memory bandwidth that is currently underutilized, and that the idea they are poorly suited for decoding is a misconception.
Kog’s approach has already generated considerable interest, with the company reporting 200 business leads after its initial preview. The first likely use case is software engineering, where developers often face long waits for results from AI coding assistants. The company is also working with design partners who generate games and apps, where faster output would directly increase revenue.
However, Kog acknowledges the market is still developing. It has learned that potential customers are not yet ready to fine-tune small models, so the company is now focusing on accelerating the development of larger models to meet the observed demand.
Delalleau’s unique background, combining solid-state physics from École Polytechnique with offensive cybersecurity, informs the company’s deep-level approach. This method involves a hands-on, time-consuming process of GPU engineering research for each new chip, which currently limits the number of hardware platforms the 11-person team can support. In the long term, Kog plans to automate this methodology.
The startup is part of a broader movement exploring software-based acceleration, with other players like ZML also working on hardware-agnostic solutions. Kog sees its work as similar to that of Stanford’s Hazy Research lab, but with an even deeper focus on GPU acceleration.
Backed by Bpifrance and the French Tech 2030 program, and supported by Scaleway, Kog’s next milestone is to demonstrate its technology on a major model. The CEO expects to achieve a 10x speedup on such a model by September, which would be key to securing a Series A round.
#AI, #GPUs, #Inference, #Startup
It is a long established fact that a reader will be distracted by the readable content of a page when looking at its layout. The point of using Lorem Ipsum is that it has a more-or-less normal distribution
Copyright PopularTechNews. 2024