Kog Goes Deeper to Squeeze More Inference Out of GPUs

The race for faster AI inference is heating up, and markets gave Cerebras and its purpose-built chips a warm welcome during its IPO debut in May. But French startup Kog is betting that there's far more performance to be unlocked from conventional GPUs.

In May, the startup hit the front page of Hacker News with a tech preview designed to prove that "extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own" โ€” such as the AMD MI300X and NVIDIA H200 models used in the demo.

Some were disappointed to learn that this didn't extend to GPUs in laptops, but others saw the potential. With inference speed and costs becoming critical bottlenecks, Kog's promise to unlock new capabilities on existing hardware through software optimization attracted more than just onlookers. "We had 200 tangible business leads," CEO Gaรซl Delalleau told TechCrunch.

Based on early feedback, the solo founder expects software engineering to be the first major use case. Veteran Claude Code users are well aware that they sometimes wait hours for results. Anthropic itself understands that speed is worth money, charging a premium for Claude's Fast Mode.

Kog aims to target customers deterred by those delays, typically those relying on AI workflows for professional tasks. The startup also has design partners that let users generate games and apps with a prompt, and for whom faster outcomes thanks to the Kog Inference Engine (KIE) would translate into more revenue, Delalleau said.

The company recognizes that this market isn't fully mature yet. While observing demand, Kog learned that prospective customers aren't prepared to fine-tune small models. "And that's why since the launch, we've been fully focused on accelerating the development of larger models to meet the demand we've seen," he added.

This leaves Kog with a significant leap to make to deliver on its promise of "30x faster LLM inference." Its demo showcased an impressive 3,000 per-request tokens per second (TPS) โ€” but with a purpose-built small model of only 2 billion parameters, the now open-sourced Laneformer 2B.

Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can challenge inference chips. "GPUs have a bright future," he said. For Kog's CEO, the notion that GPUs aren't well suited for decoding has become a misconception; newer GPUs offer increasing memory bandwidth that's begging to be unlocked.

Kog isn't alone in believing that software optimization can help GPUs exceed their specs. ZML, also from France, released hardware-agnostic software that bypasses Nvidia's CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research, with an even deeper focus on GPU acceleration.

Delalleau himself isn't a researcher, and his first startup, TechCrunch50 2009 alum ...

via TechCrunch AI

Related