French startup Kog is pursuing a novel approach to boost AI inference speed by optimizing existing GPUs rather than relying on specialized hardware. The company demonstrated its Kog Inference Engine (KIE) on AMD MI300X and Nvidia H200 GPUs, achieving 3,000 tokens per second for a 2 billion parameter model. This performance, based on the open-sourced Laneformer 2B, highlights Kog’s focus on unlocking the potential of conventional datacenter GPUs for large language model (LLM) inference. The startup’s CEO, Gaël Delalleau, emphasized that the goal is to provide faster results for users who rely on AI for professional tasks, such as generating apps or games with prompts.
Kog’s strategy centers on software optimization to enhance GPU performance, a path that contrasts with companies like Cerebras, which developed purpose-built chips. Delalleau noted that while demand for faster inference is growing, many customers are not yet ready to fine-tune small models. Instead, Kog has focused on accelerating the development of larger models to meet this demand. The company acknowledges the challenge of scaling its approach to LLMs, which are significantly more complex than the smaller models used in its demo. However, Delalleau remains confident that the same optimization techniques can be applied to larger models, given the increasing memory bandwidth of modern GPUs.
Kog’s approach is part of a broader trend in the AI industry, where software optimization is being explored as a way to enhance GPU performance without requiring new hardware. The startup’s methods are similar to those of Stanford University’s Hazy Research, though Kog’s focus is more deeply rooted in low-level GPU engineering. Delalleau, who has a background in offensive cybersecurity, explained that this mindset helps the team understand how to maximize GPU capabilities. However, this hands-on approach requires significant time and resources, limiting the number of GPUs Kog can support in the near term. The company plans to expand its capabilities through agent-based pipelines in the future, which could also align with Europe’s push for greater technological sovereignty.
Source: techcrunch