The tech world is abuzz with Kog's ambitious plan to harness conventional GPUs for lightning-fast inference. The startup has already shown impressive results with their Laneformer model, achieving an astounding 3,000 tokens per second. But will this approach work with larger language models? CEO Gaël Delalleau remains optimistic.
Delalleau’s background in offensive cybersecurity and solid-state physics provides a unique insight into optimizing hardware performance. His team is working diligently to push the boundaries of what GPUs can achieve, even dedicating several months to dissect each new chip model.
The race for faster inference isn't just about speed; it's also about cost. Kog aims to unlock capabilities on existing hardware through software optimization, which could be a game-changer for enterprises and developers alike. However, the company faces challenges in scaling their methodology across multiple models and chips.
As Europe looks to strengthen its tech sovereignty, Kog’s efforts could provide a significant boost. The startup has already secured support from Scaleway and backing from France's Bpifrance, positioning it well for future growth.
The key will be demonstrating this approach works on larger language models by September. If successful, Kog could revolutionize how we use GPUs in AI applications, making fast inference a reality on the hardware we already own.







