--- title: 'Research Engineer at Kog' canonical: 'https://feeny.ai/job/research-engineer-kog-paris-5ffy8ryxsz1a' type: 'job' last_seen: '2026-09-08' --- # Research Engineer at Kog - **Company:** Kog - **Location:** Paris, France - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-06-09 - **Last confirmed live:** 2026-09-08 - **Apply:** https://jobs.ashbyhq.com/kog/364b3702-a7a9-431d-8e70-f8fe3f4d2f01 ## Job description ## About Kog Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding). We co-design the model architecture and the execution engine together. Our Laneformer model uses Delayed Tensor Parallelism (DTP), a novel architecture that restructures the Transformer dependency graph so inter-GPU communication overlaps with computation rather than blocking it. We pre-trained a 2B-parameter DTP model on 6T tokens on 256 H100 GPUs. We are a team of 11 people, including 10 engineers and 5 PhDs. Test it at playground.kog.ai http://playground.kog.ai. Read the technical details on the Kog Labs blog https://blog.kog.ai. ## What you will work on You will imagine, design, and run experiments to understand how architectural decisions propagate through inference behavior, morph existing open-weight models into architecture variants optimized for speed, and turn findings into measurable gains in generation speed and model quality. - Design new model architecture variants, including routing strategies, attention mechanisms, and MoE structure, with execution constraints as a first-order design input. - Extend the Laneformer thesis by exploring inference-aware architectural variants such as DTP, Ladder Residual, and PT-Transformer, and finding what compounds at scale. - Own the post-training pipeline across fine-tuning, evaluation methodology, and adaptation of existing open-weight models toward architecture variants optimized for inference speed. - Scale the stack to large MoE models such as DeepSeek v4 and Qwen 3, working through routing, expert parallelism, and communication patterns at inference time. - Write up findings as research papers, submit them to top venues, and present them at conferences. - Contribute to building AI agents that will perform architecture research and training experiments autonomously, starting from the research foundations we are building now. ## What we look for - You have designed or changed model architecture, where the structure itself was the object of the work. Showing that work, a paper, a repository, or a thesis, is a requirement to move forward. - You reason about model design and hardware together, tracing how communication structure and layer dependencies shape inference behavior, with fluency in Transformers and MoE deep enough to weigh trade-offs. - Stronger signals include inference-aware architectural variants such as DTP, Ladder Residual, or PT-Transformer, and post-training methods such as fine-tuning, preference optimization, or quantization, including at research scale. - A top engineering school or a PhD with concrete architecture work counts, even without industry experience. ## What we offer - Direct access to AMD and NVIDIA datacenter GPUs from day one - A team where creativity and technical judgment carry weight and where the people closest to the problem shape the key decisions - Problems that sit on the critical path of model execution speed and that directly influence what the system can become - A remote-friendly working model, with one mandatory week per month in our Paris office. Travel and accommodation covered by the company. - Compensation aligned with top AI research profiles, including equity ## About Kog ## Company Overview - **One-liner**: Kog builds a realtime AI inference engine that achieves up to 3,000 tokens per second per request on standard datacenter GPUs through low‑level GPU engineering and model‑architecture co‑design. - **Entity Type**: Private (seed‑stage startup) - **Headquarters**: Paris, France - **Founded**: 2023 - **Founders**: Gaël Delalleau ## Core Business - **Primary industry**: AI infrastructure / GPU inference - **Target customers**: B2B – GPU providers, AI labs, enterprises deploying AI agents and agentic workflows - **Mission statement**: "Power every digital experience by providing a vertically integrated, realtime AI inference engine." ## Products & Services - **Kog Inference Engine (KIE)**: A drop‑in replacement for vLLM and TensorRT‑LLM. Uses a monokernel runtime (single persistent GPU program), custom collective‑communication library (KCCL), and hand‑crafted assembly‑level GPU code (CUDA/PTX, HIP/CDNA) – no dependency on PyTorch, Triton, or vendor libraries on the hot path. - **Kog LaneFormer**: A model architecture that implements Delayed Tensor Parallelism (DTP), overlapping cross‑device communication with computation to hide latency in multi‑GPU decoding. - **Kog Communication Library (KCCL)**: A latency‑optimized (<3 μs) inter‑GPU communication layer tuned at the assembly level for each GPU architecture, replacing RCCL/NCCL. ## Market Standing - **Valuation / Market Cap**: Not disclosed - **Key metric**: Total funding – $5M raised (Seed round led by Varsity VC with BPI France Deep Tech Program) - **Notable investors**: Varsity VC, BPI France; awarded the French Tech 2030 label (October 2025) - **Growth signals**: Launched public tech preview on AMD MI300X in May 2026 reaching 3,000 tok/s per request; 2,100 tok/s on NVIDIA H200; up to 3.5× faster than vLLM on AMD Instinct; team of 14 with 10 engineers/researchers (5 PhDs); press coverage from French Tech Journal. ## Competitive Advantages - **Hardware‑software co‑design**: Every layer is optimised for uninterrupted compute – monokernel eliminates kernel‑launch overhead and host‑side scheduling. - **Custom low‑level stack**: Avoids general‑purpose frameworks (PyTorch, Triton, CUTLASS, etc.) on the critical path; uses hand‑tuned assembly for each GPU. - **Delayed Tensor Parallelism (DTP)**: Novel architecture that hides all‑reduce communication behind computation, crucial for single‑request latency. - **Speed niche**: Matches dedicated inference‑silicon speeds on standard datacenter GPUs, without lock‑in to proprietary hardware. ## Strategic Focus - **Current priorities**: Optimising single‑request decode latency for AI agents; expanding support for larger models and batching; targeting sovereign‑AI buyers and enterprise GPU fleets. - **Direction**: Prove that software optimisations can unlock the full speed of existing GPU hardware, reducing reliance on custom chips. ## Why Work Here - **Culture**: High‑density team of PhDs and GPU research engineers who “treat hardware limits as a starting point.” Emphasis on deep technical challenges at the intersection of architecture, runtime, and kernel design. - **Team size**: ~14 people (as of mid‑2026), offering outsized impact per person. - **Locations**: Headquarters in Paris, with presence in Hungary and Portugal – likely hybrid/remote‑friendly. - **Perks**: Work on cutting‑edge inference technology with top investors (Varsity VC) and government recognition (French Tech 2030). The company is small enough that engineers shape the entire stack. ## Sources 1. [Kog.ai](https://www.kog.ai/) 2. [LinkedIn](https://fr.linkedin.com/company/kogai) 3. [Kog Blog – Real‑time LLM Inference on Standard GPUs (3,000 tok/s)](https://blog.kog.ai/real-time-llm-inference-on-standard-gpus-3-000-tokens-s-per-request/) 4. [French Tech Journal – AI Industry Spent Billions Chasing Faster Chips](https://www.frenchtechjournal.com/the-ai-industry-spent-billions-chasing-faster-chips-inference-startup-kog-says-they-were-solving-the-wrong-problem/) 5. [GitHub](https://github.com/kog-ai) ## Other roles at Kog - [GPU Engineer](https://feeny.ai/job/gpu-engineer-kog-paris-c71xtrnkrb6r) — Paris, France - [Research Engineer](https://feeny.ai/job/research-engineer-neo4j-london-60zcrspx4hjf) — London, United Kingdom - [Research Engineer](https://feeny.ai/job/research-engineer-greptile-san-francisco-a4grmc0eancw) — San Francisco, CA - [Research Engineer](https://feeny.ai/job/research-engineer-tessera-labs-san-jose-2atk17ayc2pm) — San Jose, CA - [Research Engineer](https://feeny.ai/job/research-engineer-verisign-reston-k0fceehptgcr) — Reston, VA - [Research Engineer](https://feeny.ai/job/research-engineer-antimetal-hq-new-york-ny-n112cvsmbp0c) — HQ / New York, NY - [Research Engineer](https://feeny.ai/job/research-engineer-decagon-san-francisco-fm6k4wz5aj59) — San Francisco, CA - [Research Engineer](https://feeny.ai/job/research-engineer-imc-sydney-dvh6etb3b0n3) — Sydney, Australia - [Research Engineer](https://feeny.ai/job/research-engineer-imc-hong-kong-33rrwf6g8nby) — Hong Kong, Hong Kong - [Research Engineer](https://feeny.ai/job/research-engineer-superannotate-ai-san-francisco-84vwmt7042sg) — San Francisco, CA