--- title: 'Software Engineer, Inference Runtime at LM Studio' canonical: 'https://feeny.ai/job/software-engineer-inference-runtime-lm-studio-new-york-4sp2mefb88j5' type: 'job' last_seen: '2026-09-10' --- # Software Engineer, Inference Runtime at LM Studio - **Company:** LM Studio - **Location:** New York, NY - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-08-07 - **Last confirmed live:** 2026-09-10 - **Apply:** https://jobs.ashbyhq.com/lm-studio/ba7b2c29-c2ad-4ad2-98f4-a8ba373b34e0 ## Job description LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family. As a team, we work with high technical intensity and personal responsibility. We are looking for curious, self-motivated, creative, and technically excellent teammates to join us and build the future of human-AI interactions in software. ## The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on. ## Qualifications - Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure - Strong programming ability in Python and C++ - Deep understanding of transformer architectures and the mechanics of model inference - Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement - Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM - Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution - Takes personal responsibility for the correctness and performance of their work Bonus Qualifications - Past contributions to open-source inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM ## Responsibilities - Maintain and push forward our inference stack on-device and in the cloud - Bring up new model architectures and multimodal models - Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes - Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution - Benchmark and diagnose correctness and performance problems across the inference stack - Contribute upstream to open-source projects such as llama.cpp and MLX ## Benefits - Competitive salary and equity grants - Great medical, vision, dental healthcare plans - Catered team lunch / expensed dinners in the office - Flexible PTO - Flexible WFH - Sun-drenched office in SoHo in NYC ## About LM Studio ## Company Overview - **One-liner**: LM Studio provides a local AI toolkit that enables users to discover, download, and run open-source large language models (LLMs) on their own hardware, entirely offline. - **Entity Type**: Private (Venture-backed, Unattributed VC) - **Headquarters**: Brooklyn, New York, United States - **Founded**: 2023 - **Founders**: Yagil Burowski ## Core Business - **Primary industry**: Artificial Intelligence / Developer Tools - **Target customers**: B2B and B2C; individual developers, data scientists, and enterprises seeking private, local AI inference. - **Mission or purpose statement**: To make local AI models more accessible and allow users to run LLMs on their own hardware, privately and for free. ## Products & Services - **LM Studio Desktop App**: The core product, a free desktop application for macOS, Windows, and Linux that allows users to discover, download, and run open-source LLMs locally. Features include chatting with documents, a built-in API server, and support for agentic tool calling. [lmstudio.ai](https://lmstudio.ai/) - **LM Studio Enterprise**: A paid offering for organizations that need to deploy local AI at scale, with additional management and security features. [lmstudio.ai](https://lmstudio.ai/enterprise) - **LM Link**: A service launched in partnership with Tailscale that allows users to access their local models remotely from anywhere, with end-to-end encryption. [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) - **SDKs & CLI**: Developer tools including `lmstudio-js` (JavaScript SDK), `lmstudio-python` (Python SDK), and `lms` (CLI) for integrating LM Studio into applications and workflows. [lmstudio.ai](https://lmstudio.ai/) - **llmster (Headless Deployment)**: A no-GUI version of LM Studio's core for deploying on servers, cloud instances, or in CI/CD pipelines. [lmstudio.ai](https://lmstudio.ai/) ## Market Standing - **Valuation/Market Cap**: Not publicly disclosed. - **Key Metric**: Total funding of **$19.32M** raised as of a venture round in June 2025. [cbinsights.com](https://www.cbinsights.com/company/lm-studio) - **Notable Investors/Partners**: FirstMark Capital, Matrix Partners, Torch Capital, Preston-Werner Ventures. [cbinsights.com](https://www.cbinsights.com/company/lm-studio) - **Growth Signals**: The company has grown headcount by **+100% YoY** to 12 employees. Website traffic has grown **+187.9% yearly**. The company has formed key partnerships with Tailscale (for LM Link) and was featured in Apple's M5 announcement. [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) ## Competitive Advantages - **Privacy-First Local Inference**: Unlike cloud-based AI services, LM Studio runs entirely on the user's hardware, ensuring data never leaves the machine. This is a critical differentiator for privacy-conscious enterprises and developers. - **Broad Model Support**: Supports a wide range of open-source model formats (GGUF, MLX) and architectures, making it a versatile hub for the open-source LLM ecosystem. - **Cross-Platform & Apple Silicon Optimization**: Deeply optimized for Apple Silicon (MLX) and NVIDIA GPUs (CUDA), providing strong performance on consumer hardware. The company was featured in Apple's M5 announcement. [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) - **Developer-First Ecosystem**: Offers a comprehensive suite of developer tools (SDKs, CLI, REST API) and supports industry standards like OpenAI's API format, making it easy to integrate into existing workflows. ## Strategic Focus - **Headless & Enterprise Expansion**: The launch of `llmster` (headless deployment) and the Enterprise tier signals a push beyond individual developers into server and enterprise markets. [lmstudio.ai](https://lmstudio.ai/) - **Remote Access & Networking**: The LM Link partnership with Tailscale aims to make local models accessible from anywhere, bridging the gap between local privacy and remote convenience. [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) - **Ecosystem & Partnerships**: Deepening integrations with major AI players (Google DeepMind, NVIDIA, OpenAI) and hardware partners (Apple, NVIDIA DGX) to remain the go-to local inference platform. [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) ## Why Work Here - **High-Impact, Small Team**: With only 12 employees, new hires will have significant ownership and influence over the product and direction. The team is growing rapidly (+100% YoY). [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) - **Cutting-Edge AI Work**: The engineering team works directly on the frontier of local AI, contributing to open-source projects like `llama.cpp`, `mlx-lm`, and `mlx-engine`. [lmstudio.ai](https://www.lmstudio.ai/careers) - **Global & Remote-Friendly**: While headquartered in Brooklyn, NY, the company has team members in Uruguay and Turkey, suggesting a remote or hybrid-friendly culture. [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) - **Open Source Culture**: The company actively contributes to and builds on open-source AI infrastructure, making it an attractive place for engineers who value transparency and community. - **Current Open Roles**: As of the latest data, LM Studio is hiring for a **Frontend Software Engineer** and a **Systems Engineer (C++)**, indicating a focus on both the user-facing app and the core AI runtime. [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) ## Sources 1. [lmstudio.ai](https://lmstudio.ai/) 2. [lmstudio.ai/careers](https://www.lmstudio.ai/careers) 3. [cbinsights.com](https://www.cbinsights.com/company/lm-studio) 4. [linkedin.com](https://www.linkedin.com/company/lmstudio-ai) 5. [jobs.ashbyhq.com](https://jobs.ashbyhq.com/lm-studio) ## Other roles at LM Studio - [Full Stack Software Engineer](https://feeny.ai/job/full-stack-software-engineer-lm-studio-new-york-hswhkeskpnen) — New York, NY - [Software Engineer, Agent Harness](https://feeny.ai/job/software-engineer-agent-harness-lm-studio-new-york-fjvw9gwp8djy) — New York, NY - [GTM, Enterprise](https://feeny.ai/job/gtm-enterprise-lm-studio-new-york-g4p593zmzqbq) — New York, NY - [Software Engineer, Application](https://feeny.ai/job/software-engineer-application-lm-studio-new-york-nsj36x3yeen1) — New York, NY