--- title: 'Inference Engineer at Designworks Talent' canonical: 'https://feeny.ai/job/inference-engineer-designworks-talent-bellevue-4q7758cdc709' type: 'job' last_seen: '2026-09-12' --- # Inference Engineer at Designworks Talent - **Company:** Designworks Talent - **Location:** Bellevue, WA - **Employment:** full-time - **Work type:** hybrid - **Posted:** 2026-07-21 - **Last confirmed live:** 2026-09-12 - **Apply:** https://jobs.ashbyhq.com/designworkstalent/a5540625-7ad0-458f-8110-d462447b714a ## Job description Inference Engineer Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles available) Build the Inference Platform Powering Next-Generation AI Applications ## About the Opportunity A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications. Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads. We're seeking Inference Engineers to build and operate the model-serving systems behind a next-generation AI inference platform. This team focuses on delivering high-throughput, low-latency, reliable inference experiences that enable customers to consume advanced AI capabilities through production-scale APIs. The Opportunity This is a foundational engineering role focused on building the systems that bring AI models from research environments into reliable production services. You'll work on the infrastructure layer responsible for serving large models efficiently, optimizing performance, and ensuring reliability as usage scales. You'll collaborate closely with GPU performance, AI training infrastructure, platform engineering, and operations teams to solve complex challenges around model serving, latency optimization, resource efficiency, and production reliability. This opportunity is ideal for engineers who enjoy working at the intersection of distributed systems, machine learning infrastructure, GPU computing, and large-scale production systems. ## What You'll Do - Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads. - Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads. - Design systems that maximize GPU utilization while maintaining predictable performance and reliability. - Improve the scalability and operational maturity of inference platforms as customer demand grows. - Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving. - Develop monitoring, alerting, and operational practices to maintain reliable inference services. - Investigate and resolve performance, reliability, and capacity challenges across inference workloads. - Contribute to architecture decisions and engineering standards as the platform evolves. ## What We're Looking For - Experience building and operating production machine learning inference or model-serving systems at scale. - Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency. - Experience designing reliable distributed systems or production infrastructure. - Understanding of GPU-backed AI workloads and the challenges of scaling inference systems. - Strong engineering fundamentals and the ability to independently own complex technical problems. - Comfortable working in a fast-moving environment where systems and processes are being built from the ground up. ## Preferred Qualifications - Experience with modern inference-serving frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar technologies. - Experience optimizing LLM inference workloads or large-scale AI serving platforms. - Background operating API-based AI products or high-volume production services. - Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms. - Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning. - Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization. ## Compensation - Competitive base pay for Bellevue market - Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance - U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays. Location - Hybrid role based in the Bellevue, WA area. - Approximately three days per week in the office. - Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply. - U.S. work authorization is required. Visa sponsorship is not currently available. Why Join? - Build the inference platform powering the next generation of AI applications. - Work directly on large-scale model serving, GPU optimization, and production AI systems. - Solve complex challenges around latency, throughput, reliability, and cost efficiency. - Join early enough to influence architecture, tooling, and engineering practices. - Collaborate with a highly experienced team building critical AI infrastructure from the ground up. - Enjoy the ownership and technical impact of a startup environment backed by significant long-term investment. ## About Designworks Talent ## Company Overview - **One-liner**: Designworks Talent provides custom, AI-enhanced recruiting solutions and enterprise workforce strategies to help companies hire smarter, faster, and with greater precision. - **Entity Type**: Private (Privately Held) - **Headquarters**: Los Angeles, California, United States - **Founded**: 2009 - **Founders**: Not publicly available ## Core Business - **Primary industry/industries**: Staffing and Recruiting, Talent Acquisition, Enterprise Workforce Solutions - **Target customers**: B2B; high-growth startups to global enterprises, including HR and Talent Acquisition teams - **Mission or purpose statement**: "Exceptional teams build exceptional companies. We combine human expertise with AI to help you hire smarter, faster, and with greater precision." [designworkstalent.com](https://www.designworkstalent.com/) ## Products & Services - **Recruiting Process Outsourcing (RPO)**: Scalable, end-to-end recruiting that acts as an extension of the client's team, managing part or all of the hiring process to reduce costs and improve candidate experience. [designworkstalent.com/Services](https://www.designworkstalent.com/Services) - **Recruiting Strategy & Consulting**: Design of systems and strategies for hiring, including process design, interview best practices, employer branding, tech stack selection, budget planning, and workforce forecasting. [designworkstalent.com/Services](https://www.designworkstalent.com/Services) - **Executive Search**: Confidential, precision-focused search for senior-level executives (C-suite, VP, Director roles) with an emphasis on cultural alignment. [designworkstalent.com/Services](https://www.designworkstalent.com/Services) - **Direct Hire Placement**: Full-time placements for core teams across engineering, operations, finance, HR, and more, focused on skill and cultural fit. [designworkstalent.com/Services](https://www.designworkstalent.com/Services) - **Fractional & Interim Talent**: Agile, flexible recruiting support for early-stage startups, scaling companies, or those navigating key transitions. [designworkstalent.com/Services](https://www.designworkstalent.com/Services) - **Early Stage Start-Up Solution**: End-to-end recruiting function tailored to tech startups, from sourcing and interviewing to closing critical hires. [designworkstalent.com/Services](https://www.designworkstalent.com/Services) ## Market Standing - **Valuation/Market Cap**: Not disclosed - **Key Metric**: Total Funding – Not disclosed (Privately held, bootstrapped or self-funded) - **Notable Investors/Partners**: Not publicly available - **Growth Signals**: Headcount grew by 50% year-over-year (from 2 to 3 employees) [LinkedIn](https://www.linkedin.com/company/designworks-talent-llc). Website traffic grew 69% month-over-month with 708 monthly visits [LinkedIn](https://www.linkedin.com/company/designworks-talent-llc). Employer rating of 5.0/5.0 based on 3 reviews [LinkedIn](https://www.linkedin.com/company/designworks-talent-llc). ## Competitive Advantages - **AI + Human Expertise**: Combines AI-driven tools with deep recruiting expertise to cut time-to-fill and deliver top talent quickly. [designworkstalent.com](https://www.designworkstalent.com/) - **Flexible Engagement Models**: Offers full-service searches, project-based hiring, and a pay-as-you-go model, providing clients with on-demand support without long-term commitments. [designworkstalent.com](https://www.designworkstalent.com/) - **Customized Solutions**: Every search is tailored to the client's goals, culture, and growth stage, avoiding a cookie-cutter approach. [designworkstalent.com](https://www.designworkstalent.com/) - **Agile Talent Strategies**: Delivers flexibility, speed, and continuous improvement, partnering with a range of organizations from startups to global enterprises. [designworkstalent.com](https://www.designworkstalent.com/) ## Strategic Focus - Scaling the business through a flexible, on-demand recruiting model that appeals to companies of all sizes. - Expanding into new markets and verticals by leveraging AI and data-driven insights to improve hiring efficiency. - Building long-term partnerships with clients by focusing on candidate experience, employer branding, and strategic workforce planning. ## Why Work Here - **Culture**: High employee satisfaction with a 5.0/5.0 rating across work-life balance, compensation, and culture [LinkedIn](https://www.linkedin.com/company/designworks-talent-llc). - **Work Environment**: Hybrid workspace model, with offices in Los Angeles, Boston, Chicago, and Palm Springs [Built In](https://builtin.com/company/designworks-talent-llc). - **Team**: Small, agile team (3 employees) offering a close-knit, entrepreneurial environment with significant ownership and impact. - **Growth**: Company is in a growth phase (50% headcount growth YoY), presenting opportunities for early employees to shape the company’s trajectory. ## Sources 1. [Designworks Talent Website](https://www.designworkstalent.com/) 2. [Designworks Talent Services Page](https://www.designworkstalent.com/Services) 3. [Designworks Talent LinkedIn Profile](https://www.linkedin.com/company/designworks-talent-llc) 4. [Designworks Talent Careers on Built In](https://builtin.com/company/designworks-talent-llc) 5. [Designworks Talent Job Openings](https://www.designworkstalent.com/postions) ## Other roles at Designworks Talent - [Chief Information Officer (CIO SVP/IT)](https://feeny.ai/job/chief-information-officer-cio-svp-it-designworks-talent-bellevue-pcvjkqnprhq3) — Bellevue, WA - [Director of Project Development - Forensic Investigations](https://feeny.ai/job/director-of-project-development-forensic-investigations-designworks-talent-tampa-hjhxc7ezktwv) — Tampa, FL - [Director of Project Development - Forensic Investigations](https://feeny.ai/job/director-of-project-development-forensic-investigations-designworks-talent-1gwxg1921g1j) — Dallas, TX - [Infrastructure Systems Engineer](https://feeny.ai/job/infrastructure-systems-engineer-designworks-talent-indianapolis-a6sy57ct4xft) — Indianapolis, IN - [Principal Data Center Infrastructure Software Engineer](https://feeny.ai/job/principal-data-center-infrastructure-software-engineer-designworks-talent-p3yej8s4j42x) — Bellevue, WA - [IT Consultant](https://feeny.ai/job/it-consultant-designworks-talent-indianapolis-chm1sqbbvs8m) — Indianapolis, IN - [Virtualization & Orchestration Engineer](https://feeny.ai/job/virtualization-orchestration-engineer-designworks-talent-bellevue-wtp6p2exdscf) — Bellevue, WA - [GPU Performance / Kernel Engineer](https://feeny.ai/job/gpu-performance-kernel-engineer-designworks-talent-bellevue-fvqk4w1cf1x2) — Bellevue, WA - [AI Training Infrastructure Engineer](https://feeny.ai/job/ai-training-infrastructure-engineer-designworks-talent-bellevue-w25acyb6f2gr) — Bellevue, WA - [Hardware Design Engineer](https://feeny.ai/job/hardware-design-engineer-designworks-talent-bellevue-bdbze54pjjfe) — Bellevue, WA