
AI Infrastructure Engineer at NIO USA, INC (San Jose, CA)
NIO USA, INC· San Jose, CA· $192k–$250k·
Role details
Job description
About NIO
NIO Inc. is a pioneer and a leading company in the global smart electric vehicle market. Founded in November 2014, NIO aspires to shape a sustainable and brighter future with the mission of “Blue Sky Coming”. NIO envisions itself as a user enterprise where innovative technology meets experience excellence. NIO designs, develops, manufactures and sells smart electric vehicles, driving innovations in next-generation core technologies. NIO distinguishes itself through continuous technological breakthroughs and innovations, exceptional products and services, and a community for shared growth. NIO provides premium smart electric vehicles under the NIO brand, premium smart electric vehicles for families through the ONVO brand, and small smart high-end electric cars with the FIREFLY brand.
Roles and Responsibilities
- Design and implement high-performance, scalable inference systems for LLMs and VLMs across cloud, edge, and edge-cloud hybrid platforms.
- Develop and optimize custom kernels and operators for specific hardware accelerators (GPU, NPU, DSP, etc.), improving throughput, latency, and memory efficiency.
- Integrate advanced optimization techniques such as KV-cache management, tensor/model parallelism, quantization, and memory-efficient execution into production inference systems.
- Partner with system and hardware teams to ensure tight hardware-software integration and optimal performance across diverse compute environments.
- Translate architectural requirements into robust, maintainable, production-ready software that meets performance, safety, and reliability standards.
- Define and drive the evolution roadmap for LLM/VLM inference in the AIOS stack, ensuring scalability and adaptability to new workloads.
- Stay ahead of industry trends and competitor solutions, applying best practices from both AI and large-scale systems engineering.
Must Qualifications
- 5+ years of hands-on software development experience in building and optimizing AI inference systems at scale.
- Direct experience in LLM/VLM model internals, including Transformer-based architectures, inference bottlenecks, and optimization techniques.
- Strong expertise in performance engineering: kernel development, parallelism strategies, memory optimization, and distributed inference systems.
- Proficiency with GPU/NPU programming (CUDA, or vendor-specific SDKs), compiler toolchains, and deep learning frameworks (PyTorch, or TensorFlow).
- Strong programming skills in C/C++, with a track record of delivering high-performance, production-grade software.
- Solid foundation in computer architecture, systems programming (CPU/GPU pipelines, memory hierarchy, scheduling), and embedded systems.
- BS/MS in Computer Science, Computer Engineering, or related technical field.
- Excellent communication and collaboration skills, with the ability to work across cross-functional teams.
Preferred Qualifications
- Master’s or PhD degree in Computer Science, Electrical/Computer Engineering, or related fields, plus 5 years industry experience
- Experience building inference serving systems for large models, including batching, scheduling, caching, and load balancing.
- Expertise in hardware-aware model optimization (e.g., kernel fusion, mixed precision, quantization, pruning).
- Familiarity with edge and embedded AI, including real-time constraints and limited-resource optimization.
- Contributions to widely used AI frameworks, libraries, or performance-critical software (open source or proprietary).
Compensation
- Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.
- Please note that the compensation details listed in US role postings reflect the base salary only. It does not include discretionary bonus, equity, or benefits.
Benefits:
Along with competitive pay, as a full-time NIO employee, you are eligible for the following benefits on the first day you join NIO:
- Anthem Blue Cross, HSA, and Kaiser HMO medical plans with $0 for Employee Only Coverage.
- Dental (including orthodontic coverage) and vision plan. Both provide options with a $0 paycheck contribution covering you and your eligible dependents.
- Company Paid HSA (Health Savings Account) Contribution when enrolled in the High Deductible Anthem Blue Cross medical plan
- Healthcare and Dependent Care Flexible Spending Accounts (FSA)
- 401(k) with Brokerage Link option
- Company paid Basic Life, AD&D, short-term and long-term disability insurance
- Employee Assistance Program
- Sick and Vacation time
- 13 Paid Holidays a year
- Paid Parental Leave for first 8 weeks at full pay (eligible after 90 days of employment with NIO)
- Paid Disability Leave for first 6 weeks at full pay (eligible after 90 days of employment with NIO)
- Voluntary benefits including: Voluntary Life and AD&D options for you, your spouse/domestic partner and dependent child(ren), pet insurance
- Commuter benefits
- Mobile Cell Phone Credit
- Free lunch and snacks
- Onsite gym
- Employee discounts and perks program
Why work at NIO USA, INC
- Innovative environment: Work on cutting-edge EV, autonomous driving, AI, and battery technology at a company that challenges traditional automotive.
- Benefits: Health/dental/vision insurance, FSA/HSA, 401K with RSUs and annual bonus, paid vacation, sick leave, maternity/paternity leave, flexible hours, remote work options, free lunch/snacks, gym discount, commuter program, and mobile phone allowance.
- Global presence: Opportunity to work in San Jose (US), Munich, Oslo, Budapest, Oxford, Netherlands, and other locations.
- Engineering culture: Emphasis on R&D – roles include Hypervisor/Linux Kernel Developer, Simulation Software Engineer, AI Infrastructure Engineer, LLM Algorithmic Optimization Engineer, and AI Robotics Researcher.
- Award-winning: BaaS recognized by Fast Company; NOMI AI won 2021 Artificial Intelligence award; Firmware-Over-The-Air won 2021 DEVIES Award.