
Member of Technical Staff, RL Infra at Inception (Bay Area)
Inception· Bay Area· $200k–$350k·
Role details
Salary
$200k–$350k
Work type
Onsite
Employment
Full-Time
Equity
Yes
Job description
The Role
We're looking for engineers and scientists to design, optimize, and maintain the core systems that enable scalable, efficient reinforcement learning for large models. This role sits at the intersection of research and large-scale systems engineering: you'll wear many hats, from optimizing rollout and reward pipelines to enhancing reliability, observability, and orchestration, collaborating closely with researchers to make RL stable, fast, and production-ready.
Key Responsibilities
- Design, build, and optimize the infrastructure that powers large-scale reinforcement learning and post-training workloads.
- Improve the reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
- Develop shared monitoring and observability tools to ensure high uptime, debuggability, and reproducibility for RL systems.
Qualifications
- BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
- Understanding of ML frameworks (PyTorch, TensorFlow, Ray, Megatron) from a systems perspective.
- Experience working with reinforcement learning workloads (PPO, DPO, RLHF, or reward modeling).
- Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.
Preferred Skills
- Experience building and maintaining large-scale language models with tens of billions of parameters or more.
- Experience with ML workflow orchestration tools (Kubeflow, Airflow).
- Background in performance optimization and profiling of ML systems.
Why work at Inception
- Culture: A tight‑knit team of scientists, engineers, and builders focused on shipping frontier AI research. Emphasis on innovation, speed, and real‑world impact.
- Work policy: In‑office for most roles (Palo Alto HQ); a Marketing Intern position is listed as remote. Typical time on‑site is full‑time in the office.
- Notable perks / engineering culture: Opportunity to work on cutting‑edge diffusion LLMs from scratch, contribute to foundational research (papers published), and collaborate with alumni from top AI labs. Roles span from kernel engineering to RL infrastructure and product management.
- Growth: Rapidly scaling headcount (+187% YoY) with open positions across multiple disciplines, offering strong career progression in a high‑visibility startup.