Retool

Software Engineer, Agent Platform at Retool (San Francisco, CA)

Retool· San Francisco, CA· $164k–$306k·

Role details

Salary
$164k–$306k
Work type
Hybrid
Employment
Full-Time
Equity
Yes
Commission
Yes

Job description

WHY WE’RE LOOKING FOR YOU

We’re building AI-native products where model behavior is part of the product, not just an implementation detail. As LLMs become more capable, the bottleneck is no longer access to models—it’s making non-deterministic systems reliable, evaluable, and trustworthy in production.

We’re hiring AI Engineers to own that problem end to end. This is not a role for someone who simply integrates AI APIs into features. It’s for engineers who take responsibility for how probabilistic systems behave over time, how quality is measured in the presence of variance, and how capabilities improve without regressing.

If model regressions, subtle behavior drift, or edge-case failures keep you up at night—and you enjoy that kind of ownership—we want to talk.

WHAT YOU’LL DO

As an AI Engineer, you’ll own model-driven behavior in production systems, working across product, infrastructure, and evaluation layers. Your work will directly shape what users experience—and how confidently the team can ship. You might:

  • Own the behavior of AI-powered features across multiple product surfaces, including quality, safety, variance, and failure modes
  • Design and evolve prompting, retrieval, routing, and tool-use strategies that embrace non-determinism while bounding its downside
  • Build and maintain evaluation systems that measure model performance using statistical signals, distributions, and trends—not just pass/fail tests
  • Detect, diagnose, and resolve non-deterministic failures such as hallucinations, partial correctness, instruction drift, or sensitivity to context changes
  • Define and implement guardrails, fallbacks, and degradation paths that keep systems useful even when models behave unexpectedly
  • Partner with product and infra teams to decide when probabilistic behavior is “good enough” to ship—and when it isn’t
  • Influence model selection, model behavior, and tool design to balance quality, cost, latency, and robustness for real user workflows

You’ll work across the stack (e.g., TypeScript, Node.js, React), but your leverage won’t come from code volume alone—it will come from shaping runtime behavior with precision, measurement, and intent.

WHAT THIS ROLE IS (AND IS NOT)

This role is:

  • Accountable for AI behavior, not just system correctness
  • Grounded in evaluation, iteration, and regression prevention under non-determinism
  • Comfortable designing systems where outputs vary, confidence is probabilistic, and correctness is contextual
  • Focused on shipping dependable products on top of imperfect components

This role is not:

  • Adding LLM calls to existing features and moving on
  • Treating models as black boxes with undefined behavior
  • Shipping AI features without owning their long-term reliability, drift, or user trust

THE SKILLSET YOU’LL BRING

  • 6+ years of professional engineering experience, with ownership over complex systems in production
  • Demonstrated experience owning AI/LLM behavior beyond basic integration, including mitigation of variance and failure modes
  • Comfort reasoning about probabilistic systems and tradeoffs (quality vs. cost, recall vs. precision, speed vs. robustness)
  • Experience designing or maintaining evaluation frameworks, golden datasets, regression detection, or human-in-the-loop feedback loops
  • Strong product intuition—you care deeply about what “good” looks like even when outputs are non-deterministic
  • Ability to operate independently in ambiguous problem spaces and set quality standards others rely on
  • Strong opinions, weakly held—you iterate quickly and adjust based on evidence and observed runtime behavior

BONUS POINTS

  • Experience with RAG, agentic systems, or tool-using models in production
  • Familiarity with vector databases, embeddings, or retrieval pipelines
  • Exposure to fine-tuning, model routing, or post-training techniques
  • Experience building shared AI infrastructure used by multiple teams
  • History of mentoring engineers on designing for non-determinism and evaluation-driven development

WHO YOU’LL WORK WITH

You’ll join a small, senior team focused on advancing AI capabilities across the product. You’ll collaborate closely with product engineers, infra engineers, designers, and PMs—often acting as the final owner of AI behavior and quality before features reach users.

Your work will set standards that others build on. If you enjoy being the person teams rely on when AI behavior matters most—and certainty is never guaranteed—you’ll thrive here.

READY TO BUILD RELIABLE AI SYSTEMS?

If you’re excited to move beyond demos and take real ownership of non-deterministic behavior in production—defining quality, preventing regressions, and turning variability into a strength—we’d love to meet you.

Why work at Retool

  • Culture: The company values ambition, intense curiosity, energy, and deep care. They emphasize moving fast, acting like an owner, and being both demanding and supportive.
  • Work Environment: Hybrid model with offices in San Francisco, New York, Salt Lake City, and London. The company believes great work happens when teams collaborate in person.
  • Product Impact: Employees are close to the product and use Retool internally, giving them a direct voice in improving it. The mission is to enable a wider community of builders to create production-grade software safely.
  • Hiring Process: Structured, with initial recruiter/hiring manager chats, a technical evaluation or practical exercise, and a final round of 3–5 interviews. The company aims for transparency and candidate preparation.
  • Employee Sentiment: Employer rating of 3.4/5.0 (127 reviews). Strengths include Work-Life Balance (4.1) and Culture (3.7). Compensation (3.4) and Career (3.4) are areas for potential improvement.

Application questions