Elicit

ML Research Resident at Elicit (Oakland, CA)

Elicit· Oakland, CA·

Role details

Work type
Remote
Employment
Contract

Job description

Elicit is building a research agent that can use an unlimited amount of test-time compute while keeping its reasoning transparent and verifiable.

The residency Transformers do a fixed amount of computation per token, and the quality of work degrades rapidly when they are applied iteratively. As research resident, you'll work with us for 3 months on developing computational procedures (operators) that can reliably improve a knowledge state over thousands of iterations. What is a knowledge state? A knowledge state consists of structured information - for example, a scientific paper might be represented as a set of claims supported by evidence and connected through logical reasoning; this might be combined with scratchpads, evergreen “notes to self”, search trees, and other information. What counts as improvement? Like scientists, we want LLMs to make genuine progress in understanding - separating inferences from raw evidence, finding connections between ideas, building clearer explanations, and identifying gaps in reasoning. But unlike typical ML systems that are often trained to do “whatever works”, we need improvements that are epistemically sound - each step should make the knowledge state more useful while remaining human-readable. An improvement might reorganize information to better answer a question, find an implicit assumption in an argument, or connect evidence across multiple sources. As research resident, your work will focus on designing and testing improvement operators that maintain stability over 1000+ iterations while making genuine progress. You'll start with simple cases (e.g., shallow refactoring of scientific papers) and demonstrate reliable iteration before scaling to more complex reasoning tasks. Developing systems that perform legible reasoning over long horizons addresses core challenges in AI transparency and scalable reasoning.

About you

Strong candidates will have experience with LLMs, good intuitions about what makes reasoning systematic and verifiable, and care about AI transparency. The best applicants will additionally have a strong software engineering background and concrete examples of how they've applied this background to come up with novel abstractions that push the frontiers of automated reasoning.

Logistics

  • 3-month contract role
  • Compensation: $12-15k/month depending on experience
  • Location: In-person (Oakland) or remote (US)
  • Potential of full-time offer for exceptional candidates

Location and travel We have a great office in Oakland, CA, and we'd love to see you there if you're local. That said, we're just as happy for you to work remotely. We do get the whole team together for a quarterly retreat somewhere fun, because in-person time matters to us.

Why work at Elicit

  • Culture: High‑agency, low‑bureaucracy environment where “everyone takes ownership and avoids egos and status games.” The team values transparency, truth‑seeking, and continuous growth. Cross‑team collaboration is the norm elicit.com.
  • Remote/Hybrid Policy: Flexible – team members can work from the Oakland office or remotely within time zones between GMT and GMT‑8. Quarterly in‑person company offsites are expected (travel 4+ times per year). Minimal meetings and no micromanagement elicit.com.
  • Compensation & Benefits:
    • Competitive salaries (e.g., Senior Software Engineer $185‑270K + equity).
    • Full coverage health, dental, vision, and life insurance for employee and generous family coverage.
    • 20 weeks paid parental leave (12 weeks fully paid).
    • 401K with 6% employer match.
    • $1,000 quarterly AI experimentation & learning budget.
    • Flexible PTO with recommended minimum 20 days/year.
    • New Mac + $1,000 workstation budget in first year, $500 annually thereafter.
    • Team administrative assistant available for personal/work tasks elicit.com.
  • Engineering Culture: Engineers work directly with LLM experts; the company predates the AI bubble, offering deep learning opportunities. The stack is ML‑heavy, and engineers ship user‑facing features weekly. Roles include ML Engineer, Frontend Engineer, AI Engineer, and Evaluation Engineer.
  • Potential Downsides: High ambiguity as an early‑stage startup; projects shift frequently. Self‑direction is expected; close mentorship is limited. Limited management track (few direct reports) – the team is intentionally kept small.

Application questions