Title: Founding Research Engineer - RL and Post Training
Location: San Francisco, CA (Onsite)
Company Description
Our client, a leader in the AI-native healthcare and drug discovery space, is looking to hire a Founding Research Engineer for RL and Post Training.
They're building the data layer that turns messy real-world clinical workflows into AI-ready products for leading AI labs, human data companies, and frontier biotech teams. The long-term vision is linking real-world healthcare data with genomics, imaging, biomarkers, and experimental data to make high-quality healthcare accessible to everyone and radically accelerate drug discovery.
They're backed by PeakXV, Y Combinator, Afore Capital, SV Angel, plus angels from OpenAI, Meta, and Google DeepMind.
Responsibilities:
As an RL Engineer, you'll build reinforcement learning environments and post-training systems for healthcare AI. This means sourcing high-value clinical data, turning it into model-ready workflows, and building tasks, rewards, verifiers, benchmarks, and agent environments where models can learn against meaningful and measurable outcomes.
You'll work across the full RL loop, from environment and reward design to training, evaluation, and iteration. Projects may span clinical reasoning, longitudinal patient care, diagnostic decision-making, chronic disease management, and biomedical research. This is a hands-on engineering role: you'll build environments, run experiments, train models and agents, analyze failures, and improve the data and feedback signals that determine what models learn.
- Build healthcare-specific RL environments, including tasks, action spaces/tool interfaces, reward functions, verifiers, and evaluation harnesses
- Run post-training experiments on language models and agents using techniques like SFT, RLVR, RLHF/RLAIF, and reward modeling
- Turn clinical and biomedical datasets into training environments with measurable, verifiable outcomes
- Design rewards and verifiers that capture correctness across clinical reasoning and longitudinal decision-making tasks
- Train and evaluate multi-step agents operating across patient histories, clinical tools, and structured/unstructured medical data
- Build scalable pipelines for rollouts, training, evaluation, experiment tracking, and dataset iteration
- Analyze model failures and use them to improve environments, rewards, datasets, and subsequent training runs
Is this you?
- Excited to apply frontier RL methods to healthcare, medicine, and biological data
- Experience with reinforcement learning, LM post-training, agent environments, reward modeling, evaluation, or related ML systems
- Strong judgment on data quality (signal, label fidelity, coverage, longitudinal depth, clinical relevance)
- Can move fast from research concept to working prototype, then iterate on empirical results
- Comfortable designing controlled experiments, building baselines, and drawing trustworthy conclusions from noisy real-world data
- Comfortable in large ML codebases, debugging training runs, data pipelines, eval harnesses, and model behavior
- A self-starter who owns ambiguous problems and drives projects to completion
- Thrives where research, engineering, product, and customer needs intersect
Preferred Qualifications:
- 2+ years of experience in reinforcement learning, post-training, agents, or related ML systems
- Prior healthcare experience not required
Perks:
- Base salary: $200,000–$350,000
- Equity: 0.1%–1.0%
- 401(k) with company match
- Medical, dental, and vision insurance
- Complimentary lunch daily at the office
- Unlimited budget for AI tools and software
- Relocation and joining bonus available based on role and circumstances
Equal Opportunity:
Our client is an equal opportunity employer, committed to building a diverse and inclusive team. All qualified applicants will receive consideration for employment regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or any other characteristic protected by applicable law.