All open roles

Research Engineer

Engineering · Posted 18 days ago

The largest obstacle to useful current AI, and well-targeted ASI, is misalignment. To that end, we develop high-quality evals for misaligned behavior. Frontier labs use our evals to test & train their models propensities for reward hacking, and measure the efficacy of their own AI control measures.

Way of working
Hybrid
Location
San Francisco, CA, Open to relocation
Pay range
$200,000 to $400,000
Level
Junior
Experience
1+ years
Type
Full time
Visa sponsorship
Not offered for this role
The company
Computer and Network Security · under 20 people

Skills that matter here

Python

The full description

About Goodhart Labs

The largest obstacle to useful current AI, and well-targeted ASI, is misalignment. To that end, we develop high-quality evals for misaligned behavior. Frontier labs use our evals to test & train their models propensities for reward hacking, and measure the efficacy of their own AI control measures.

Goodhart Labs is the applied AI arm of ZeroPath. ZeroPath is itself an application security company. It was a Top 10 finalist in RSAC’s 2026 Innovation Sandbox and has raised $12.5M from Y Combinator, HOF Capital, SurgePoint Capital, Crosspoint Capital, and Paul Graham. The team is 12 people and based in San Francisco.

What you’ll do

Research Engineers at Goodhart independently author environments that elicit misbehavior from frontier models, and manage & improve the base of software we use to isolate these behaviors. A single Research Engineer owns the entire environment engineering process, including initial ideation, hack identification, grader iteration, measurement, and refinement.

Increasingly, authoring these environments primarily involves interacting with LLMs. Apart from specific parts of the process, you’ll spend most of your time prompting agents to perform tasks on your behalf, reviewing their work, and making qualitative design decisions that frontier agents are incapable of performing. The environments we author are diverse, and to make these decisions effectively, you must be able to pick up new domains relatively quickly. Practically, the most important skill is interacting with agents efficiently and directing them according to their capabilities.

Because many of the behaviors we look for are revealed only at the frontier of model capabilities, the job involves both authoring high-quality, long-horizon environments and thinking critically about what kinds of misalignment you can elicit from frontier models.

A good fit tends to:

- Be serious about making recursive self-improvement go well

- Have some traditional SWE experience

- Display strong conceptual and moral reasoning abilities

- Pick up new domains relatively quickly

- Direct agents efficiently and according to their capabilities, and tell when their work is subtly wrong

- Be comfortable owning an environment end to end, with the accountability that implies

How we work

While authoring an environment is largely solo work, we make a point of spending time together to share learnings and context regularly, with daily team meals, meetings, and weekly check-ins. If you’re remote, you’ll get feedback on every environment you submit through review, and we check in daily for new hires. If you’re in San Francisco, you’d be working at the office for at least 40 hours a week, split however suits you. We care most about the quality of what you ship.

We do not create evals or RL environments that we think differentially advance AI capabilities. If you are interested in doing such work, there are many other companies that do that.

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you