What You’ll Do
- Develop risk monitoring for conversational, recommender, and tool-using AI systems.
- Design safety measurements for single- and muspolti-turn features, including production monitoring metrics, risk-specific decision criteria, and model performance metrics.
- Build reusable Python pipelines, LLM-as-a-judge workflows, and reporting dashboards.
- Work directly with Product, Engineering, and Trust & Safety to translate insights into safety mitigations and improvements.
- Communicate results clearly to technical and non-technical stakeholders.
- Full-time availability is preferred, although part-time arrangements may be considered.
Who You Are
- You have personally delivered safety evaluations or mitigations for a real AI or machine-learning product.
- You have strong agentic coding and data analysis skills, with sufficient Python and SQL to independently judge code and queries.
- You have experience designing evaluation datasets, rubrics, and metrics.
- You are comfortable structuring ambiguous problems and defining success criteria, methodologies, and tradeoffs with stakeholders.
- You have experience working cross-functionally across multiple domains, including research, engineering, product, policy, or Trust & Safety.
- You communicate clearly in writing and have a record of turning research findings into action.
It’s a Plus If You Have
- Experience calibrating LLM judges or building human-in-the-loop evaluations.
- Experience evaluating multi-turn or tool-using agents.
- Experience with multilingual evaluation.
Target profile
More of a traditional Data Scientist / Data Analyst profile with AI safety exposure, rather than a hardcore AI engineer or research scientist.
Candidates who have worked specifically on AI safety, measurement, monitoring, or risk-related analytics would be particularly relevant.
Core focus
Focused on measurement and monitoring of AI safety.
Defining what is safe/unsafe, training LLM judges to detect specific risks, calculating metrics, and turning those metrics into dashboards/reports.
The ultimate goal is to provide stakeholders with data to support decision-making around the platform.
Python / SQL
Primarily used for research, analysis, reporting, and troubleshooting existing pipelines.
Candidates should be able to understand and debug existing pipelines, but are not expected to design data pipelines from scratch.
Claude/Codex may be used to support some data engineering tasks, but this is not a core data engineering role.
Target backgrounds
Strong preference for candidates from companies already operating AI systems at scale, such as Google or Meta.
AI startups specifically focused on safety are also relevant, although these profiles may be less common.
Candidates coming from companies with little/no exposure to AI safety are less likely to be a strong match.
Seniority
Flexible on seniority, with no upper limit.
More senior candidates can be considered if they are still interested in hands-on work.
The scope can be adjusted based on the candidate’s level and experience.
Location / Time zone
The team primarily works on Eastern Time due to significant collaboration with Europe.
Pacific Time candidates can be considered if they are willing to work East Coast hours.
For Pacific candidates, the preference would be for someone more senior who can operate with greater independence due to the limited overlap.
Distinction from AI Safety Engineer / Research roles
This role is more focused on measurement, monitoring, metrics, and reporting.
The other AI Safety roles are more focused on AI/safety research and engineering, working on more open-ended and ambiguous problems.
AI Engineers can still be a fit if their actual experience aligns with the measurement/analytics focus of this role.