We are looking for AI Safety Practitioner candidates for a project delivered through Mercor.
What you'll do
- Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
- Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains.
- Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking.
- Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
- Provide structured feedback to improve model alignment and safety performance.
- Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.
What you need
- Excellent written English, critical thinking, and analytical reasoning skills.
- Ability to consistently evaluate nuanced and policy-sensitive scenarios.
Nice to have
- Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation.
- Familiarity with safety policies, content moderation, or evaluation rubric development.
- Experience reviewing complex, high-risk, or ambiguous content.
Who you work with
Project and contracting process: Mercor. Applications continue on the provider's website.

