We are looking for AI Safety Experts — English & Dutch candidates for a project delivered through Mercor.
What you'll do
- Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
- Apply structure by following taxonomies, benchmarks, and playbooks to ensure testing consistency.
- Document reproducibly by producing reports, datasets, and attack cases that customers can act on.
What you need
- Prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing).
- Curiosity and an adversarial mindset to push systems to breaking points.
- Structured approach using frameworks or benchmarks.
- Communication skills to explain risks clearly to technical and non-technical stakeholders.
- Adaptability to thrive on moving across projects and customers.
Nice to have
- Specialty in Adversarial ML, including jailbreak datasets, prompt injection, RLHF/DPO attacks, and model extraction.
- Specialty in Cybersecurity, including penetration testing, exploit development, and reverse engineering.
- Specialty in Socio-technical risk, including harassment/disinformation probing, abuse analysis, and conversational AI testing.
- Specialty in Creative probing, utilizing psychology, acting, or writing for unconventional adversarial thinking.
Expertise
Who you work with
Project and contracting process: Mercor. Applications continue on the provider's website.

