We are looking for Software Engineer – AI Code Evaluation & Benchmarking (US candidates only) candidates for a project delivered through Turing.
What you'll do
- Review and evaluate AI-generated code for correctness, efficiency, maintainability, and adherence to requirements.
- Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.
- Debug code, reproduce issues, and verify fixes across different programming environments.
- Assess model-generated explanations, reasoning, and implementation approaches for technical accuracy.
- Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
- Identify edge cases, failure modes, and areas where AI systems struggle with software engineering problems.
- Document findings clearly and provide structured feedback to improve evaluation quality and consistency.
- Collaborate with project teams to establish quality standards and evaluation methodologies.
What you need
- Strong understanding of data structures, algorithms, software design principles, and debugging methodologies.
- Experience performing code reviews and evaluating code quality in production or large-scale codebases.
- Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
- Familiarity with version control systems (e.g., Git) and modern software development workflows.
- Strong written communication skills and attention to detail.
Nice to have
- Experience with AI/ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects.
- Experience evaluating AI-generated code, benchmark creation, or software quality assessment.
Expertise
Who you work with
Project and contracting process: Turing. Applications continue on the provider's website.

