We are looking for AI Quality Analyst (Personalization) - Vietnamese candidates for a project delivered through Turing.
What you'll do
- Evaluate a new personalization feature for Gemini, assessing how well the model uses information from past conversations, Gmail, Google Search, and YouTube activity.
- Design prompts from the perspective of personal experiences to test the model's capabilities.
- Assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness.
- Design and execute multi-turn conversational prompts (typically 1-5 turns) requiring the AI to utilize personal information and experiences.
- Evaluate model responses based on prompt intent, checking if personalization was appropriately applied.
- Analyze responses for Grounding issues, ensuring claims about the user are supported by evidence and not flawed inferences or hallucinations.
- Assess Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating".
- Rigorously evaluate and stack-rank two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.
- Write clear, defensible rationales for comparisons, explicitly referencing where issues or positive aspects occurred in the conversation.
- Extract and verify "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.
- Maintain strict data hygiene by deleting evaluation conversations to prevent them from polluting future chat history.
What you need
- Ability to read and write in Vietnamese with a high degree of competency.
- Willingness to use a primary personal Google account and enable personal data sources for a genuine assessment.
- Full-time availability in local time zone.
- Exceptional analytical thinking to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality.
- Creative prompt engineering experience, including designing creative, multi-turn starting prompts based on personal context.
- Strong evaluation acumen, including understanding personalization concepts and identifying incorrect personalization, poor inferences, and forced connections.
- Meticulous attention to detail, including spotting subtle differences in naturalness and overnarrating in side-by-side model responses.
- Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers.
- Ability to provide constructive feedback and detailed annotations.
- Excellent communication and collaboration skills.
- Self-motivated and able to work independently in a remote setting.
- Desktop/Laptop setup with a good internet connection.
Nice to have
- Experience in data annotation, AI quality evaluation, content moderation, or a related role.
Expertise
Who you work with
Project and contracting process: Turing. Applications continue on the provider's website.

