In this role, you will analyze structured data, create challenging prompts, evaluate AI-generated responses for factual accuracy and reasoning quality, and provide evidence-backed feedback to improve model performance.
We are looking for curious, detail-oriented professionals who enjoy solving complex problems and evaluating AI systems. The ideal candidate can analyze data, think critically, validate responses against evidence, and clearly articulate why a model's output is correct or incorrect. Experience working with AI models and creating challenging evaluation prompts is a strong advantage.
Employment type: Contractor assignment (no medical/paid leave).
Benefits
Opportunity to work on cutting-edge AI projects.
Competitive compensation.
Flexible working hours and remote work environment.
Responsibilities
Create challenging prompts that evaluate an LLM's ability to retrieve, analyze, and reason over structured data.
Assess AI-generated responses for factual accuracy, logical reasoning, and completeness.
Identify model failures, inconsistencies, hallucinations, and reasoning gaps.
Validate model outputs using provided datasets and supporting evidence.
Document findings with clear, evidence-based explanations.
Consistently follow annotation guidelines and maintain high-quality standards.
Preferred skills
Experience working with Large Language Models (LLMs) or Generative AI.
Familiarity with prompt engineering, AI evaluation, data annotation, or model testing.
Experience working with structured datasets (CSV, Excel, databases, etc.).
Ability to identify edge cases and design prompts that expose model limitations.
Experience and education
Master's degree or higher in any discipline.
Minimum 3 years of professional, research, or teaching experience.
Strong analytical and critical thinking skills.
Exceptional attention to detail and ability to validate information against source data.
Language requirements
Excellent written English communication skills.
Assessment requirements
Shortlisting based on qualifications and assessment scores.
Schedule details
Commitments Required: 40, 30 or 20 hours per week with at least 4 hours PST overlap.
Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.