Evaluate multilingual AI reference responses and scoring rubrics for Project Meteor. You’ll use detailed guidelines to identify quality issues, correct critical errors, and help improve how advanced AI systems are assessed.
Meteor focuses on the reference answers and evaluation rubrics used to train and assess multilingual AI models. The work includes both open-ended and closed-ended tasks and requires careful, consistent judgment.
Responsibilities
Review AI-generated reference responses for quality and accuracy.
Evaluate and improve scoring rubrics.
Identify and correct critical P0 issues.
Follow detailed annotation guidelines and meet project quality requirements.
Required skills
Native-level fluency in one of the available project languages.
Strong reading comprehension, written communication, and attention to detail.
Basic familiarity with medical, financial, and legal topics.
Ability to follow detailed instructions carefully.
A computer with a reliable internet connection.
Schedule details
Schedule: Flexible; task availability may vary.
Eligibility
Work arrangement: Fully remote
Applicant eligibility: 7 eligible countries
View all 7 eligible countries
Eligible countries: Brazil, Indonesia, Malaysia, Mexico, Thailand, United Kingdom, Vietnam
Compensation details
Compensation is calculated at a fixed hourly rate. The rate shown on the OneForma platform depends on the language and location you select.
Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.