AI Evaluators: Assessing a Shopping Assistant
We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant.
Hiring companyTerac
About the role
This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.
You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.
Responsibilities
- Review real user interaction traces with an AI shopping assistant.
- Identify logical failures, inaccuracies, or poor recommendations in the text.
- Create structured rubrics and verifiers to judge response quality.
- Commit to 20+ hours per week of evaluation work on our internal platform.
Experience and education
This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.
- Experience in data evaluation, quality assurance, or AI training.
- Strong analytical skills with the ability to spot subtle errors in text.
- Familiarity with e-commerce search and digital shopping experiences.
- Ability to commit to a sustained workload of 20+ hours per week.
Schedule details
This is a fully remote opportunity that can be completed on your own schedule.
Opportunities can be extended, shortened, or concluded early depending on needs and performance.
Eligibility
Work arrangement: Fully remote
Eligible countries: United States, United Kingdom, Canada
- For US, UK, or Canadian AI Evaluators (18+).
- Know a US, UK, or Canadian degree-holder (18+) who regularly uses LLMs for complex workflows and has 20+ hours a week for a project?
- Your participation will not involve access to confidential or proprietary information from any employer, client, or institution.
- We are unable to support H1-B or STEM OPT candidates at this time.
Compensation details
$30 per hour.
- Payments are processed weekly based on services rendered.
More about the hiring companyTerac

