Thriveth

The AI job market, made clearer

Find work worth
your expertise.

Back to all jobs

SWE-Bench Task Auditor

Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models.

United States
Hiring companyMercor
Application processN/ANo approved reviews yet
Work experienceN/ANo approved reviews yet

About the role

Contract and Payment Terms

  • You will be engaged as an independent contractor.
  • This is a fully remote role that can be completed on your own schedule.
  • Projects can be extended, shortened, or concluded early depending on needs and performance.
  • Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution.

Responsibilities

  • You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.

Required skills

  • 3+ years professional software engineering.
  • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles).
  • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking.
  • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++).

Preferred skills

  • Familiarity with SWE-Bench (Verified) or similar repository benchmarks.
  • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.).
  • Prior code-review or task-grading experience.

Eligibility

Work arrangement: Fully remote

Eligible countries: United States

  • Please note: We are unable to support H1-B or STEM OPT candidates at this time.

Compensation details

  • Payments are weekly on Stripe or Wise based on services rendered.
More about the hiring companyMercor
Find AI-training work. Know what to expect.

Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.