Thriveth

The AI job market, made clearer

Find work worth
your expertise.

Back to all jobs

ML Challenge Task Auditor

Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models.

United States
Hiring companyMercor
Application processN/ANo approved reviews yet
Work experienceN/ANo approved reviews yet

About the role

Note: this role evaluates applied/experimental ML rigor — it is not an LLM-application-building or MLOps role.

Contract and Payment Terms

  • You will be engaged as an independent contractor.
  • This is a fully remote role that can be completed on your own schedule.
  • Projects can be extended, shortened, or concluded early depending on needs and performance.
  • Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution.

Responsibilities

  • You'll assess experiment design, model-selection reasoning, and evaluation methodology — and provide clear, rubric-based written feedback.

Required skills

  • 3+ years hands-on applied/experimental ML (experiment design, model selection, hyperparameter tuning, evaluation methodology).
  • Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene.
  • Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost).
  • Ability to critique ML claims against evidence and reproduce results.

Preferred skills

  • Competition / benchmark experience (e.g., Kaggle).
  • Graduate research or publication record in applied ML.
  • Prior task-grading or peer-review experience.

Eligibility

Work arrangement: Fully remote

Eligible countries: United States

  • Please note: We are unable to support H1-B or STEM OPT candidates at this time.

Compensation details

  • Payments are weekly on Stripe or Wise based on services rendered.
More about the hiring companyMercor
Find AI-training work. Know what to expect.

Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.