Thriveth

The AI job market, made clearer

Find work worth
your expertise.

Back to all jobs

AI Safety Experts — English & Danish

We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers.

Remote
Hiring companyMercor
Application processN/ANo approved reviews yet
Work experienceN/ANo approved reviews yet

About the role

At Mercor, we believe the safest AI is the one that’s already been attacked — by us.

This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated.

What Success Looks Like

  • You uncover vulnerabilities automated tests miss.
  • You deliver reproducible artifacts that strengthen customer AI systems.
  • Evaluation coverage expands: more scenarios tested, fewer surprises in production.
  • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary.

Why Join Mercor

  • Build experience in human data-driven AI red teaming at the frontier of safety.
  • Play a direct role in making AI systems more robust, safe, and trustworthy.

Contract and Payment Terms

  • You will be engaged as an independent contractor.
  • This is a fully remote role that can be completed on your own schedule.
  • Projects can be extended, shortened, or concluded early depending on needs and performance.
  • Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution.

Responsibilities

  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation.
  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent.
  • Document reproducibly: produce reports, datasets, and attack cases customers can act on.

Required skills

  • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing).
  • You’re curious and adversarial: you instinctively push systems to breaking points.
  • You’re structured: you use frameworks or benchmarks, not just random hacks.
  • You’re communicative: you explain risks clearly to technical and non-technical stakeholders.
  • You’re adaptable: thrive on moving across projects and customers.

Preferred skills

  • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
  • Cybersecurity: penetration testing, exploit development, reverse engineering.
  • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.
  • Creative probing: psychology, acting, writing for unconventional adversarial thinking.

Language requirements

  • Fluent Language Skills Required: English & Danish. Native fluency in English and Danish is required for this position.

Eligibility

Work arrangement: Fully remote

  • Please note: We are unable to support H1-B or STEM OPT candidates at this time.

Compensation details

  • Payments are weekly on Stripe or Wise based on services rendered.
More about the hiring companyMercor
Find AI-training work. Know what to expect.

Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.