Thriveth

The AI job market, made clearer

Find work worth
your expertise.

Back to all jobs

Machine Learning Engineers: Scenario Building for Reinforcement Learning

We're hiring AI researchers and machine learning engineers to participate in building worlds within a reinforcement learning platform.

United States$90 per hour
Hiring companyTerac
Application processN/ANo approved reviews yet
Work experienceN/ANo approved reviews yet

About the role

This work directly influences how agents interact with complex, simulated environments during their training cycles. Your technical expertise will help us refine the tools and interfaces used to create robust testing scenarios.

You will connect to our remote platform to design and construct specific scenarios for reinforcement learning agents. Throughout the session, you will configure environmental parameters, define spatial constraints, and run preliminary agent interactions to test your setup. You will document your workflow and note any friction points encountered while structuring the environment. Finally, you will participate in an interview to share your feedback on the platform's overall usability.

Responsibilities

  • Design and build specific scenarios within a remote reinforcement learning platform.
  • Configure environmental parameters and define agent interaction rules.
  • Test initial agent behaviors to validate your scenario structure.
  • Walk us through your workflow and highlight areas for platform improvement.

Experience and education

This study targets professionals with hands-on experience in simulation design and reinforcement learning environments. We welcome machine learning engineers, AI researchers, simulation developers, and technical game designers accustomed to RL frameworks. Candidates should be highly comfortable configuring complex platform interfaces and defining structured agent scenarios.

  • Professional experience in machine learning or artificial intelligence research.
  • Hands-on background in building simulations or reinforcement learning environments.
  • Familiarity with configuring platform interfaces and defining reward structures.
  • Comfortable articulating technical feedback during a remote interview.

Schedule details

This is a fully remote opportunity that can be completed on your own schedule.

Opportunities can be extended, shortened, or concluded early depending on needs and performance.

Eligibility

Work arrangement: Fully remote

Eligible countries: United States

  • For US Codex and Claude Coders (18+).
  • Know a US-based coder over 18 who uses Codex or Claude for coding on a daily basis?
  • Your participation will not involve access to confidential or proprietary information from any employer, client, or institution.
  • We are unable to support H1-B or STEM OPT candidates at this time.

Compensation details

$90 per hour.

  • Payments are processed weekly based on services rendered.
More about the hiring companyTerac
Find AI-training work. Know what to expect.

Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.