Thriveth

The AI job market, made clearer

Find work worth
your expertise.

Back to all jobs
Posted 21 Aug 2026

Project Cursa - Robot Manipulation Video Annotator

We are looking for detail-oriented annotators to help label robot manipulation videos for AI training purposes.

Philippines
Hiring companyWelo Data
Application processN/ANo approved reviews yet
Work experienceN/ANo approved reviews yet

About the role

About the Role.

You'll watch short videos of robots performing manipulation tasks (filmed from three synchronized camera angles) and produce precise, structured, natural-language descriptions of the actions taking place. This work directly supports the development of robotics AI models and requires strong written English, sharp observational skills, and the discipline to follow a detailed style guide consistently.

What We're Looking For.

Responsibilities

  • Watch short robot manipulation videos, each filmed from three synchronized camera views (an overhead view and views from each of the robot's two wrist-mounted cameras).
  • Break each video into time segments and write clear, natural-language descriptions for each segment.
  • Apply labels at three levels of detail for each applicable segment:
  • Atomic motion (a few seconds) — a single small movement (e.g., "close fingers around the red handle").
  • Skill / subtask (several seconds to ~20 seconds) — a complete, meaningful action (e.g., "pick up the red block by its edge").
  • Task / goal (up to ~1 minute) — the overall purpose of a sequence of skills (e.g., "place all blocks in the container").
  • Ensure every moment of video is covered by a label at two or more of these levels — no gaps, including idle or pause moments.
  • Accurately describe exactly what happens, including when something doesn't go as planned (a dropped object, a failed grasp, a slipped grip). Precision matters more than making the robot look successful.
  • Cross-reference all three camera angles: use the overhead view to understand the overall scene and object identity, and the close-up wrist views to confirm exact contact and grasp details.
  • Follow a detailed style guide covering vocabulary for actions, spatial relationships, object descriptions, and manner of movement, applying it consistently across many episodes.
  • Participate in periodic calibration sessions to align your labeling with the team and the client's reference examples.

Required skills

  • Strong written English — you'll write dozens of short, precise descriptive sentences per video and need to vary your language rather than repeating the same phrases.
  • Sharp attention to detail — able to distinguish small differences (a successful grasp vs. a fumble, a push vs. a drag, which specific object part is being touched).
  • Comfort following a detailed, structured style guide and applying it consistently, even in ambiguous or edge-case scenarios.
  • Basic comfort with spatial/mechanical description (left/right, above/below, naming object parts like handles, lids, or edges).
  • Reliable, self-directed work habits — this is often heads-down work with periodic check-ins rather than close supervision.

Preferred skills

  • Prior experience with video annotation, data labeling, transcription, or QA work.
  • Familiarity with robotics terminology (grippers, end-effectors, manipulation) — helpful but not necessary, as the style guide is self-contained.
  • Experience with annotation tools such as Label Studio.

Eligibility

Work arrangement: Fully remote

Eligible countries: Philippines

More about the hiring companyWelo Data
Find AI-training work. Know what to expect.

Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.