The audio you will work with includes natural conversations and scripted speech across a wide range of topics, speakers, and acoustic conditions.
This role requires a trained ear and careful attention to everything happening in a recording, not just the words being said. Our transcription standard captures the full richness of human communication, and you will be asked to label audio that you hear, including words, noises, and non-speech audio.
You will work from clear guidelines and reference materials that define how each tag type should be used. Accuracy, consistency, and attention to detail are the core requirements of this role.
Experience and education
Strong listening skills with the ability to parse overlapping speech, accented speakers, and varied recording quality.
Comfortable working with structured annotation guidelines, including tagging conventions beyond plain text.
Methodical and consistent, able to apply a defined style across many hours of audio.
Experienced in transcription, captioning, subtitling, or a related field (preferred but not required).
Language requirements
Native-level fluency in Japanese.
Equipment requirements
A reliable computer and stable internet connection.
Headphones or earphones recommended for accurate audio review.
A quiet environment for focused listening.
Schedule details
Flexible Hours.
Eligibility
Work arrangement: Fully remote
Compensation details
up to $110 USD / audio hour transcribed.
Additional performance-based incentives may be available.
Rate applies to audio time transcribed for this specific project. Applications may also be considered for other Babel Audio projects, which may carry different rates. Please review our Terms of Use and Privacy Policy for full details.
Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.