Your work will shape how models learn, reason, and perform through high-quality, real-world input.
Responsibilities
Design, build, and maintain scalable big data pipelines and architectures to support robust data solutions.
Collaborate with cross-functional teams to understand and deliver on data requirements and business objectives.
Implement data integration, transformation, and processing solutions using Python and relevant big data technologies.
Develop, manage, and optimize distributed databases and storage systems for efficiency and reliability.
Monitor, troubleshoot, and enhance data systems to ensure high availability and performance.
Enforce data quality, security, and governance standards across all solutions.
Document solutions and communicate complex technical concepts effectively, both in writing and verbally.
Required skills
Proven expertise in big data engineering with hands-on experience building and maintaining large-scale data pipelines.
Advanced proficiency in Python for data processing, automation, and integration.
Deep understanding of relational and NoSQL databases, including optimization and management techniques.
Experience with distributed data processing frameworks (e.g., Hadoop, Spark, Flink).
Strong foundation in data modeling, ETL processes, and data warehousing principles.
Excellent written and verbal communication skills, with the ability to convey technical ideas clearly to both technical and non-technical stakeholders.
Detail-oriented, proactive, and self-motivated, thriving in remote and autonomous work environments.
Preferred skills
Prior experience in fast-paced or startup-like environments supporting global teams.
Expertise with cloud-based big data platforms (e.g., AWS, GCP, Azure).
Familiarity with machine learning operations and data science workflows is a plus.
Thriveth makes AI data-training work easier to find, understand, and navigate. We replace uncertainty with clear opportunities, realistic expectations, and insights from real application journeys.