Synthetic Data Engineer (AI Data/Training)
Hyphen Connect Limited
Apply to this job Boston, US Until 8/21/2026 First posted April 25, 2026 Last posted April 25, 2026
Job description
We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops. Your expertise will drive the success of data processing and model training within the organization.
Responsibilities:
- Design domain-specific synthetic data generation (SDG) pipelines via self-instruct and constitutional prompting.
- Implement automated quality scoring and de-duplication systems.
- Manage data pipelines that feed directly into SFT and DPO training loops.
Qualifications:
- Proven experience building large-scale data pipelines (Airflow, Spark, Ray).
- Deep knowledge of prompt engineering for data generation.
- Familiarity with dataset distillation and bias mitigation.
About this role
Summary
Design and implement synthetic data pipelines for AI training and quality management.
Job title
Synthetic Data Engineer
Experience level
Industry
software
Location requirements
Located in Boston, USA, remote work not specified
Salary
Not specified
Management role
No
Skills & keywords
Required skills
data pipelinestraining dataquality scoringde-duplication
Preferred skills
prompt engineeringdataset distillationbias mitigationAirflowSparkRay
Specializations
AIdata generationtraining datadata pipelines
Locations
Structured locations inferred from the posting.
Boston, MA, USA
Work arrangement unknown City