Synthetic Data Engineer (AI Data/Training)
Hyphen Connect Limited
Apply to this job Oregon, US Until 8/21/2026 First posted April 25, 2026 Last posted April 25, 2026
Job description
We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops. Your expertise will drive the success of data processing and model training within the organization.
Responsibilities:
- Design domain-specific synthetic data generation (SDG) pipelines via self-instruct and constitutional prompting.
- Implement automated quality scoring and de-duplication systems.
- Manage data pipelines that feed directly into SFT and DPO training loops.
Qualifications:
- Proven experience building large-scale data pipelines (Airflow, Spark, Ray).
- Deep knowledge of prompt engineering for data generation.
- Familiarity with dataset distillation and bias mitigation.
About this role
Summary
Design and implement data generation pipelines for AI training and quality systems.
Job title
Synthetic Data Engineer (AI Data/Training)
Experience level
Industry
software
Location requirements
Oregon, USA; on-site work required
Salary
Not specified
Management role
No
Skills & keywords
Required skills
data pipelinesAirflowSparkRaydata generation
Preferred skills
prompt engineeringdataset distillationbias mitigation
Specializations
AIdata engineeringmachine learning
Locations
Structured locations inferred from the posting.
Unknown location
On-site