Synthetic Data Engineer (AI Data/Training)

Hyphen Connect Limited

Apply to this job
China Until 8/21/2026 First posted April 25, 2026 Last posted April 25, 2026
Job description

We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops. Your expertise will drive the success of data processing and model training within the organization.

 

Responsibilities:

  • Design domain-specific synthetic data generation (SDG) pipelines via self-instruct and constitutional prompting.
  • Implement automated quality scoring and de-duplication systems.
  • Manage data pipelines that feed directly into SFT and DPO training loops.

Qualifications:

  • Proven experience building large-scale data pipelines (Airflow, Spark, Ray).
  • Deep knowledge of prompt engineering for data generation.
  • Familiarity with dataset distillation and bias mitigation.
About this role

Summary

Design and implement synthetic data pipelines for AI training, ensuring quality and bias mitigation.

Job title

Synthetic Data Engineer (AI Data/Training)

Experience level

Industry

software

Location requirements

must be in China, remote work not specified

Salary

Not specified

Management role

No

Skills & keywords

Required skills

data pipelinesPythondata management

Preferred skills

prompt engineeringdataset distillationbias mitigationAirflowSparkRay

Specializations

AIdata generationmachine learningdata pipelines
Locations

Structured locations inferred from the posting.

China

Work arrangement unknown Country