Synthetic Data Engineer (AI Data/Training)

Hyphen Connect Limited

Apply to this job
Australia Until 8/21/2026 First posted April 25, 2026 Last posted April 25, 2026
Job description

We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops. Your expertise will drive the success of data processing and model training within the organization.

 

Responsibilities:

  • Design domain-specific synthetic data generation (SDG) pipelines via self-instruct and constitutional prompting.
  • Implement automated quality scoring and de-duplication systems.
  • Manage data pipelines that feed directly into SFT and DPO training loops.

Qualifications:

  • Proven experience building large-scale data pipelines (Airflow, Spark, Ray).
  • Deep knowledge of prompt engineering for data generation.
  • Familiarity with dataset distillation and bias mitigation.
About this role

Summary

Design and implement synthetic data pipelines for AI training and data management.

Job title

Synthetic Data Engineer (AI Data/Training)

Experience level

Industry

software

Location requirements

Australia, remote work not specified

Salary

Not specified

Management role

No

Skills & keywords

Required skills

large-scale data pipelinesdata generationquality scoringde-duplication

Preferred skills

prompt engineeringdataset distillationbias mitigationAirflowSparkRay

Specializations

AIdata engineeringsynthetic data
Locations

Structured locations inferred from the posting.

Australia

Work arrangement unknown Country