Site Reliability Engineer — ETL Platform, IS&T Ai & Data Platforms

Shanghai Until 10/7/2026 4+ years exp First posted August 8, 2026 Last posted August 8, 2026
Job description

We are hiring an SRE to own reliability, performance, data freshness, and operational readiness for our production ETL platform. The platform supports data ingestion, transformation, and loading workflows across Kubernetes-based environments — including Airflow-based loader jobs and Spark-on-EKS jobs that load data into a Datalake/Lakehouse.

Description

You will operate and triage production pipelines end-to-end: extractors, loaders, batch jobs, streaming ingestion (Kafka), Spark workloads, and Airflow DAGs. You will tune Kubernetes and Spark for stability, build observability tooling, drive root cause analysis, write automation to reduce toil, and manage configuration and secrets through GitOps-style processes. The ideal candidate troubleshoots distributed systems from logs, metrics, and infrastructure signals, and drives permanent fixes through engineering partnership.

Minimum Qualifications

4+ years of experience in SRE, DevOps, platform engineering, data infrastructure, or production operations with strong Linux troubleshooting skills.
Strong Kubernetes/EKS operations experience (kubectl, deployments, pods, resource limits, service accounts, secrets, workload debugging) and hands-on experience supporting Apache Spark on Kubernetes — including tuning, log analysis, memory issues, shuffle failures, and performance bottlenecks.
Production experience with Apache Airflow (DAG operations, task failures, retries, scheduling, SLA misses) and Kafka or similar streaming platforms (consumer groups, lag, offsets, partitions, secure client connectivity).
Familiarity with Datalake/Lakehouse architectures, object storage (S3 or equivalent), monitoring/logging tools (Splunk, Prometheus, Grafana, CloudWatch, ELK, Datadog), and solid SQL skills.
Scripting proficiency in Python and Bash, with strong incident management, RCA, change management, and operational documentation skills.

Preferred Qualifications

Experience supporting metadata-driven ETL platforms or internal ETL frameworks.
Experience with Lakehouse technologies (Spark, Iceberg, IRC, Hive Metastore, Parquet) and data loading performance tuning.
Experience with GitOps or source-of-truth configuration management, and infrastructure-as-code tools (Terraform, Helm, Argo CD, Ansible).
Experience operating multi-region or region-specific data platforms.
Experience with certificate/PKI/TLS management, JKS/truststore, Kerberos, or key rotation processes.
Familiarity with Spark performance tuning at scale, and platform dependencies such as API gateways, RabbitMQ, Redis, or Cassandra.

About this role

Summary

Manage reliability and performance of ETL platform with Kubernetes, Spark, Kafka, and Airflow.

Job title

Site Reliability Engineer — ETL Platform, IS&T Ai & Data Platforms

Experience level

4+ years

Minimum experience

4+ years exp

Industry

software

Location requirements

Shanghai, on-site role, no remote work allowed

Salary

Not specified

Management role

No

Skills & keywords

Required skills

kubernetessparkairflowkafkalinux troubleshootingpythonbashmonitoring toolssql

Preferred skills

lakehousegitopsterraformhelmregion-specific data platformsTLS managementspark performance tuning

Specializations

kubernetessparkairflowkafkadata infrastructure
Locations

Structured locations inferred from the posting.

Shanghai, China

On-site City