Site Reliability Engineer - Cloud Infrastructure

Dublin, Ireland Until 8/21/2026 H-1B sponsor history First posted March 29, 2026 Last posted March 29, 2026
Job description

Description

The Technical Infrastructure SRE team is responsible for managing the whole infrastructure and applications. Our mission is to ensure all production systems can support our fast growing world-wide user base as well as keep the entire systems stable, efficient and cost effective. We manage deployments, system capacity, traffic scheduling, fault tolerance, disaster recovery, emergency response, automations, operation platforms development, etc. 

Our team is full of diversity. We have team members in Singapore and China. Now we are extending our teams to Ireland. We are looking forward to seeing new talents joining our team and together helping TikTok grow.

- Reliability: Ensure the stability of the company's core infrastructure (system high availability and reliability), focus on system performance and capacity, establish O&M (Operation & Maintenance) standards and SOP processes.
- Reliability: Troubleshooting and locating technical issues, collaborate with the technical team to develop and implement system capacity planning, performance testing, anomaly analysis, and fault diagnosis and resolution strategies.
- Efficiency: Research and evaluate large-scale system architectures and technologies, use new tools and technologies to improve existing systems and processes to support business development.
- Efficiency: Design and implement O&M platforms to achieve efficient, automated, and intelligent system maintenance.
- Cost: Develop delivery standards for mass production system scales, from budgeting to resource delivery, to online system capacity assessments, to help the company optimize IT costs.
- Compliance: Design and establish new IDC, design and implement data protection plans to meet standard requirements.

Requirements

Minimum Qualifications:
- Bachelor's / Master's Degree in Computer Science or related major.
- Solid basic knowledge of computer software, understanding of Linux operating system, storage, network IO and other related principles.
- Familiar with one or more programming languages, such as Python, Go, and Java. Knowledge of design patterns and coding principles is necessary.
- Familiar with Elastic Search

Preferred Qualifications:
- Experience with storage, and relevant system experience with the following: KV, Table, Graph, Redis, MySQL, MongoDB, MQ, and Kafka.
- Experience with computing & big data, and system experience with the following: Kubernetes, Docker/Containers, AIops, Spark, Flink, Function as a service, RPC Framework, and Service Mesh.

About this role

Summary

Manage infrastructure, ensure system reliability, optimize performance, and automate operations.

Job title

Site Reliability Engineer - Cloud Infrastructure

Experience level

null

Industry

software

Location requirements

Dublin, Ireland; remote not specified

Salary

Not specified

Visa sponsorship

H-1B sponsor history

Management role

No

Skills & keywords

Required skills

LinuxPythonGoJavaElastic Search

Preferred skills

storageKVRedisMySQLMongoDBKafkaKubernetesDockerFlinkSparkAIopsService Mesh

Specializations

cloud infrastructuresystem stabilityautomationcost optimizationdisaster recovery
Locations

Structured locations inferred from the posting.

Dublin, Ireland

Work arrangement unknown City