External_Site Reliability Engineer_SG

Singapore Until 8/22/2026 3+ years exp First posted June 8, 2026 Last posted June 8, 2026
Job description

At IBM, work is more than a job - it's a calling: To build. To design. To code. To consult. To think along with clients and sell. To make markets. To invent. To collaborate. Not just to do something better, but to attempt things you've never thought possible. Are you ready to lead in this new era of technology and solve some of the world's most challenging problems? If so, lets talk.


As an SRE in IBM Consulting, you'll serve as a leader in defining solutions for clients. You'll identify insights and tasks that can be automated. You'll have the opportunity to identify points of improvement in technical processes and propose new ways to do it through automation, help our customer to resolve their pain points and, through co-creation, define solutions that allow improving the efficiency of their operations. Your primary responsibilities include: Strategic Design and Analysis of Distributed Systems: Design, analyze, and troubleshooting large-scale distributed systems. Proactive Reliability Management and Incident Response: Participate in on-call rotation, engage with product teams to fix production outages, and carry forward action items to improve ongoing reliability. Empowering Tools and Automation for Enhanced Reliability: Develop effective tooling, alerts, and response to both identify and address reliability risks including automatic problem detection and mitigation.

  • 3 or more years’ experience in a systems engineering infrastructure role, or related experience as a Software Engineer / other technical role
  • Detailed, practical knowledge of systems administration practices for Linux / Windows (mainly Linux). Ability to work with little-or-no supervision on business-as-usual SRE / DevSecOps systems administration tasks
  • Command line configuration of Linux and Windows hosts, as well as familiarity with common GUI tools and x-server applications
  • Enterprise development experience in high-level programming languages such as Java, C, C++, Python, R, etc.
  • Has a passion for working within an operational environment and to take work through to its conclusion
  • Experience with Agile methods and working within a Sprint-based setting
  • Experience with IBM Cloud, AWS, Azure or similar cloud environments
  • Experience providing networks for / within container environments – such as Kubernetes, OpenShift or Docker Swarm
  • Experience with Akamai WAF technology.
About this role

Summary

Design, analyze, and troubleshoot large-scale distributed systems, automate, and improve reliability.

Job title

Site Reliability Engineer

Experience level

3+ years

Minimum experience

3+ years exp

Industry

software

Location requirements

Singapore-based role, remote work allowed

Salary

Not specified

Management role

No

Skills & keywords

Required skills

LinuxWindowscommand lineJavaCC++PythonAgilecloud environmentsKubernetesOpenShiftDocker SwarmAkamai WAF

Preferred skills

None specified

Specializations

distributed systemsreliability managementautomationcloud environmentscontainer networks
Locations

Structured locations inferred from the posting.

Singapore

Remote Country