Site Reliability Engineer

Singapore Until 8/21/2026 First posted June 8, 2026 Last posted June 8, 2026
Job description

A career in IBM Consulting is rooted by long-term relationships and close collaboration with clients across the globe. You'll work with visionaries across multiple industries to improve the hybrid cloud and AI journey for the most innovative and valuable companies in the world. Your ability to accelerate impact and make meaningful change for your clients is enabled by our strategic partner ecosystem and our robust technology platforms across the IBM portfolio; including Software and Red Hat. Curiosity and a constant quest for knowledge serve as the foundation to success in IBM Consulting. In your role, you'll be encouraged to challenge the norm, investigate ideas outside of your role, and come up with creative solutions resulting in ground breaking impact for a wide network of clients. Our culture of evolution and empathy centers on long-term career growth and development opportunities in an environment that embraces your unique skills and experience.

As an SRE in IBM Consulting, you'll serve as a leader in defining solutions for clients. You'll identify insights and tasks that can be automated. You'll have the opportunity to identify points of improvement in technical processes and propose new ways to do it through automation, help our customer to resolve their pain points and, through co-creation, define solutions that allow improving the efficiency of their operations. Your primary responsibilities include: Strategic Design and Analysis of Distributed Systems: Design, analyze, and troubleshooting large-scale distributed systems. Proactive Reliability Management and Incident Response: Participate in on-call rotation, engage with product teams to fix production outages, and carry forward action items to improve ongoing reliability. Empowering Tools and Automation for Enhanced Reliability: Develop effective tooling, alerts, and response to both identify and address reliability risks including automatic problem detection and mitigation.

Create up to 5 bullets max

Create up to 3 bullets max (encouraging then to focus on required skills)

About this role

Summary

Design, analyze, and improve reliability of large-scale distributed systems through automation and incident management.

Job title

Site Reliability Engineer

Experience level

null

Industry

software

Location requirements

Singapore, remote work not specified

Salary

Not specified

Management role

No

Skills & keywords

Required skills

distributed systemsautomationmonitoringincident response

Preferred skills

None specified

Specializations

distributed systemsautomationreliabilityincident responsemonitoring
Locations

Structured locations inferred from the posting.

Singapore

Work arrangement unknown Country