Site Reliability Engineer
A career in IBM Consulting is rooted by long-term relationships and close collaboration with clients across the globe. You'll work with visionaries across multiple industries to improve the hybrid cloud and AI journey for the most innovative and valuable companies in the world. Your ability to accelerate impact and make meaningful change for your clients is enabled by our strategic partner ecosystem and our robust technology platforms across the IBM portfolio; including Software and Red Hat. Curiosity and a constant quest for knowledge serve as the foundation to success in IBM Consulting. In your role, you'll be encouraged to challenge the norm, investigate ideas outside of your role, and come up with creative solutions resulting in ground breaking impact for a wide network of clients. Our culture of evolution and empathy centers on long-term career growth and development opportunities in an environment that embraces your unique skills and experience.
As an SRE in IBM Consulting, you'll serve as a leader in defining solutions for clients. You'll identify insights and tasks that can be automated. You'll have the opportunity to identify points of improvement in technical processes and propose new ways to do it through automation, help our customer to resolve their pain points and, through co-creation, define solutions that allow improving the efficiency of their operations. Your primary responsibilities include: Strategic Design and Analysis of Distributed Systems: Design, analyze, and troubleshooting large-scale distributed systems. Proactive Reliability Management and Incident Response: Participate in on-call rotation, engage with product teams to fix production outages, and carry forward action items to improve ongoing reliability. Empowering Tools and Automation for Enhanced Reliability: Develop effective tooling, alerts, and response to both identify and address reliability risks including automatic problem detection and mitigation.
Create up to 5 bullets max
Create up to 3 bullets max (encouraging then to focus on required skills)
Summary
Design, analyze, and improve reliability of large-scale distributed systems through automation and incident management.
Job title
Site Reliability Engineer
Experience level
null
Industry
software
Location requirements
Singapore, remote work not specified
Salary
Not specified
Management role
No
Required skills
Preferred skills
Specializations
Structured locations inferred from the posting.
Singapore