Lead SRE / Platform Engineer - APP/PROD Support

AT&T Mobility Services LLC

Apply to this job
IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City Until 9/27/2026 10+ years exp First posted July 29, 2026 Last posted July 29, 2026
Job description

Lead SRE / Platform Engineer - APP/PROD Support

Job Summary / About the Role

AT&T is looking for a Lead SRE / Platform Engineer to own the technical direction, reliability roadmap, and platform governance for enterprise integration, messaging, and event-driven ecosystems. This is the senior-most IC role in the support organization: the candidate is the final technical escalation point for high-severity incidents, mentors Tier 2 and Tier 3 engineers, and drives cross-team architecture and reliability decisions that shape how the platform evolves.

Key Responsibilities

  • Own end-to-end platform reliability strategy for availability, resilience, latency, and operational efficiency across the support organization.

  • Serve as the final technical escalation and decision-making authority for high-severity/complex incidents, driving resolution and post-incident reliability improvements.

  • Provide technical leadership and mentoring to Tier 2 and Tier 3 engineers, including skill development, code/config review, and troubleshooting coaching.

  • Drive DevOps and automation strategy including Golden Image improvements and support automation use cases across teams.

  • Define and enforce GitHub Actions pipelines and CI/CD reliability standards platform-wide.

  • Lead JFROG Helm chart automation and JFROG images/ACR migration initiatives.

  • Own microservices deployment enablement strategy and platform/tooling upgrade roadmap.

  • Own and govern monitoring, alerting, observability, and logging stack architecture:

Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.

  • Own health-check framework strategy including Airflow health-check requirements.

  • Partner with architecture and delivery leadership on reliability, scalability, and platform evolution decisions.

  • Own cloud infrastructure governance: creation, maintenance, access controls, and policy enforcement.

  • Own capacity planning, DR planning/exercises, and platform best-practice documentation sign-off.

  • Own cost management governance, role enforcement, and license management decisions.

  • Set and maintain SOP standards for alerts and incident patterns across Tier 1-3.

Required Qualifications / Must-Have Skills

  • 10+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles, including demonstrated technical leadership.

  • Proven experience mentoring or technically leading engineers, with the ability to set standards and review the work of others.

  • Deep hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations at scale.

  • Advanced experience with CI/CD engineering and GitHub Actions, including defining organization-wide standards.

  • Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks, including architecture-level ownership.

  • Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.

  • End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).

  • Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).

  • Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.

  • Experience owning governance controls: access management, role enforcement, and separation of duties.

  • Proven track record leading high-severity incident response and driving post-incident reliability improvement programs.

  • Strong stakeholder communication skills to represent platform reliability decisions to architecture and leadership audiences.

Good-to-Have / Nice-to-Have

  • Postgres performance and reliability operations.

  • Telecom-scale high-availability systems experience.

  • Prior people-management or formal team-lead experience.

What We Offer

  • Highest IC-level ownership over platform reliability strategy and standards.

  • Direct influence on architecture, governance, and cross-team technical decisions.

  • Mentorship scope across Tier 2 and Tier 3 engineering teams.

  • Enterprise-scale impact across observability, automation, and resilience engineering.

Weekly Hours:

40

Time Type:

Regular

Location:

IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City

It is the policy of AT&T to provide equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state or local law. In addition, AT&T will provide reasonable accommodations for qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.

About this role

Summary

Lead platform reliability, automation, incident response, and cross-team architecture

Job title

Lead SRE / Platform Engineer - APP/PROD Support

Experience level

10+ years

Minimum experience

10+ years exp

Industry

telecommunications

Location requirements

Hyderabad, India; on-site work required, no remote option

Salary

Not specified

Management role

No

Skills & keywords

Required skills

kubernetesazure kubernetes serviceci/cdgithub actionsprometheusgrafanaalertmanagerpythonconfluent kafkaconfluent cloudazure event hubapache flinkgovernance controls

Preferred skills

postgreshigh-availability systemspeople management

Specializations

srekubernetescloud-nativedevopsobservability
Locations

Structured locations inferred from the posting.

Hyderabad, Telangana, India

On-site City