Jobs Companies AT&T Senior App/Prod Support (Tier 3 SRE for event-driven ecosystems)

À propos de ce poste Senior App/Prod Support (Tier 3 SRE for event-driven ecosystems) chez AT&T

AT&T · Sur site · IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City

Key Responsibilities

- Own platform reliability practices for availability, resilience, latency, and operational efficiency.

- Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.

- Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.

- Lead JFROG Helm chart automation and JFROG images/ACR migration work.

- Support microservices deployment enablement and platform/tooling upgrades.

- Own and optimize monitoring, alerting, observability, and logging stack components:

  Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.

- Support health-check frameworks including Airflow health-check requirements.

- Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.

- Collaborate with architecture and delivery teams on reliability and scalability patterns.

- Lead cloud infrastructure creation, maintenance, governance, and access controls.

- Drive capacity planning, DR planning/exercises, and platform best-practice documentation.

- Support cost management, role enforcement, and license management governance.

- Maintain SOP documentation for established alerts and incident patterns.

Required Qualifications / Must-Have Skills

- 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.

- Strong hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations.

- Advanced experience with CI/CD engineering and GitHub Actions.

- Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.

- Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.

- End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).

- Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).

- Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.

- Experience in governance controls: access management, role enforcement, and separation of duties.

- Proven high-severity incident leadership and post-incident reliability improvement execution.

Good-to-Have / Nice-to-Have

- Postgres performance and reliability operations.

- Telecom-scale high-availability systems experience.

Experience Level

Senior to Lead IC (typically 10 to 17 years)

Location / Work Mode

Onsite (Hyderabad / Bangalore or designated AT&T location)

What We Offer

- Opportunity to define and scale platform reliability standards.

- High technical ownership and strong cross-functional influence.

- Enterprise-scale impact across observability, automation, and resilience engineering.

Weekly Hours:

40

Time Type:

Regular

Location:

IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City, IND:KA:Bangalore / Intl Tech Park, Navigator Bldg, Whitefield Road: Intl Tech Park, Navigator Bldg:Intl Tech Park, Navigator Bldg

AT&T and its subsidiaries are committed to equal employment opportunity. All hiring, promotion, and other employment decisions remain merit-based and free from discrimination on the basis of race, color, religion, religious creed, national origin, ancestry, age, sex, sexual orientation, gender, gender identity, gender expression, physical disability, mental disability, pregnancy, medical condition, genetic information, marital status, citizenship status, military status, veteran status, or any other characteristic protected by federal, state, or local laws. In addition, AT&T will provide reasonable accommodations to qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made. Click here to learn more or request an application accommodation here.

Prêt à postuler chez AT&T ?
Postuler chez AT&T

À propos de AT&T

We are pioneers of making connections and have been ever since Alexander Graham Bell invented the telephone and founded our company. That was nearly 150 years ago, and we haven’t stopped innovating since. At our core, we help bring families, communities, and businesses together with the products and services they need to thrive every day. From the widespread and growing availability of 5G and Fiber to working on things we once only dreamed of—at AT&T, we create connections that change the world.

Voir tous les emplois chez AT&T →

Emplois similaires

General Matter
Reliability Engineer
General Matter
⚡ Postuler tôt Los Angeles, CA Sur site $140,000–$200,000
● Nouveau 👁 Vu ✓ Postulé il y a 2 h
General Matter
DevOps / Site Reliability Engineer
General Matter
⚡ Postuler tôt Los Angeles, CA Sur site $100,000–$200,000
● Nouveau 👁 Vu ✓ Postulé il y a 2 h
Wikimedia Foundation
Senior Site Reliability Engineer, Wikimedia Enterprise
Wikimedia Foundation
⚡ Postuler tôt Remote · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 2 h
MS
Senior Reliability Engineer, Technical Lead
Muon Space
⚡ Postuler tôt San Jose, CA Hybride $216,000–$236,000
● Nouveau 👁 Vu ✓ Postulé il y a 3 h
Fingerprint
Senior Site Reliability Engineer
Fingerprint
⚡ Postuler tôt Remote · lieu restreint $152,000–$205,000
● Nouveau 👁 Vu ✓ Postulé il y a 5 h
Stone - Linkedin
SITE RELIABILITY ENGINEER II
Stone - Linkedin
⚡ Postuler tôt Remoto Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 5 h
Incode Technologies
Senior Site Reliability Engineer (Public Sector)
Incode Technologies
⚡ Postuler tôt United States Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 5 h
Zscaler
Staff Site Reliability Engineer
Zscaler
⚡ Postuler tôt San Jose, California, USA Hybride $122,500–$175,000
● Nouveau 👁 Vu ✓ Postulé il y a 6 h
Zscaler
Sr. Staff Site Reliability Engineer-Federal, Security Clearance
Zscaler
⚡ Postuler tôt Crystal City, Virginia, USA Sur site $140,000–$200,000
● Nouveau 👁 Vu ✓ Postulé il y a 6 h

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez AT&T

Voir tous les emplois chez AT&T →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit