Sobre este puesto de Senior Production Support Engineer(4+ Years in Azure and AWS infrastructure support and SQL Server administration) en FIS
Job Title
Senior Production Support Engineer
Intro:
We are FIS. Our technology powers the world’s economy and our teams bring innovation to life. We champion diversity to deliver the best products and solutions for our colleagues, clients and communities. If you’re ready to start learning, growing and making an impact with a career in fintech, we’d like to know: Are you FIS?
About the Role
As a Senior Production Support Engineer, you will ensure the stability, reliability, and performance of business-critical applications across cloud and hybrid environments. You will apply a Site Reliability Engineering (SRE) mindset to support production platforms, improve observability, reduce operational toil through automation, and resolve complex technical issues. Working across infrastructure, applications, databases, and container platforms, you will partner with Engineering, Infrastructure, Security, and Product teams to maintain service excellence. Success in this role is measured through service reliability, operational efficiency, incident resolution effectiveness, and continuous improvement outcomes.
What You Will Be Doing
- Provide L2/L3 production support for business-critical applications and services.
- Lead troubleshooting across applications, databases, cloud infrastructure, networking, and Kubernetes platforms.
- Participate in incident response, major incident management, and on-call support activities.
- Conduct detailed Root Cause Analysis (RCA) and drive permanent corrective actions.
- Develop and maintain operational runbooks, recovery procedures, and support documentation.
- Monitor and improve availability, MTTR, service reliability, and operational performance metrics.
- Support and troubleshoot Azure and AWS environments, including networking connectivity, DNS, routing, firewalls, security groups, load balancers, hybrid connectivity, VPNs, and private network integrations.
- Support AKS and Kubernetes-hosted applications, including workloads, namespaces, ingress controllers, and containerized services.
- Support CI/CD deployment pipelines using GitHub and ArgoCD and investigate deployment and configuration issues.
- Design and enhance monitoring and observability solutions using Dynatrace, Azure Monitor, Azure Log Analytics, and AWS CloudWatch.
- Develop dashboards, alerts, SLIs, and SLOs while leveraging logs, metrics, traces, and distributed tracing data.
- Identify automation opportunities, reduce manual operational effort, and contribute to platform reliability and resilience initiatives.
- Collaborate with Development and DevOps teams to improve production readiness, platform stability, and operational maturity.
- Promote SRE best practices including observability, automation, reliability engineering, blameless post-mortems, and proactive service management.
Required Qualifications
- Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience.
- Proven experience in Production Support, Site Reliability Engineering, or a similar operational support function.
- Strong Azure and AWS infrastructure support experience.
- Strong understanding of DNS, TCP/IP networking, routing, firewalls, and connectivity troubleshooting.
- Hands-on experience with Microsoft SQL Server administration, troubleshooting, performance tuning, locking, blocking, and query optimization.
- Experience supporting AKS and Kubernetes environments.
- Experience supporting GitHub and ArgoCD operational processes and deployment pipelines.
- Strong monitoring and observability expertise with Dynatrace, Azure Monitor, Azure Log Analytics, and AWS CloudWatch.
- Experience with distributed tracing, application performance monitoring, and log analysis.
- Strong understanding of cloud-native architectures and microservices.
- Demonstrated experience conducting Root Cause Analysis and managing complex production incidents.
- Experience supporting high-availability, mission-critical production systems.
- Strong knowledge of SRE principles, including observability, automation, reliability engineering, error budgets, Service Level Objectives (SLOs), and operational excellence.
- Strong analytical, problem-solving, stakeholder management, and communication skills.
Preferred Qualifications
- Experience supporting financial services or highly regulated environments.
- Knowledge of Infrastructure as Code tools including Terraform, ARM, Bicep, or CloudFormation.
- Experience with PowerShell, Bash, Python, or automation scripting.
- Experience implementing operational dashboards and automated remediation solutions.
- Understanding of DevOps practices and CI/CD methodologies.
- Industry certifications in Azure, AWS, Kubernetes, or Site Reliability Engineering disciplines.
- Ability to troubleshoot complex issues across multiple technology layers.
- Customer-focused mindset with a passion for automation, reliability, and continuous improvement.
What We Offer you
At FIS, we are as committed to growing our employees’ careers as our own business. We offer:
Opportunities to innovate in fintech
Inclusive and diverse team atmosphere
Professional and personal development
Resources to contribute to your community
Competitive salary and benefits
Privacy Statement
FIS is committed to protecting the privacy and security of all personal information that we process in order to provide services to our clients. For specific information on how FIS protects personal information online, please see the Online Privacy Notice.
Sourcing Model
Recruitment at FIS works primarily on a direct sourcing model; a relatively small portion of our hiring is through recruitment agencies. FIS does not accept resumes from recruitment agencies which are not on the preferred supplier list and is not responsible for any related fees for resumes submitted to job postings, our employees, or any other part of our company.
#pridepass