Sobre este puesto de Senior Manager, Site Reliability Engineering - Cloud Infrastructure en Futurex
About Futurex
Futurex is a Texas-based leader in hardware security modules (HSMs), enterprise key management, and data protection, trusted by global financial institutions, payment processors, and cloud providers for more than four decades. Futurex Cloud Infrastructure delivers secure, highly available services to customers around the world.
Role overview
The Senior Manager, Site Reliability Engineering – Cloud Infrastructure owns the availability, security, and operational excellence of Futurex Cloud Infrastructure worldwide. You will lead a globally distributed SRE team that runs production infrastructure across multiple regions, cloud providers, and datacenters.
This role sits at the intersection of cloud operations, security, and compliance. Our customers depend on services that cannot go down, so reliability here is measured in customer trust as much as uptime.
You will report to executive leadership and partner closely with Engineering, Product, Security, Compliance, and Customer Support.
Key responsibilities
Team leadership
• Lead, hire, and develop a global SRE team across multiple time zones, with clear ownership and follow-the-sun coverage.
• Own the global on-call program, including rotation design, response time standards, escalation paths, and coverage across all regions.
• Hold the team accountable for incident response, remediation follow-through, and adherence to operational standards.
• Set the SRE roadmap, staffing plan, and operating budget.
Reliability and service levels
• Define and own SLIs, SLOs, and error budgets for Cloud Infrastructure services, and use them to balance stability against release velocity.
• Partner with Sales and Legal so customer SLA commitments are consistently met.
• Track SLA compliance and error budget consumption across all services.
Operational reporting
• Deliver regular reporting to executive leadership on uptime, SLA performance, incident trends, and on-call metrics such as response times and escalation volume.
• Maintain dashboards that give leadership and stakeholders real-time visibility into service health.
• Produce customer-facing incident reports and root cause analyses for significant events.
Incident management
• Own the incident response process, including severity definitions, escalation paths, and customer communications.
• Lead postmortems and drive remediation items to closure.
• Serve as the senior escalation point during major incidents.
Infrastructure and operations
• Review monitoring, observability, and alerting standards.
• Maintain and regularly test disaster recovery and business continuity plans against defined RTO and RPO targets.
Security and compliance
• Operate within the controls required by industry security and compliance frameworks.
• Oversee secure operational procedures, including dual control and chain of custody where required.
• Enforce strict tenant isolation, least-privilege access, and change management.
• Support internal and external audits with accurate operational evidence.
Requirements
Required qualifications
• 10+ years in SRE, DevOps, or production infrastructure roles, including 5+ years managing engineering teams.
• Experience leading geographically distributed teams across multiple time zones.
• Proven ownership of a production SaaS or cloud platform with contractual SLAs and 24x7 operations.
• Hands-on depth in at least one major cloud provider (AWS, GCP, or Azure), with working knowledge of the others.
• Strong background in Linux, networking, infrastructure as code (Terraform or similar), and CI/CD.
• Experience operating under compliance frameworks such as PCI DSS, SOC 2, or ISO 27001.
• A track record of building incident management, observability, and DR programs.
• Clear written and verbal communication with executives, customers, and auditors.
• Bachelor's degree in Computer Science, Engineering, or equivalent experience.
Preferred qualifications
• Experience operating security-sensitive or cryptographic infrastructure.
• Familiarity with security certification standards such as FIPS 140.
• Experience running hybrid environments that combine public cloud with physical datacenter hardware.
• Experience leading datacenter migrations or regional expansions.
• Background in financial services or another highly regulated industry.
Benefits
- Health, dental, vision, life, and short/long-term disability insurance
- Paid vacation, holidays, and sick leave
- Competitive compensation and opportunities for advancement
- Retirement plan with employer contribution match
- Welcoming, family-style corporate culture uniquely suited to fast-paced, entrepreneurial, and motivated individuals
- One of San Antonio’s “Best Places to Work” for nine consecutive years