About this Lead Software & Production Engineer role at CAI
Req number:
R8507Employment type:
Full timeWorksite flexibility:
RemoteWho we are
CAI is a global services firm with over 9,000 associates worldwide and a yearly revenue of $1.3 billion+. We have over 40 years of excellence in uniting talent and technology to power the possible for our clients, colleagues, and communities. As a privately held company, we have the freedom and focus to do what is right—whatever it takes. Our tailor-made solutions create lasting results across the public and commercial sectors, and we are trailblazers in bringing neurodiversity to the enterprise.
Job Summary
As a Lead Software & Production Engineer, you will report to the Sr. Manager of Software Engineering and is responsible for ensuring the stability, reliability, performance, and operational excellence of critical business applications and services. This role provides technical leadership for production support activities, major incident management, problem resolution, root cause analysis, service reliability improvements, and operational readiness across the product portfolio.Job Description
We are looking for a Lead Software & Production Engineer serves as the primary technical escalation point for complex production issues and works closely with Engineering, Product Management, Infrastructure, Security, QA, and Business stakeholders to ensure high system availability and customer satisfaction. This role requires strong expertise in application support, troubleshooting, software engineering, DevOps practices, observability tools, incident response, and continuous improvement methodologies. This position will be full-time and remote.
What You'll Do
Lead production support activities for business-critical applications and services
Ensure high availability, performance, reliability, and operational stability of supported systems
Act as the technical escalation point for critical incidents and complex production issues
Coordinate cross-functional teams during incident response and service restoration efforts
Drive effective communication during incidents, outages, and service disruptions
Lead Major Incident Management (MIM) activities and support timely issue resolution
Conduct root cause analysis (RCA) and post-incident reviews to identify corrective and preventative actions
Develop and maintain incident management processes, support runbooks, and operational procedures
Track recurring issues and implement permanent fixes to reduce incident volume
Monitor support metrics and service trends to proactively identify risks and improvement opportunities
Establish and enforce production support standards, operational procedures, and service-level objectives
Drive application resiliency, scalability, performance, and reliability improvements
Implement monitoring, alerting, observability, and automated recovery capabilities
Lead initiatives focused on reducing technical debt and increasing platform stability
Ensure systems meet security, compliance, audit, and governance requirements
Provide technical leadership and guidance to production support engineers
Mentor team members and promote knowledge sharing across the organization
Collaborate with development teams to improve application supportability and operational readiness
Participate in architecture reviews to ensure operational requirements are considered throughout the software lifecycle
Define best practices for support transition, deployment readiness, and production acceptance
Drive automation initiatives to improve efficiency and reduce manual operational tasks
Improve incident detection, diagnosis, remediation, and reporting processes
Partner with engineering teams to implement self-healing and proactive support capabilities
Promote continuous improvement through metrics, analysis, and operational reviews
Partner with Product Owners, Engineering teams, Infrastructure teams, Vendors, and Business stakeholders
Provide operational status reporting and executive-level communication on system health and service performance
Participate in change management processes to assess operational risk and production readiness
Ensure support priorities align with business and operational objectives
Participate in production support on-call and escalation processes as required
Perform other job-related duties as assigned by management
Budget and Resource Planning
Understand accounting rules related to expense and capital activities
Identify opportunities for operational cost optimization
Monitor production support efforts and resource utilization
Support planning and prioritization of maintenance, reliability, and technical debt initiatives
What You'll Need
Required:
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline
8+ years of experience in Software Engineering, Application Support, Site Reliability Engineering (SRE), or Production Operations
3+ years of experience leading technical teams or production support organizations
Experience supporting mission-critical enterprise applications in a 24x7 environment
Demonstrated success leading major incident management and service restoration efforts
Experience working within globally distributed teams and complex enterprise environments
Deep understanding of Incident, Problem, Change, and Release Management processes
Experience leading critical incident response and coordinating cross-functional resolution teams
Strong analytical and troubleshooting capabilities for complex application and infrastructure issues
Proven ability to perform detailed root cause investigations for system failures and performance degradation
Establish corrective and preventative actions to improve overall service reliability
Experience implementing and supporting monitoring, alerting, logging, and observability platforms
Ability to define service health indicators, operational dashboards, and support metrics
Strong understanding of modern application architectures, APIs, microservices, databases, and cloud technologies
Working knowledge of Java, React, NextJS, SQL, and related enterprise application technologies
Ability to review code, identify supportability improvements, and collaborate effectively with development teams
Experience with CI/CD pipelines, deployment automation, and operational tooling
Ability to automate repetitive support tasks and operational processes
Familiarity with Infrastructure as Code (IaC), scripting, and cloud-native operational practices
Strong understanding of system performance, scalability, resiliency, and disaster recovery principles
Experience improving application uptime and reducing service disruptions
Ensure compliance with enterprise standards, security requirements, and operational controls
Participate in audits, operational reviews, and risk management activities
Lead, mentor, and develop production support engineers
Foster a culture of accountability, ownership, innovation, and continuous improvement
Promote knowledge sharing and operational best practices across teams
Ability to communicate technical issues and operational risks to both technical and executive audiences
Strong collaboration skills with Product, Engineering, Infrastructure, QA, Security, and Business teams
Preferred:
Experience supporting enterprise applications in cloud environments such as Azure or AWS
Experience with ServiceNow, Jira, Dynatrace, Splunk, AppDynamics, Datadog, Grafana, or similar observability platforms
Experience managing on-call support rotations and operational support models
Understanding of ITIL processes and Service Management best practices
Experience implementing SRE practices, service-level indicators (SLIs), service-level objectives (SLOs), and operational maturity frameworks
Physical Demands
Ability to safely and successfully perform the essential job functions
Sedentary work that involves sitting or remaining stationary most of the time with occasional need to move around the office to attend meetings, etc.
Ability to conduct repetitive tasks on a computer, utilizing a mouse, keyboard, and monitor
Reasonable accommodation statement
If you require a reasonable accommodation in completing this application, interviewing, completing any pre-employment testing, or otherwise participating in the employment selection process, please direct your inquiries to [email protected] or (888) 824 – 8111.