Jobs › Companies › Systems Engineering Solutions Corporation › Reliability Engineer

Sobre este puesto de Reliability Engineer en Systems Engineering Solutions Corporation

Systems Engineering Solutions Corporation · Presencial · Hanscom Air Force Base, Massachusetts, United States

This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a Reliability Engineer. The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of mission‑critical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities.

 

 

Location: This position will be hybrid remote. Candidates will be required to work onsite as needed. Candidates preferred to be located near Hanscom AFB (Boston, MA).

Requirements

System Reliability & Availability

  • Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments
  • Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
  • Identify reliability risks and implement mitigation strategies across the system lifecycle
  • Conduct capacity planning and performance modeling to ensure systems scale to meet demand

Monitoring, Observability & Alerting

  • Implement and manage monitoring, logging, and tracing solutions to provide full system observability
  • Define actionable alerting thresholds that minimize noise and enable rapid incident detection
  • Analyze trends and metrics to proactively identify potential reliability issues

Incident Response & Problem Management

  • Participate in on‑call rotations and lead incident response activities for production systems
  • Coordinate troubleshooting efforts across development, infrastructure, and security teams
  • Conduct post‑incident reviews (PIRs) and develop corrective and preventive action plans
  • Track recurring issues and ensure root causes are resolved

Automation & Engineering Excellence

  • Automate operational tasks to reduce manual intervention and operational risk
  • Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR)
  • Promote “automation over toil” and standardize operational workflows

Reliability‑Focused Engineering

  • Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability
  • Validate disaster recovery (DR) and business continuity plans; test failover mechanisms
  • Support chaos engineering, fault injection testing, and resilience validation where appropriate

Collaboration & Governance

  • Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives
  • Document system reliability standards, runbooks, and operational procedures
  • Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls)

 

Required Skills:

·       Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree.  

·       Active Secret clearance at a minimum required to start  

·       US citizenship required 

·       Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services

·       Experience with containerized environments (Docker, Kubernetes)

·       Familiarity with CI/CD pipelines and deployment automation

·       SLOs and error budgets

·       Capacity modeling and performance testing

·       Strong understanding of:

·       Distributed systems and high‑availability architectures

·       Linux/Windows system administration

·       Networking fundamentals (DNS, TCP/IP, load balancing)

·       Hands-on experience with:

·       Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)

·       Infrastructure as Code (Terraform, ARM, CloudFormation)

·       Scripting or programming languages (Python, Bash, Go, PowerShell, or similar)

·       Experience supporting incident management and on‑call operations

 

Preferred Skills

  • Experience with USAF Cloud One or Platform 1. 
  • Experience with Zero Trust Architecture 
  • Cloud certifications in AWS, Azure, Google, or Oracle clouds 

Benefits

SES provides a competitive salary and the following benefits:

  • Medical
  • Dental
  • Vision
  • AD&D
  • STD
  • LTD
  • Company paid Life Insurance
  • 401k with employer contribution
  • Paid Time Off
  • Pet Insurance
¿Listo para postularte en Systems Engineering Solutions Corporation?
Postúlate en Systems Engineering Solutions Corporation

Sobre Systems Engineering Solutions Corporation

SES Corporation is a rapidly growing Small Business and Federal Contractor. We are an information technology consulting services firm that partners with its customers in every aspect of their engineering lifecycle.

We pride ourselves on being able to seamlessly integrate the most talented technical resources with our client’s team and mission. Our resources are experts in Systems Engineering, Systems Integration, System Test, Infrastructure Development, Deployment, and Lab Support. Within these disciplines, we apply lessons learned from several successful mission critical programs to effectively manage technical risk exposure to our clients.

Ver todos los empleos en Systems Engineering Solutions Corporation →

Empleos similares

Cloudlinux
Senior Database Reliability Engineer (DBRE) (remote work)
Cloudlinux
⚡ Postúlate pronto Warsaw, Masovian Voivodeship,... · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 4h
Rowan Digital Infrastructure
Reliability Engineer - Commissioning
Rowan Digital Infrastructure
⚡ Postúlate pronto Temple, TX Híbrido $130,000–$150,000
● Nuevo 👁 Visto ✓ Postulado hace 5h
Camunda
Senior Site Reliability Engineer
Camunda
⚡ Postúlate pronto Remote · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 6h
Megaport
Senior Site Reliability Engineer
Megaport
⚡ Postúlate pronto Sao Paulo Híbrido
● Nuevo 👁 Visto ✓ Postulado hace 6h
Vast
Thermal Fluids Design Reliability Engineer
Vast
⚡ Postúlate pronto Long Beach, California, United... Presencial $162,360–$265,392
● Nuevo 👁 Visto ✓ Postulado hace 6h
Vast
Structures Design Reliability Engineer
Vast
⚡ Postúlate pronto Long Beach, California, United... Presencial $162,360–$265,392
● Nuevo 👁 Visto ✓ Postulado hace 6h
Vast
Propulsion Design Reliability Engineer
Vast
⚡ Postúlate pronto Long Beach, California, United... Presencial $162,360–$265,392
● Nuevo 👁 Visto ✓ Postulado hace 6h
Rent the Runway
Site Reliability Engineer
Rent the Runway
⚡ Postúlate pronto Galway, Ireland Híbrido
● Nuevo 👁 Visto ✓ Postulado hace 8h
NL
Senior Cloud Platform & Site Reliability Engineering Lead
National Life Insurance Company
⚡ Postúlate pronto Addison, TX; Montpelier, VT Presencial $136,875–$200,750
● Nuevo 👁 Visto ✓ Postulado hace 8h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Systems Engineering Solutions Corporation

Ver todos los empleos en Systems Engineering Solutions Corporation →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis