À propos de ce poste Systems Development Engineer (SRE/DevOps) chez Modeln
Job Responsibilities:
• Develop and maintain infrastructure and configuration as code using CloudFormation, Terraform, Ansible, and related automation tools.
• Administer and optimize AWS environments, including core services, networking, security, and architecture for availability, performance, and cost.
• Manage and support Kubernetes clusters and containerized workloads, including configuration, scaling, and upgrades.
• Design, implement, and evolve end‑to‑end monitoring and observability frameworks using tools such as Open Telemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic, or similar platforms.
• Create and maintain dashboards, logs, traces, SLIs/SLOs, and automated alerting systems to ensure reliability and rapid detection of anomalies.
• Embed observability, CI/CD best practices, and operational readiness into all stages of the software development lifecycle in partnership with engineering teams.
• Lead or participate in incident response, troubleshooting, and root cause analysis for production incidents, using observability data to drive fast resolution.
• Automate operational tasks, runbooks, and incident remediation workflows to reduce toil and improve service reliability.
• Contribute to risk mitigation, backup, and disaster recovery strategies, including periodic testing and continuous improvement.
• Participate in shared after hours support and project work as needed.
Job Qualification:
• 2-4 years of experience designing, implementing, and maintaining CI/CD pipelines (e.g., Harness, GitHub Actions, ArgoCD or similar tools).
• Hands‑on experience with automation tools and Infrastructure as Code / Configuration as Code (CloudFormation, Terraform, Ansible).
• Strong understanding of Infrastructure as Code and Configuration as Code principles and patterns.
• Solid grasp of the software development lifecycle and modern SRE/DevOps practices.
• AWS administration and architecture experience, including networking, security, IAM, and core services.
• Experience operating Kubernetes clusters (EKS or other distributions) and containerized workloads.
• Deep experience with monitoring and observability tools such as OpenTelemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic, or equivalent, including metrics, logs, and traces.
• Ability to define and track SLIs/SLOs and use them to guide reliability improvements.
• Proficiency in Linux administration, including system configuration, troubleshooting, and performance tuning.
• Programming/scripting skills in at least one language such as Python, Go, or Rust for automation, tooling, and observability integrations.
• Solid understanding of networking, load balancing, and performance tuning.
• Experience troubleshooting complex distributed systems, supporting incident response, and driving root cause analysis.
• Familiarity with risk mitigation, backup, and disaster recovery concepts.
Preferred
• Experience building unified observability platforms or standardized dashboards for multiple services/teams.
• Experience with GitOps workflows and tools for declarative infrastructure and application delivery.
• Background in incident command and post‑mortem frameworks.
• Experience integrating observability and reliability practices into microservices and/or serverless architectures.
• Experience integrating testing, security and compliance checks into CI/CD pipelines.