Sobre esta vaga de IND - Senior Staff Engineer, Reliability na The Hartford
We’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals – and to help others accomplish theirs, too. Join our team as we help shape the future.
Job Summary:**
We are seeking an experienced and highly skilled Senior Staff Engineer, Reliability to join our team at Hartford Global Services Private Limited. In this critical role, you will be responsible for leading initiatives to enhance the reliability, availability, and performance of our systems and applications. You will drive the adoption of best practices in Site Reliability Engineering (SRE), collaborate with development and operations teams, and mentor junior engineers to build a robust and resilient infrastructure.
Job Responsibilities:**
* Lead the design, development, and implementation of robust reliability solutions for complex distributed systems.
* Drive the adoption of SRE principles and practices across engineering teams, including incident response, post-mortems, SLO/SLA definition, and error budget management.
* Proactively identify potential failure points and performance bottlenecks within our infrastructure and applications, and propose effective mitigation strategies.
* Develop and implement advanced monitoring, alerting, and logging solutions to ensure high visibility into system health and performance.
* Automate operational tasks and build tools to improve efficiency and reduce manual effort.
* Collaborate closely with development teams to ensure reliability is designed into new features and services from inception.
* Conduct root cause analysis for production incidents and implement preventative measures to avoid recurrence.
* Mentor and guide junior engineers on reliability best practices, SRE principles, and effective troubleshooting techniques.
* Stay abreast of industry trends and emerging technologies in reliability engineering, cloud platforms, and distributed systems.
* Participate in on-call rotations to support critical production systems.
Job Qualifications:**
* Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.
* 8+ years of experience in software development, DevOps, or Site Reliability Engineering roles, with a strong focus on system reliability and performance.
* Proven experience in designing, building, and operating highly available and scalable distributed systems.
* Expertise in at least one major cloud platform (AWS, Azure, GCP) and its associated services.
* Strong programming skills in languages such as Python, Go, Java, or C++.
* In-depth knowledge of monitoring tools (e.g., Prometheus, Grafana, ELK stack) and observability best practices.
* Solid understanding of containerization technologies (Docker, Kubernetes) and microservices architectures.
* Experience with CI/CD pipelines and infrastructure as code (Terraform, Ansible).
* Excellent problem-solving, analytical, and communication skills.
* Demonstrated ability to lead technical initiatives and mentor other engineers.
* Experience with incident management and post-mortem processes.