About this Lead Software Engineer role at The Hartford
We’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals – and to help others accomplish theirs, too. Join our team as we help shape the future.
Job Summary
Hartford is seeking a motivated Reliability & Observability Engineer to support production reliability and application-level observability for critical technology enablement platforms. This is an early-career role designed for engineers building foundational expertise in monitoring, alerting, and operational stability.
The engineer will work closely with senior reliability engineers, application development teams, and platform partners to improve system visibility, identify issues early, and support incident response. This role focuses on observability and operational reliability, not performance or load testing.
Key Responsibilities
- Observability Platform Support
Support the configuration and day-to-day operation of observability capabilities for Agile Enablement SaaS platforms (such as Rally, Clarity, and Apptio) and their integrations, in collaboration with senior engineers, SRE teams, and vendors. - Application-Level Monitoring on AWS
Monitor application health, logs, and key performance indicators for applications running on AWS, with a focus on application behavior and user experience rather than infrastructure administration. - Dashboard, Visualization & Alert Development
Build and maintain dashboards, metrics, and alert rules using Splunk, ensuring alerts are actionable, relevant, and continuously refined. - Performance Baseline & Trend Analysis
Maintain baseline metrics, identify early signs of degradation, document observations, and escalate findings to senior engineering partners. - Incident Support & RCA Participation
Act as an observability point of contact during incidents by gathering logs and metrics from Splunk and contributing supporting data for root-cause analysis and problem-management activities. - Automation & Scripting Assistance
Write or modify basic automation scripts (e.g., Python or shell) to assist with monitoring, log processing, or alert tuning, under the guidance of senior engineers. - Documentation & Team Enablement
Maintain documentation for dashboards, alerts, and monitoring configurations, and create simple guides to help application teams use observability tools effectively. - Proactive Monitoring & Reliability Improvements
Assist in identifying opportunities for proactive alerting and contribute to foundational remediation or self-healing initiatives that reduce manual intervention. - Continuous Learning & Skill Development
Actively develop skills in observability, reliability engineering, scripting, and cloud-native monitoring practices, applying new learning to improve system stability.
Qualifications & Skillset
Must Have
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field
- 2–4 years of experience in reliability engineering, observability, production support, or a related role
- Hands-on experience with Splunk for log analysis, dashboards, alerts, and operational monitoring
- Experience supporting application-level monitoring for applications running on AWS
- Familiarity with incident support processes and collaboration across distributed or remote teams
- Strong analytical, troubleshooting, and problem-solving skills
Nice to Have
- Exposure to SaaS or Agile Enablement platforms (e.g., Rally, Jira, Clarity, Apptio, TargetProcess, ServiceNow)
- Basic familiarity with Agile, DevOps, or Site Reliability Engineering (SRE) practices
- Introductory scripting experience (Python, shell, or similar)
Preferred Certifications (Optional)
- Splunk Core Certified Power User
- AWS Cloud Practitioner