AHEAD builds platforms for digital business. By weaving together advances in cloud infrastructure, automation and analytics, and software delivery, we help enterprises deliver on the promise of digital transformation.
At AHEAD, we prioritize creating a culture of belonging, where all perspectives and voices are represented, valued, respected, and heard. We create spaces to empower everyone to speak up, make change, and drive the culture at AHEAD.
We are an equal opportunity employer, and do not discriminate based on an individual's race, national origin, color, gender, gender identity, gender expression, sexual orientation, religion, age, disability, marital status, or any other protected characteristic under applicable law, whether actual or perceived.
We embrace all candidates that will contribute to the diversification and enrichment of ideas and perspectives at AHEAD.
AHEAD is seeking a Principal Observability & Reliability Architect to join our Observability practice within Intelligent Operations. You are the senior technical authority and client advisor for observability and reliability: an expert in Dynatrace at enterprise scale, fluent in open standards and adjacent platforms, and accountable for the outcomes your programs deliver in reliability, efficiency, cost, and adoption. You will architect enterprise observability solutions, lead the SRE practices that turn telemetry into dependable services, advise executive stakeholders, serve as escalation point for delivery teams, and grow the practice through pre-sales, offerings, reusable content, and mentorship. This is a full-time, remote role based in the United States, with occasional travel (0 to 15%) based on engagement needs.
Responsibilities
Architect Dynatrace at enterprise scale: multi-tenant and hybrid designs, ActiveGate topology, OneAgent and OpenTelemetry instrumentation strategy, consumption and licensing governance, tagging and ownership models, and AI-workload observability.
Design end-to-end observability architectures across monitoring, logging, metrics, tracing, telemetry pipelines, alerting, event correlation, and service visibility in hybrid and multi-cloud environments.
Establish and mature SRE practices with clients: SLIs and SLOs, error budgets, production readiness, incident response and postmortems, and reliability roadmaps tied to business impact.
Lead assessment and advisory workshops that define use cases, maturity roadmaps, operating models, and adoption strategies, including AIOps and automation with Davis AI, alerting profiles, and workflows.
Define standards for telemetry onboarding, naming, tagging, service ownership, access, dashboards, alert governance, runbooks, and operational handoff, and advise on telemetry governance: data quality, retention, sampling, cardinality, and cost.
Lead modernization initiatives: tool and alert rationalization, telemetry strategy, migration to Dynatrace from legacy APM and monitoring platforms, and integration with ITSM, CMDB, event management, and automation platforms.
Lead complex programs, own solution design and architectural review, articulate trade-offs, and act as escalation point for delivery teams.
Advise client executives on platform strategy and value realization, and report on program health and outcomes.
Provide architecture and quality oversight across engagements, intervening early where outcomes are at risk.
Support pursuits as the technical expert: scoping, positioning, demonstrations, estimate validation, and client-facing technical narratives.
Build reusable assets such as Dynatrace reference architectures, governance models, accelerators, and points of view, and contribute thought leadership through content, partner material, and conference speaking.
Mentor architects and consultants across the practice, and maintain Dynatrace Professional certification plus professional-level certification on at least one additional platform.
Qualifications
8 or more years of hands-on experience in observability, APM, SRE, or related disciplines, including architecting enterprise-scale solutions across distributed systems and multi-cloud estates.
4 or more years of hands-on enterprise Dynatrace experience, including architecture and governance, OneAgent and Kubernetes deployment, Smartscape and PurePath, Grail and DQL, Davis AI, SLOs, workflows and automation, and ITSM integration; Dynatrace Professional certification held or attainable within six months.
Applied SRE experience defining SLIs and SLOs, operating error budgets, running production readiness and incident reviews, and leading reliability programs that measurably reduce incidents and time to resolve.
Working expertise in OpenTelemetry, Prometheus, and the Grafana ecosystem, and in public cloud monitoring services on AWS, Azure, or GCP.
Strong knowledge of telemetry governance (routing, transformation, enrichment, retention, access, cost) and experience defining enterprise standards for dashboards, alerts, tagging, and service ownership.
Expert knowledge of platform architecture, API integration patterns, and automation frameworks (Terraform, Ansible, Python, or similar).
Strong consultative and executive-facing presence, with experience leading workshops and translating business needs into architecture and delivery plans.
Demonstrated leadership mentoring technical teams; familiarity with ITIL, ITSM, and DevOps principles and with scoping, estimating, and change control in consulting delivery.
Strong communication skills, attention to detail, and a self-starting work style; able to travel occasionally (0% to 15%).
Preferred Qualifications
Both Dynatrace Professional certifications, or Dynatrace partner program experience.
Migration experience from AppDynamics, New Relic, Datadog, Splunk, or legacy monitoring platforms to Dynatrace; LogicMonitor or other infrastructure monitoring experience is a plus.
Telemetry pipeline tools such as OpenTelemetry Collector, Grafana Alloy, Fluent Bit, Kafka, Cribl, or Vector, plus Kubernetes, CI/CD, and infrastructure as code.
Integration with ServiceNow, Jira Service Management, PagerDuty, Opsgenie, BigPanda, or xMatters.
Published thought leadership, conference speaking, or ownership of a named offering or accelerator; relevant cloud, SRE, ITIL, or FinOps certifications are a plus.
The compensation range indicated in this posting reflects the On-Target Earnings (“OTE”) for this role, which includes a base salary and any applicable target bonus amount. This OTE range may vary based on the candidate’s relevant experience, qualifications, and geographic location.
Why AHEAD:
Through our daily work and internal groups like Moving Women AHEAD and RISE AHEAD, we value and benefit from diversity of people, ideas, experience, and everything in between.
We fuel growth by stacking our office with top-notch technologies in a multi-million-dollar lab, by encouraging cross department training and development, sponsoring certifications and credentials for continued learning.
USA Employment Benefits include:
- Medical, Dental, and Vision Insurance
- 401(k)
- Paid company holidays
- Paid time off
- Paid parental and caregiver leave
Use of AI:
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, assessing responses, or to capture recordings and create transcriptions or summaries during interviews. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans.
You may opt-out of the review or analysis of your application and resume by AI tools by using the
General Application. Please include the role you wish to apply for in the Additional Information field. You may also choose to opt-out of recording and transcription at any time, including after joining an interview. Candidates will not be penalized for choosing to opt-out.