Mission
Own the functional design and capability roadmap for an enterprise-scale observability platform — the trusted signal foundation that enables AIOps across a large, complex technology estate. The platform creates the telemetry, health models, SLOs, dashboards, and detection signals required by event management, AI-guided diagnosis, and automated remediation, and sits at the start of the detect-to-correct value chain.
You will help advance a modern, vendor-agnostic, composable observability platform built on open industry standards including OpenTelemetry and Prometheus, competing with the best commercial tools in a rapidly evolving market. Working at the intersection of SRE, user experience, architecture, and engineering, you will translate product strategy into coherent capabilities — trusted telemetry, real-time monitoring, dashboards, synthetic monitoring, RUM, alerting, and SLO management — that strengthen proactive detection, signal quality, and continuous reliability across highly complex, interdependent application and platform landscapes.
What You Will Enable
- Earlier, more reliable detection — Give SRE and operations teams continuous visibility into application health, dependencies, and end-user experience, with measurable SLOs and detection signals that identify degradation before users report it.
- A trusted signal foundation for AIOps — Deliver consistent, contextual, and governable telemetry that can be correlated, enriched, and acted on by event management, AI-guided diagnosis, automation, and future agentic capabilities.
- Faster cross-domain diagnosis — Ensure telemetry carries the identifiers and context needed to map signals to separately managed topology and service models, while supporting topology discovery from telemetry such as distributed traces.
- Coherent self-service observability — Provide clear integration patterns, reusable dashboards, monitoring experiences, alerting, synthetic monitoring, RUM, and SLO capabilities that teams can adopt consistently and at scale.
- Continuous reliability improvement — Turn operational evidence into better health models, SLOs, monitoring, dashboards, and alerting so recurring weaknesses are identified and observability improves over time.
Responsibilities
- Own end-to-end functional design — Define how platform capabilities work individually and together across telemetry ingestion, real-time monitoring, dashboards, synthetic monitoring and RUM, alerting, and SLO management; eliminate functional gaps, duplication, and inconsistent experiences.
- Own the capability roadmap — Translate product strategy, SRE needs, and the competitive observability landscape into a clear capability-level roadmap, defining what each capability must deliver to strengthen the detect-to-correct value chain and enable AIOps.
- Lead deep technical discovery — Work directly with SREs, operations teams, platform engineers, application owners, and architects to understand operational pain points, current workflows, telemetry needs, failure modes, and the complex technical setups of observed applications and platforms.
- Define rigorous product requirements — Turn discovery into functional specifications, user journeys, telemetry and integration requirements, non-functional expectations, user stories, and precise acceptance criteria that engineering teams can implement with confidence.
- Guide engineering design — Partner closely with engineers and architects on APIs, telemetry models, integration patterns, configuration models, scalability, security, and extensibility; challenge design choices and make functional trade-offs grounded in technical reality.
- Enable topology integration — Define how telemetry signals are consistently identifiable and mappable to topology and service context managed by the responsible platform, and how telemetry such as traces can support topology discovery and enrichment.
- Create a coherent, state-of-the-art product experience — Ensure consistent concepts, workflows, defaults, and self-service patterns across capabilities and personas, benchmarked against leading observability platforms, OpenTelemetry evolution, SRE practices, and emerging AIOps patterns.
- Assure functional quality and measurable value — Validate that delivered capabilities match design intent, work coherently end to end, and contribute to SLO coverage, system-detected incidents, MTTD, MTTR, automation readiness, adoption, and user satisfaction.
About us
We produce globally recognized brands and we grow the best business leaders in the industry. With a portfolio of trusted brands as diverse as ours, it is paramount our leaders can lead with courage the vast array of brands, categories and functions. We serve consumers around the world with one of the strongest portfolios of trusted, quality, leadership brands, including Always®, Ariel®, Gillette®, Head & Shoulders®, Herbal Essences®, Oral-B®, Pampers®, Pantene®, Tampax® and more. Our community includes operations in approximately 70 countries worldwide. Visit http://www.pg.com to know more.