About LivePerson
LivePerson (NASDAQ: LPSN) is the global leader in enterprise conversations. Hundreds of the world’s leading brands — including HSBC, Chipotle, and Virgin Media — use our award-winning Conversational Cloud platform to connect with millions of consumers. We power nearly a billion conversational interactions every month, providing a uniquely rich data set and safety tools to unlock the power of Conversational AI for better customer experiences. At LivePerson, we foster an inclusive workplace culture that encourages meaningful connection, collaboration, and innovation.
Role Overview
LivePerson transforms customer care from voice calls to mobile messaging. Our cloud-based software platform, LiveEngage, allows brands with millions of customers and tens of thousands of care agents to deliver digital experiences at scale. The Cloud DevOps team at LivePerson is looking for a Principal Site Reliability Engineer (Principal SRE) to provide technical leadership across the organization and help shape the reliability, scalability, security, and operational excellence of our cloud platforms and services. The ideal candidate is a highly experienced engineer who can solve complex technical problems, influence engineering teams without direct authority, and drive large-scale initiatives from strategy through implementation.
Responsibilities
- Provide technical leadership and direction for reliability and platform engineering across multiple teams.
- Design and evolve highly available, scalable, secure, and resilient systems, with a strong focus on Google Cloud Platform (GCP).
- Lead complex, cross-team initiatives across cloud infrastructure, Kubernetes, networking, observability, security, and software delivery.
- Define and drive SRE practices including SLOs, SLIs, error budgets, reliability reviews, capacity planning, and operational readiness.
- Lead technical response to complex production incidents and drive long-term corrective and preventative actions.
- Develop automation and infrastructure-as-code solutions using Python, Terraform, Ansible, Bash, and other modern engineering tools.
- Provide technical leadership for Kubernetes platforms and containerized workloads, including architecture, scalability, performance, and reliability.
- Establish and evolve GitOps deployment practices using Kubernetes, Helm, and FluxCD.
- Define and improve CI/CD practices using GitLab CI/CD, focusing on reliability, security, scalability, and developer experience.
- Drive observability improvements using metrics, logs, traces, dashboards, and actionable alerting.
- Identify systemic reliability risks, technical debt, and architectural weaknesses and drive sustainable solutions.
- Create reusable platforms, tooling, and engineering patterns that enable teams to operate reliable services independently.
- Mentor Senior SREs, Team Leads, and other technical leaders while raising engineering standards across the organization.
- Influence architecture and technical decisions across teams and communicate complex technical concepts and trade-offs to engineering leadership.