About Delinea

Delinea is a pioneer in securing human and machine identities through intelligent, centralized authorization, empowering organizations to seamlessly govern their interactions across the modern enterprise. Leveraging AI-powered intelligence, Delinea’s leading cloud-native Identity Security Platform applies context throughout the entire identity lifecycle – across cloud and traditional infrastructure, data, SaaS applications, and AI. It is the only platform that enables you to discover all identities – including workforce, IT administrator, developers, and machines – assign appropriate access levels, detect irregularities, and respond to threats in real-time. With deployment in weeks, not months, 90% fewer resources to manage than the nearest competitor, and a 99.995% uptime, Delinea delivers robust security and operational efficiency without compromise. Learn more about Delinea on Delinea.com, LinkedIn, X, and YouTube.

Site Reliability Engineer Summary

Our growing technology company is seeking an experienced Site Reliability Engineer with deep Azure expertise to help maintain the availability, performance, and reliability of our critical SaaS applications. In this role, you will own and drive automation, monitoring, incident response, and infrastructure improvements across our multi-cloud, multi-region environment, working closely with senior engineering and cross-functional teams.

What You’ll Do

  • Own the availability and performance of production SaaS applications running on Azure (AKS, App Service, Redis, SQL, Service Bus etc), across multiple geographic regions.
  • Lead troubleshooting and resolution of cloud infrastructure and application issues, including AKS pod/node failures, deployment rollbacks, ingress and networking issues, and resource/autoscaling problems.
  • Participate in an on-call rotation (including weekends) and drive incident response from detection through resolution, with a primary focus on customer experience and minimizing impact.
  • Drive improvements to disaster recovery, failover, and incident management processes across multi-region deployments.
  • Build and maintain automation scripts and monitoring tools to reduce manual toil and streamline operational tasks.
  • Author post-incident reviews (RCAs), identify root causes, and drive preventive action items to closure.
  • Partner with senior engineers and cross-functional teams to implement and improve reliability, observability, and performance best practices.
  • Contribute to continuous improvement initiatives across infrastructure, tooling, and process.
  • Communicate clearly with customer-facing stakeholders when incidents require external status updates or written incident summaries.

Why work at Delinea?

  • We're passionate problem-solvers helping the world's largest organizations protect what matters most: their human and machine identities.
  • We invest in people who are smart, self-motivated, and collaborative.
  • What we offer in return is meaningful work, a culture of innovation and great career progression.
  • We take care of our employees. We offer competitive salaries, a meaningful bonus program, and excellent benefits, including healthcare insurance, as well as pension/retirement matching, comprehensive life insurance, an employee assistance program, time off plans, and paid company holidays.

What You’ll Need

  • 5+ years of relevant experience in Site Reliability Engineering, DevOps, or Cloud Administration, with demonstrated ownership of production systems.
  • Hands-on experience administering Azure environments, including AKS (Kubernetes), core Azure services, cloud networking, and cloud security fundamentals.
  • Solid understanding of monitoring, logging, and alerting practices (e.g., Datadog, Azure Monitor, ELK stack), including hands-on troubleshooting with log analysis and stack traces using Datadog APM.
  • Familiarity with networking fundamentals: firewalls, load balancers, VPNs, DNS, and routing.
  • Experience with automation and scripting (PowerShell, Python, or similar).
  • Practical understanding of backup, redundancy, and disaster recovery strategies in cloud environments, including geo-redundant / multi-region deployments.
  • Strong ownership mindset across the full incident lifecycle, from detection through post-mortem, with a customer-first approach.
  • Comfort communicating clearly and professionally in written form, including incident status updates and post-incident summaries.

We’d Love to See

  • Experience with AWS Cloud Platform.
  • Experience with CI/CD tools such as Azure DevOps.
  • Experience with infrastructure-as-code tools such as Terraform or ARM templates.
  • Prior experience operating SaaS products with regional tenant architectures (e.g., multiple geo-specific production environments).