Senior DevOps & Site Reliability Engineer
Executive Placements · North Riding
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →We are looking for an experienced Senior DevOps & Site Reliability Engineer (SRE) to design, build and operate highly available, secure, scalable and automated enterprise technology platforms.
This is a senior hands-on engineering role spanning DevOps, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and DevSecOps .
The successful candidate will work across engineering and delivery teams to improve platform reliability, deployment velocity, resilience, automation, operational efficiency and production performance , while supporting mission-critical enterprise applications.
Key Responsibilities
DevOps & Platform Engineering
Design, build and maintain cloud-native infrastructure and platform services .
Develop and maintain Infrastructure as Code (IaC) solutions.
Automate infrastructure provisioning, configuration and operational processes.
Build reusable engineering tools, deployment templates and platform components.
Establish and standardise platform engineering practices across multiple delivery teams.
Identify opportunities to reduce manual intervention and increase engineering automation.
CI/CD & Release Automation
Design, implement and maintain enterprise-grade CI/CD pipelines for application and infrastructure deployments.
Implement automated testing, security scanning, code-quality controls and release automation.
Enable automated deployments, rollback and recovery processes.
Improve deployment frequency while reducing change and deployment risk.
Continuously optimise software delivery and release-management processes.
Site Reliability Engineering
Implement and mature Site Reliability Engineering practices across production environments.
Define, monitor and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs) .
Improve application and platform availability, scalability, resilience and performance .
Lead production incident response, troubleshooting, problem management and Root Cause Analysis (RCA) .
Drive proactive reliability improvements and reduction of technical debt.
Improve Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR) .
Azure Cloud Engineering
Design, implement and operate enterprise Microsoft Azure environments.
Work extensively with technologies such as
Azure Kubernetes Se
.special-hidden { display: none; }