Senior DevOps & Site Reliability Engineer at Datonomy Solutions
Datonomy Solutions (Cape Town) · Sandton, Gauteng · R90,000 - R100,000 per month
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →We are looking for an experienced Senior DevOps & Site Reliability Engineer (SRE) to design, build and operate highly available, secure, scalable and automated enterprise technology platforms. This is a senior hands-on engineering role spanning DevOps, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and DevSecOps .
The successful candidate will work across engineering and delivery teams to improve platform reliability, deployment velocity, resilience, automation, operational efficiency and production performance , while supporting mission-critical enterprise applications.
Key Responsibilities DevOps & Platform Engineering
Design, build and maintain cloud-native infrastructure and platform services .
- Develop and maintain Infrastructure as Code (IaC) solutions.
- Automate infrastructure provisioning, configuration and operational processes.
- Build reusable engineering tools, deployment templates and platform components.
- Establish and standardise platform engineering practices across multiple delivery teams.
- Identify opportunities to reduce manual intervention and increase engineering automation.
CI/CD & Release Automation
Design, implement and maintain enterprise-grade CI/CD pipelines for application and infrastructure deployments.
- Implement automated testing, security scanning, code-quality controls and release automation.
- Enable automated deployments, rollback and recovery processes.
- Improve deployment frequency while reducing change and deployment risk.
- Continuously optimise software delivery and release-management processes.
Site Reliability Engineering
Implement and mature Site Reliability Engineering practices across production environments.
- Define, monitor and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs) .
- Improve application and platform availability, scalability, resilience and performance .
- Lead production incident response, troubleshooting, problem management and Root Cause Analysis (RCA) .
- Drive proactive reliability improvements and reduction of technical debt.
- Improve Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR) .
Azure Cloud Engineering
Design, implement and operate enterprise Microsoft Azure environments.
- Work extensively with technologies such as:
Azure Kubernetes Service ( AKS )
- Azure App Services
- Azure Networking
- Azure Monitor
- Azure Storage
- Azure Identity Services
- Design highly available and disaster-recovery-capable environments.
- Optimise cloud environments for performance, resilience, security and cost .
- Support hybrid-cloud and multi-cloud environments where required.
Containers & Kubernetes
Build, deploy and support containerised applications using Docker and Kubernetes .
- Manage Kubernetes environments, particularly Azure Kubernetes Service (AKS) .
- Develop and maintain deployment configurations using Helm .
- Support container-platform reliability, scalability and operational performance.
- OpenShift experience would be advantageous.
Infrastructure as Code & Automation Hands-on experience with technologies such as:
Terraform
- Bicep
- ARM Templates
- Ansible
Candidates should be comfortable using Infrastructure as Code to build repeatable, scalable and governed enterprise infrastructure.
Monitoring & Observability
Implement comprehensive logging, monitoring, metrics, tracing and alerting .
- Build operational dashboards and platform insights.
- Establish enterprise observability standards.
- Implement proactive and predictive monitoring.
- Use observability information to improve application and infrastructure reliability.
Relevant technologies may include
Dynatrace
- Grafana
- Prometheus
- Elastic Stack / ELK
- Splunk
- Azure Monitor
- OpenTelemetry
DevSecOps & Security
Embed DevSecOps practices throughout the software-delivery lifecycle.
- Integrate security scanning and controls into CI/CD pipelines.
- Support vulnerability identification, remediation and risk reduction.
- Ensure cloud and platform environments comply with enterprise security and regulatory requirements.
- Work closely with information-security teams to continuously improve platform security.
Technical Leadership
Provide technical leadership across DevOps, Cloud, Platform and SRE teams.
- Mentor and coach junior and intermediate engineers.
- Contribute to architecture decisions and technology roadmaps.
- Promote engineering standards and operational best practice.
- Lead cross-functional initiatives aimed at improving engineering productivity and reliability.
Minimum Experience
8+ years' experience across software engineering, infrastructure engineering, cloud engineering, DevOps or platform engineering.
- 5+ years' hands-on DevOps engineering experience.
- 3+ years' Site Reliability Engineering or production-operations experience.
- Proven experience supporting mission-critical production systems .
- Experience operating large-scale enterprise technology platforms .
- Strong exposure to highly available and business-critical environments.
Essential Technical Skills Cloud
Microsoft Azure
- Azure Kubernetes Service (AKS)
- Azure Networking
- Azure App Services
- Azure Monitor
- Azure Storage
- Azure Identity
DevOps / CI/CD
Azure DevOps
- GitHub / Git
- Jenkins
- SonarQube
- Artifactory and/or Nexus
Infrastructure Automation
Terraform
- Bicep
- ARM Templates
- Ansible
Containers
Kubernetes
- Docker
- Helm
Observability
Dynatrace
- Grafana
- Prometheus
- Elastic Stack
- Splunk
- Azure Monitor
- OpenTelemetry
Scripting / Development Strong scripting or programming ability using technologies such as:
Python
- PowerShell
- Bash
- C#
- Java
Go experience would be advantageous.
Core Technical Competencies
DevOps Engineering
- Site Reliability Engineering
- Azure Cloud Engineering
- Platform Engineering
- Infrastructure as Code
- Kubernetes / Container Orchestration
- CI/CD
- Infrastructure Automation
- DevSecOps
- Cloud Architecture
- Observability
- Continuous Delivery
- Systems Integration
- Capacity Planning
- Performance Optimisation
- Incident & Problem Management
- Root Cause Analysis
Behavioural Competencies
Strong technical problem-solving ability
- Strategic thinking
- Strong decision-making skills
- Collaboration across engineering disciplines
- Stakeholder management
- Continuous-improvement mindset
- Coaching and mentoring capability
- Accountability and ownership
- Customer-centric approach
Qualifications A Bachelor's Degree or equivalent technical qualification in one of the following areas is preferred:
Computer Science
- Information Technology
- Software Engineering
- Information Systems
Preferred Certifications Relevant certifications would be advantageous, including:
Microsoft Certified: Azure DevOps Engineer Expert
- Microsoft Certified: Azure Solutions Architect Expert
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
- HashiCorp Terraform Associate
- AWS Certified DevOps Engineer
- ITIL Foundation
- SRE Foundation Certification
Ideal Candidate The ideal candidate is a senior, hands-on engineer who can bridge software development, cloud infrastructure, DevOps, platform engineering and production operations .
They should have deep experience building and running highly available enterprise environments and possess a strong automation-first and reliability-focused mindset .
This person should be equally comfortable troubleshooting a critical production issue, building Terraform infrastructure, improving a Kubernetes platform, designing a CI/CD pipeline, implementing observability, defining SLOs and mentoring other engineers.
Desired Skills
- Senior DevOps Engineer
- Site Reliability Engineer
- SRE