Senior DevOps & Site Reliability Engineer at Datonomy Solutions

Datonomy Solutions (Cape Town) · Sandton, Gauteng · R90,000 - R100,000 per month

Stop applying one at a time.

JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.

Start free — we apply for you →

We are looking for an experienced  Senior DevOps & Site Reliability Engineer (SRE)  to design, build and operate highly available, secure, scalable and automated enterprise technology platforms. This is a senior hands-on engineering role spanning  DevOps, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and DevSecOps .

The successful candidate will work across engineering and delivery teams to improve  platform reliability, deployment velocity, resilience, automation, operational efficiency and production performance , while supporting mission-critical enterprise applications.

Key Responsibilities DevOps & Platform Engineering

Design, build and maintain  cloud-native infrastructure and platform services .

  • Develop and maintain  Infrastructure as Code (IaC)  solutions.
  • Automate infrastructure provisioning, configuration and operational processes.
  • Build reusable engineering tools, deployment templates and platform components.
  • Establish and standardise platform engineering practices across multiple delivery teams.
  • Identify opportunities to reduce manual intervention and increase engineering automation.

CI/CD & Release Automation

Design, implement and maintain enterprise-grade  CI/CD pipelines  for application and infrastructure deployments.

  • Implement automated testing, security scanning, code-quality controls and release automation.
  • Enable automated deployments, rollback and recovery processes.
  • Improve deployment frequency while reducing change and deployment risk.
  • Continuously optimise software delivery and release-management processes.

Site Reliability Engineering

Implement and mature  Site Reliability Engineering practices  across production environments.

  • Define, monitor and manage  Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs) .
  • Improve application and platform  availability, scalability, resilience and performance .
  • Lead production incident response, troubleshooting, problem management and  Root Cause Analysis (RCA) .
  • Drive proactive reliability improvements and reduction of technical debt.
  • Improve  Mean Time to Detect (MTTD)  and  Mean Time to Recover (MTTR) .

Azure Cloud Engineering

Design, implement and operate enterprise  Microsoft Azure  environments.

  • Work extensively with technologies such as:

Azure Kubernetes Service ( AKS )

  • Azure App Services
  • Azure Networking
  • Azure Monitor
  • Azure Storage
  • Azure Identity Services
  • Design highly available and disaster-recovery-capable environments.
  • Optimise cloud environments for  performance, resilience, security and cost .
  • Support hybrid-cloud and multi-cloud environments where required.

Containers & Kubernetes

Build, deploy and support containerised applications using  Docker  and  Kubernetes .

  • Manage Kubernetes environments, particularly  Azure Kubernetes Service (AKS) .
  • Develop and maintain deployment configurations using  Helm .
  • Support container-platform reliability, scalability and operational performance.
  • OpenShift experience would be advantageous.

Infrastructure as Code & Automation Hands-on experience with technologies such as:

Terraform

  • Bicep
  • ARM Templates
  • Ansible

Candidates should be comfortable using Infrastructure as Code to build repeatable, scalable and governed enterprise infrastructure.

Monitoring & Observability

Implement comprehensive  logging, monitoring, metrics, tracing and alerting .

  • Build operational dashboards and platform insights.
  • Establish enterprise observability standards.
  • Implement proactive and predictive monitoring.
  • Use observability information to improve application and infrastructure reliability.

Relevant technologies may include

Dynatrace

  • Grafana
  • Prometheus
  • Elastic Stack / ELK
  • Splunk
  • Azure Monitor
  • OpenTelemetry

DevSecOps & Security

Embed  DevSecOps  practices throughout the software-delivery lifecycle.

  • Integrate security scanning and controls into CI/CD pipelines.
  • Support vulnerability identification, remediation and risk reduction.
  • Ensure cloud and platform environments comply with enterprise security and regulatory requirements.
  • Work closely with information-security teams to continuously improve platform security.

Technical Leadership

Provide technical leadership across DevOps, Cloud, Platform and SRE teams.

  • Mentor and coach junior and intermediate engineers.
  • Contribute to architecture decisions and technology roadmaps.
  • Promote engineering standards and operational best practice.
  • Lead cross-functional initiatives aimed at improving engineering productivity and reliability.

Minimum Experience

8+ years' experience  across software engineering, infrastructure engineering, cloud engineering, DevOps or platform engineering.

  • 5+ years' hands-on DevOps engineering experience.
  • 3+ years' Site Reliability Engineering or production-operations experience.
  • Proven experience supporting  mission-critical production systems .
  • Experience operating  large-scale enterprise technology platforms .
  • Strong exposure to highly available and business-critical environments.

Essential Technical Skills Cloud

Microsoft Azure

  • Azure Kubernetes Service (AKS)
  • Azure Networking
  • Azure App Services
  • Azure Monitor
  • Azure Storage
  • Azure Identity

DevOps / CI/CD

Azure DevOps

  • GitHub / Git
  • Jenkins
  • SonarQube
  • Artifactory and/or Nexus

Infrastructure Automation

Terraform

  • Bicep
  • ARM Templates
  • Ansible

Containers

Kubernetes

  • Docker
  • Helm

Observability

Dynatrace

  • Grafana
  • Prometheus
  • Elastic Stack
  • Splunk
  • Azure Monitor
  • OpenTelemetry

Scripting / Development Strong scripting or programming ability using technologies such as:

Python

  • PowerShell
  • Bash
  • C#
  • Java

Go experience would be advantageous.

Core Technical Competencies

DevOps Engineering

  • Site Reliability Engineering
  • Azure Cloud Engineering
  • Platform Engineering
  • Infrastructure as Code
  • Kubernetes / Container Orchestration
  • CI/CD
  • Infrastructure Automation
  • DevSecOps
  • Cloud Architecture
  • Observability
  • Continuous Delivery
  • Systems Integration
  • Capacity Planning
  • Performance Optimisation
  • Incident & Problem Management
  • Root Cause Analysis

Behavioural Competencies

Strong technical problem-solving ability

  • Strategic thinking
  • Strong decision-making skills
  • Collaboration across engineering disciplines
  • Stakeholder management
  • Continuous-improvement mindset
  • Coaching and mentoring capability
  • Accountability and ownership
  • Customer-centric approach

Qualifications A Bachelor's Degree or equivalent technical qualification in one of the following areas is preferred:

Computer Science

  • Information Technology
  • Software Engineering
  • Information Systems

Preferred Certifications Relevant certifications would be advantageous, including:

Microsoft Certified:  Azure DevOps Engineer Expert

  • Microsoft Certified:  Azure Solutions Architect Expert
  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD)
  • HashiCorp Terraform Associate
  • AWS Certified DevOps Engineer
  • ITIL Foundation
  • SRE Foundation Certification

Ideal Candidate The ideal candidate is a senior, hands-on engineer who can bridge  software development, cloud infrastructure, DevOps, platform engineering and production operations .

They should have deep experience building and running highly available enterprise environments and possess a strong  automation-first and reliability-focused mindset .

This person should be equally comfortable troubleshooting a critical production issue, building Terraform infrastructure, improving a Kubernetes platform, designing a CI/CD pipeline, implementing observability, defining SLOs and mentoring other engineers.

Desired Skills

  • Senior DevOps Engineer
  • Site Reliability Engineer
  • SRE
Auto-apply to this jobView original posting ↗