DevOps / SRE Cloud Engineer (Site Reliability Engi
Executive Placements · Bo-Kaap, Western Cape
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Location: Remote Employment Type: Full-Time Industry: Cloud Infrastructure | DevOps | Site Reliability Engineering | Data Technology
WatersEdge Solutions is partnering with a growing technology business to appoint a DevOps / SRE Cloud Engineer (Site Reliability Engineer) to take end-to-end ownership of its Azure-based Kubernetes platform.
This is a senior, hands-on engineering role focused on building and operating resilient cloud infrastructure, strengthening CI/CD, and creating the observability needed to run production workloads reliably at scale. You’ll work across Azure, Kubernetes, Terraform, Docker, GitHub Actions and SRE practices , with the autonomy to own the environment and continuously improve its reliability, security and performance.
About the Role
As DevOps / SRE Cloud Engineer, you’ll take ownership of the production Kubernetes and Azure environment, managing infrastructure as code and ensuring workloads are secure, observable and dependable.
You’ll operate Kubernetes clusters, maintain modular Terraform infrastructure, build and improve deployment pipelines, manage cloud identity and secrets, and establish effective monitoring, alerting and SLOs.
This role requires someone who is comfortable going beyond maintaining infrastructure. You’ll be expected to understand how the platform behaves under real production load, troubleshoot complex issues alongside engineering teams, respond effectively to incidents and make practical improvements that prevent problems from recurring.
Key Responsibilities
Operate and maintain production Kubernetes clusters and Azure cloud infrastructure.
Manage infrastructure through modular Terraform code across multiple environments.
Own Kubernetes deployments, configuration, upgrades and operational troubleshooting.
Design and maintain Helm charts and environment-specific deployment configurations.
Manage Kubernetes node pools, resource requests and limits, and capacity requirements.
Maintain secure and reliable CI/CD pipelines using GitHub Actions .
Implement linting, testing and build quality gates for production deployments.
Manage container image build and publishing workflows.
Maintain branch protection and required-status-check processes.
Own platform observability across dashboards, metrics, alerting and structured logging.
Define and maintain meaningful SLOs and alerting strategies.
Participate in and improve incident response processes.
Perform capacity planning based on measured production workloads.
Manage cloud secrets, identity and network security across environments.
.special-hidden { display: none; }