DevOps / SRE Cloud Engineer (Site Reliability Engi
Executive Placements · Kensington
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Role purpose & context
Our client is looking for a senior DevOps / SRE Cloud Engineer to own the companys Azure-based Kubernetes platform end-to-end resilient infrastructure, hardened CI/CD, and strong observability so the companys workloads run reliably at scale.
Key roles & responsibilities
· Operate production Kubernetes clusters and Azure infrastructure as Terraform code.
· Maintain CI/CD pipelines (GitHub Actions) with quality gates for reliable deployments.
· Own observability: dashboards, alerting, SLOs, and incident response.
· Manage secrets, identity, and network security across environments.
· Partners with engineering teams to troubleshoot and improve reliability.
Must-have technical skills / experience
· Azure cloud platform production experience with AKS, ADLS Gen2 (blob storage), Azure Key Vault, Entra ID (incl. workload identity / managed identities), Azure networking (VNets, private endpoints, firewall rules), and Azure Service Bus or an equivalent message broker.
· Kubernetes in production deploying and operating self-managed (non-PaaS) workloads: Helm chart authoring, environment overlays, K8s operators, node-pool design, resource requests/limits and capacity sizing, pod troubleshooting, upgrades.
· Terraform authoring and maintaining modular IaC for cluster, identity, storage, and secrets provisioning; state management and plan/apply discipline across environments.
· Docker image builds (multi-arch), Compose-based local development stacks, container health checks and memory-limit tuning.
· CI/CD with GitHub Actions building and maintaining pipelines for lint/test/build gates and container image publishing; branch-protection / required-status-check workflows.
· Observability and SRE practice Prometheus + Grafana (dashboards, alert rules), SLO/alerting design, incident response, capacity planning from measured load; structured logging.
· Linux and shell scripting strong Bash; comfortable owning operational scripts as maintained, tested code.
· Secrets management provider-chain patterns (Key Vault in prod, env/file locally); keeping secrets out of repos and images.
Preferred / nice-to-have technical skills
· Apache Spark on Kubernetes operations Spark Operator, cluster sizing, executor/driver tuning, Spark Connect; Jupyter Hub on K8s.
· Open data-platform stack Hive Metastore, Trino, Apache Ranger, Delta Lake, S3-compatible object storage (MinIO); understanding of how governed query surfaces are wired.
· Open Telemetry SDK-based metrics/traces instrumentation and collector topology; Open Lineage/Marquez lineage.
· Python enough to maintain operational tooling an
.special-hidden { display: none; }