Senior DevOps / SRE Cloud Engineer
Hire Resolve
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Job Description
- A company that develops, owns, and operates utility-scale power generation and energy storage plants across the African continent is seeking a Senior DevOps / SRE Cloud Engineer who will own the company's Azure-based Kubernetes platform end-to-end - remote or hybrid based on location.
Responsibilities
- Infrastructure as Code: Provision and operate production Kubernetes AKS and Azure infrastructure using modular Terraform.
- CI/CD Automation: Maintain GitHub Actions pipelines featuring quality gates, security checks, and automated deployments.
- Observability & SRE: Own monitoring, alerting, SLOs, capacity planning, and incident response across environments.
- Security & Access: Manage platform secrets, network controls VNets, private endpoints, and Entra ID identity governance.
- Developer Enablement: Partner with engineering teams to resolve operational bottlenecks and drive reliability standards.
Minimum Requirements
- Experience: 5+ years in DevOps/SRE/Platform Engineering including 2–3 years operating K8s in production and 2+ years on Azure.
- Education: Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- Certifications: CKA, CKAD, or Azure Solutions Architect / DevOps Engineer certifications are advantageous.
- Soft Skills: Strong problem-solving ability, clear technical communication, and the capacity to operate autonomously within a hybrid/remote team.
- Azure Platform: Production experience with AKS, ADLS Gen2, Key Vault, Entra ID Workload/Managed Identities, VNets/Private Endpoints, and Service Bus or equivalent broker.
- Production Kubernetes: Advanced operational depth in Helm chart authoring, operator deployment, node-pool sizing, pod troubleshooting, and cluster upgrades.
- Terraform: Proven ability to author modular IaC, manage state safely, and maintain strict plan/apply disciplines.
- Containerization & CI/CD: Hands-on Docker multi-arch builds, optimization and GitHub Actions pipeline development with required status checks.
- Observability: Practical experience configuring Prometheus + Grafana, structured logging, and SLO-based alerting.
- Linux & Automation: Strong Bash scripting with clean, code-maintained operational tooling.
Preferred Experience
- Data Platforms: Apache Spark on K8s Spark Operator/Connect, JupyterHub, Delta Lake, Trino, Hive Metastore, or Apache Ranger.
- Advanced Telemetry: OpenTelemetry SDKs/Collector topology and OpenLineage/Marquez integration.
- Languages & Databases: Intermediate Python infrastructure testing/pytest and basic DBA management for SQL Server and PostgreSQL.
- Compliance & Isolation: Multi-tenant architecture design, dependency auditing, image provenance, and POPIA/GDPR/ISO 27001 controls.