DevOps / SRE Cloud Engineer (Site Reliability Engineer)

Recru-it · R Undisclosed

Stop applying one at a time.

JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.

Start free — we apply for you →

Role purpose & context

Our client is looking for a senior DevOps / SRE Cloud Engineer to own the company’s Azure-based Kubernetes platform end-to-end — resilient infrastructure, hardened CI/CD, and strong observability so the company’s workloads run reliably at scale.

Key roles & responsibilities

· Operate production Kubernetes clusters and Azure infrastructure as Terraform code.

· Maintain CI/CD pipelines (GitHub Actions) with quality gates for reliable deployments.

· Own observability: dashboards, alerting, SLOs, and incident response.

· Manage secrets, identity, and network security across environments.

· Partners with engineering teams to troubleshoot and improve reliability.

Must-have technical skills / experience

· Azure cloud platform — production experience with AKS, ADLS Gen2 (blob storage), Azure Key Vault, Entra ID (incl. workload identity / managed identities), Azure networking (VNets, private endpoints, firewall rules), and Azure Service Bus or an equivalent message broker.

· Kubernetes in production — deploying and operating self-managed (non-PaaS) workloads: Helm chart authoring, environment overlays, K8s operators, node-pool design, resource requests/limits and capacity sizing, pod troubleshooting, upgrades.

· Terraform — authoring and maintaining modular IaC for cluster, identity, storage, and secrets provisioning; state management and plan/apply discipline across environments.

· Docker — image builds (multi-arch), Compose-based local development stacks, container health checks and memory-limit tuning.

· CI/CD with GitHub Actions — building and maintaining pipelines for lint/test/build gates and container image publishing; branch-protection / required-status-check workflows.

· Observability and SRE practice — Prometheus + Grafana (dashboards, alert rules), SLO/alerting design, incident response, capacity planning from measured load; structured logging.

· Linux and shell scripting — strong Bash; comfortable owning operational scripts as maintained, tested code.

· Secrets management — provider-chain patterns (Key Vault in prod, env/file locally); keeping secrets out of repos and images.

Preferred / nice-to-have technical skills

· Apache Spark on Kubernetes operations — Spark Operator, cluster sizing, executor/driver tuning, Spark Connect; Jupyter Hub on K8s.

· Open data-platform stack — Hive Metastore, Trino, Apache Ranger, Delta Lake, S3-compatible object storage (MinIO); understanding of how governed query surfaces are wired.

· Open Telemetry — SDK-based metrics/traces instrumentation and collector topology; Open Lineage/Marquez lineage.

· Python — enough to maintain operational tooling and pytest-based environment/infrastructure test tiers.

· SQL Server and PostgreSQL — operational administration (backups, container deployments, EF-migration-driven schemas).

· Multi-tenant / regulated-data environments — tenant isolation patterns, POPIA/GDPR/ISO 27001-adjacent controls; secret scanning and supply-chain hygiene (dependency audit, image provenance).

Seniority and experience

· Senior (mid-senior acceptable with strong K8s depth). 5+ years DevOps/SRE/platform engineering, including 2–3 years running Kubernetes workloads in production and 2+ years on Azure. Must be able to own the environment independently.

Required qualifications

· Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

· Cloud/Kubernetes certification (e.g. Azure, CKA) advantageous.

· Proven experience operating production-grade cloud infrastructure at scale; strong communication skills and ability to work independently in a hybrid/remote team.

Why you’ll love working for the company

The company believes in taking care of their team and creating an environment where you can thrive. As part of their company, you’ll enjoy:

· Flexible working arrangements: Whether you’re a night owl or an early bird, they offer hybrid and remote options to suit your lifestyle

· Comprehensive benefits: From a wellness program to home office reimbursements and continuous learning opportunities, they have got you covered.

· Team culture: Fun team-building activities, regular socials, and a supportive, inclusive culture that values transparency, accountability, and work-life balance.

· Performance incentives: Competitive salaries, ESOP, and recognition for your hard work. Role purpose & context:

Our client is looking for a senior DevOps / SRE Cloud Engineer to own the company’s Azure-based Kubernetes platform end-to-end — resilient infrastructure, hardened CI/CD, and strong observability so the company’s workloads run reliably at scale.

Key roles & responsibilities

· Operate production Kubernetes clusters and Azure infrastructure as Terraform code.

· Maintain CI/CD pipelines (GitHub Actions) with quality gates for reliable deployments.

· Own observability: dashboards, alerting, SLOs, and incident response.

· Manage secrets, identity, and network security across environments.

· Partners with engineering teams to troubleshoot and improve reliability.

Must-have technical skills / experience

· Azure cloud platform — production experience with AKS, ADLS Gen2 (blob storage), Azure Key Vault, Entra ID (incl. workload identity / managed identities), Azure networking (VNets, private endpoints, firewall rules), and Azure Service Bus or an equivalent message broker.

· Kubernetes in production — deploying and operating self-managed (non-PaaS) workloads: Helm chart authoring, environment overlays, K8s operators, node-pool design, resource requests/limits and capacity sizing, pod troubleshooting, upgrades.

· Terraform — authoring and maintaining modular IaC for cluster, identity, storage, and secrets provisioning; state management and plan/apply discipline across environments.

· Docker — image builds (multi-arch), Compose-based local development stacks, container health checks and memory-limit tuning.

· CI/CD with GitHub Actions — building and maintaining pipelines for lint/test/build gates and container image publishing; branch-protection / required-status-check workflows.

· Observability and SRE practice — Prometheus + Grafana (dashboards, alert rules), SLO/alerting design, incident response, capacity planning from measured load; structured logging.

· Linux and shell scripting — strong Bash; comfortable owning operational scripts as maintained, tested code.

· Secrets management — provider-chain patterns (Key Vault in prod, env/file locally); keeping secrets out of repos and images.

Preferred / nice-to-have technical skills

· Apache Spark on Kubernetes operations — Spark Operator, cluster sizing, executor/driver tuning, Spark Connect; Jupyter Hub on K8s.

· Open data-platform stack — Hive Metastore, Trino, Apache Ranger, Delta Lake, S3-compatible object storage (MinIO); understanding of how governed query surfaces are wired.

· Open Telemetry — SDK-based metrics/traces instrumentation and collector topology; Open Lineage/Marquez lineage.

· Python — enough to maintain operational tooling and pytest-based environment/infrastructure test tiers.

· SQL Server and PostgreSQL — operational administration (backups, container deployments, EF-migration-driven schemas).

· Multi-tenant / regulated-data environments — tenant isolation patterns, POPIA/GDPR/ISO 27001-adjacent controls; secret scanning and supply-chain hygiene (dependency audit, image provenance).

Seniority and experience

· Senior (mid-senior acceptable with strong K8s depth). 5+ years DevOps/SRE/platform engineering, including 2–3 years running Kubernetes workloads in production and 2+ years on Azure. Must be able to own the environment independently.

Required qualifications

· Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

· Cloud/Kubernetes certification (e.g. Azure, CKA) advantageous.

· Proven experience operating production-grade cloud infrastructure at scale; strong communication skills and ability to w

Auto-apply to this jobView original posting ↗