Site Reliability Engineer
Sun International · Johannesburg, Gauteng
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Job Description
- The Site Reliability Engineer SRE is a technical contributor responsible for improving the reliability, resilience and operational performance of Sun International's technology services, platforms and infrastructure.
- The role focuses on observability, monitoring, alerting, event management, reliability analytics, automation and auto-remediation, proactively identifying potential issues before they impact the business.
- Working across software engineering, infrastructure, enterprise applications and service management teams, the SRE implements monitoring and automation solutions that strengthen service reliability and reduce operational.
Core behavioural & Technical / proficiency competencies
- Cloud platform management across Azure, AWS and/or GCP.
- Containerisation technologies, including Docker and Kubernetes.
- Monitoring, observability and alerting tools such as Prometheus, Grafana and ELK.
- Scripting and operational automation using Python, Bash and/or PowerShell.
- Development of automation, auto-remediation capabilities and operational runbooks.
- Incident management and proactive identification of service reliability risks.
- Root Cause Analysis RCA and analysis of recurring operational failures.
- Linux and Windows system administration.
- Networking fundamentals.
- Git version control.
- Troubleshooting and problem-solving across technology services and infrastructure.
- Analysis of operational telemetry, event data and service performance trends.
- Application of reliability standards, security requirements, operational controls and governance practices.
- Cross-functional collaboration with software engineering, infrastructure, enterprise applications and service management teams.
- Analytical thinking and evidence-based decision-making.
- Continuous improvement and operational excellence.
- Collaboration and knowledge sharing.
- Operational excellence and accountability.
Job Requirements
Qualifications
- Degree in Computer Science, Engineering, Information Technology or a related discipline required
- Cloud platform certification, e.g. Azure Administrator Associate or AWS Certified SysOps Administrator preferred
- ITIL Foundation Certification preferred
Experience
- 2–5 years' experience in infrastructure operations, technology operations, monitoring platforms, cloud operations, Site Reliability Engineering, observability tooling or a related technology discipline.
- Practical experience with monitoring and observability, cloud platforms, automation/scripting, incident management and troubleshooting aligned to an SRE or technology operations environment