Site Reliability Engineer (AI Factory)
Liquidtech · South Africa
Posted 19 August 2026
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Role Purpose Cassava Technologies is a pan-African digital services leader. We are building a state-of-the-art AI Factory in Cape Town, South Africa, based on the NVIDIA NCP (NVIDIA Cloud Partner) program. This facility will serve as the foundation for high-performance AI workloads across the continent. The Site Reliability Engineer (SRE) - Compute is responsible for engineering reliability into the Cassava AI Factory platform. The role focuses on reducing operational risk, improving service availability, and eliminating repetitive operational effort (toil) through automation, observability, and reliability-focused engineering. The SRE works closely with AI Cloud Operations Engineers and AI Engineering teams to ensure that the AI Factory becomes more stable, predictable, and scalable over time.You ensure the GPUs are healthy, updated, and never starved for data. You will manage the extreme-throughput storage layers and the full lifecycle of the NVIDIA compute stack. Deep Expertise: High-Performance Storage (HPS) architectures (specifically Weka or similar) and deep knowledge of the NVIDIA GPU stack.AI Specifics: Managing NVIDIA drivers, firmware (SBIOS/VBIOS), and the CUDA software layer. Familiarity with local high-end NVMe disk pooling. Key Responsibilities: Executing hardware acceptance testing (fio, cluster validation), managing firmware/driver parity across the cluster, troubleshooting GPU health (NVLink errors, thermal throttling), and optimizing data ingress/egress for massive AI datasets. What You Will Do Day-to-Day ● Monitor & Respond: Act as the first line of defense for the AI Factory, monitoring the health of the 2,048 GPU cluster and responding to complex, multi-layered incidents. ● Automate: Ruthlessly eliminate manual toil by building robust automation for provisioning, configuration management, and self-healing. ● Collaborate: Work seamlessly alongside hardware vendors (HPE, NVIDIA), the network engineering team, and software architects to bridge physical data center realities with cloud-native workflows. ● Scale: Assist in capacity planning, acceptance testing (UAT) for new hardware deliveries, and continuously tuning the environment for maximum throughput. Role Description Engineer platform reliability across AI Factory services (BMaaS, GPUaaS, LLM Training, AIFaaS, AISaaS).Define, measure, and improve Service Level Objectives (SLOs), and reliability targets. Identify systemic weaknesses that contribute to outages, performance degradation, or instability. Full Stack Oversight: Oversee the daily health of the environment, from the physical layer (Power/Cooling/Cabling) up through Compute/Storage (NVIDIA H200s/Infiniband Fabric, WEKA Storage) and the Orchestration Layer (Rafay).Act as a Tier 2 escalation point for complex incidents, working closely with the AI Centre of excellence engineering teams. Support Tier 1 Engineers during high severity incidents with deep technical analysis. Participate in on call rotation for critical escalations. Drive post incident reviews focused on learning and prevention. Work closely with vendors – ADC, WEKA, HPe, Nvidia, Rafay, LIT C2 for troubleshooting and incident resolution. Identify repetitive, manual, or high effort operational activities that do not add long term value. Identify opportunities for automation to reduce recurring incidents and manual intervention. Improve tooling, scripts, self-healing mechanisms, and guardrails that reduce operational load. Design and maintain observability dashboards for availability and performance of the AI Factory. Improve monitoring, alerting, and dashboards to reduce noise and improve signal quality. Ensure alerts are actionable, severity aligned and mapped to operational response. Proactive Monitoring: Utilize observability tools to detect bottlenecks (thermal throttling, packet loss) before they disrupt customer training runs or inferencing requests. Collaborate with AI Engineering and vendors on platform design changes that improve reliability and recovery. Contribute to capacity modelling and scaling strategies. Identify reliability risks introduced by new features, changes, or capacity growth. Operational Security: Enforce network security policies to ensure tenant isolation and data sovereignty. Maintenance Windows: Schedule and execute patching cycles (Firmware, NVIDIA Drivers, OS) with extreme care, ensuring coordination with customers to avoid killing active training jobs. Work closely with all levels of support tier’s to understand the incident behaviours and how to reduce repeat incidents. Compliance: Maintain strict operational adherence to NVIDIA NCP standards for availability and performance. Work closely with facility providers to ensure facility reliability is in place, ensure optimal uptime of cooling, power and security systems of the facility provider. Collaborate with Security teams on patching strategies and remediation workflows. Ensure secure configuration of infrastructure components (OS, Kubernetes, networking layers).Contribute to incident response for security-related events (breaches, compromise, data exposure)