Site Reliability Engineer (AI Factory)
Liquidtech · South Africa
Posted 31 July 2026
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Cassava Technologies is a pan-African digital services leader. We are building a state-of-the-art AI Factory in Cape Town, South Africa, based on the NVIDIA NCP (NVIDIA Cloud Partner) program. This facility will serve as the foundation for high-performance AI workloads across the continent. The Site Reliability Engineer – Network is responsible for engineering reliability into the network and interconnect layers of the Cassava AI Factory platform. The role focuses on reducing operational risk, ensuring predictable high‑performance connectivity, and eliminating repetitive operational effort (toil) through automation, observability, and reliability‑focused engineering across InfiniBand fabric, Ethernet networking, and inter‑node communication layers. The SRE works closely with AI Cloud Operations Engineers, AI Engineering teams, and vendors to ensure that network performance does not constrain AI workloads. You are the guardian of the data path. You will ensure the high-speed Ethernet and Infiniband fabrics deliver uncompromising performance and strict multi-tenant isolation. Deep Expertise: Arista WAN routing, Cumulus Linux (NVIDIA Spectrum) switching, and complex BGP/EVPN/-VXLAN architecture and protocols. AI Specifics: Managing and troubleshooting Infiniband fabrics for East-West GPU traffic. Key Responsibilities: Automating network configurations, managing AS numbers and IPv4 prefix allocations, securing OOB networks, and maintaining perimeter/internal firewalls (Fortigate). What you will do day-to-day: Monitor & Respond: Act as the first line of defense for the AI Factory, monitoring the health of the 2,048 GPU cluster and responding to complex, multi-layered incidents. Automate: Ruthlessly eliminate manual toil by building robust automation for provisioning, configuration management, and self-healing. Collaborate: Work seamlessly alongside hardware vendors (HPE, NVIDIA), the network engineering team, and software architects to bridge physical data center realities with cloud-native workflows. Scale: Assist in capacity planning, acceptance testing (UAT) for new hardware deliveries, and continuously tuning the environment for maximum throughput. Key Responsibilities: Reliability Engineering • Engineer reliability across AI Factory network infrastructure, including: InfiniBand Fabric, Data centre Ethernet networks, GPU node interconnects, Storage networks (FCOE) Incident Engineering and escalation support Diagnose and resolve issues such as: Fabric instability, Packet loss / retransmissions, Congestion or oversubscription, GPU node communication failures. Toil identification and reduction Identify repetitive, manual, or high‑effort operational activities that do not add long‑term value. Observability and monitoring engineering Design and maintain observability dashboards for availability and performance of the AI Factory focusing on : Latency, Throughput, Packet Loss, Network Congestion. Platform and capacity improvement Collaborate with AI Engineering and vendors on platform design changes that improve reliability and recovery Operational Involment Contribute to incident response for security-related events (breaches, compromise, data exposure)