Systems Operations Engineer (Permanent)
Recruit · Western Cape , CPT - Northern Suburbs · Market Related Monthly Cost To Company (Market related)
Posted 6 August 2026
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Job Title: Systems Operations Engineer Start Date: 2026-08-06 - 2026-09-05 Vacancy Type: Permanent PE011687 Sectors: Information Technology Location: Western Cape , CPT - Northern Suburbs Salary: Market Related Monthly Cost To Company (Market related) Brief: Job purpose Our client is looking for an experienced Systems Operations Engineer to own the foundation their platform and their business run on. In this role you keep their servers, storage, networking, and identity systems up, current, and secure - and you are the person accountable for the uptime of that infrastructure. You will work closely with the company's DevOps Engineer (with deliberately overlapping skills, so the two of you can cover for one another) and alongside their internal IT Systems Administrator. The company's infrastructure already spans multiple territories - your focus is continuous improvement: securing and simplifying the estate while keeping it current, resilient, and well documented. This is a hands-on senior role with real ownership and the autonomy to set direction in your area. Key responsibility areas: Drive continuous improvement of the existing estate - securing, simplifying, consolidating, and reducing complexity and risk over time. Maintain and manage internal and external infrastructure for software development and production - servers, switches, firewalls, routing, virtual machines, and storage. Keep operating systems and platforms patched, current, and hardened across the server estate. Manage virtualised environments (Proxmox / KVM; VMware experience welcome) - VM lifecycle, high availability, and capacity. Own storage and data protection - ZFS/NFS pools, snapshots, replication, and documented, tested backups. Manage network and edge security - firewall policy, VPN, segmentation, and gateway platforms (OPNsense, FortiGate). Operate the identity platform (Authentik / SSO) and wire services into it securely. Operate the MariaDB/MySQL high-availability database cluster (Galera/MaxScale) - replication, failover, and recovery. Maintain the infrastructure source of truth (DCIM/IPAM) so the estate is accurately documented. Ensure redundancy across systems and continuous monitoring of critical services. Produce and maintain design and systems documentation, runbooks, and recovery procedures. Continuously improve operational practices, processes, and infrastructure. Participate in an after-hours and weekend support rotation shared across the infrastructure team. Required skills: Excellent Linux system administration (Ubuntu, CentOS/RHEL) and associated technologies. Strong Linux command-line proficiency and a genuine patching / currency discipline. Working knowledge of network services (NFS, iSCSI, SMB/Samba, LDAP, DNS). Strong understanding of the OSI model and networking fundamentals. Working knowledge of package management (rpm, deb). Basic version control familiarity (Git). Required experience: Configuring and maintaining virtualised environments (Proxmox, VMware, or Xen). Storage administration with real depth - pool design, replication, and recovery (ZFS a strong plus). High-availability or load-balancing environments - implementing and operating them. Relational databases (PostgreSQL, MySQL/MariaDB) and HA clustering concepts. Backup and recovery tooling (Veeam, Bacula, or similar). Monitoring tooling (Zabbix, Prometheus/Grafana, Nagios, or similar). Scripting for automation (Bash, Python). Working within agile delivery (Jira, Bitbucket) and a change-control discipline. Advantageous Experience: Cloud platforms (AWS, Azure, GCP) - relevant to the company's international expansion. Identity and access management platforms (Authentik, Keycloak, WSO2). Networking depth and troubleshooting (TCP/IP, DNS, VPN, routing/switching, firewalls). Server hardening and defensive systems (file integrity, IDS/IPS, Fail2Ban, auditing). PKI and certificate lifecycle management (Let's Encrypt, commercial CAs). Security audits and standards exposure (PCI-DSS, ISMS, ISO 27000) - important in the company's regulated industry. Configuration management (Ansible). Attributes: Leadership that comes from experience; able to set and execute direction in your area. Self-motivated, proactive, and comfortable working independently with little supervision. Strong team spirit and able to collaborate across technical and non-technical staff. Logical, objective, and able to think beyond the obvious solution. Comfortable working under pressure and juggling multiple initiatives. Clear written and verbal communication, up and down the organisation. Qualifications: Degree/Diploma in Information Systems or related field. At least 5 years in IT operations and/or systems administration. At least 2 years in a senior or lead capacity. Detail: Job purpose Our client is looking for an experienced Systems Operations Engineer to own the foundation their platform and their business run on. In this role you keep their servers, storage, networking, and identity systems up, current, and secure - and you are the person accountable for the uptime of that infrastructure. You will work closely with the company's DevOps Engineer (with deliberately overlapping skills, so the two of you can cover for one another) and alongside their internal IT Systems Administrator. The company's infrastructure already spans multiple territories - your focus is continuous improvement: securing and simplifying the estate while keeping it current, resilient, and well documented. This is a hands-on senior role with real ownership and the autonomy to set direction in your area. Key responsibility areas: Drive continuous improvement of the existing estate - securing, simplifying, consolidating, and reducing complexity and risk over time. Maintain and manage internal and external infrastructure for software development and production - servers, switches, firewalls, routing, virtual machines, and storage. Keep operating systems and platforms patched, current, and hardened across the server estate. Manage virtualised environments (Proxmox / KVM; VMware experience welcome) - VM lifecycle, high availability, and capacity. Own storage and data protection - ZFS/NFS pools, snapshots, replication, and documented, tested backups. Manage network and edge security - firewall policy, VPN, segmentation, and gateway platforms (OPNsense, FortiGate). Operate the identity platform (Authentik / SSO) and wire services into it securely. Operate the MariaDB/MySQL high-availability database cluster (Galera/MaxScale) - replication, failover, and recovery. Maintain the infrastructure source of truth (DCIM/IPAM) so the estate is accurately documented. Ensure redundancy across systems and continuous monitoring of critical services. Produce and maintain design and systems documentation, runbooks, and recovery procedures. Continuously improve operational practices, processes, and infrastructure. Participate in an after-hours and weekend support rotation shared across the infrastructure team. Required skills: Excellent Linux system administration (Ubuntu, CentOS/RHEL) and associated technologies. Strong Linux command-line proficiency and a genuine patching / currency discipline. Working knowledge of network services (NFS, iSCSI, SMB/Samba, LDAP, DNS). Strong understanding of the OSI model and networking fundamentals. Working knowledge of package management (rpm, deb). Basic version control familiarity (Git). Required experience: Configuring and maintaining virtualised environments (Proxmox, VMware, or Xen). Storage administration with real depth - pool design, replication, and recovery (ZFS a strong plus). High-availability or load-balancing environments - implementing and operating them. Relational databases (PostgreSQL, MySQL/MariaDB) and HA clustering concepts. Backup and recovery tooling (Veeam, Bacula, or similar). Monitoring tooling (Zabbix, Prometheus/Grafana, Nagios, or similar). Scripting for automation (Bash, Python). Working within agile delivery (Jira, Bitb