Principal Telco SysOps Engineer
Comsol Networks (Pty) Ltd · Centurion, Gauteng
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Introduction
The Principal Telco SysOps Engineer is the senior technical authority responsible for the architecture, reliability, automation, security and lifecycle management of the infrastructure platforms supporting the organisation's telecommunications network. The role has particular responsibility for the infrastructure underlay supporting the 5G SA Core and associated OSS platforms , including Linux, compute, virtualisation, Kubernetes/container platforms, storage, IP networking, databases, observability and infrastructure automation. This is a senior hands-on technical role , not primarily a management position. The Principal Engineer leads complex infrastructure troubleshooting, establishes engineering standards and provides technical leadership to the Telco SysOps team.
Duties & Responsibilities
Telco Infrastructure
- Own the technical architecture and operational integrity of infrastructure supporting critical telecommunications services.
- Design and operate highly available carrier-grade infrastructure.
- Own the infrastructure underlay supporting 5G SA Core network functions and OSS platforms.
- Ensure appropriate redundancy across compute, networking, storage and supporting services.
- Define infrastructure capacity, performance and resilience requirements.
- Lead infrastructure lifecycle management, upgrades and technology refreshes.
- Establish infrastructure standards across production environments and data centres.
Linux & Operating Systems
- Act as technical authority for production Linux environments.
- Diagnose complex operating-system and application-platform problems.
- Manage system services, kernel parameters, networking, storage, filesystems and resource utilisation.
- Develop standardised OS builds and configuration baselines.
- Manage patching, vulnerability remediation and lifecycle management.
- Optimise Linux platforms for high-throughput and latency-sensitive telecommunications workloads.
Kubernetes, Containers & Telco Cloud
- Architect and operate Kubernetes platforms supporting CNFs and OSS applications.
- Troubleshoot Kubernetes networking, scheduling, storage and application dependencies.
- Manage cluster capacity, resilience and availability.
- Understand CNI, CSI, ingress/load balancing and service networking.
- Manage Helm-based and containerised application deployments.
- Support CNF lifecycle operations without requiring ownership of the 5GC application itself.
- Understand NUMA, CPU pinning, hugepages, SR-IOV and other performance considerations relevant to telecommunications workloads where required.
Virtualisation & Compute
- Architect and support virtualised and bare-metal infrastructure.
- Manage hypervisor, VM and compute-cluster environments.
- Troubleshoot resource contention and infrastructure performance.
- Define CPU, RAM and storage capacity requirements.
- Support hardware lifecycle, firmware and platform upgrades.
Infrastructure Networking
- Design and troubleshoot the IP underlay supporting 5GC and OSS platforms.
- Strong understanding of:
- TCP/IP
- VLANs
- LACP
- MLAG/MC-LAG
- VRFs
- BGP
- OSPF/IS-IS
- ECMP
- VXLAN/EVPN where deployed
- Troubleshoot packet loss, latency, MTU and asymmetric routing.
- Understand data-centre leaf/spine architectures.
- Troubleshoot Linux and Kubernetes networking end-to-end.
- Work closely with IP/MPLS and security teams on external connectivity.
Storage & Databases
- Design and support highly available storage platforms.
- Understand block, file and object storage.
- Troubleshoot storage latency, capacity and performance.
- Support database infrastructure used by OSS and telecommunications applications.
- Understand database clustering, replication, backup and recovery concepts.
- Coordinate with application/database specialists where deep database engineering is required.
OSS Infrastructure
Provide infrastructure engineering and operational support for OSS platforms including:
- Network management systems
- Element management systems
- Fault management
- Performance management
- Configuration management
- Inventory systems
- Mediation platforms
- Log management
- Telemetry platforms
- Network analytics systems
Own the infrastructure and platform layers while working with OSS/application engineers on application functionality.
Observability
- Define the observability architecture for critical telecommunications platforms.
- Implement monitoring across:
- Hardware
- Linux
- Kubernetes
- VMs
- Storage
- Network
- Databases
- Applications
- Develop service and infrastructure dashboards.
- Implement centralised logging and telemetry.
- Establish meaningful alert thresholds and reduce alert noise.
- Define SLIs/SLOs for infrastructure services.
- Ensure infrastructure faults can be correlated with network-service impact.
Automation & DevOps
- Drive automation of infrastructure operations.
- Develop automation using:
- Python
- Ansible
- Bash
- REST APIs
- Establish Infrastructure-as-Code practices.
- Maintain infrastructure configuration through Git.
- Develop automated health checks and validation.
- Automate provisioning and configuration.
- Integrate infrastructure operations with OSS/workflow systems.
- Promote CI/CD practices for infrastructure configuration and automation.
Reliability & Incident Management
- Act as senior technical escalation point for major infrastructure incidents.
- Lead cross-domain troubleshooting involving compute, storage, network, Linux, Kubernetes and OSS.
- Perform detailed root-cause analysis.
- Identify systemic failure modes and drive permanent remediation.
- Conduct resilience and failure testing.
- Improve MTTR and reduce repeat incidents.
- Lead technical post-incident reviews.
Technical Leadership
- Define Telco SysOps engineering standards.
- Review infrastructure designs.
- Mentor Senior and intermediate engineers.
- Provide technical guidance during major changes.
- Lead infrastructure acceptance into production.
- Challenge vendor designs and RCAs where appropriate.
Establish operational documentation and runbook standards.
Desired Experience & Qualification
Minimum Experience
- 10+ years in systems, infrastructure, telecommunications or service-provider engineering.
- Significant experience operating highly available production infrastructure.
- Advanced Linux experience.
- Advanced networking knowledge.
- Strong experience with virtualisation and/or Kubernetes.
- Demonstrated automation capability.
- Experience supporting geographically distributed production environments.
Direct 5G Core experience is not required .
Experience supporting mobile, telecommunications, ISP or other carrier-grade infrastructure is strongly advantageous.
Core Technical Skills
Candidates should demonstrate advanced capability across most of
Linux: RHEL/Rocky/Ubuntu or equivalent; systemd; networking; storage; performance troubleshooting.
Containers: Kubernetes, Docker/containerd, Helm, CNI/CSI.
Networking: TCP/IP, VLAN, VRF, BGP, OSPF/IS-IS, ECMP, load balancing, DNS, NTP and firewalls.
Automation: Python, Bash, Ansible, Git, REST APIs, YAML/JSON and Infrastructure-as-Code.
Observability: Prometheus, Grafana, ELK/OpenSearch or equivalent.
Infrastructure: VMware/KVM/OpenStack or equivalent, bare metal, SAN/NAS/object storage, HA and clustering.
Advantageous
- Red Hat RHCE/RHCA
- CKA/CKS
- CCNP/CCIE, JNCIP/JNCIE or equivalent
- VMware/OpenStack experience
- Ceph or distributed storage
- GitLab/Jenkins/CI-CD
- Terraform
- PostgreSQL/MySQL/Oracle operational experience
- Telecom OSS experience
- Telco Cloud/CNF infrastructure experience
- Experience supporting 4G/5G infrastructure without necessarily being a Core engineer