Senior Telco SysOps Engineer

Comsol Networks (Pty) Ltd · Centurion, Gauteng

Stop applying one at a time.

JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.

Start free — we apply for you →

Introduction

The Senior Telco SysOps Engineer is responsible for the engineering, operation, maintenance and troubleshooting of production infrastructure supporting telecommunications services. The position is a hands-on infrastructure engineering role spanning Linux, compute, virtualisation, Kubernetes, storage, networking, monitoring and automation.

Duties & Responsibilities

Infrastructure Operations

  • Operate production infrastructure supporting telecommunications applications.
  • Maintain highly available compute, network and storage platforms.
  • Monitor infrastructure health and capacity.
  • Troubleshoot hardware, OS, network and application-platform issues.
  • Perform infrastructure upgrades and maintenance.
  • Conduct pre- and post-change validation.
  • Maintain infrastructure configuration and documentation.

Linux

  • Administer production Linux servers.
  • Troubleshoot CPU, memory, filesystem, disk and network problems.
  • Manage services and processes.
  • Analyse system and application logs.
  • Perform patching and OS maintenance.
  • Manage users, permissions, certificates and system configuration.
  • Troubleshoot DNS, NTP and other infrastructure services.

Kubernetes & Containers

  • Operate Kubernetes clusters.
  • Deploy and troubleshoot containerised applications.
  • Troubleshoot pods, nodes, services, storage and networking.
  • Monitor cluster capacity and performance.
  • Support CNF and OSS application teams at infrastructure level.

Networking

  • Troubleshoot server and application connectivity.
  • Work with VLANs, routing, VRFs and firewalls.
  • Understand BGP and dynamic routing fundamentals.
  • Troubleshoot DNS and load-balancing problems.
  • Diagnose MTU, latency, packet-loss and connectivity problems.
  • Capture and analyse traffic using tcpdump/Wireshark.
  • Troubleshoot virtual and host networking.

OSS Platform Support

Provide infrastructure support for telecommunications OSS platforms including:

  • NMS/EMS
  • Fault management
  • Performance management
  • Configuration management
  • Inventory
  • Monitoring
  • Logging
  • Telemetry and analytics platforms

Work with application and vendor teams to distinguish application faults from infrastructure faults .

Monitoring & Observability

  • Maintain infrastructure monitoring.
  • Develop and maintain dashboards.
  • Investigate alarms.
  • Maintain centralised logging.
  • Improve monitoring coverage.
  • Identify performance trends before they cause outages.
  • Assist in defining meaningful operational thresholds.

Automation

  • Develop operational scripts using Python and Bash.
  • Interact with systems using REST APIs.
  • Maintain scripts and configurations in Git.
  • Automate recurring operational activities.
  • Develop automated infrastructure validation and health checks.

Incident & Problem Management

  • Respond to production incidents.
  • Independently diagnose complex infrastructure problems.
  • Participate in engineering escalation/on-call.
  • Produce technical incident reports.
  • Perform RCA.
  • Work with equipment/software vendors on defects.
  • Implement corrective and preventative actions.

Desired Experience & Qualification

Minimum Experience

  • 5–8+ years in systems, infrastructure, ISP or telecommunications engineering.
  • Strong production Linux experience.
  • Strong IP networking fundamentals.
  • Experience with virtualisation and/or container platforms.
  • Experience operating highly available production systems.
  • Practical scripting or automation experience.

Core Technical Skills

Strong capability in

  • Linux
  • TCP/IP
  • DNS/NTP
  • VLANs
  • Routing
  • Firewalls
  • Virtualisation
  • Storage
  • Bash
  • Git
  • Wireshark/tcpdump

Working knowledge of

  • Docker/containerd
  • OpenShift
  • Python
  • REST APIs
  • Monitoring and logging platforms
  • High-availability architectures

Advantageous

  • Telecommunications/ISP environment experience
  • OSS/NMS platform experience
  • Telco Cloud/CNF infrastructure
  • Red Hat certification
  • CKA
  • CCNP/JNCIP
  • VMware/OpenStack
  • Ceph
  • Prometheus/Grafana
  • ELK/OpenSearch
  • CI/CD
  • Database administration/troubleshooting
Auto-apply to this jobView original posting ↗