Supervisor: System Monitoring (2988)
The South African National Roads Agency SANRAL · Gauteng
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →MINIMUM REQUIREMENTS
- A Bachelors degree or an equivalent qualification at an NQF level 7: Information Technology, Electrical Engineering, Electronic Engineering or equivalent.
- A minimum of five 5 years relevant experience in systems or networking monitoring.
- A minimum of three 3 years supervisory or team lead experience.
Certification
- ITIL Foundation
- CCNA and/or
- CCNP
The successful candidate will be required to provide after hours support during major network incidents or infrastructure emergencies.
ADVANTAGEOUS
- Fibre Optic Network Design / Telecommunications Engineering certifications.
- Network +.
TECHNICAL COMPETENCIES
- Advanced knowledge of enterprise monitoring operations, including real-time monitoring of applications, infrastructure, networks, databases, and digital services in a 24/7 environment.
- Strong understanding of server infrastructure, WAN/LAN connectivity, network devices, storage systems, and service availability monitoring.
- Understanding of ITIL-based Incident, Problem, and Change Management processes, SLA management, escalation protocols, operational continuity practices and Root Cause Analysis RCA methodologies.
- Knowledge of system health indicators, capacity monitoring, performance analytics, trend analysis, and proactive fault detection.
KEY RESPONSIBILITIES
Monitoring Operations Supervision: 24/7 Environment
- Supervise and coordinate 24/7 monitoring operations within SANRAL's Central Operations Centre COC, ensuring uninterrupted operational oversight of all critical ICT services.
- Oversee continuous monitoring of enterprise systems, including, business applications, network infrastructure, server environments, databases, Intelligent Transportation Systems ITS, toll operational platforms, payment processing systems and website and customer-facing digital services.
- Ensure real-time visibility of system performance, service health, alarms, and infrastructure availability through integrated monitoring dashboards.
- Monitor and validate operational dashboards, event consoles, and automated alerts generated by monitoring tools.
- Ensure proper shift handover processes, including transfer of unresolved incidents, operational updates, and outstanding risks.
- Verify that all monitoring activities comply with approved Standard Operating Procedures SOPs, escalation protocols, and service continuity processes.
- Coordinate operational response during service disruptions, system failures, and infrastructure incidents.
- Ensure proactive identification of operational risks before they escalate into service outages.
- Maintain oversight of all mission-critical systems to ensure operational readiness across SANRAL's digital services environment.
- Escalate critical service degradation affecting tolling operations, financial systems, and traffic management services.
Incident Detection, Escalation and Response Management
- Ensure rapid identification, validation, and classification of incidents affecting monitored systems.
- Assess system alerts to determine operational severity, business impact, and urgency.
- Distinguish between genuine incidents, false positives, and informational events.
- Coordinate timely escalation of incidents to relevant technical teams, including Infrastructure support, Network support, Application support, Database administrators, Security teams and external service providers.
- Ensure incidents are accurately logged, categorised, prioritised, and tracked in IT service management platforms.
- Monitor incident response and resolution times against approved service level agreements SLAs.
- Ensure all critical incidents are escalated within defined timelines.
- Provide operational oversight during Priority 1 and Priority 2 incidents, ensuring immediate response and structured communication.
- Coordinate bridge calls and operational communications during major incidents.
- Track incident lifecycle from detection to closure, ensuring completeness of incident records.
- Review recurring incidents and identify escalation trends requiring permanent corrective action.
- Support post-incident reviews and Root Cause Analysis RCA processes.
System Health Monitoring and Performance Oversight
- Monitor overall health, performance, and availability of ICT infrastructure, applications, databases, networks, and operational systems.
- Ensure proactive identification of performance bottlenecks, capacity constraints, and service degradation risks.
- Analyse operational metrics, trends, and system performance indicators to support service optimisation.
- Coordinate with technical teams to address infrastructure, application, and network performance issues.
- Monitor capacity utilisation and recommend improvements to support business growth and operational resilience.
- Ensure operational monitoring thresholds, baselines, and performance indicators remain aligned with business requirements.
- Validate system availability targets and ensure compliance with agreed uptime requirements.
- Support proactive maintenance activities aimed at improving system stability and service reliability.
Monitoring, Administration and Optimisation Tools
- Oversee the administration and effective operation of enterprise monitoring tools used within the COC.
- Ensure monitoring systems are configured to provide accurate event detection, threshold management, and alerting.
- Validate integration between monitoring platforms and service management tools.
- Review and optimise alert configurations to reduce false positives and alert fatigue.
- Ensure proper maintenance of monitoring dashboards, event correlation engines, and reporting systems.
- Coordinate tuning of alerts based on operational trends and system changes.
- Recommend enhancements to monitoring capabilities, including automation and predictive analytics.
- Support implementation of new monitoring technologies and tools.
- Ensure monitoring coverage aligns with changes in SANRAL's ICT environment.
- Coordinate dashboard enhancements to improve operational visibility for management and technical teams.
Reporting, Compliance and Audit Support
- Produce daily, weekly, monthly, and ad hoc operational monitoring reports.
- Report on System availability, Infrastructure uptime, Incident volumes, SLA performance, Major service disruptions, Operational risks and Capacity trends
- Maintain accurate operational records for governance, audit, and compliance purposes.
- Support internal and external ICT audits by providing monitoring evidence, incident records, and operational logs.
- Ensure all monitoring activities comply with SANRAL governance standards and ICT operational policies.
- Maintain updated monitoring documentation, SOPs, incident procedures, and escalation matrices.
- Assist in preparation of management reports for ICT governance forums.
- Track compliance with service management and operational reporting standards.
- Support risk assessments related to operational service continuity.
Team Leadership and Operational Coordination
- Supervise, guide, and support System Monitoring Analysts and operational monitoring personnel.
- Allocate shift responsibilities and ensure adequate operational coverage across all monitoring functions.
- Manage shift schedules, standby arrangements, and operational resource planning.
- Conduct regular performance reviews and operational coaching sessions.
- Identify skills gaps and support staff development initiatives.
- Ensure staff adherence to monitoring procedures, escalation protocols, and reporting requirements.
- Foster a high-performance operational culture focused on.
- Support onboarding and training of new monitoring personnel.
- Promote collaboration between monitoring teams and other ICT support units.
- Ensure operational discipline and consistency across all monitoring shifts.