Event Streaming Engineer
Falcorp Resourcing · Fairlands · Market Related
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →Introduction
The Event Streaming Engineer is responsible for designing, deploying, operating, and scaling high-throughput, low-latency, real-time data streaming platforms and event-driven architectures. Operating at the core of enterprise event distribution, real-time integration platforms, and distributed streaming engines, this professional acts as a key technical contributor who builds, hardens, and maintains the core infrastructure underpinning event-driven ecosystems. The position requires a systems and platform engineer with strong problem-solving skills, a deep structural understanding of distributed log systems, broker cluster topology, and partitioning strategies, alongside a performance-oriented approach to platform reliability, network throughput, and data availability.
Duties & Responsibilities
- Design, deploy, automate, and manage scalable, enterprise-grade event streaming clusters (Apache Kafka, Confluent Platform, AWS Kinesis, or Apache Pulsar) across hybrid cloud environments.
- Own event platform infrastructure reliability, broker configuration, cluster sizing, partition placement, and replication topologies to guarantee fault tolerance and high availability.
- Manage and scale event platform ecosystem components, including Schema Registries, Kafka Connect clusters, REST Proxies, and Stream Processing engines (Kafka Streams, Flink, or ksqlDB).
- Implement automated cluster management, multi-region replication (e.g., MirrorMaker 2, Confluent Cluster Linking), disaster recovery strategies, and zero-downtime rolling upgrades.
- Enforce end-to-end event platform security standards, including TLS/SSL encryption, SASL/Kerberos authentication, Fine-Grained Access Control (ACLs), and integration with enterprise IAM systems.
- Establish comprehensive real-time telemetry, alerting, and observability frameworks for distributed brokers, connectors, and consumer groups using tools like Prometheus, Grafana, and OpenTelemetry.
- Partner closely with application development teams to advise on consumer group scaling, partition key design, schema evolution, dead-letter queue (DLQ) patterns, and optimal producer/consumer configurations.
- Diagnose, debug, and resolve platform-level incidents, broker degradation, network bottlenecks, disk I/O constraints, and consumer lag spikes.
- Drive Infrastructure-as-Code (IaC) practices using Terraform, Ansible, Docker, and Kubernetes (e.g., Strimzi or Confluent Operators) to automate streaming platform provisioning and GitOps pipelines.
- Participate actively in Agile ceremonies, providing precise technical estimations, infrastructure story point sizing, and clear operational execution pathways during sprint planning.
Desired Experience & Qualification
- National Diploma or Bachelor’s Degree in Computer Science, Computer Engineering, Information Technology, or a related technical discipline.
- A minimum of 5 years of commercial engineering experience, with at least 3 years dedicated strictly to engineering, operating, and maintaining production-grade event streaming platforms and distributed message brokers.
- Relevant industry certifications such as Confluent Certified Administrator for Apache Kafka (CCAAK), Confluent Certified Developer (CCDAK), or AWS Certified Data Engineer / DevOps Engineer are highly preferred.
- Experience in Telecommunications, Media, FinTech, or Large-Scale Enterprise domains will be advantageous.
- Deep structural knowledge of distributed systems internals, consensus mechanisms (ZooKeeper / KRaft), page cache management, JVM tuning, and disk I/O optimization.
- Advanced proficiency in Linux systems administration, network engineering (TCP/IP, socket tuning, load balancing), and shell scripting (Bash, Python, or Go).
- Proven hands-on experience with container orchestration (Kubernetes), Helm charts, and cloud-native operator models for streaming platforms.
- Solid grasp of database integration patterns, Change Data Capture (CDC) with Debezium, and schema governance (Avro, Protobuf, JSON Schema).
- Advanced proficiency navigating, customizing, and generating insights from enterprise tracking ecosystems.
- Exceptional verbal and written communication capability, with a proven track record of translating complex platform infrastructure constraints into clear summaries for technical and non-technical stakeholders.