Data & AI Engineer (JHB Hybrid)
Datafin · Johannesburg, Gauteng · R Undisclosed
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →ENVIRONMENT
A provider of tailored Financial Solutions seeks a Data & AI Engineer to build, maintain, and improve reliable pipelines and curated datasets that feed the company’s analytics, reporting, and AI workloads, and you will help run the governed in-house AI platform that sits on top of them. This is one combined role, not a Data Engineer with an AI paragraph bolted on. Data flows from SQL Server and Azure SQL sources into Fabric, through bronze, silver, and gold layers, and out to the semantic, reporting, and AI consumption layers the business depends on. Expect the work to sit at roughly two-thirds Data Engineering and one-third applied AI in the first year, shifting further towards AI as adoption grows. Both halves run inside the same established standards, which cover metadata-driven, incremental, idempotent ELT, source control, gated deployment, and human review of anything AI-generated. You will grow into full ownership of production data and AI assets end-to-end: build it, document it, support it.
DUTIES
Data Engineering: Governed pipelines and the estate -
- Develop and maintain batch and incremental ELT pipelines that ingest, transform, and deliver trusted datasets across our SQL Server / Azure SQL / Microsoft Fabric estate.
- Write and optimise T-SQL transformation logic, stored procedures, views, and reconciliation queries, following our standards (CTEs over subqueries, idempotent and metadata-driven patterns).
- Build and maintain Fabric pipelines, Dataflow Gen2 flows, and Lakehouse / Warehouse data assets.
- Support data modelling: dimensional models, medallion (bronze / silver / gold) layer design, partitioning, and performance optimisation for analytics workloads.
- Operate the estate: orchestration, scheduling, monitoring, logging, alerting, and methodical incident resolution for data jobs and platform processes.
- Implement data validation, reconciliation, row-count and parity checks so reporting outputs can be trusted and prove it with evidence.
Applied AI: The governed AI platform -
- Build and maintain the curated, documented datasets, retrieval indexes, and interfaces that governed AI platform and AI-assisted use cases consume.
- Help run the in-house AI platform: content ingestion, indexing, configuration, access control, and routine health checks.
- Develop Python automation for AI workloads, covering data preparation, embedding and indexing jobs, evaluation harnesses, and approved API integrations.
- Build and evaluate AI use cases such as search, summarisation, classification, extraction, and workflow automation, measuring each against defined success criteria with human-review checkpoints built in.
- Enforce data minimisation and anonymisation on every AI path: no personal information, credentials, or production data into unapproved tools, and no uncontrolled egress of data.
- Record AI evaluation evidence, including what was tested, what the model produced, what a human reviewer changed, and why, so that AI-assisted output stays traceable and auditable.
Shared: Governance, delivery, and documentation -
- Work the governed delivery flow: Git / Azure DevOps source control, feature branches, peer-reviewed pull requests, and gated environment promotion (Dev ? UAT ? production).
- Produce and maintain documentation: pipeline design notes, source-to-target mappings, runbooks, AI use-case and configuration notes, and support procedures.
- Collaborate with Analysts, the Head of Data & AI, and business stakeholders to translate requirements into governed, supportable data assets, without business logic drifting into unmanaged reports.
What you’ll work with –
- Azure SQL and on-premises SQL Server, which form the core transactional and reporting estate.
- The wider Fabric surface, including mirroring and CDC ingestion, deployment pipelines, and capacity management.
- Python for transformation, automation, and tooling; DAX exposure on the BI side.
- Azure DevOps for repos, boards, CI/CD pipelines, and gated deployments.
- Power BI and SSRS as the consumption layer you feed.
- The governed in-house AI platform, which covers agent tooling, retrieval and indexing, and the controls that make it safe to use on sensitive data. You will help run it, not just use it.
- AI-assisted engineering under policy, meaning sanitised inputs, human review of generated code, and traceable output. It is how the whole team works, and how you will work too.
- Evaluation and human-review workflows, which are how we decide whether an AI output is good enough to use and how that decision is evidenced.
REQUIREMENTS
- Minimum 1-year hands-on experience in Data Engineering, ETL / ELT, analytics engineering, database development, applied AI, or a closely related production-oriented role.
- Solid SQL Server and T-SQL. You should be confident with joins, aggregations, CTEs, and multi-step transformation logic, and able to debug and reason about query performance.
- Demonstrable Microsoft Fabric experience. You have built, supported, or substantially prototyped at least one of the following: a Fabric pipeline, a Lakehouse or Warehouse solution, a Dataflow Gen2 flow, or a Fabric notebook. Work, internship, or serious project evidence all count.
- Understanding of data warehousing and pipeline design concepts (staging, incremental loads, dimensional modelling).
- Working Python. You can read, write, and modify scripts for data preparation, automation, and integration, because the AI half of the role runs on it.
- Practical, sceptical familiarity with modern AI tooling. You have used LLM or AI-assisted tools on real work and can explain where they got things wrong, not just that they were useful. No formal Machine-Learning or Data-Science background is required.
- Familiarity with source control (Git or similar) and structured, reviewed ways of releasing changes.
- Genuine care for data quality, security, least privilege, and confidentiality, including the discipline never to put sensitive data into an unapproved tool.
Advantageous –
- Python or PySpark depth for data transformation at scale.
- Dataflow Gen2, mirroring / CDC, and Fabric deployment pipelines.
- Power BI semantic model and DirectLake awareness; SSRS report development.
- Experience with APIs, file feeds, or enterprise source systems.
- Retrieval-augmented generation, vector or semantic search, and embedding concepts.
- Azure AI services, prompt and evaluation frameworks, agent tooling, or MLOps exposure.
- DP-700 (Fabric Data Engineer Associate) or another relevant Microsoft certification. These are valued rather than required and are achievable after appointment.
- Relevant qualification (Computer Science, Information Systems, Engineering, or related) or equivalent practical experience.
- Experience in a regulated financial-services or similarly controlled environment.
ATTRIBUTES
- You want to learn, so you ask good questions, absorb feedback, and close the loop.
- You verify before you trust, and that includes AI output above all.
- You take ownership end-to-end: build it, document it, support it.
- You have a quality mindset, so you check your work and raise issues early rather than hoping.
- You are comfortable with sensitive data and disciplined governance; process is a feature to you, not friction.
- You keep composure in a delivery-focused environment and communicate honestly about progress and blockers.
- Integrity and values, with the ability to handle sensitive and confidential information.
- Attention to detail and a results-driven quality mindset.
- Strong problem-solving; composure in a fast-moving, dynamic environment.
- Works well in a team and independently; communicates openly.
Growth mindset, actively wanting to learn and develop
Desired Skills
- Data
- AI
- Engineer