AI Platform Engineers (AI/GenAI | Cloud | Kubernetes) – Contract – Onsite – Sandton
Datafin Recruitment · Johannesburg, Gauteng
Stop applying one at a time.
JobAlertsZA auto-applies to South African jobs like this one for you, overnight. Upload your CV once — we do the applying.
Start free — we apply for you →ENVIRONMENT
- Our client is seeking highly specialised AI Platform Engineers to design, build, operate and optimise enterprise-grade AI infrastructure within a complex, regulated environment.
- This is not a general cloud engineering, IT infrastructure or data science role. The successful candidates must have hands-on experience supporting production AI workloads across multi-cloud environments and must be comfortable working across AI infrastructure, model serving, agentic AI, security, observability, infrastructure-as-code and AI cost governance.
DUTIES
- Design, deploy and optimise scalable multi-cloud AI platform infrastructure.
- Build reusable platform components for AI gateways, model serving, vector databases, data pipelines and GPU workloads.
- Develop and maintain infrastructure-as-code using Terraform, Pulumi, CloudFormation or equivalent technologies.
- Design infrastructure supporting agentic AI, including orchestration environments, tool-calling, agent memory, state management and multi-agent communication.
- Implement cloud-agnostic model-serving patterns that support workload portability.
- Define and manage AI platform SLAs covering availability, inference latency, throughput and reliability.
- Implement platform observability, monitoring, incident management, release management and operational runbooks.
- Design and implement zero-trust security controls for AI platforms.
- Manage AI compute expenditure through cost attribution, chargeback/showback, workload optimisation and usage reporting.
- Maintain technical documentation, architectural decision records and governance evidence.
- Mentor engineers and contribute to platform engineering standards and delivery practices.
REQUIREMENTS
- Senior level: approximately 5–8 years of relevant cloud and AI platform engineering experience.
- Lead/Principal level: approximately 8–12 years of relevant experience, including technical leadership and responsibility for engineering teams or platform squads.
Mandatory Technical Experience
Candidates must demonstrate meaningful production experience in most of the following:
At least two of the following AI ecosystems
- AWS Bedrock or SageMaker
- Microsoft Azure AI Foundry or Azure OpenAI
- Databricks AI
- Enterprise Hugging Face deployments
- Kubernetes, Docker, Helm and containerised platform services.
- Terraform, Pulumi, AWS CDK, CloudFormation or equivalent infrastructure-as-code.
- CI/CD and automated deployment of cloud or AI platform components.
- Production model-serving infrastructure, AI gateways or inference endpoints.
- Platform observability using tools such as Prometheus, Grafana, Datadog, OpenTelemetry or Databricks Lakehouse Monitoring.
- Cloud security, identity and access management, including OAuth/OIDC, JWT, RBAC or ABAC.
- Production incident management, SLAs, release management and operational readiness.
- Experience within banking, financial services or another highly regulated enterprise environment.
Specialist AI Experience
- Candidates should demonstrate practical experience in one or more of the following:
- Agent orchestration frameworks such as LangGraph, AutoGen, AWS Bedrock Agents or Microsoft Foundry Agent Service.
- Model Context Protocol, tool-calling APIs and agent state or memory management.
- Retrieval-augmented generation and vector database infrastructure.
- Cloud-agnostic model serving using tools such as ONNX, BentoML, Triton Inference Server or vLLM.
- MLOps platforms such as MLflow, Kubeflow or Airflow.
- GPU cluster management and inference or training workload optimisation.
- Prompt-injection prevention, output filtering, data-exfiltration controls and AI threat modelling.
AI Finops Experience
Candidates should have experience with some combination of
- AI or cloud cost attribution and tagging.
- Chargeback and showback models.
- Token, GPU, DBU or provisioned-throughput cost management.
- Rightsizing, workload scheduling and reserved or spot-instance optimisation.
- Cost dashboards, anomaly detection and cost-per-use-case reporting.
- Communicating technical cost trade-offs to senior technology, business or finance stakeholders.
Qualifications
- Postgraduate qualification in Computer Science, Information Technology, Data Science, Mathematics, Statistics, Engineering or a related quantitative field.
- A Master's degree is preferred and may be required for certain senior appointments.
Relevant certifications are strongly preferred, including
- AWS Solutions Architect Professional or AWS Machine Learning
- Microsoft Azure AI Engineer
- FinOps Certified Practitioner
- Certified Cloud Security Professional or equivalent
- HashiCorp Terraform Associate
- Kubernetes certification