Start Date
Immediate
Expiry Date
16 Dec, 26
Salary
30000.0
Posted On
17 Sep, 26
Experience
5 year(s) or above
Remote Job
Yes
Telecommute
Yes
Sponsor Visa
No
Skills
Industry
Electrical Equipment Manufacturing
Principal Evals Engineer - AI & Agentic Systems
About the Organisation
Our client is building one of the region’s most ambitious government AI programmes, developing the AI infrastructure, applications and platforms required to operate AI-native public services at scale.
The systems being built must operate reliably across Arabic and English, within strict requirements around data sovereignty, privacy, security and regulated infrastructure.
These are production systems supporting consequential workflows and senior decision-makers. Whether an AI system works cannot simply be a matter of opinion.
It has to be measurable.This role owns how that measurement happens.
The Role
We are hiring a Principal Evals Engineer to own how the organisation determines whether its AI systems actually work.
You will set the direction for evaluation across a portfolio spanning AI assistants, retrieval systems, agentic workflows, voice applications and document intelligence, while building the infrastructure that turns “it seems better” into measurable evidence.
You own measurement and release evidence, partnering closely with Principal-level engineers responsible for AI architecture, product engineering, platform development and forward deployment.
Evaluating AI systems is fundamentally different from conventional software testing. The same input can produce different outputs, correctness can be subjective, and failures such as hallucination, poor grounding, behavioural drift and prompt injection cannot be captured through conventional assertions alone.
At Principal level, you will be expected to have already tackled these problems in production environments.
This is primarily an individual contributor role. Your influence comes from the infrastructure you build, the quality bar you establish and the engineering decisions your evidence enables.
What You Own
Evaluation strategy and architecture.Define what gets measured, at which layer, using which methodology and how results feed into product and release decisions. Establish the reference architecture engineering teams build against.
The shared evaluation platform.Build evaluation harnesses, golden-set management, dataset versioning, automated grading, behavioural regression detection and reporting infrastructure.
Grading you can trust.Own judge-model selection, rubric design and calibration against human labels — including understanding when automated grading cannot be trusted.
Retrieval, agents and multilingual evaluation.Measure grounding, citation correctness, tool use, multi-step reasoning and failure recovery. Build dedicated Arabic evaluation datasets and ensure judges are properly calibrated rather than assuming English evaluation methods transfer directly.
Evidence behind engineering decisions.Model swaps, prompt changes, framework migrations and infrastructure decisions should be supported by defensible evaluation evidence before changing production behaviour.
Production and adversarial evaluation.Build online evaluation, sampling, human review, drift detection and alerting alongside adversarial testing for prompt injection, jailbreak resistance and data leakage.
Quality gates and escapes.Integrate AI evaluation with conventional test automation and CI/CD, turning production quality escapes into evaluations capable of detecting the same failure before release.
Evaluation culture.Enable engineering teams to run rigorous evaluations independently and ensure evaluation is designed into systems from the beginning rather than added before launch.
How To Apply:
Incase you would like to apply to this job directly from the source, please click here
Principal Evals Engineer - AI & Agentic Systems
About the Organisation
Our client is building one of the region’s most ambitious government AI programmes, developing the AI infrastructure, applications and platforms required to operate AI-native public services at scale.
The systems being built must operate reliably across Arabic and English, within strict requirements around data sovereignty, privacy, security and regulated infrastructure.
These are production systems supporting consequential workflows and senior decision-makers. Whether an AI system works cannot simply be a matter of opinion.
It has to be measurable.This role owns how that measurement happens.
The Role
We are hiring a Principal Evals Engineer to own how the organisation determines whether its AI systems actually work.
You will set the direction for evaluation across a portfolio spanning AI assistants, retrieval systems, agentic workflows, voice applications and document intelligence, while building the infrastructure that turns “it seems better” into measurable evidence.
You own measurement and release evidence, partnering closely with Principal-level engineers responsible for AI architecture, product engineering, platform development and forward deployment.
Evaluating AI systems is fundamentally different from conventional software testing. The same input can produce different outputs, correctness can be subjective, and failures such as hallucination, poor grounding, behavioural drift and prompt injection cannot be captured through conventional assertions alone.
At Principal level, you will be expected to have already tackled these problems in production environments.
This is primarily an individual contributor role. Your influence comes from the infrastructure you build, the quality bar you establish and the engineering decisions your evidence enables.
What You Own
Evaluation strategy and architecture.Define what gets measured, at which layer, using which methodology and how results feed into product and release decisions. Establish the reference architecture engineering teams build against.
The shared evaluation platform.Build evaluation harnesses, golden-set management, dataset versioning, automated grading, behavioural regression detection and reporting infrastructure.
Grading you can trust.Own judge-model selection, rubric design and calibration against human labels — including understanding when automated grading cannot be trusted.
Retrieval, agents and multilingual evaluation.Measure grounding, citation correctness, tool use, multi-step reasoning and failure recovery. Build dedicated Arabic evaluation datasets and ensure judges are properly calibrated rather than assuming English evaluation methods transfer directly.
Evidence behind engineering decisions.Model swaps, prompt changes, framework migrations and infrastructure decisions should be supported by defensible evaluation evidence before changing production behaviour.
Production and adversarial evaluation.Build online evaluation, sampling, human review, drift detection and alerting alongside adversarial testing for prompt injection, jailbreak resistance and data leakage.
Quality gates and escapes.Integrate AI evaluation with conventional test automation and CI/CD, turning production quality escapes into evaluations capable of detecting the same failure before release.
Evaluation culture.Enable engineering teams to run rigorous evaluations independently and ensure evaluation is designed into systems from the beginning rather than added before launch.