Data Engineer at Sutherland
Dubai, Dubai, United Arab Emirates -
Full Time


Start Date

Immediate

Expiry Date

27 Dec, 26

Salary

3000.0

Posted On

28 Sep, 26

Experience

3 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Industry

Information Technology & Services

Description

About The Role


We are looking for a specialized Data Intelligence Machine Learning Engineer to design and implement in-house tools that automate our data labelling pipelines. Your primary goal will be to reduce our reliance on manual annotation by leveraging techniques like Active Learning, Weak Supervision, and Synthetic Data Generation. You will bridge the gap between raw data collection and model-ready datasets, ensuring high-quality labels at scale.


Key Responsibilities


  • Architect Labelling Pipelines: Design and deploy end-to-end automated labelling systems using frameworks like Snorkel, Cleanlab, or custom active learning loops.
  • Develop "Human-in-the-Loop" (HITL) Systems: Build interfaces and workflows where models pre-label data and humans only intervene on high-uncertainty samples.
  • Quality Assurance & Denoising: Implement algorithmic checks to identify and correct mislabelled or "noisy" data within existing datasets.
  • Tooling & Integration: Collaborate with software engineers to integrate labelling tools with our existing data lakes and ML training infrastructure.
  • Model Optimization: Fine-tune "teacher" models to generate high-quality pseudo-labels for "student" models.
  • Set up and maintain robust data preparation infrastructure—optimising for data quality, speed, and seamless integration with downstream MLOps pipelines.
  • Perform data visualization and in-depth analysis using advanced data and feature engineering techniques. You’ll help transform raw data into actionable insight, supporting both research and deployment.
  • Work closely with Data Scientists, Software Engineers, and Product teams to ensure high data quality and usability across products and projects.


How To Apply:

Incase you would like to apply to this job directly from the source, please click here

Responsibilities

About The Role


We are looking for a specialized Data Intelligence Machine Learning Engineer to design and implement in-house tools that automate our data labelling pipelines. Your primary goal will be to reduce our reliance on manual annotation by leveraging techniques like Active Learning, Weak Supervision, and Synthetic Data Generation. You will bridge the gap between raw data collection and model-ready datasets, ensuring high-quality labels at scale.


Key Responsibilities


  • Architect Labelling Pipelines: Design and deploy end-to-end automated labelling systems using frameworks like Snorkel, Cleanlab, or custom active learning loops.
  • Develop "Human-in-the-Loop" (HITL) Systems: Build interfaces and workflows where models pre-label data and humans only intervene on high-uncertainty samples.
  • Quality Assurance & Denoising: Implement algorithmic checks to identify and correct mislabelled or "noisy" data within existing datasets.
  • Tooling & Integration: Collaborate with software engineers to integrate labelling tools with our existing data lakes and ML training infrastructure.
  • Model Optimization: Fine-tune "teacher" models to generate high-quality pseudo-labels for "student" models.
  • Set up and maintain robust data preparation infrastructure—optimising for data quality, speed, and seamless integration with downstream MLOps pipelines.
  • Perform data visualization and in-depth analysis using advanced data and feature engineering techniques. You’ll help transform raw data into actionable insight, supporting both research and deployment.
  • Work closely with Data Scientists, Software Engineers, and Product teams to ensure high data quality and usability across products and projects.


Loading...