Start Date
Immediate
Expiry Date
18 Nov, 26
Salary
0.0
Posted On
20 Aug, 26
Experience
0 year(s) or above
Remote Job
Yes
Telecommute
Yes
Sponsor Visa
Yes
Skills
Industry
Information Services
Job Description
Roles & Responsibilities
Role Purpose :
The Data Engineer prepares the Group’s enterprise data and knowledge sources so that AI, Generative AI, and Agentic AI solutions operate on information that is accurate, current, governed, and fit for purpose.
Spanning structured data in core business systems and unstructured content across documents, policies, and correspondence, the role builds and maintains the pipelines and knowledge repositories that underpin retrieval-augmented AI. It is the control point that ensures agents draw only on trusted, approved sources — making this role the single greatest determinant of whether the Group’s AI outputs can be relied upon in business decisions
▶ Data Preparation for AI Use Cases
– Prepare enterprise data and knowledge sources for AI use cases, working from prioritised business requirements defined with the AI / Agentic AI Lead.
– Support data extraction, cleansing, classification, tagging, and indexing across structured, semi-structured, and unstructured sources.
– Build and maintain ingestion pipelines from ERP, CRM, HRMS, procurement systems, the data lake, and document repositories.
– Design chunking, metadata, and enrichment strategies that materially improve retrieval relevance and answer quality.
– Handle multi-format content — PDF, Office documents, scanned material, and email — including OCR and text extraction where required.
▶ Knowledge Repositories & RAG Enablement
– Build and maintain knowledge repositories for RAG-based AI solutions, including embedding generation, vector store management, and index refresh cycles.
– Implement versioning and change detection so that repositories remain synchronised with authoritative source systems.
– Define and apply access controls at the data layer so that retrieval respects existing entitlement and confidentiality boundaries.
– Measure and tune retrieval performance, working with AI engineers to diagnose grounding failures and improve recall and precision.
▶ Data Quality & Master Data Readiness
– Work with functional and technical teams to improve data quality and master data readiness across customer, vendor, product, employee, and asset domains.
– Profile source data to quantify completeness, consistency, duplication, and timeliness, and report readiness objectively to initiative sponsors.
– Implement automated data quality rules, validation checks, and exception reporting within pipelines.
– Support remediation of root-cause data issues with business data owners rather than correcting symptoms downstream.
▶ Governance, Trust & Security
– Ensure AI agents use trusted, approved, and governed data sources, and that unapproved or unclassified content is excluded from AI consumption.
– Apply data classification, retention, and privacy requirements in line with Group policy and UAE data protection regulation.
– Maintain lineage and cataloguing so that any AI output can be traced back to its underlying source with confidence.
– Collaborate with Cybersecurity and Compliance on access reviews, data residency, encryption, and audit evidence.
▶ Platform Operations & Collaboration
– Operate and optimize cloud data platforms and pipelines for reliability, performance, and cost efficiency. – Monitor pipeline health, resolve failures, and maintain documentation, runbooks, and operational handover materials. – Partner with AI engineers, application teams, and business analysts throughout the delivery cycle, from discovery to production support. – Contribute to Group data standards, reusable pipeline patterns, and shared engineering practice.
Desired Candidate Profile
▶ Education – Bachelor’s degree in Computer Science, Information Systems, Data Engineering, Statistics, or a related discipline. – Postgraduate qualification in Data Science, Analytics, or Computer Science is an advantage. ▶ Professional Certifications – Cloud data certification such as AWS Certified Data Engineer or Data Analytics, Microsoft Azure Data Engineer Associate, Google Cloud Professional Data Engineer, Databricks Data Engineer, or SnowPro. – Certification or formal training in data governance, data management (for example DAMA CDMP), or data privacy is advantageous. – Training in AI/GenAI data preparation, vector databases, or RAG architecture is an asset. ▶ Experience – 5–8 years of experience in data engineering, data platforms, analytics, ETL/ELT, data warehousing, or cloud data solutions. – Hands-on experience in preparing structured and unstructured data for analytics, AI, GenAI, and Agentic AI use cases. – Experience with tools such as AWS S3, Glue, Redshift, Athena, Azure Data Factory, Synapse, Fabric, Databricks, Snowflake, BigQuery, or similar. – Demonstrated experience improving data quality and master data readiness in partnership with business functions. – Experience building and maintaining knowledge repositories or search/retrieval indexes is strongly preferred. – Exposure to multi-entity or group environments with heterogeneous source systems is an advantage. ▶ Key Skills & Attributes – Strong SQL and Python capability with a disciplined, production-grade engineering approach. – Rigorous attention to data accuracy, lineage, and reproducibility. – Ability to assess and communicate data readiness honestly, including when a use case should not yet proceed. – Effective collaboration with business data owners to resolve issues at source. – Sound understanding of data privacy, classification, and security obligations. – Pragmatic balance between speed of delivery and long-term maintainability of data assets.
Employment Type
Company Industry