Design, develop, and maintain end to end data pipelines (batch and near real time) on AWS Data Platform
Build and manage ETL/ELT workflows using AWS services (e.g., AWS Glue, S3, Redshift, Athena, EMR), dbt and orchestration tools such as Airflow
Implement data ingestion patterns from diverse sources (databases, APIs, files, event streams) into lake/warehouse layers such as raw, cleansed, and curated data layers
Develop transformation logic using SQL and Python/PySpark for cleansing, enrichment, and standardisation
Implement robust data quality checks, reconciliation controls, and monitoring/alerting for failures and anomalies
Collaborate with data analysts/data scientists to model datasets for analytics and machine learning consumption.
Contribute to DataOps/DevOps practices: version control, CI/CD, automated testing, release management, and operational support.
Produce and maintain technical documentation (data flows, mappings, job schedules, runbooks, and operational procedures)
Optimise Data Pipeline performance and Support workflow orchestration and scheduling
Support production deployments and operations
Required Skills & Experience
Extensive experience as a Data Engineer
Advanced SQL skills
Hands on experience working with Teradata and Siebel CRM data sets
Experience delivering data pipelines in a large-scale enterprise data platform environment
Strong hands-on AWS experience with common data services such as : Amazon S3, AWS Glue, Amazon Redshift, Amazon Athena, Amazon EMR and dbt
Strong programming capability in Python and strong data transformation experience using PySpark (preferred) and/or Spark.