Data Engineer at Thermo Fisher Scientific
United Arab Emirates, Dubai, United Arab Emirates -
Full Time


Start Date

Immediate

Expiry Date

29 Nov, 26

Salary

0.0

Posted On

31 Aug, 26

Experience

0 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Industry

Information Technology & Services

Description

About the job

We help the world run better


At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.


What You Will Build


As a Data Engineer on the Service & Support Data Lake team, you will help modernize our data engineering capabilities. We run an in-house data lake that powers AI services for SAP customer support and generates insights through project and agent mining. You will join a collaborative team evolving from isolated pipelines to a unified intelligent platform for self-service analytics and autonomous data agents. Data governance underpins every solution we deliver end to end.


Responsibilities


  • Design and maintain scalable pipelines with clear architecture and high quality standards.
  • Build real-time and batch ingestion from diverse sources with automated validation.
  • Write production-grade code with strong testing discipline (unit/integration tests), code review quality, and maintainable modular design.
  • Troubleshoot complex data issues end to end, including root-cause analysis, performance bottlenecks, and reliability incidents.
  • Implement anonymization and PII controls aligned with governance requirements.
  • Develop metadata pipelines for schema profiling, business context extraction, and lineage tracking.
  • Contribute to semantic layers that map technical fields to business terminology for self-service and natural-language analytics use cases.
  • Use AI-assisted development practices to accelerate delivery and maintenance; prior autonomous-agent experience is a plus, not a requirement.
  • Establish monitoring, alerting, and data quality controls to ensure secure, reliable analytical assets, evaluate their AI/ML readiness based on data science requirements.
  • Partner with the Tech Lead and Architect to convert requirements into production-ready systems.

What You Bring


  • Bachelor's degree or equivalent practical experience.
  • 3+ years of experience coding in Python (pandas, pytest) and SQL.
  • 3+ years of experience with Spark / Big Data processing (PySpark): transformations, partitioning, performance optimization.
  • 3+ years designing and deploying data pipelines, including managing data schemas and processing high-volume workflows.
  • Strong software engineering fundamentals: clean code, modular design, debugging, version control, and maintainable documentation.
  • Strong SQL and data modeling capabilities (normalized and denormalized patterns, data contracts, schema evolution).
  • Hands-on testing and release practices: unit/integration testing, CI/CD pipelines, and safe production rollout.
  • Experience in observability and operations: metrics, logging, alerting, and on-call friendly troubleshooting.
  • Experience with SQL databases (PostgreSQL preferred, others acceptable) and NoSQL databases (Elasticsearch and Delta Lake required).
  • Experience with real-time and batch ingestion (APIs - polling & push, Kafka streaming).
  • Experience with workflow orchestration (Kubeflow Pipelines, Airflow, Prefect, or Dagster).
  • Experience with data governance: data redaction, anonymization, and PII handling.
  • Proficiency in Git workflows: branching strategies, code review, CI/CD integration.

Preferred Qualifications


  • 5+ years designing enterprise-scale data platforms and analytics infrastructure.
  • Familiarity with LLM-enabled data applications, including RAG, embeddings, vector search, and evaluation from a data platform perspective.
  • Experience building data services and data APIs that support AI-related applications and analytics products.
  • Experience productionizing data and feature pipelines that support machine learning and intelligent applications.
  • Strong interest in AI-native data engineering and agentic data workflows, including orchestration, tool integration, evaluation, and workflow automation.
  • Understanding of MLOps/LLMOps principles to ensure scalable and reliable deployment of text processing and redaction pipelines.
  • Experience with monorepo or shared library architecture patterns.
  • Ability to operate across ambiguity and influence cross-functional technical decisions.
  • Demonstrated learning agility in adopting emerging data concepts (for example, data agents and semantic layer patterns).
Responsibilities
Loading...