About the job
We help the world run better
At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.
What You Will Build
As a Data Engineer on the Service & Support Data Lake team, you will help modernize our data engineering capabilities. We run an in-house data lake that powers AI services for SAP customer support and generates insights through project and agent mining. You will join a collaborative team evolving from isolated pipelines to a unified intelligent platform for self-service analytics and autonomous data agents. Data governance underpins every solution we deliver end to end.
Responsibilities
- Design and maintain scalable pipelines with clear architecture and high quality standards.
- Build real-time and batch ingestion from diverse sources with automated validation.
- Write production-grade code with strong testing discipline (unit/integration tests), code review quality, and maintainable modular design.
- Troubleshoot complex data issues end to end, including root-cause analysis, performance bottlenecks, and reliability incidents.
- Implement anonymization and PII controls aligned with governance requirements.
- Develop metadata pipelines for schema profiling, business context extraction, and lineage tracking.
- Contribute to semantic layers that map technical fields to business terminology for self-service and natural-language analytics use cases.
- Use AI-assisted development practices to accelerate delivery and maintenance; prior autonomous-agent experience is a plus, not a requirement.
- Establish monitoring, alerting, and data quality controls to ensure secure, reliable analytical assets, evaluate their AI/ML readiness based on data science requirements.
- Partner with the Tech Lead and Architect to convert requirements into production-ready systems.
What You Bring
- Bachelor's degree or equivalent practical experience.
- 3+ years of experience coding in Python (pandas, pytest) and SQL.
- 3+ years of experience with Spark / Big Data processing (PySpark): transformations, partitioning, performance optimization.
- 3+ years designing and deploying data pipelines, including managing data schemas and processing high-volume workflows.
- Strong software engineering fundamentals: clean code, modular design, debugging, version control, and maintainable documentation.
- Strong SQL and data modeling capabilities (normalized and denormalized patterns, data contracts, schema evolution).
- Hands-on testing and release practices: unit/integration testing, CI/CD pipelines, and safe production rollout.
- Experience in observability and operations: metrics, logging, alerting, and on-call friendly troubleshooting.
- Experience with SQL databases (PostgreSQL preferred, others acceptable) and NoSQL databases (Elasticsearch and Delta Lake required).
- Experience with real-time and batch ingestion (APIs - polling & push, Kafka streaming).
- Experience with workflow orchestration (Kubeflow Pipelines, Airflow, Prefect, or Dagster).
- Experience with data governance: data redaction, anonymization, and PII handling.
- Proficiency in Git workflows: branching strategies, code review, CI/CD integration.
Preferred Qualifications
- 5+ years designing enterprise-scale data platforms and analytics infrastructure.
- Familiarity with LLM-enabled data applications, including RAG, embeddings, vector search, and evaluation from a data platform perspective.
- Experience building data services and data APIs that support AI-related applications and analytics products.
- Experience productionizing data and feature pipelines that support machine learning and intelligent applications.
- Strong interest in AI-native data engineering and agentic data workflows, including orchestration, tool integration, evaluation, and workflow automation.
- Understanding of MLOps/LLMOps principles to ensure scalable and reliable deployment of text processing and redaction pipelines.
- Experience with monorepo or shared library architecture patterns.
- Ability to operate across ambiguity and influence cross-functional technical decisions.
- Demonstrated learning agility in adopting emerging data concepts (for example, data agents and semantic layer patterns).
Incase you would like to apply to this job directly from the source, please click here