About the job
We help the world run better
At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.
What You Will Build
As a Data Engineer on the Service & Support Data Lake team, you will help modernize our data engineering capabilities. We run an in-house data lake that powers AI services for SAP customer support and generates insights through project and agent mining. You will join a collaborative team evolving from isolated pipelines to a unified intelligent platform for self-service analytics and autonomous data agents. Data governance underpins every solution we deliver end to end.
Responsibilities
- Design and maintain scalable pipelines with clear architecture and high quality standards.
- Build real-time and batch ingestion from diverse sources with automated validation.
- Write production-grade code with strong testing discipline (unit/integration tests), code review quality, and maintainable modular design.
- Troubleshoot complex data issues end to end, including root-cause analysis, performance bottlenecks, and reliability incidents.
- Implement anonymization and PII controls aligned with governance requirements.
- Develop metadata pipelines for schema profiling, business context extraction, and lineage tracking.
- Contribute to semantic layers that map technical fields to business terminology for self-service and natural-language analytics use cases.
- Use AI-assisted development practices to accelerate delivery and maintenance; prior autonomous-agent experience is a plus, not a requirement.
- Establish monitoring, alerting, and data quality controls to ensure secure, reliable analytical assets, evaluate their AI/ML readiness based on data science requirements.
- Partner with the Tech Lead and Architect to convert requirements into production-ready systems.
What You Bring
- Bachelor's degree or equivalent practical experience.
- 3+ years of experience coding in Python (pandas, pytest) and SQL.
- 3+ years of experience with Spark / Big Data processing (PySpark): transformations, partitioning, performance optimization.
- 3+ years designing and deploying data pipelines, including managing data schemas and processing high-volume workflows.
- Strong software engineering fundamentals: clean code, modular design, debugging, version control, and maintainable documentation.
- Strong SQL and data modeling capabilities (normalized and denormalized patterns, data contracts, schema evolution).
- Hands-on testing and release practices: unit/integration testing, CI/CD pipelines, and safe production rollout.
- Experience in observability and operations: metrics, logging, alerting, and on-call friendly troubleshooting.
- Experience with SQL databases (PostgreSQL preferred, others acceptable) and NoSQL databases (Elasticsearch and Delta Lake required).
- Experience with real-time and batch ingestion (APIs - polling & push, Kafka streaming).
- Experience with workflow orchestration (Kubeflow Pipelines, Airflow, Prefect, or Dagster).
- Experience with data governance: data redaction, anonymization, and PII handling.
- Proficiency in Git workflows: branching strategies, code review, CI/CD integration.
Preferred Qualifications
- 5+ years designing enterprise-scale data platforms and analytics infrastructure.
- Familiarity with LLM-enabled data applications, including RAG, embeddings, vector search, and evaluation from a data platform perspective.
- Experience building data services and data APIs that support AI-related applications and analytics products.
- Experience productionizing data and feature pipelines that support machine learning and intelligent applications.
- Strong interest in AI-native data engineering and agentic data workflows, including orchestration, tool integration, evaluation, and workflow automation.
- Understanding of MLOps/LLMOps principles to ensure scalable and reliable deployment of text processing and redaction pipelines.
- Experience with monorepo or shared library architecture patterns.
- Ability to operate across ambiguity and influence cross-functional technical decisions.
- Demonstrated learning agility in adopting emerging data concepts (for example, data agents and semantic layer patterns).