Senior Machine Learning Engineer at Zalando
Berlin, Berlin, Germany -
Full Time


Start Date

Immediate

Expiry Date

18 Nov, 26

Salary

0.0

Posted On

20 Aug, 26

Experience

0 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

Yes

Skills

Industry

Information Technology & Services

Description

WHAT WE’D LOVE YOU TO DO (AND LOVE DOING)



  • Design & Architecture Play a key role in the design, architecture, and development of end-to-end data engineering and MLOps solutions with full operational responsibilities on cloud infrastructure (AWS, Databricks, Kubernetes).
  • Streaming Pipelines Gather requirements and design high-throughput, low-latency batch and real-time feature pipelines using Apache Flink (Java) and Spark (Python) to provision production-grade features to our central Hopsworks Feature Store.
  • System Operationalization Drive the operationalization, model serving, and maintenance (MLOps/MLaaS) of our real-time inference sponsored prediction system and new Ad candidate retrieval systems.
  • Science Collaboration Collaborate closely with Applied Scientists to optimize data pipeline runtime, data quality, and model performance, latency, and memory usage.
  • Operational Excellence Take ownership of the operational excellence of our AI systems, implementing robust CI/CD pipelines, continuous monitoring, and automated alerting for distributed systems to maximize scalability and reliability.
  • Communication & Roadmapping Communicate effectively with product managers, data scientists, and engineering peers, translating complex engineering concepts into actionable roadmaps.
  • Workflow Automation Strive to continuously improve and automate the time-to-market for the team's experimentation-to-production workflows.


WE’D LOVE TO MEET YOU IF



  • Solid Foundation You hold a degree in Computer Science, a related technical field, or have equivalent practical experience showcasing strong software engineering fundamentals.
  • Streaming & Data Engineering You have significant hands-on experience designing, building, and maintaining high-throughput, low-latency data streaming applications. Practical experience with Apache Flink and Spark is highly required.
  • MLOps & Model Serving You possess professional experience in machine learning operationalization, model serving (e.g., Triton, SageMaker), data version control, and workflow orchestration (e.g., Airflow or Databricks workflows).
  • Strong Programming Skills You are proficient in Java and/or Python, with a strong passion for writing clean, testable, and maintainable production code. Familiarity with ML libraries (e.g., PyTorch, TensorFlow) is a major plus.
  • Modern Practices You are well-versed in Agile methodologies, CI/CD pipelines, and establishing effective metrics and monitoring for large-scale distributed systems.
  • Collaboration & Mentorship You have experience working closely with applied scientists and mentoring other engineers, with excellent verbal and written communication skills to bridge technical gaps across stakeholders.


Responsibilities
Loading...