Site Reliability Engineer at Zalando
Germany, Vienna, Germany -
Full Time


Start Date

Immediate

Expiry Date

18 Nov, 26

Salary

0.0

Posted On

20 Aug, 26

Experience

0 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

Yes

Skills

Industry

Information Technology & Services

Description

THE ROLE AND THE TEAM

The team serves as the bridge between Customer Operations, Software Development and Infrastructure teams and are responsible for ensuring stability, high-performance, and scalability across our global application landscape. 

Central to the role is proactive prevention: monitoring and alerting to catch issues before impact, proactive issue detection and rigorous Root Cause Analysis (RCA). You will support our 1st and 2nd level support teams with deep-dive code/environment level troubleshooting and automation of repetitive tasks. 

INCLUSIVE BY DESIGN

If you think you have what it takes, we encourage you to apply even if you don't meet every single requirement. You may just be the right candidate for this or other roles!

At Zalando, our vision is to be the leading European technology platform for fashion and lifestyle – one that thrives on diversity and is truly inclusive by design. We believe that diverse teams fuel innovation and creativity, and we actively seek out talent from all backgrounds.

We actively seek to reduce bias in our hiring and employment processes, focusing on your qualifications, skills, and contributions. To support this, we kindly ask that you refrain from including personal details such as your photo, age, or marital status in your CV, ensuring a fair and equitable evaluation based solely on your abilities and potential.

We are committed to providing an exceptional and accessible candidate experience for everyone. If you require any accommodations to support you throughout the hiring process, please let us know – we are here to assist you.

Discover more about our commitment to creating a diverse and inclusive workplace:https://jobs.zalando.com/en/our-culture/diversity-and-inclusion

WHAT WE’D LOVE YOU TO DO (AND LOVE DOING) 

  • Monitoring, Alerting & Observability: Implementing tools and dashboards to track system health and business processes in real-time, ensuring the team is proactively notified of any technical or functional anomalies.
  • Incident & Problem Management (L3-Support): Providing expert-level troubleshooting to resolve complex technical issues, performing root-cause analysis and proposing solutions to prevent recurring failures.
  • Performance Optimisation & Load Testing: Optimizing system performance and scalability by assessing responsiveness under load and fine-tuning configurations to ensure a seamless experience for our users.
  • Security Management & Vulnerability Scanning: Proactively identifying and mitigating security vulnerabilities within the application and its environment, following a standardized framework.
  • Compliance & Audit Readiness (ISO 27001, SOC2, GDPR): Ensuring all operations meet legal and industry standards and providing the necessary evidence for internal or external audits.

Responsibilities
Loading...