Planning, building and operating scalable Kubernetes clusters with GPU support as the basis of an AI service platform
Optimization of cluster resources for AI/ML workloads (scheduling, GPU management, scaling)
Conception, implementation and review of security guidelines
Development and implementation of an update, patch and recovery strategy
Implementation and operation of monitoring and observability solutions (e.g., Prometheus, Grafana)
Close collaboration in an agile, interdisciplinary team of data scientists, ML/AI engineers, and full-stack software developers.
Support and mentoring of the development team regarding questions about deployment, operation, and efficient use of Kubernetes
Your profile:
Completed studies (Master's degree or equivalent) in computer science, mathematics or a related field
Several years of practical experience (at least 5 years) in operating container and virtualization environments with a focus on Kubernetes (ideally for on-premises or hybrid cloud)
Solid knowledge of network and security concepts in the context of Kubernetes and experience with the implementation of security policies.
Experience with the deployment and administration of CI/CD pipelines (e.g., GitLab CI, Tekton, ArgoCD)
A good understanding of DevOps principles and Infrastructure-as-Code.
Advantageous, but not a requirement: Development and customization of Kubernetes operators as well as experience with MLOps technologies (Kubeflow, MLflow, vLLM, NVIDIA Triton, DeepSpeed)
Knowledge of SQL, NoSQL and vector database administration
Fluent German skills and very good written and spoken English skills