MLOps Engineer at EPAM Systems, Inc
Utrecht, Utrecht, Netherlands -
Full Time


Start Date

Immediate

Expiry Date

18 Dec, 26

Salary

0.0

Posted On

19 Sep, 26

Experience

7 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Industry

Information Technology & Services

Description

Responsibilities

  • Build and maintain platform components for ML model training, deployment, serving and monitoring
  • Develop and optimize CI/CD pipelines for machine learning workflows
  • Implement and support model lifecycle management, including registries and observability tooling
  • Design and manage scalable, secure deployments using containerization and Kubernetes
  • Enable secure, reusable and automated workflows to enhance ML developer productivity
  • Extend platform capabilities to support LLMOps, RAG and agentic AI workloads
  • Collaborate with engineering teams to improve reliability, automation and operational maturity
  • Apply governance and compliance standards across AI operations
  • Participate in presales and client-facing sessions to translate requirements into scalable solutions
  • Advocate cloud best practices for reliability, scalability and cost optimization


Responsibilities

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering or related discipline
  • Experience in delivering machine learning or MLOps systems into production environments
  • Proficiency in Python for building services, APIs, scripts and CI/CD automation
  • Working knowledge of modern MLOps stacks including experiment tracking and artifact management
  • Hands-on experience with orchestration tools (e.g., Kubeflow, Apache Airflow, Metaflow or Prefect)
  • Demonstrated skills with Docker, Kubernetes and distributed deployments
  • Practical knowledge of Infrastructure-as-Code (Terraform) and a major cloud provider (AWS, Azure or GCP)
  • Familiarity with ML model serving, scaling and monitoring frameworks in production
  • Strong communications skills to convey technical decisions and engage with clients effectively

Nice to have

  • Background deploying Generative AI solutions, LLM inference pipelines or agentic AI systems
  • Experience with feature stores, vector databases and retrieval-augmented generation approaches
  • Knowledge of AI governance, security and compliance for regulated sectors
  • Familiarity with advanced observability and tracing solutions, such as OpenTelemetry or Langfuse
  • Consulting or enterprise architecture experience in large-scale AI programs
  • Understanding of FinOps strategies for managing GPU/CPU costs in cloud environments
  • Certifications in cloud technologies (AWS, Azure, GCP) or Kubernetes (CKA/CKAD)
  • Expertise in securing and operationalizing ML/LLM/agent-based systems for enterprise readiness


Loading...