Software Engineer, ML Ops and Platform at Tools for Humanity
berlin, Berlin, Germany -
Full Time


Start Date

Immediate

Expiry Date

03 Dec, 26

Salary

0.0

Posted On

04 Sep, 26

Experience

0 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Industry

Information Technology & Services

Description

About the Team


We are building a planet‑scale biometric recognition system that will serve more than a billion users and enable them to become part of the World protocol. We use cutting-edge Machine Learning models deployed on custom hardware to enable high-quality image acquisition, identification, and fraud prevention, all while requiring minimal user interaction.

We are building a biometric recognition and fraud detection engine that works on the 1bn people scale. Therefore, its performance needs to outperform all the current recognition technologies. We leverage our powerful custom-made iris recognition and presentation attack detection device, the Orb, combined with the latest research from the field of AI and Deep Learning.


To reach our next milestone—continuous, trustworthy ML innovation across millions of edge devices—we’re hiring a Senior ML Platform & Ops Engineer to own the ML lifecycle from data to device. You’ll design and operate production-grade pipelines that transform state-of-the-art ML research into deployed models with clear telemetry, rollback, and reproducibility. If you thrive on building self‑service platforms that turn research ideas into reliable, observable production systems, we’d love to meet you.

Key Responsibilities

  • Design, build, and operate reliable, observable infrastructure for training, evaluation, telemetry ingestion, and deployment.
  • Maintain CI/CD workflows and automated pipelines.
  • Edge‑aware rollout services with staged deployment, A/B experimentation and instant rollback across Orbs, Orb Mini and Mobile Apps.
  • Develop secure APIs and backend services that expose governed datasets and model artefacts at scale.
  • Implement automated checks, drift detection, and alerting for real-time model monitoring.
  • Champion best practices in data lineage, reproducibility, privacy‑by‑design, security and secure edge delivery.
  • Collaborate across ML research, product, and firmware teams to streamline delivery and feedback loops.

About You

  • 5+ years building ML infrastructure, data platforms, or production ML systems at scale.
  • Track record of delivering platforms and CI/CD pipelines used daily by ML or data teams.
  • Hands-on experience running large-scale training on multi-tenant GPU clusters to maximize throughput and reliability.
  • You’ve built versioned dataset & lineage systems with slice-level provenance and governed access, making every model reproducible to the exact data, features, code, and config used.
  • Deep understanding of containerisation (Docker) and orchestration (Kubernetes/EKS) plus Infrastructure‑as‑Code (Terraform/CDK/Cloudformation).
  • Strong backend engineering skills in Python and/or Go; you value clean, maintainable code.
  • Deep understanding of modern CI/CD, model packaging, and observability practices.
  • Comfortable operating production systems, defining SLAs, and handling rollout or incident workflows.
  • Comfortable using modern Agentic AI development

Responsibilities
Loading...