Data Engineer at Taod Consulting GmbH
Berlin, Berlin, Germany -
Full Time


Start Date

Immediate

Expiry Date

08 Jan, 27

Salary

65000.0

Posted On

10 Oct, 26

Experience

2 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Industry

Information Technology & Services

Description
  • Skills & QualificationsA bachelor's degree in a relevant field (e.g., computer science, computer engineering, software engineering) is required.
  • 5+ years of experience designing, implementing, and managing large-scale distributed data processing systems, or working within large-scale distributed ML data frameworks, with recent experience using e.g. Ray, Apache Spark, workflow orchestrators, Apache Arrow, and/or Parquet.
  • Demonstrated ownership of a data processing system under real throughput, reliability, and cost pressure. The specific frameworks matter less to us than evidence that you have had to reason about where a large pipeline breaks and why.
  • Experience profiling and optimizing throughput and cost across heterogeneous workloads, including GPU-accelerated stages.
  • Experience with dataset versioning, lineage, and reproducibility tooling.
  • Ability to collaborate effectively with cross-functional teams, document best practices, and stay updated with the latest advancements in large-scale data processing and software development.
  • Experience with workload managers (e.g., Ray, Kubernetes, Slurm).
  • Familiarity with containerization tools (e.g., Docker, Enroot).
  • Familiarity with data infrastructures and platforms (e.g., vector databases).


Responsibilities
  • Skills & QualificationsA bachelor's degree in a relevant field (e.g., computer science, computer engineering, software engineering) is required.
  • 5+ years of experience designing, implementing, and managing large-scale distributed data processing systems, or working within large-scale distributed ML data frameworks, with recent experience using e.g. Ray, Apache Spark, workflow orchestrators, Apache Arrow, and/or Parquet.
  • Demonstrated ownership of a data processing system under real throughput, reliability, and cost pressure. The specific frameworks matter less to us than evidence that you have had to reason about where a large pipeline breaks and why.
  • Experience profiling and optimizing throughput and cost across heterogeneous workloads, including GPU-accelerated stages.
  • Experience with dataset versioning, lineage, and reproducibility tooling.
  • Ability to collaborate effectively with cross-functional teams, document best practices, and stay updated with the latest advancements in large-scale data processing and software development.
  • Experience with workload managers (e.g., Ray, Kubernetes, Slurm).
  • Familiarity with containerization tools (e.g., Docker, Enroot).
  • Familiarity with data infrastructures and platforms (e.g., vector databases).


Loading...