AI Engineer at Harnham
Utrecht, Uttaradit, Netherlands -
Full Time


Start Date

Immediate

Expiry Date

24 Dec, 26

Salary

50000.0

Posted On

25 Sep, 26

Experience

1 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Industry

Information Technology & Services

Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Engineer, LLM Inference Optimization based in Netherlands.

How To Apply:

Incase you would like to apply to this job directly from the source, please click here

Responsibilities
  • Own optimization initiatives for specific model families, customer endpoints, and inference serving backends.
  • Evaluate inference engines and recommend practical serving configurations based on workload requirements.
  • Diagnose and resolve model quality, performance, and reliability regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token.
  • Deploy, configure, benchmark, and extend modern inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or equivalent technologies.
  • Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
  • Implement or integrate advanced inference techniques such as speculative decoding, draft models, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
  • Develop reproducible benchmark harnesses covering TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory usage, reliability, and cost per token.
  • Partner with GPU kernel and platform engineers to identify bottlenecks across model code, kernels, runtimes, schedulers, gateways, and cluster infrastructure.


Loading...