Start Date
Immediate
Expiry Date
17 Nov, 26
Salary
0.0
Posted On
19 Aug, 26
Experience
0 year(s) or above
Remote Job
Yes
Telecommute
Yes
Sponsor Visa
Yes
Skills
Industry
Information Technology & Services
Main Responsibilities
• Improve and optimise LLM inference performance across distributed, multi‑chip and multi‑node environments
• Apply strong understanding of transformer architectures, including dense and Mixture‑of‑Experts (MoE) models
• Benchmark leading LLMs (LLaMA, Mistral, Qwen, DeepSeek) across varied hardware stacks
• Design and implement attention‑level optimisations (Flash Attention, grouped‑query, sliding‑window)
• Deliver model‑level optimisation including quantisation (INT8/FP8), KV‑cache strategies, batching and parallelism
• Work closely with hardware, systems and compiler teams to co‑design efficient inference pipelines
• Build and maintain benchmarking frameworks to measure latency, throughput and scaling behaviour
• Evaluate architectural trade‑offs and contribute to deployment strategies for large‑scale environments
How To Apply:
Incase you would like to apply to this job directly from the source, please click here