This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Engineer, LLM Inference Optimization based in Switzerland.
As a Senior Machine Learning Engineer, you will drive the optimization of large language and vision-language model inference from model artifacts through production deployment. You will work across model internals, inference engines, serving architectures, and benchmarking to improve latency, throughput, memory efficiency, GPU utilization, reliability, and cost per token. This is a hands-on role focused on solving complex performance challenges and delivering measurable improvements to production systems. You will collaborate closely with kernel, platform, infrastructure, research, product, and customer-facing teams. Your work will involve evaluating serving configurations, diagnosing performance and quality regressions, and implementing advanced inference optimization techniques. You will also establish reproducible benchmarks and safe rollout practices for high-throughput AI workloads.