Senior Applied Scientist, Efficient LLM Inference & Model Optimization

Nebius · London

<div class="content-intro"><p><strong>About Nebius:</strong></p> <p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p> <p>Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.</p> <p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&amp;D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&amp;D.</p></div><h3><strong><span data-contrast="auto">The role</span></strong><span data-ccp-props="{}">&nbsp;</span></h3> <p><span data-ccp-props="{}">Nebius Token Factory needs scientists who can turn frontier inference bottlenecks into research problems, publish credible work, and then help ship the results into production. This is not a papers-only research role. The Applied Scientist is expected to design rigorous experiments, write strong code, collaborate with engineers, and convert research into deployed inference capabilities.</span></p> <p>A Senior Applied Scientist owns well-scoped research and production optimization projects. They can publish or prepare high-quality technical work while also producing code, experiments, and prototypes that engineers can use.</p> <p><strong><span data-contrast="auto"><span data-ccp-charstyle="Strong">Your responsibilities</span><span data-ccp-charstyle="Strong">:</span></span></strong><span data-ccp-props="{"134233117":true,"134233118":true}">&nbsp;</span></p> <ul> <li> <p data-renderer-start-pos="2682" data-local-id="06f359694ab7">Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff.</p> </li> <li> <p data-renderer-start-pos="2796" data-local-id="7c017788c1e5">Prepare internal reports, technical blogs, or papers when the work is externally credible.</p> </li> <li> <p data-renderer-start-pos="2890" data-local-id="7efe020884ba">Partner directly with MLEs to ensure research prototypes become usable production components.</p> </li> <li> <p data-renderer-start-pos="2987" data-local-id="1481cbfe2c41">Define and execute research programs in efficient LLM and <span data-highlighted="true" data-vc="highlighted-text">VLM</span> inference with measurable production impact.</p> </li> <li> <p data-renderer-start-pos="3097" data-local-id="87a2ac19f5af">Invent, evaluate, and productionize methods for quantization, <span data-highlighted="true" data-vc="highlighted-text"><span class="_kqswh2mm"><span class="_5pioz8co _189e1dm9 _1il9buyh _19lc184f _d0altlke" data-testid="definition-highlighter">QAT</span></span></span>, distillation, speculative decoding, <span data-highlighted="true" data-vc="highlighted-text">KV</span>-cache reuse, <span data-highlighted="true" data-vc="highlighted-text">KV</span>-cache compression, long-context inference, <span data-highlighted="true" data-vc="highlighted-text">MoE</span> routing, and model/runtime co-optimization.</p> </li> <li> <p data-renderer-start-pos="3313" data-local-id="01d7f1451adb">Build high-quality prototypes in PyTorch, Triton, <span data-highlighted="true" data-vc="highlighted-text">CUDA</span>-adjacent tooling, or inference-serving frameworks, then work with MLEs and platform engineers to productionize them.</p> </li> <li> <p data-renderer-start-pos="3488" data-local-id="7252b6ce1513">Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.</p> </li> <li> <p data-renderer-start-pos="3642" data-local-id="c70959baf3b3">Publish papers, technical reports, blog posts, and open-source artifacts that build external credibility for Nebius Token Factory.</p> </li> <li> <p data-renderer-start-pos="3776" data-local-id="ccb9fc6eec1b">Collaborate with MLE, GPU kernel, backend infrastructure, product, and customer teams to choose high-leverage research bets.</p> </li> <li> <p data-renderer-start-pos="3904" data-local-id="b653633dbf1f">Mentor engineers and scientists on experimental design, scientific rigor, and model/system tradeoffs.</p> </li> </ul> <p><strong><span data-contrast="auto"><span data-ccp-charstyle="Strong">Must-haves</span><span data-ccp-charstyle="Strong">:</span></span></strong><span data-ccp-props="{}">&nbsp;</span></p> <ul> <li> <p data-renderer-start-pos="4036" data-local-id="d0838f7fe31a"><span data-highlighted="true" data-vc="highlighted-text">PhD</span> in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field.</p> </li> <li> <p data-renderer-start-pos="4201" data-local-id="d6a3c2d74b2c">Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.</p> </li> <li> <p data-renderer-start-pos="4385" data-local-id="5f311fbd5e78">Strong hands-on coding ability in Python and PyTorch; ability to move from idea to experiment to prototype quickly.</p> </li> <li> <p data-renderer-start-pos="4504" data-local-id="12c5979cb37e">Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs.</p> </li> <li> <p data-renderer-start-pos="4652" data-local-id="2a3641ca38bf">Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis.</p> </li> <li> <p data-renderer-start-pos="4776" data-local-id="5ad1594b4f5b">Excellent written and verbal communication.</p> </li> </ul> <p><strong><span data-contrast="auto"><span data-ccp-charstyle="Strong">Nice</span><span data-ccp-charstyle="Strong">-</span><span data-ccp-charstyle="Strong">to</span><span data-ccp-charstyle="Strong">-</span><span data-ccp-charstyle="Strong">have</span><span data-ccp-charstyle="Strong">s</span><span data-ccp-charstyle="Strong">:</span></span></strong><span data-ccp-props="{}">&nbsp;</span></p> <ul> <li> <p data-renderer-start-pos="4851" data-local-id="7b1fc9cd775f">First-author publications in NeurIPS, <span data-highlighted="true" data-vc="highlighted-text">ICML</span>, <span data-highlighted="true" data-vc="highlighted-text">ICLR</span>, MLSys, <span data-highlighted="true" data-vc="highlighted-text">ACL</span>, <span data-highlighted="true" data-vc="highlighted-text">EMNLP</span>, <span data-highlighted="true" data-vc="highlighted-text">ASPLOS</span>, <span data-highlighted="true" data-vc="highlighted-text">OSDI</span>, <span data-highlighted="true" data-vc="highlighted-text">SOSP</span>, <span data-highlighted="true" data-vc="highlighted-text">ISCA</span>, <span data-highlighted="true" data-vc="highlighted-text">HPCA</span>, or comparable venues.</p> </li> <li> <p data-renderer-start-pos="4977" data-local-id="dfddd25a2b60">Experience deploying ML models or inference optimizations in production.</p> </li> <li> <p data-renderer-start-pos="5053" data-local-id="389d84a803be">Experience with vLLM, SGLang, TensorRT-LLM, <span data-highlighted="true" data-vc="highlighted-text">NVIDIA</span> Dynamo, FlashAttention, FlashInfer, Triton, <span data-highlighted="true" data-vc="highlighted-text">CUDA</span>, or PyTorch internals.</p> </li> <li> <p data-renderer-start-pos="5179" data-local-id="a23b78c77076">Experience with post-training, <span data-highlighted="true" data-vc="highlighted-text"><span class="_kqswh2mm"><span class="_5pioz8co _189e1dm9 _1il9buyh _19lc184f _d0altlke" data-testid="definition-highlighter">SFT</span></span></span>, <span data-highlighted="true" data-vc="highlighted-text"><span class="_kqs

View application details