Research Engineer (LLM Training and Performance)
hace 4 días
Madrid
ppAt JetBrains, code is our passion. Ever since we started back in 2000, we have been striving to make the strongest, most effective developer tools on earth. By automating routine checks and corrections, our tools speed up production, freeing developers to grow, discover, and create. /ppWe’re looking for a Research Engineer who will own the training stack and model architecture for our Mellum LLM family. Your job is easier said than done: make training faster, cheaper, and more stable at a large scale. You’ll profile, design, and implement changes to the training pipeline – from architecture to custom GPU kernels, as needed. /ph3As Part Of Our Team, You Will /h3ulliBe responsible for improving end-to-end performance for multi-node LLM pre‑training and post‑training pipelines. /liliProfile hotspots (Nsight Systems/Compute, NVTX) and fix them using compute/comm overlap, kernel fusion, scheduling, etc. /liliDesign and evaluate architecture choices (depth/width, attention variants including GQA/MQA/MLA/Flash‑style, RoPE scaling/NTK, and MoE routing and load‑balancing). /liliImplement custom ops (Triton and/or CUDA C++), integrate via PyTorch extensions, and upstream when possible. /liliPush memory/perf levers: FSDP/ZeRO, activation checkpointing, FP8/TE, tensor/pipeline/sequence/expert parallelism, NCCL tuning. /liliHarden large runs by building elastic and fault‑tolerant training setups, ensuring robust checkpointing, strengthening reproducibility, and improving resilience to preemption. /liliKeep the data path fast using streaming and sharded data loaders and tokenizer pipelines, as well as improve overall throughput and cache efficiency. /liliDefine the right metrics, build dashboards, and deliver steady improvements. /liliRun both pre‑training and post‑training (including SFT, RLHF, and GRPO‑style methods) efficiently across sizable clusters. /li /ulpWe’ll be happy to bring you on board if you have: /pulliStrong PyTorch and PyTorch Distributed experience, having run multi‑node jobs with tens to hundreds of GPUs. /liliHands‑on experience with Megatron‑LM/Megatron‑Core/NeMo, DeepSpeed, or serious FSDP/ZeRO expertise. /liliReal profiling expertise (Nsight Systems/Compute, nvprof) and experience with NVTX‑instrumented workflows. /liliGPU programming skills with Triton and/or CUDA, and the ability to write, test, and debug kernels. /liliA solid understanding of NCCL collectives, as well as topology and fabric effects (IB/RoCE), and how they show up in traces. /li /ulh3Our Ideal Candidate Would Have Experience With /h3ulliFlashAttention‑2 and 3, CUTLASS and CuTe, TransformerEngine and FP8, Inductor, AOTAutograd, and torch.compile. /liliMoE at scale (expert parallel, router losses, capacity management) and long‑context tricks (ALiBi/YaRN/NTK scaling). /liliKubernetes or SLURM at scale, placement and affinity tuning, as well as AWS, GCP, and Azure GPU fleets. /liliWeb‑scale data plumbing (streaming datasets, Parquet and TFRecord, tokenizer perf), eval harnesses, and benchmarking. /liliSafety and post‑training methods, such as DPO, ORPO, GRPO, and reward models. /liliInference ecosystems such as vLLM and paged KV. /li /ulh3We are an equal opportunity employer /h3pWe know great ideas can come from anyone, anywhere. That’s why we do our best to create an open and inclusive workplace – one that welcomes everyone regardless of their background, identity, religion, age, accessibility needs, or orientation. /ppWe process the data provided in your job application in accordance with the Recruitment Privacy Policy. /p /p #J-18808-Ljbffr