Senior Applied Research Engineer | Barcelona | Up to €150k
9 days ago
Madrid
ph3Senior Applied Research Engineer | Barcelona, Spain /h3pWe are partnered with a cutting- edge AI company shaping the future of enterprise decision- making. Founded by experienced technologists from leading research environments, the firm has developed a market- leading platform purpose- built for the structured data that underpins critical business decisions. Backed by top-tier investors and trusted by some of the world’s largest organisations, the company helps enterprises unlock significant value by enabling more accurate, forward-looking decision-making. You will work on novel technical challenges in large-scale model development and contribute to technology that is changing how major organisations operate. This is an opportunity to join a category-defining company at an early stage and help shape its trajectory. /ph3Location compensation /h3pLocation: Barcelona, SpainbrSalary: Up to €150,000 (plus equity)brIndustry: Technology /ph3Key responsibilities /h3ulliProfile end-to-end distributed training runs to identify bottlenecks across compute, GPU memory, and inter-GPU communication. /liliInfluence architectural decisions to improve efficiency and reliability of large-scale training jobs, including developing Triton/CUDA kernels when needed. /liliDesign and implement model scaling, parallelisation, and memory optimisation techniques for training workloads with very large context sizes. /liliCollaborate closely with ML Researchers to diagnose architectural inefficiencies, ensure new research ideas scale efficiently in practice, and share internal knowledge on optimisation. /liliDrive productionisation and serving of models from the research side, including improving inference efficiency via techniques such as quantisation. /li /ulh3Qualifications /h3ulliMust have strong understanding of modern ML architectures and large-scale training pipelines. /liliHands-on experience running distributed training jobs on multi-GPU systems. /liliAdvanced profiling and debugging across CPU, GPU, memory usage, latency, and inter-GPU communication. /liliStrong programming skills in Python. /liliExperience with model scaling and parallelisation strategies, including tensor and pipeline parallelism. /li /ulh3Nice to have /h3ulliFamiliarity with NCCL, MPI, and distributed communication primitives. /liliKnowledge of PyTorch and Triton internals. /liliProgramming experience with C++ and CUDA. /li /ulh3Benefits /h3pCompetitive compensation with salary and equity and comprehensive benefits. Relocation support for employees moving to join the team in an office location. /ppA mission- driven, low-ego culture valuing diversity of thought, ownership, and bias towards action. /p /p #J-18808-Ljbffr