Senior Research Engineer (Agentic Behavior)
hace 12 días
Madrid
ppAt JetBrains, code is our passion. Ever since we started, back in 2000, we've been striving to make the strongest, most effective developer tools on earth. Today, AI-powered coding agents are becoming a core part of how developers write Kotlin – and we want to make sure they write it well. /p pThe Kotlin AI Value Stream team is responsible for how AI agents understand, generate, and improve Kotlin code across all platforms: Android, Kotlin Multiplatform, server‑side, web, desktop, and others. We build the evaluation infrastructure, error analysis tools, and post‑training pipelines that measure and improve agent behavior on real Kotlin developer tasks. /p pAs a Research Engineer on this team, you'll own the end‑to‑end loop: Analyze how agents fail on Kotlin → build evals that capture those failures → research and implement methods to fix them → measure the improvement. Your work will directly shape how millions of developers experience Kotlin through AI coding agents. /p h3As Part Of Our Team, You Will /h3 h3Build tools for agentic error analysis /h3 ul liDesign and implement tooling to systematically capture, classify, and analyse errors that AI coding agents make when generating Kotlin code. /li liBuild observability pipelines over agentic traces – mining patterns from agent sessions in JetBrains IDEs, Junie, Claude Code, Cursor, and other coding agents. /li /ul h3Build evaluation pipelines /h3 ul liDesign, implement, and maintain evaluation pipelines that measure Kotlin code generation quality across dimensions, including correctness, idiomaticity, build success, framework usage, and test coverage. /li liBuild simulation environments where coding agents can be measured on realistic Kotlin developer tasks – from greenfield KMP projects and Gradle dependency management to migrating Spring applications from Java to Kotlin. /li liOwn evaluation infrastructure: metrics, experiment tracking, automated regression checks, and reproducible benchmarking. /li /ul h3Research methods for improving agent and model behavior on Kotlin /h3 ul liExperiment with post‑training techniques (SFT, DPO, GRPO) to improve how models handle Kotlin‑specific patterns, idioms, and frameworks. /li liInvestigate context‑engineering approaches: CLAUDE.md/AGENTS.md files, compiler‑as‑verifier feedback loops, Kotlin LSP integration, and MCP‑based tooling. /li liRun experiments to measure impact: A/B comparisons, benchmark suites, and before/after analyses on real codebases. /li liCollaborate with model providers (Anthropic, OpenAI, and Google) to translate Kotlin‑specific findings into model improvements. /li /ul h3Build public Kotlin benchmarks /h3 ul liDesign and build open‑source benchmarks that measure AI coding agent performance on Kotlin tasks and eventually become the standard reference for the ecosystem. /li liCreate task datasets covering the breadth of Kotlin usage: the server side (Spring, Ktor), multiplatform projects (KMP), build systems (Gradle), Android, library development, and others. /li liInclude both mined real‑world tasks and carefully designed synthetic tasks that test specific Kotlin capabilities. /li liMaintain and evolve benchmarks as models improve, ensuring they remain challenging, relevant, and contamination‑resistant. /li /ul h3We'll be happy to have you on board if you have: /h3 ul liHands‑on experience building evaluation or analysis pipelines for LLMs or AI coding agents in a research or production setting. /li liStrong Python engineering skills (at least three years), with the ability to write clean, maintainable code in data‑heavy and ML‑adjacent codebases. /li liExperience with data analysis at scale: querying large datasets (SQL/Athena), building data pipelines, and performing statistical analysis of experimental results. /li liThe ability to own projects end to end – from identifying a problem in agent traces to designing an eval, running experiments, and shipping a fix. /li liA product‑aware mindset: You care about how agents are actually used by developers and can translate real failure modes into evaluation and training work. /li liFamiliarity with Kotlin or a strong willingness to develop deep Kotlin expertise (you'll be living in Kotlin codebases daily). /li /ul h3Our Ideal Candidate Would Also Have Experience With /h3 ul liPost‑training LLMs: SFT, RLHF, DPO, GRPO – either hands‑on training or designing the data and reward pipelines that feed into training. /li liModern deep learning frameworks (PyTorch) and LLM training stacks (TRL, verl, Megatron, or similar). /liliAI agent development: tool‑using agents, multi‑step coding workflows, agentic frameworks. /li liEvaluation frameworks and tools: Inspect AI, Promptfoo, LM‑evaluation‑harness, or custom eval pipelines. /li liExperiment tracking and observability: Weights Biases, MLflow, Langfuse, or similar. /li liThe Kotlin ecosystem: Android, Gradle, KMP, Spring, Ktor – with an understanding of the developer workflows that agents need to support. /li liContributing to or maintaining open‑source projects, especially benchmarks or evaluation tools. /li /ul pDon't check every box? That's okay – if you're excited about this work and bring strong fundamentals, we'd love to hear from you. We're happy to talk and provide the training you need to grow into the role. /p h3Why join JetBrains? /h3 ul liStrong base salary. We offer competitive pay that reflects your skills and experience. /li liFlexible work location. Enjoy the freedom to work from home or from the office. /li liRemote work. Spend up to 30 days per year working remotely from abroad. /li liExtra time off. More days to relax, recharge, and do the things you love. /li liMedical insurance allowance. Enjoy peace of mind for you and your family /li liLearning and development opportunities. Access to conferences, courses, and language classes. /li liRelocation support. We help make your move as smooth and stress‑free as possible. /li liLanguage classes. Pick up the local language or sharpen your English skills. /li liFuel your day. Enjoy a hot meal or receive a lunch allowance on workdays. /li liMental health support. To help you feel your best, we provide easy access to professional mental health services. /li liSports benefit. Enjoy an on‑site gym or sports club stipend. /li liInternal events. Join company‑wide celebrations and team gatherings. /li liSome benefits may vary depending on location. /li /ul h3We are an equal opportunity employer /h3 pWe know great ideas can come from anyone, anywhere. That’s why we do our best to create an open and inclusive workplace – one that welcomes everyone regardless of their background, identity, religion, age, accessibility needs, or orientation. /p pWe process the data provided in your job application in accordance with the Recruitment Privacy Policy. /p /p #J-18808-Ljbffr