Mid AI Engineer
3 days ago
Barcelona
ph3About Valeria /h3pValeria is building the future of HR and payroll in Spain. We’re an AI-native platform that automates contracts, payroll, and compliance for companies with high employee turnover (hospitality, delivery, events, agriculture). We’re rethinking how an entire industry works—moving from manual, error‑prone processes to intelligent automation. /ppWe’re a fast‑growing startup backed by top investors, disrupting a €5B industry that is still stuck in spreadsheets and legacy software. /ph3Your Role /h3pYou’ll work closely with the AI Lead, designing product AI features and transforming internal processes with AI, helping build a cross‑functional AI team with impact across every department. You’ll be a core member of our AI team, building the agents that power Valeria in production—talking to real customers and handling real payroll and legal processes. You’ll own AI features end‑to‑end: designing agent architectures, engineering context, building evals that hold up, and shipping reliable systems into a domain where correctness genuinely matters. /ph3Responsibilities /h3ulliPrototype the complex: build proofs of concept that solve hard problems in innovative ways, then take them to production /liliTranslate business into AI: understand the business problem deeply and land it into a solid technical solution /liliDesign agents end‑to‑end: multi‑agent architectures, tools, tool‑calling, function calling, memory, orchestration, and state management /liliMaster context engineering: decide what goes into the context window and how (system prompts, few‑shot, retrieval, memory, compaction, token management), understanding why behaviour changes and anticipating failure modes (hallucinations, edge cases, prompt injection) /liliBuild the LLM harness: the layer around the model—tool interfaces, output parsing and validation, retries, fallbacks, guardrails, scaffolding, and flow control—that turns a model into a reliable production agent /liliEnsure reliability: solid evals and observability (datasets, metrics, regressions, production tracing) before every release—never on a single happy‑path /liliBuild high‑quality RAG systems: embeddings, vector stores, chunking, retrieval, re‑ranking, and grounding in a compliance‑heavy context /liliPick the right model: integrate and compare GPT, Gemini, and Claude, reasoning about cost, latency, reliability, context window, and fallback /liliIntegrate systems: build MCP servers and integrations with external systems /liliShip production code: solid Python, APIs, tests, CI/CD, and the team’s best practices /liliStay on the frontier: keep up with the latest models and technologies and test them to spot opportunities /liliOwn features end‑to‑end: from technical design to deployment, monitoring, and iteration based on customer feedback /liliMentor interns and evangelise AI across other departments as we scale the team /li /ulh3Required /h3ulli3 years of professional software engineering experience building production systems /liliStrong Python skills and solid backend fundamentals (APIs, SQL, Git, testing) /liliHands‑on experience with LLMs / agents in production: LangChain / LangGraph, RAG, prompting, tool‑calling, or equivalents /liliContext engineering and evaluation mindset: you reason about why models behave the way they do, and you validate with evals instead of a single test /liliAbility to design and break down medium‑complexity solutions autonomously, communicating progress, blockers, and trade‑offs clearly /liliStartup mindset: comfortable with ambiguity, high autonomy, and fast iteration cycles /liliStrong communication skills and ability to collaborate across product, design, and business teams /liliFluent in English and/or Spanish /li /ulh3Highly Valued /h3ulliCloud experience: Azure (Azure OpenAI / AI Foundry) and GCP (Vertex AI / Gemini) /liliHands‑on practice / familiarity with AI coding tools such as Claude Code, Cursor, Codex, and similar /liliReact and basic frontend notions for full‑stack contributions /liliExperience deploying agents/models in production at scale /liliMCP, advanced function calling, and evaluation frameworks (LangSmith, RAGAS, or similar) /liliFine‑tuning / model optimization techniques /liliBackground in FinTech, HR‑tech, or regulated industries (compliance‑heavy products, government integrations) /li /ulh3What makes you a great fit /h3pYou’re the kind of engineer who treats LLMs as systems to be understood, not black boxes to be prompted once. You care about reliability, you anticipate how agents fail, and you build the harness and evals that keep them honest in production. You’re pragmatic but principled, comfortable moving fast in a startup, and excited to work in a domain where correctness matters—getting payroll wrong affects real people’s lives. Bonus points if you love being on the frontier of applied AI. /ph3Our Stack /h3ulliLanguage: Python, async APIs, SQL, Git /liliAI frameworks: LangChain / LangGraph /liliLLMs agents: prompting, context engineering, agent harness/scaffolding, tool‑calling, multi‑step agents, RAG, embeddings, vector databases /liliModels Cloud: Azure OpenAI, Gemini (GCP), Claude /liliQuality Observability: evals, testing, LLM tracing (Datadog) /liliChannel: WhatsApp API /li /ulh3Our technical philosophy /h3pAs an AI‑native product, we build intelligence into every layer—automating altas, bajas, payroll, and compliance through agents that run in production, not demos. We care about reliable, testable systems over framework magic, and we treat evals and observability as first‑class. If you’re excited about applying AI to solve real business problems (not building AI for AI’s sake), you’ll love working here. /ph3What We Offer /h3ulliCompetitive compensation: €40,000‑45,000 gross salary Equity /liliFree lunch when you’re at the office thanks to Kombo Nora /liliFlexible remuneration with Coverflex /liliFlexibility: Hybrid setup (HQ in Barcelona), 60 days/year remote work from anywhere /liliUnlimited vacation days—take the time you need, no counting days /li /ulh3Our hiring process /h3ulliIntro call with People (30 min) /liliInterview with the Hiring Manager (45 min) /liliTech Assessment – Onsite at the office (1 hour) /liliFounders interview (45 min) /liliOffer /li /ulh3Why Join Valeria Now? /h3ulliTiming: We’re past the “idea stage” with real customers and revenue, but early enough that you’ll define how we scale our AI /liliMarket opportunity: €5B market in Spain, every company with employees needs payroll, and current solutions are outdated and painful /liliReal AI ownership: you won’t assist on AI projects—you’ll build them, decide on them, and see their impact on thousands of people /liliCareer growth: be a critical AI hire, build the playbook, and grow as we scale the team /li /ul /p #J-18808-Ljbffr