Machine Learning Systems Engineer
hace 10 días
Madrid
ph3About the job /h3 pWe are developing a highly scalable media intelligence platform that processes, analyzes, and structures large volumes of multimedia content across text, image, video, and audio. As a Senior Applied ML Engineer, you will architect and build the core backend systems that power media ingestion, processing workflows, metadata generation, AI-based analysis, semantic search, and retrieval across large media libraries. /p pWe are looking for a Senior Applied ML Engineer who can design, implement, optimize, and evaluate a production‑grade moderation pipeline using open‑source models. /p pThis role requires deep backend engineering expertise, strong system design capability, and practical experience integrating AI/ML systems into production workflows. You will work on complex media‑processing pipelines, video/audio analysis, OCR, speech‑to‑text, embedding generation, vector search, multimodal model integrations, and high‑throughput asynchronous workloads. You will collaborate closely with engineering leadership to define backend architecture, improve reliability and scalability, and guide other engineers in delivering secure, observable, and high‑performance systems. /p h3Responsibilities /h3 h3Backend Architecture System Ownership /h3 ul liArchitect, build, and operate scalable backend services for a media intelligence platform, with a focus on clean, maintainable, and production‑ready systems. /li liOwn critical backend components end to end, from system design and API contracts through implementation, deployment, monitoring, and iteration. /li liDrive architectural decisions across APIs, processing pipelines, distributed compute, storage, search, observability, cloud infrastructure, and model‑serving workflows. /li liDesign data models and storage patterns for media assets, generated metadata, embeddings, processing jobs, model outputs, search indexes, and audit trails. /li liDesign high‑throughput media ingestion and processing pipelines for large volumes of video, audio, image, and text content. /li liBuild distributed, event‑driven workflows for media processing using queues and pub/sub systems such as SQS, Kafka, Pub/Sub, or equivalent technologies. /li liImplement reliable asynchronous processing patterns, including retries, idempotency, dead‑letter queues, backpressure handling, and fault‑tolerant job execution. /li /ul h3AI/ML Integration Model Workflows /h3 ul liLead the development and optimization of metadata extraction, content analysis, scene detection, transcription, embedding generation, and multimodal AI inference workflows. /li liIntegrate and optimize AI/ML services within backend workflows, including model APIs, embedding pipelines, OCR, speech‑to‑text, scene analysis, multimodal inference, batching, caching, and fallback strategies. /li liCollaborate with ML engineers, data scientists, or external model providers to benchmark models, compare quality/latency trade‑offs, and safely roll out model upgrades. /li /ul h3Model Serving Performance Optimization /h3 ul liOptimize AI/ML inference workflows for latency, throughput, reliability, and cost across both real‑time and batch‑processing paths. /li liWork with model‑serving systems such as vLLM, Triton, TGI, SageMaker, Vertex AI, or custom inference services to improve batching, concurrency, warmup behavior, timeout handling, autoscaling, and GPU utilization. /li liEvaluate and apply practical model optimization techniques such as quantization, model distillation, batching, caching, prompt optimization, and routing to smaller or cheaper models where appropriate. /li liDesign and maintain vector search and indexing systems using technologies such as Pinecone, Weaviate, Qdrant, Elastic Vectors, FAISS, pgvector, or similar tools. /li liBuild retrieval workflows that support semantic search, similarity matching, duplicate detection, media discovery, and structured metadata search. /li liMonitor model and system performance in production, including API latency, queue depth, processing time, model error rates, GPU utilization, confidence distributions, drift signals, and cost per processed item. /li /ul h3Infrastructure, Reliability Observability /h3 ul liDeploy and operate systems on AWS, GCP, Azure, or equivalent cloud platforms, including compute, storage, networking, queues, model‑serving infrastructure, and monitoring systems. /li liEnsure system reliability through logging, metrics, tracing, alerting, dashboards, operational runbooks, and incident‑response best practices. /li /ul h3Collaboration Engineering Leadership /h3 ul liCollaborate with product, design, data, and ML teams to deliver media‑rich, AI‑powered product features. /li liMentor junior and mid‑level engineers, support technical planning, review designs, and raise engineering quality across the team. /li liParticipate in code reviews, documentation, technical planning, and continuous improvement of engineering practices. /li liEnsure code quality through testing, peer review, clear documentation, and maintainable implementation patterns. /li /ul h3Education Experience /h3 ul liBachelor's degree in Computer Science, Engineering, or equivalent practical experience. /li li5–7+ years of backend engineering experience, ideally building scalable distributed systems, media platforms, data pipelines, or high‑throughput backend services. /li liPrior experience owning major backend modules end to end, including architecture, implementation, deployment, monitoring, and production operations. /li li3+ years of experience integrating AI/ML inference systems into backend workflows, including model APIs, embedding pipelines, OCR, speech‑to‑text, scene detection, or multimodal model outputs. /li liHands‑on experience creating AI‑powered processing pipelines for image, video, audio, or text analysis. /li liPractical experience with production model optimization, especially for image, video, embedding, or multimodal models, including batching, caching, quantization, prompt optimization, routing strategies, latency reduction, and cost optimization. /li liPrior experience with vector search, semantic search, media retrieval, or similarity‑matching systems is strongly preferred. /li liExperience mentoring engineers, leading technical discussions, and influencing architectural decisions across backend, infrastructure, and AI/ML workflows. /li /ul h3Technical Skills /h3 ul liStrong expertise in Python and/or Node.js with deep understanding of building scalable RESTful APIs and backend architectures. /li liExperience with HuggingFace transformers ecosystem and deep learning frameworks such as PyTorch and TensorFlow. /li liStrong experience with SQL/NoSQL databases, schema design, and data modeling. /li liPreferred exposure to distributed systems, microservices, asynchronous processing, and event‑driven patterns with SQS, Pub/Sub, Kafka, or other queueing/pub‑sub systems. /li liExperience deploying production systems on AWS, GCP, or similar cloud platforms. /li liKnowledge of infrastructure patterns (compute, storage, networking, observability). /li /ul h3AI/ML Integration /h3 ul liExperience orchestrating embedding generation, scene detection, OCR, speech‑to‑text, image classification, video analysis, and multimodal model integrations. /li liExperience optimizing inference workflows for latency, throughput, reliability, and cost. /li liExperience working with scalable and optimized inference settings, including tuning sampling parameters, managing output‑length formats, and configuring reasoning‑related behaviors. /li liFamiliarity with practical model optimization techniques such as batching, caching, quantization, model distillation, prompt optimization, fallback routing, and use of smaller models where appropriate. /li liExperience working with model‑serving systems such as vLLM, Triton, TGI, SageMaker, Vertex AI, or custom inference services is preferred. /li liExperience working with LLM and multi‑modal evaluation and benchmarking frameworks and domain‑specific benchmarks with the ability to interpret results and optimize model performance accordingly. /li /ul h3System Design Architecture /h3 ul liPreferred understanding of distributed systems, scaling patterns, and performance engineering. /li liAbility to design modular, maintainable, and efficient architectures. /li liExperience with API versioning, modularization, and designing long‑running workflows. /li liUnderstanding of performance bottlenecks and low‑latency backend patterns. /li /ul /p #J-18808-Ljbffr