Site Reliability Engineer
1 day ago
Madrid
ph3About Tinybird /h3pAt Tinybird, we help developers and data teams unlock the power of real‑time data. We enable the rapid building of data pipelines and innovative data products by ingesting multiple data sources at scale, querying them with familiar SQL, and publishing low‑latency, high‑concurrency APIs for applications. Developers can create fast APIs in minutes, transforming what once took hours or days. /ph3About the Platform Team /h3pThe Platform team builds, operates, and continually improves the technical foundations that Tinybird relies on. We ensure safety at scale, predictable evolution, and product and customer growth with minimal operational friction. Our focus is on reliability, observability, performance, infrastructure, cost efficiency, CI/CD, development environments, and critical backend services. We are both an infrastructure and operations team, making Tinybird’s foundations reliable, observable, scalable, and easier to evolve. /ph3What We Are Looking For /h3pbRole /b: Site Reliability Engineer who loves keeping large‑scale distributed systems reliable and adaptable as they grow. You should understand hardware and software, grasp our product, and tackle real challenges faced by customers and internal teams. /pulliStrong experience designing, building, and running distributed cloud architectures and large‑scale web‑based production systems. /liliDeep knowledge of Kubernetes – essential. Comfortable designing and operating production‑grade clusters, writing custom controllers or operators, and tuning autoscaling mechanisms (e.g., KEDA, Karpenter). /liliKnowledge of how Kubernetes manages networking, storage, scheduling, and resources; able to analyze performance and failure scenarios at scale. /liliSkilled in AWS and GCP. /liliBaseline coding skills (Python or C++). Able to explore our codebase, ClickHouse source code, and understand how things work. /liliExperience operating close to production: debugging incidents, analyzing system behavior, improving observability, and enhancing service reliability. /liliSystems thinking with attention to edge cases, failure modes, and implementation details. Careful about performance, reliability, cost efficiency, and operational simplicity. /liliAction orientation – quick decisions, iterative delivery, and ownership over issues that need fixing. /liliCuriosity about data and SQL; comfortable querying our own data. Experience with ClickHouse or launching databases at scale is a plus. /liliFamiliarity with Traefik, Varnish, Redis, Terraform, or Ansible is helpful but not mandatory. /liliStrong written communication for asynchronous work and documentation. /liliFluency in English and Spanish (English is primary; Spanish is common in the Platform team). /liliWillingness to participate in on‑call rotations and located in an EU timezone. /li /ulh3Our Stack /h3ulliKubernetes – foundation of our infrastructure. /liliClickHouse – primary data store. /liliPython – backend; some performance‑critical components in C++. /liliVarnish – load balancing and caching. /liliTraefik – ingress and routing. /liliRedis – metadata store. /liliZookeeper – coordinating ClickHouse replicas. /liliArgoCD – GitOps‑based continuous delivery. /liliGrafana, Loki, Mimir, OpenTelemetry – monitoring, alerting, telemetry. /li /ulh3What You Will Work On /h3ulliEnhance high‑availability and elasticity for automatic, efficient scaling. /liliImprove observability across resource usage, service metrics, telemetry, dashboards, and alerts. /liliStrengthen disaster recovery tools, incident discovery, and on‑call experience. /liliHandle Kubernetes lifecycle tasks, cluster infrastructure, autoscaling, and safe deployments. /liliUnderstand ClickHouse internals and extract peak performance. /liliIdentify bottlenecks and improve performance across storage, networking, and compute. /liliReduce operational burden by automating and managing processes. /liliAssist with incident prevention, post‑reliability reviews, and follow‑up tasks. /liliStrengthen CI/CD foundations for greater confidence in deployments. /liliCollaborate with product and backend teams to design architecture, optimize resources, and increase platform autonomy. /li /ulh3Typical Day /h3pYour focus remains on the Platform, but priorities often come from product, customers, and engineering teams. You might design autoscaling behavior, develop Kubernetes infrastructure, investigate production issues, improve dashboards, enable safer deployments, optimize ClickHouse, or help reduce platform friction for other teams. /ph3How We Work /h3pWe are a remote‑first company with occasional in‑person gatherings in Madrid to foster alignment and problem solving. We value ownership, transparency, clear communication, and documented decisions. Work is closely aligned with product, support, customer success, and other engineering teams to support the entire company. /ph3Benefits /h3ulli22 days of holiday a year plus birthday and public holidays. /liliFreedom to work from wherever suits you best. /liliUp to €2,800 for home‑workspace setup. /liliRemote‑first culture with flexibility. /li /ul /p #J-18808-Ljbffr