Database Systems Engineer Madrid-Hybrid Engineering
il y a 5 jours
Madrid
ppWe’re looking for a Database Systems Engineer who thrives at the frontier of distributed databases, storage engines, transactional systems, and analytical query execution. /p h3Database Systems Engineer — Distributed Storage Transactional Query Engines /h3 h3Role Summary /h3 pFastS3 is building AI-native data infrastructure from the ground up. We are looking for an engineer to help design the database layer that sits above PixelDB, our distributed key/value object storage engine, and works with TEMPANO, our custom Iceberg catalog and compaction API. /p pThis role is not about building a conventional application database, a Postgres extension, or a simple lakehouse connector. It is about helping design the foundation for a system that can support transactional, analytical, and vector workloads over the same live data layer. /p pYou’ll work closely with engineering leadership to prototype and build the core components of a distributed table engine: commit logs, MVCC, stable Row IDs, secondary indexes, snapshot publication, compaction, query execution, and integration with Iceberg-style table metadata. /p h3What You’ll Be Doing /h3 ul liDesign and prototype a transactional table engine on top of PixelDB. /li liWork on MVCC, commit logs, transaction visibility, snapshot isolation, and serializable consistency. /li liDesign stable Row ID abstractions so indexes point to logical rows rather than physical storage locations. /li liBuild metadata and manifest flows that connect live transactional writes to TEMPANO and Iceberg-compatible analytical snapshots. /li liHelp design compaction, version cleanup, and garbage collection that do not block long-running analytical queries. /li liExplore query execution paths for point reads, range scans, analytical scans, joins, and vector retrieval. /li liWork on distributed storage layout, partitioning, object placement, erasure‑coded data, and recovery semantics. /li liCollaborate with systems and AI engineers to support agentic workloads that require transactional, analytical, and semantic access to live data. /li liBenchmark latency, throughput, consistency behavior, write amplification, and query performance under concurrent workloads. /li /ul h3What We Need to See /h3 ul liStrong systems programming experience in Rust, C++, Go, or C. /li liPractical knowledge of database internals, storage engines, distributed databases, or query engines. /li liExperience with MVCC, WAL/commit logs, transaction processing, indexing, compaction, or recovery. /li liUnderstanding of distributed systems concepts such as consensus, replication, failure recovery, consistency, and concurrency control. /li liFamiliarity with analytical formats or engines such as Apache Iceberg, Delta Lake, Parquet, Arrow, Trino, Spark, DuckDB, or ClickHouse. /li liAbility to reason deeply about tradeoffs between OLTP, OLAP, and vector workloads. /li liComfort moving from architectural design to prototype code, benchmarks, and production‑quality implementation. /li liStrong debugging, profiling, and performance‑analysis skills. /li /ul h3Ways to Stand Out /h3 ul liExperience building or contributing to database kernels, storage engines, query planners, transaction managers, or distributed SQL systems. /li liHands‑on work with PostgreSQL internals, CockroachDB, YugabyteDB, TiDB, FoundationDB, ScyllaDB, Cassandra, ClickHouse, DuckDB, Neon, or similar systems. /li liExperience with Apache Iceberg, Delta Lake, table catalogs, manifest generation, compaction, or lakehouse metadata. /li liKnowledge of vector indexes such as HNSW, IVF, PQ, ANN search, or hybrid vector/SQL execution. /li liExperience designing systems around stable logical IDs, append‑only storage, object storage, LSM trees, B‑trees, or log‑structured architectures. /li liBackground in high‑performance networking, distributed storage, RDMA, erasure coding, or object storage. /li liOpen‑source contributions, papers, or deep technical writing related to databases, storage engines, or distributed systems. /li /ul h3Why Join Us? /h3 pFastS3 is building infrastructure for the next generation of AI‑native data systems. Our vision is to move beyond fragmented pipelines, replicated warehouses, and bolt‑on vector extensions by designing a data layer where transactional, analytical, and semantic workloads can operate over the same live data. /p pYou’ll join a team working at the intersection of distributed storage, database internals, lakehouse architecture, vector retrieval, and AI infrastructure. If you’re excited by the idea of building a new database architecture instead of extending yesterday’s systems, this is the role. /p pWe’re growing our team in Madrid and looking for engineers who want to build core infrastructure from first principles. /p pFastS3 is proud to be an inclusive, equal opportunity employer committed to diversity, equity, and accessibility for all. /p /p #J-18808-Ljbffr