4+ years of professional software engineering experience building and operating production systems.
Engineering background in Computer Science, Mathematics, Software Engineering, Physics, or a similar field, or equivalent practical experience.
Strong coding skills with demonstrated proficiency in one or more languages such as Java, Rust, Scala, or C++.
Experience developing database engines, distributed data processing systems, storage systems, or comparable infrastructure, with depth in areas such as query planning, execution, or performance optimization.
Strong foundations in algorithms, data structures, and concurrency, with experience diagnosing correctness and performance problems in complex systems.
Strong written and verbal communication skills and the ability to work effectively across teams, incorporate feedback, and hold a high bar for quality.
What We Value
Ownership mindset and a high bar for correctness. Our systems support decisions and operations that customers depend on.
Curiosity about how data systems work, from query optimizers and execution operators to table formats and distributed processing.
Strong debugging skills and motivation to follow a problem across languages, services, and layers of the stack.
A practical approach to performance, grounded in profiling, representative workloads, and measurable improvements.
Interest in applying deep systems engineering to real-world problems, with empathy for the people who use and depend on our software.
Experience building or extending systems such as Spark, DataFusion, Iceberg, or comparable technologies, and an interest in learning across the stack.
Ability to collaborate across teams and work effectively with the open-source projects we build on. Experience contributing to open-source projects is valued, but not required.
Core Responsibilities
Designing and implementing query planning and optimization capabilities that turn complex computations into efficient execution plans
Extending execution engines with new capabilities and improving query operators, parallelism, memory management, and data movement
Developing Foundry’s Iceberg catalog and engine integrations, including table metadata, transactions, and efficient reads and writes
Building shared transformation semantics and execution interfaces for workloads across Foundry’s products and platform services
Improving incremental processing so pipelines can reuse previous results and process new data efficiently while preserving correctness
Evaluating and integrating advances in open-source data systems, validating their behavior and performance against real-world workloads
Investigating correctness and performance issues across planning, execution, and storage, and building tests and benchmarks that prevent regressions
Working with product teams and customers to translate operational needs into engine capabilities that integrate with Foundry’s security, data management, and build infrastructure
Technologies We Use
Java, Scala, Rust, and Python
Apache Spark, Apache DataFusion, Apache Comet, and Velox for data processing and query execution
Apache Iceberg for table management and catalog interoperability
Apache Arrow and Apache Parquet for in-memory data processing and columnar storage
Industry-standard build tooling, including Gradle, Cargo, and GitHub.