Builds performant data processing pipelines, manages relational databases, and constructs cloud-portable data lake environments using modern open storage formats and embedded analytical engines.
Responsibilities:
Pipeline Design: Construct resilient data processing pipelines in Python using Medallion architecture principles (Bronze/Silver/Gold) to transform raw data into analytics-ready assets.
Open Storage Management: Standardize data storage using open formats including Parquet, Apache Iceberg, and Delta Lake to ensure cross-cloud compatibility.
Embedded Analytics: Leverage DuckDB integrated with Python for fast, lightweight, and performant localized data processing within ETL workflows.
Database Operations: Architect and optimize production PostgreSQL schemas while utilizing SQLite for embedded or lightweight localized storage needs.
Distributed Compute: Scale heavy ETL,
data transformation, and analytics workloads using Databricks or Apache Spark
Requirements:
Advanced SQL proficiency with deep PostgreSQL expertise and working knowledge of SQLite.
Hands-on experience building robust analytical pipelines using Python and DuckDB.
Mastery of open table and storage formats (Parquet, Apache Iceberg, Delta Lake).
Proven track record designing data lakes using the Medallion data processing architecture.
Familiarity with the Databricks / PySpark ecosystem for distributed processing.