The Role:
Builds performant data processing pipelines, manages relational databases, and constructs cloud-portable data lake environments using modern open storage formats and embedded analytical
engines.
- Responsibilities:
Pipeline Design: Construct resilient data processing pipelines in Python using Medallion
architecture principles (Bronze/Silver/Gold) to transform raw data into analytics-ready assets.
- Open Storage Management: Standardize data storage using open formats including Parquet,
Apache Iceberg, and Delta Lake to ensure cross-cloud compatibility.
- Embedded Analytics: Leverage DuckDB integrated with Python for fast, lightweight, and
performant localized data processing within ETL workflows.
- Database Operations: Architect and optimize production PostgreSQL schemas while utilizing
SQLite for embedded or lightweight localized storage needs.
- Distributed Compute:
Scale heavy ETL, data transformation, and analytics workloads using
Databricks or Apache Spark
Requirements:
- Advanced SQL proficiency with deep PostgreSQL expertise and working knowledge of SQLite.
- Hands-on experience building robust analytical pipelines using Python and DuckDB.
- Mastery of open table and storage formats (Parquet, Apache Iceberg, Delta Lake).
- Proven track record designing data lakes using the Medallion data processing architecture.
- Familiarity with the Databricks / PySpark ecosystem for distributed processing.