Summary
Data Engineer with 4+ years of experience in building scalable ETL pipelines,
data quality frameworks,
and modern lakehouse/DWH architectures. Proficient in Python, SQL, and dbt, with hands-on experience in
Apache Airflow,
Spark, ClickHouse, and DevOps practices (CI/CD, Kubernetes). Experienced in dimensional modeling
and designing reliable data architectures.
Work Experience
- Contributing to a config-driven ETL engine for large-scale ingestion of international
statistics (SDMX, REST APIs) into S3: plugin architecture, incremental stateful collection.
- Designed a lakehouse DWH MVP with a unified dimensional model (dbt, DuckDB/ClickHouse),
harmonizing data from multiple international statistical sources.
- Maintaining Airflow orchestration and introducing data quality practices.
- Joined a newly formed Data Quality team and architected a configuration-driven ecosystem to
ensure platform-wide data integrity:
- Developed a Python/SQL framework for business logic validation (replacing Great
Expectations).
- Built a scalable Spark-based reconciliation service for detecting discrepancies between
event streams and S3 storage.
- Engineered a dynamic DAG generator in Airflow: SQL/YAML configs are parsed into
ClickHouse and auto-converted into scheduled pipelines, scaling to
2,000+ active data quality checks across the platform.
- Managed OpenMetadata infrastructure on Kubernetes (ArgoCD, custom Helm charts) and established
CI/CD pipelines (GitLab CI, Tox, Pre-commit) with database version control via Liquibase.
- Pioneered GenAI initiatives by developing the first AI agent prototype using LangGraph.
- Developed multi-gigabyte data marts for a team of 4 Data Scientists using a Spark-based internal
ETL platform
within the Hadoop ecosystem, accelerating the training and feature engineering for ~10 ML
models.
- Extended ETL capabilities by writing custom UDFs in Scala and optimizing complex SQL logic for
high-load processing.
- Maintained the backend of a real-time scoring service (PostgreSQL) and orchestrated regular data
workflows using Apache Airflow.
- Led the data migration of 4 regional services to the SMEV4 standard, managing metadata
registries and ensuring continuous data integrity.
- Designed an end-to-end reporting system: formed API requirements, built Python/SQL ETL
pipelines on a self-hosted Apache Airflow instance, and created an inter-departmental
dashboard to monitor and track the percentage of overdue citizen applications.
Skills
Programming: Python, SQL.
Big Data & ETL: Apache Airflow, Apache Spark,
dbt, ClickHouse, Hadoop (HDFS), Pandas.
DevOps & Infrastructure: Docker, Kubernetes,
ArgoCD, Helm, GitLab CI/CD, Linux, Elasticsearch.