π MΓΌnster, Germany
π Data Engineering Weiterbildung at DCI
πΌ Open to Junior Data Engineer and Internship opportunities
I build reliable batch, streaming, and cloud data pipelines with a focus on data quality, testing, orchestration, and scalable storage.
βοΈ AWS Glue Books Data Lake
S3, AWS Glue, Glue Catalog, Athena, Terraform, and Parquet data lake pipeline.
π§± dbt Banking Analytics
dbt models, dimensional data warehouse, Athena, data tests, and Terraform.
π¬οΈ Airflow CDC Books Pipeline
Dockerized Airflow pipeline with PostgreSQL upserts and CDC reporting.
π§ Delta Lakehouse Pipeline
Spark, Delta Lake, incremental processing, merge, time travel, Databricks, and Dagster.
π¨ Kafka Streaming Orders
Kafka producer and consumer with schema validation, bounded reads, offsets, and tests.
π€ Customer Risk ML Platform
MLflow Registry, batch inference, Spark MLlib, Feast, validation, DLQ, and drift monitoring.
β
LoanScope Data Quality
Data validation, quality metrics, alert records, rejected data, and dashboard output.
π©Ί Diabetes Logistic Regression
A small machine learning project using Python, pandas, and scikit-learn.
- Python, SQL, Bash, Linux
- AWS S3, Glue, Athena, IAM, Terraform
- Apache Airflow, Dagster, Kafka
- dbt, PostgreSQL, CDC
- Spark, PySpark, Delta Lake, Databricks
- Docker, pytest, Ruff, GitHub Actions
- Machine Learning, scikit-learn, MLflow Model Registry, batch inference, Feast, drift monitoring
- Data quality, validation, testing, and observability