Skip to content
View mahdisaemi-tech's full-sized avatar

Block or report mahdisaemi-tech

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
mahdisaemi-tech/README.md

Hi, I'm Mahdi Saemi πŸ‘‹

πŸš€ Junior Data Engineer

πŸ“ MΓΌnster, Germany
πŸŽ“ Data Engineering Weiterbildung at DCI
πŸ’Ό Open to Junior Data Engineer and Internship opportunities

I build reliable batch, streaming, and cloud data pipelines with a focus on data quality, testing, orchestration, and scalable storage.

🧰 Data Engineering Toolbox

Python SQL AWS Terraform Apache Airflow Dagster Kafka Spark dbt Docker GitHub Actions

πŸ—οΈ Featured Projects

☁️ AWS Glue Books Data Lake
S3, AWS Glue, Glue Catalog, Athena, Terraform, and Parquet data lake pipeline.

🧱 dbt Banking Analytics
dbt models, dimensional data warehouse, Athena, data tests, and Terraform.

🌬️ Airflow CDC Books Pipeline
Dockerized Airflow pipeline with PostgreSQL upserts and CDC reporting.

🧊 Delta Lakehouse Pipeline
Spark, Delta Lake, incremental processing, merge, time travel, Databricks, and Dagster.

πŸ“¨ Kafka Streaming Orders
Kafka producer and consumer with schema validation, bounded reads, offsets, and tests.

πŸ€– Customer Risk ML Platform
MLflow Registry, batch inference, Spark MLlib, Feast, validation, DLQ, and drift monitoring.

βœ… LoanScope Data Quality
Data validation, quality metrics, alert records, rejected data, and dashboard output.

πŸ“Œ Additional Project

🩺 Diabetes Logistic Regression
A small machine learning project using Python, pandas, and scikit-learn.

πŸ” Core Skills

  • Python, SQL, Bash, Linux
  • AWS S3, Glue, Athena, IAM, Terraform
  • Apache Airflow, Dagster, Kafka
  • dbt, PostgreSQL, CDC
  • Spark, PySpark, Delta Lake, Databricks
  • Docker, pytest, Ruff, GitHub Actions
  • Machine Learning, scikit-learn, MLflow Model Registry, batch inference, Feast, drift monitoring
  • Data quality, validation, testing, and observability

Pinned Loading

  1. airflow-cdc-books-pipeline airflow-cdc-books-pipeline Public

    Dockerized Apache Airflow pipeline for book ingestion, PostgreSQL upserts, CDC tracking, and reporting.

    Python

  2. aws-glue-books-data-lake aws-glue-books-data-lake Public

    AWS Glue data lake pipeline with S3, Glue Catalog, Athena, Terraform, and Parquet transformations.

    HCL

  3. dbt-banking-analytics dbt-banking-analytics Public

    dbt banking analytics project with Amazon Athena, Glue Catalog, dimensional models, tests, and Terraform.

    HCL

  4. delta-lakehouse-pipeline delta-lakehouse-pipeline Public

    Spark and Delta Lake pipeline demonstrating merge, idempotent incremental processing, time travel, and Databricks validation.

    Jupyter Notebook

  5. kafka-streaming-orders kafka-streaming-orders Public

    Kafka order streaming project with schema validation, bounded batch consumption, manual offsets, and Python tests.

    Python

  6. customer-risk-ml-platform customer-risk-ml-platform Public

    Tested customer-risk ML platform with MLflow Registry, batch inference, Spark MLlib, Feast materialization, event validation, DLQ routing, PSI drift monitoring, and GitHub Actions scheduling.

    Python