Data Analyst & Data Engineer — SQL · Python · Power BI. I turn raw logistics and supply-chain operations into decisions.
📍 Egypt · Open to remote & hybrid
9 yrs logistics · 549K rows → $298M revenue · 10.6M vessel tracks · 2.2M listings · 84.2% model accuracy
Data Analyst and Data Engineer with 9 years of domain expertise in international trade, logistics, and supply-chain operations before pivoting into data. I build end-to-end analytics — raw operational data → clean warehouse → dashboard — with SQL Server, Python, and Power BI. My edge: I don't just model logistics data; I know what a detention hotspot costs and which question management will ask next.
| Category | Technologies |
|---|---|
| Databases | SQL Server · T-SQL · SQLite · GeoPackage |
| Languages | Python · TypeScript · HTML · CSS |
| Data & ML | Pandas · NumPy · scikit-learn · XGBoost · Matplotlib · Seaborn · SHAP |
| BI & Viz | Power BI · Jupyter · Excel |
| Architecture | Medallion (Bronze → Silver → Gold) · Star Schema · ETL/ELT · Data Quality |
| Tools | Git · Docker · Playwright · uv |
📦 Supply Chain Delay Analysis · → Repository
15,549 orders · 152 countries · 4 shipping modes → First Class delays 98.5% of shipments.
E-commerce delivery analysis isolating which mode, region, and route drives late delivery, with verified KPIs.
Python Pandas EDA E-commerce
🚢 AIS Vessel Tracks 2025 · → Repository · → Case Study
10.6M vessel tracks across 6 US regions → 45% recreational / 26% commercial, 2.6× seasonal amplitude.
GeoPackage spatial database profiled with SQL + Python: fleet composition, traffic distribution, and vessel-type behavior.
Python SQL Geospatial Maritime AIS Matplotlib
🚚 Smart Logistics Delay Analysis · → Repository
Predictive delay model at 84.2% accuracy; traffic & waiting time identified as top drivers.
15 engineered features across XGBoost / Random Forest / Gradient Boosting, packaged with an Excel dashboard and standalone SQL script.
Python XGBoost scikit-learn Pandas SQL Server Power BI
🏥 Hospital Beds Management · → Repository
13,493 bed requests · 52 weeks → 56.6% of demand refused.
Capacity-and-demand analysis of a year of hospital operations, surfacing where refusal concentrates and why.
Python Pandas Capacity Planning EDA
🏠 USA Real Estate Analysis · → Repository
2.2M listings · 171MB CSV → $325K median price, $201 median per sq ft.
Market and data-quality analysis of a large public housing dataset — profiling, cleaning, and outlier handling at scale.
Python Pandas Big Data Data Quality EDA
📊 Telecom Churn — Data Quality Audit & Model · → Repository
140,200 rows · 5-pass audit → 11 defects found, cleaned, and documented before modeling.
A competition submission in two halves: a rigorous data-quality audit, then a 66-feature / 5-model churn pipeline with SHAP explainability — including a null-model baseline that showed the churn label carried little predictive signal.
Python scikit-learn SHAP Data Quality Churn
🏆 Logistics Intelligence Platform · → Repository · → Documentation
549,706 rows → Bronze → Silver → Gold → 11 business views → $298.6M revenue explained.
Production medallion pipeline on SQL Server 2025: 14 tables into a conformed star schema with surrogate keys and referential integrity. Every KPI cross-validated — 0 rows dropped, 0 orphaned references.
T-SQL SQL Server Medallion Star Schema ETL Data Quality
🏢 Enterprise Data Warehouse · → Repository
Full medallion warehouse integrating CRM + ERP sources, built entirely in T-SQL.
Stored procedures, ROW_NUMBER() surrogate keys, MDM business rules with CRM/ERP fallback, and data-quality tests — documented with draw.io diagrams.
T-SQL Data Warehousing Stored Procedures MDM Data Modeling
🌱 Supply Chain Carbon Emissions · → Repository
GHG emissions across 1,016 industries → Bronze → Silver → Gold → Power BI dashboard.
End-to-end pipeline (CSV → SQL Server → PyODBC EDA → Power BI) with right-skew analysis, outlier detection, and industry ranking.
Python SQL Server Medallion PyODBC Power BI
📦 Enterprise Shipping Data Simulator · → Repository
Relational shipping data generator (ports → vessels → shipments → BOL → tracking), tested to 100M+ rows.
95% clean / 5% anomalous records, weighted trade lanes, peak-season simulation — Parquet or CSV output.
Python Pandas Parquet Data Generation Multiprocessing
🔎 Egyptian Customs Tariff Scraper · → Repository
Playwright toolkit extracting HS-code classifications from Egyptian Customs.
Modular scraper with legal disclaimers, sample output, and a test suite.
Python Playwright Web Scraping Customs
Currently open to Data Analyst & Data Engineer roles — logistics and supply chain a plus.



