04 / PROJECTData Engineering · Machine Learning
Flight Delay Data Pipeline
Six stages from raw rows to a served model
A six-stage pipeline that generates, cleans, models and serves flight-delay predictions. Every stage is idempotent, so a rerun converges instead of duplicating.
Architecture
Source
Generate
Produces the 100,000-row synthetic dataset from a known formula.
Generate → Clean → Feature Engineer → Load → Train → Evaluate → Flask API / Streamlit
The problem
A model is only as reproducible as the pipeline feeding it. Reruns that double-insert rows or silently change feature definitions make evaluation meaningless.
Engineering challenge
Making a multi-stage pipeline safe to rerun, so that evaluation reflects the data and features actually intended rather than accumulated duplicates.
Approach
Split the pipeline into six discrete stages with a 4-table Postgres model separating raw, cleaned and feature data. Loads are idempotent, so any stage can be re-executed without corrupting downstream state. The trained XGBoost model is served through a Flask REST API with a Streamlit dashboard for exploration.
- PostgreSQL
- A 4-table model keeps raw, cleaned and feature data separable and queryable.
- XGBoost
- Gradient boosting handles the mixed categorical and numeric feature set well.
- Flask
- Minimal surface for serving a single prediction endpoint.
- Streamlit
- Fast exploratory interface over the same model without building a frontend.
Engineering detail
- 016-stage pipeline: generate → clean → feature engineer → load → train → evaluate
- 02100,000-row dataset
- 034-table PostgreSQL model separating raw, cleaned and feature data
- 04Idempotent reruns — re-executing a stage converges rather than duplicating rows
- 05Data cleaning and feature engineering as discrete, inspectable stages
- 06XGBoost classifier, ROC-AUC 0.710 on the generated dataset
- 07Flask REST API serving predictions
- 08Streamlit dashboard for exploration
Engineering practices
- Idempotent stage reruns.
- Cleaning and feature engineering kept separate from training.
- Dataset provenance stated alongside the metric.
Features
- Six-stage reproducible pipeline.
- Four-table Postgres model.
- Flask prediction API.
- Streamlit dashboard.
Measured results
- ROC-AUC 0.710 on the synthetic evaluation set.