04 / PROJECTData Engineering · Machine Learning

Flight Delay Data Pipeline

Six stages from raw rows to a served model

A six-stage pipeline that generates, cleans, models and serves flight-delay predictions. Every stage is idempotent, so a rerun converges instead of duplicating.

Year
2026
Status
Live
Technology
Python · Pandas · PostgreSQL · Scikit-Learn · XGBoost · Flask · Streamlit
Data note
Built on a synthetic 100,000-row dataset whose delay labels come from a known generator formula, so the reported score describes model fit on generated data — not real-world operational performance.
01 / SYSTEM MAP

Architecture

Data Engineering · Machine Learning / flow6 stages

Source

Generate

Produces the 100,000-row synthetic dataset from a known formula.

Generate → Clean → Feature Engineer → Load → Train → Evaluate → Flask API / Streamlit

02 / CONTEXT

The problem

A model is only as reproducible as the pipeline feeding it. Reruns that double-insert rows or silently change feature definitions make evaluation meaningless.

Engineering challenge

Making a multi-stage pipeline safe to rerun, so that evaluation reflects the data and features actually intended rather than accumulated duplicates.

03 / DESIGN DECISION

Approach

Split the pipeline into six discrete stages with a 4-table Postgres model separating raw, cleaned and feature data. Loads are idempotent, so any stage can be re-executed without corrupting downstream state. The trained XGBoost model is served through a Flask REST API with a Streamlit dashboard for exploration.

PostgreSQL
A 4-table model keeps raw, cleaned and feature data separable and queryable.
XGBoost
Gradient boosting handles the mixed categorical and numeric feature set well.
Flask
Minimal surface for serving a single prediction endpoint.
Streamlit
Fast exploratory interface over the same model without building a frontend.
04 / IMPLEMENTATION

Engineering detail

  1. 016-stage pipeline: generate → clean → feature engineer → load → train → evaluate
  2. 02100,000-row dataset
  3. 034-table PostgreSQL model separating raw, cleaned and feature data
  4. 04Idempotent reruns — re-executing a stage converges rather than duplicating rows
  5. 05Data cleaning and feature engineering as discrete, inspectable stages
  6. 06XGBoost classifier, ROC-AUC 0.710 on the generated dataset
  7. 07Flask REST API serving predictions
  8. 08Streamlit dashboard for exploration
05 / OPERATING PRINCIPLES

Engineering practices

  • Idempotent stage reruns.
  • Cleaning and feature engineering kept separate from training.
  • Dataset provenance stated alongside the metric.
06 / CAPABILITIES

Features

  • Six-stage reproducible pipeline.
  • Four-table Postgres model.
  • Flask prediction API.
  • Streamlit dashboard.
07 / EVIDENCE

Measured results

  • ROC-AUC 0.710 on the synthetic evaluation set.