← All projects
Applied ML · Research

BiLSTM Air Quality Forecasting

An end-to-end ML pipeline using Bidirectional LSTM networks to forecast multi-pollutant air quality from a 5-year NASA dataset for Baghdad, and to compute US-EPA Air Quality Index metrics — with seasonal feature engineering, uncertainty quantification, and a standalone interactive dashboard.

BiLSTM air quality dashboard overview with R² and MAPE charts per pollutant

Project Overview

An applied deep-learning research project: a Bidirectional LSTM (BiLSTM) network trained on a 5-year NASA air-quality dataset for Baghdad (2020–2025), forecasting five pollutants — CO, SO₂, SO₄, PM2.5, and NO₂ — and mapping predictions to US-EPA Air Quality Index health categories. Results are served through a standalone, single-file interactive dashboard rather than a static notebook.

Problem

Air-quality pollutant levels are noisy, seasonal, and interdependent, which makes reliable short- and medium-term forecasting genuinely difficult — small time series make standard train/test splits unreliable, outlier spikes distort ordinary loss functions, and different pollutants respond very differently to the same modeling approach.

Solution

A BiLSTM sequence model per pollutant, trained with Huber loss for robustness to concentration outliers, circular harmonic (sine/cosine) month encoding to capture seasonality, and walk-forward cross-validation appropriate for a small time series. 95% confidence bounds come from Monte Carlo Dropout. Every result is backed by a full statistical battery — MSE, RMSE, MAE, MAPE, R², Durbin-Watson, Pearson and Spearman correlation — plus ADF/KPSS stationarity testing on the underlying series before modeling.

Key Features

  • Overview: at-a-glance best-performing pollutant, MAPE, Pearson r, and training configuration (60 epochs, early stopping).
  • Pollutant Predictions: observed-vs-predicted charts on the held-out test set, per pollutant.
  • AQI Results: predicted Air Quality Index mapped to US-EPA health categories (Good / Moderate / Unhealthy for Sensitive Groups / Unhealthy), with observed-vs-predicted comparison.
  • 12-Month Forecast: forward-looking pollutant forecasts with Monte-Carlo-Dropout uncertainty bounds and a month-by-month value table.
  • Full Metrics Table: the complete evaluation matrix for all five pollutants side by side.
  • Stationarity Tests: ADF and KPSS test results per pollutant, checked before modeling.
NO2 observed vs predicted chart with R², MAPE, MAE, Pearson r, and DW stat
NO₂ observed vs. predicted — test set, Apr 2024–Jan 2025
12-month NO2 forecast chart, Feb 2025 to Jan 2026
12-month forecast — NO₂, Feb 2025–Jan 2026
NO2 AQI observed vs predicted with health category bands
AQI classification — observed vs. predicted, with health-category bands
ADF and KPSS stationarity test results per pollutant
Stationarity testing (ADF + KPSS) per pollutant

Architecture

A two-script pipeline, run in sequence: airquality_lstm_publication.py handles data loading, stationarity testing, seasonal feature engineering, per-pollutant BiLSTM training/evaluation, and forecasting — writing results out as CSVs and figures. generate_dashboard.py then compiles those outputs into a single standalone HTML5 file with Chart.js, so the dashboard has no server dependency at all — it's a static artifact that happens to be interactive.

Technology Stack

Python TensorFlow (BiLSTM) scikit-learn statsmodels (ADF / KPSS) pandas / NumPy Standalone HTML5 + Chart.js

Engineering Highlights

  • Huber loss over MSE: chosen specifically for robustness to concentration outliers in pollutant readings, rather than defaulting to a squared-error loss that outliers would dominate.
  • Seasonal encoding done right: month is encoded as circular harmonic sine/cosine features rather than a raw integer, so the model doesn't see December and January as maximally far apart.
  • Walk-forward cross-validation: used instead of a single random train/test split, which matters for a genuinely small time series where a lucky or unlucky split can badly mislead the evaluation.
  • Uncertainty quantification: 95% confidence bounds on the 12-month forecast come from Monte Carlo Dropout, so the forecast isn't presented as a single false-precision line.
  • Rigorous, honest evaluation: the dashboard reports full metrics for every pollutant — including the ones where the model underperforms — rather than only showcasing the best result.
  • Standard-compliant AQI: AQI is computed per US-EPA methodology for NO₂ and CO, with an explicit unit-conversion note in the code for the remaining species (the raw data ships as mol/mol mixing ratios, not the μg/m³ or ppm the standard expects).
  • Zero-dependency delivery: the dashboard is a single generated HTML file — no backend, no build step — which is also why it deploys as a static Render site with no server-side moving parts to keep running.

Results

Model performance varies meaningfully by pollutant. NO₂ forecasting performed strongly: R² = 0.9515, MAPE = 1.70%, Pearson r = 0.977 on the test set, with the derived NO₂ AQI achieving 0.7% MAPE and correct health-category classification across all 10 test months. PM2.5 showed moderate fit (R² = 0.561), while CO, SO₂, and SO₄ were harder to model reliably in this run (negative R² on the test set — reported transparently rather than omitted).

0.9515
NO₂ R² (test set)
1.70%
NO₂ MAPE
0.977
NO₂ Pearson r
0.561
PM2.5 R²
Reported honestly: CO, SO₂, and SO₄ forecasts underperformed in this run (negative R² on the test set) despite fitting the same architecture — a reminder that not every pollutant responds equally well to the same model, and a clear direction for further work.

Links