gavinbowden.me — home

Cooling Tower Predictive Maintenance

A machine learning approach to predicting remaining useful life (RUL) for NASA Langley cooling towers, joining eight years of vibration sensor data, maintenance work orders, and local weather to forecast how many days remain until the next unplanned repair. The model never beat the naive benchmark, and working out why was the most useful result.

How long until it breaks?

Fixing a cooling tower after it fails costs far more than servicing it beforehand, and that is before anyone counts the downtime. Which makes "is this unit healthy?" the wrong question to ask it. The answer is almost always yes, right up until the morning it isn't.

The useful question is how long it stays that way. I framed that as a regression problem: predict Remaining Useful Life, the number of days until the next reactive work order, for a Langley cooling tower and all of its associated pumps.

Three sources, one timeline

Three sources had to be lined up on one timeline before any modeling could happen.

  • AVEVA PI sensor data: hourly vibration readings (PeakVue and overall) from the gearboxes, bearings, and pumps across the cooling towers, from 2017 to 2025.
  • Maximo work orders: timestamped maintenance records from 2017 to 2022. Two work types, TC and REPR, mark reactive work (something broke or alarmed), and those events are the prediction target. Since the work orders stop in 2022 while the sensors keep going, only 2017 through 2022 has labels to learn from.
  • Langley AFB METAR weather: hourly local conditions from the airfield next door. The theory was that weather is what wears a cooling tower down unexpectedly.

Predicting forward, never backward

I used XGBoost regression, with the evaluation designed around this being time series data. Folds were split chronologically, with the newest year held out, rather than shuffled, so the model is always predicting forward.

A random split would let it learn from the future. That isn't learning, it's memorizing.

  • Three-fold cross-validation, ordered oldest to newest, with a final holdout year.
  • Compared grid search, random search, and Bayesian optimization for hyperparameter tuning. Bayesian won on cost, getting close to the best configuration in about 60 trials where the full grid would've taken tens of thousands.
  • Scored against a naive benchmark rather than against zero, so any improvement had to be real.

It never beat a naive guess

The naive benchmark averaged an error of 39.43 across folds (27.81, 29.04, 61.43). My first model came in around 41. Tuning never got it under the benchmark.

The model never learned anything a naive guess didn't already know.

That is a real result, not just a failed one. It says the vibration, maintenance, and weather features, as I built them, don't carry enough signal about when the next reactive repair is coming.

The third fold is where I would start over. Its error is double the other two, so whatever changed in those later years (new equipment, different maintenance habits, or just a bad year) is something the features don't capture at all. That's where the future work is.

Presenting it to thirty engineers

I presented this work to around 30 NASA engineers and staff at a Jam Session (usually led by my mentor, Charles Liles), walking through cross-validation, leakage, and hyperparameter tuning for time series maintenance data. The goal was partly to share the method and partly to recruit. Attendees with domain knowledge, like the maintenance folks, were invited to help refine the model through a shared Google Cloud Jupyter notebook, which started an ongoing collaboration.