gavinbowden.me — home

Turbine Vane Defect Detection

Howmet Aerospace inspects thousands of cast turbine vanes by having operators read X-ray images for anomalies. Through Purdue's Data Mine, our team of six built an image pipeline and a CNN classifier meant to do that flagging automatically. I owned the data splits, the model and its tuning, and the evaluation, which makes me the one who found out the honest test numbers were nowhere near the headline accuracy.

The defect you find out about later

Howmet casts the turbine vanes that sit in the hot section of jet engines. A hidden void or a non-metallic inclusion inside one is the kind of defect you find out about the expensive way, at altitude.

So every part gets X-rayed and a person reads the image. That read is about 87% accurate, and it does not scale.

Howmet brought the problem to Purdue's Data Mine asking for a model that could do the flagging automatically. Our team of six picked up where the previous year's group left off and built both halves: the image processing pipeline, and the classifier on the end of it.

Two filtered X-ray images of turbine vanes side by side, the left one showing a mottled defect region near the top edge
Anomalous (left) and normal (right) after filtering. The defect is the mottled patch below the leading edge. That's the signal the model has to learn.

Making the defect visible first

The raw input is a DICOM X-ray of the whole part, and the defect is a low-contrast texture change inside a region that is already nearly black. Feed that to a classifier directly and it spends most of its capacity learning where the vane is.

So my teammates built a pipeline that takes a DICOM straight off the machine and hands back something a network can actually separate: crop to the vane, run a horizontal Sobel filter to pull out edge structure, invert, sharpen with an unsharp mask, then push the contrast. Every image the model ever sees goes through it.

A row of the same vane X-ray at each processing stage: original DICOM, cropped, Sobel filtered, inverted, unsharp-mask sharpened, and contrast enhanced
One image through every stage, from the original DICOM on the left to classifier input on the right.

Where my part starts

This is where my part starts. We had roughly 3,000 human-flagged images, and defective parts are (fortunately for Howmet, unfortunately for us) rare. Augmentation added about 2,000 more. I split anomalous and normal separately at 7:2:1 into train, validation, and test so the class balance held in every split, and oversampled the anomalous class in training.

Diagram of the anomalous and normal sets splitting into training, validation, and testing sets
Splitting each class separately, then oversampling the anomalous side of the training set.

Borrowing features that already work

I built the classifier with transfer learning: VGG and ResNet backbones pretrained on ImageNet, with a custom classification head, fine-tuned on vane data. With a few thousand images and two classes, training from scratch was never going to beat borrowing features that already know what edges and textures look like from millions of samples. Then I swept the hyperparameters that mattered most: convolutional layers (2 to 5), batch size (16 to 32), and training length (20 to 50 epochs). The best configurations hit around 94% validation accuracy, which is what our poster leads with.

3D scatter plot of epochs, batch size, and number of convolutional layers, colored by accuracy from 90-94%.
The hyperparameter sweep. The best runs reach ~94%. The problem is that accuracy is the wrong thing to optimize here, which the test set made very clear.

94%, and the 23 it was hiding

That 94% does not survive the test set. On 159 held-out images, the model caught 4 of the 27 defective vanes.

It passed the other 23 through as normal.

Overall test accuracy is about 71%, and the dataset is about 83% normal, which means a model can answer "normal" to everything and still score in the eighties without learning a single thing. We optimized for accuracy when the number that matters in inspection is how many defects you miss. We missed most of them.

Confusion matrix on the test set: 4 anomalous correctly flagged, 23 anomalous missed, 23 normal falsely flagged, 109 normal correct
The number that actually matters is the 23 in the top right: defective vanes the model passed as normal.

The training curves say the same thing a different way. Training accuracy gets past 97% while validation stalls around 91%, training AUC reaches almost 1.0 while validation flattens in the mid-eighties, and validation loss starts rising after about the fifth epoch while training loss keeps falling. That gap is the model memorizing a small number of defect examples instead of learning what a defect looks like.

Four plots over 25 epochs (loss, accuracy, AUC, and false negatives) each showing training and validation curves diverging.
Loss, accuracy, AUC, and false negatives per epoch. Train and validation diverge early and never reconverge.

Looking back, I also don't fully trust the 94% itself. If augmented copies of an image land in both training and validation (augment first, split second), validation is partly grading the model on pictures it has already seen with a filter on, and a 94% validation vs 71% test gap is exactly what that looks like. X-rays of the same physical part landing on both sides of the split would do the same thing. It's the first thing I'd check.

What I would do differently

We concluded the model wasn't ready for the factory floor, and I still think that was the right call. But the conclusion we wrote, that more defect images would improve performance, was only part of the story. Labeled defects were exactly what nobody had, so it was an easy thing to blame. If I did it again, I would:

  • Split by physical part before augmenting anything, and only augment the training set, so validation can't flatter the model.
  • Optimize for recall, not accuracy. On an 83/17 split, accuracy rewards laziness. A weighted loss or a tuned decision threshold would've made the sweep chase a metric we actually cared about.
  • Treat it as anomaly detection rather than binary classification. We had normal vanes in abundance, so training on those alone and flagging whatever doesn't match sidesteps the class imbalance entirely.
  • Localize, don't just classify. An operator handed a yes/no from a model that misses defects has no reason to trust it. A heatmap over the suspect region gives them something to check, which is a much easier ask.