Why Doesn’t My Model Work?

An article in The Gradient examines why machine learning models that appear to perform well during development can fail when applied to real-world data, describing common pitfalls and possible prevention methods.
The author, who says they have worked in machine learning for about 20 years, wrote a guide titled "How to avoid machine learning pitfalls: a guide for academic researchers." The article points to the recently introduced REFORMS checklist for ML-based science as one prevention method.
Reported examples of failures include hundreds of models developed during the Covid pandemic that "simply don't work," and a water quality system deployed in Toronto that regularly told people it was safe to bathe in dangerous water, according to the article. Many such cases are documented in the AIAAIC repository, it states. The article also says it has been suggested these missteps are causing a reproducibility crisis in science and a lack of trust in published scientific results.
On data, the article cites a review of Covid prediction models by Roberts et al. Public datasets were later found to contain misleading signals, including overlapping records, mislabellings and hidden variables. In many Covid chest imaging datasets, body orientation served as a hidden variable: sick people were more likely to have been scanned lying down, while those standing tended to be healthy. As a result, many models became good at predicting posture but bad at predicting Covid, it says.
On leakage, the article describes pre-term birth prediction papers that applied data augmentation before splitting off the test data. When reviewers corrected this, performance dropped from near perfect to not much better than random, the article states. It also warns that repeatedly evaluating a model on the same test set, and using the same community benchmarks such as MNIST, CIFAR and ImageNet, can lead to overfitting those benchmarks.
On metrics, the article gives the example of accuracy with imbalanced data: a model that always predicts one label gets 50% accuracy if half of test samples carry that label and 90% if 90% do. It notes that a review of time series forecasting pitfalls found an autoformer can be beaten by a trivial model predicting no change at the next time step.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Have you ever trained a model you thought was good, but then it failed miserably when applied to real world data? If so, you’re in good company.