How a machine learning fuel a reproducibility crisis in science?

From biomedicine to political science, researchers are increasingly using machine learning as a tool for making model-based predictions in their data. But the claims in many of these studies are likely exaggerated, according to a pair of researchers from Princeton University in New Jersey. They want to sound a warning about what they call a "looming reproducibility crisis" in machine learning science.Machine learning is being sold as a tool that researchers can pick up in hours and use on their own, and many are heeding this advice, says Sayash Kapoor, a machine learning researcher at Princeton. "But you wouldn't expect a chemist to be able to learn how to run a lab using an online course," he says. And few scientists realize that the problems they encounter when applying artificial intelligence (AI) algorithms are common to other fields, says Kapoor, co-author of a preprint on the "crisis" 1. Peer reviewers don't have time to analyze these models, so the academy currently doesn't have mechanisms to eliminate irreproducible articles, he says. Kapoor and his co-author Arvind Narayanan created guidelines for scientists to avoid such pitfalls, including an explicit checklist to submit with each article. What is reproducibility? The Kapoor and Narayanan definition of reproducibility is broad. He says other teams should be able to replicate the results of a model, given all the details of the data, code and conditions, often called computational reproducibility, something that is already a concern for machine learning scientists. The pair also define an irreproducible model when researchers make mistakes in analyzing the data that mean the model isn't as predictive as it claims.However, the team's battle cry has struck a chord. More than 1,200 people signed up for what was initially a small online workshop on reproducibility on July 28, organized by Kapoor and his colleagues, designed to find and disseminate solutions. "Unless we do something like this, every field will continue to run into these problems over and over again," he says. Over-optimism about the powers of machine learning models could prove detrimental when applying algorithms to areas such as health and justice, says Momin Malik, a data scientist at the Mayo Clinic in Rochester, Minnesota, who will speak at the seminar. . Unless the crisis is addressed, the reputation of machine learning could suffer, he says. “I am somewhat surprised that there has not been a collapse in the legitimacy of machine learning. But I think it could come very soon. ”Machine learning problems Kapoor and Narayanan say there are similar pitfalls in applying machine learning to multiple sciences. The pair analyzed 20 reviews across 17 research fields and counted 329 Research papers whose findings could not be fully replicated due to problems in the way machine learning was applied.1 Narayanan himself is not immune: a 2015 cybersecurity paper he is a co-author of3 is among 329. " It really is a problem that the entire community must tackle collectively, "says Kapoor. Failures are not the fault of any single researcher, he adds. Instead, the fault lies with a combination of hype about AI and inadequate checks and balances.Judging such errors is subjective and often requires deep knowledge of the field in which machine learning is being applied. Some researchers whose work has been critiqued by the team disagree that their papers are flawed, or say Kapoor’s claims are too strong. In social studies, for example, researchers have developed machine-learning models that aim to predict when a country is likely to slide into civil war. Kapoor and Narayanan claim that, once errors are corrected, these models perform no better than standard statistical techniques. But David Muchlinski, a political scientist at the Georgia Institute of Technology in Atlanta, whose paper2 was examined by the pair, says that the field of conflict prediction has been unfairly maligned and that follow-up studies back up his work. The problem most important highlighted by Kapoor and Narayanan is "data leakage", when the information in the dataset that a model becomes aware of includes data that they are then evaluated. If these aren't completely separate, the model has already actually seen the answers and its predictions look much better than they actually are. The team identified eight main types of data breaches that researchers can be vigilant against.Broader issues include training models on datasets that are more limited than the population they ultimately intend to reflect, Malik says. For example, an AI that detects pneumonia on chest X-rays that has only been trained on older people might be less accurate on younger people. Another problem is that algorithms often end up relying on shortcuts that don't always work, says Jessica Hullman, a computer scientist at Northwestern University in Evanston, Illinois, who will speak at the seminar. For example, a computer vision algorithm could learn to recognize a cow from the grassy background in most cow images, then fail when it encounters the image of the animal on a mountain or beach. The high accuracy of predictions in tests often leads people to think that the models capture the "true structure of the problem" in a human-like way, he says. The situation is similar to the replication crisis in psychology, where people rely too much on statistical methods, he adds. The hype about machine learning capabilities made its findings too easy for researchers to accept, says Kapoor. The word "prediction" itself is problematic, Malik says, as most predictions are, in effect, backtests and have nothing to do with forecasting the future.

Enjoyed this article? Stay informed by joining our newsletter!

Comments

You must be logged in to post a comment.

About Author