As artificial intelligence improves by leaps and bounds, basic data science access has become increasingly democratized. Traditional entry barriers to the field, such as a lack of data and computing power, have been swept aside by a steady stream of new data startups (some offering access for as little as a cup of coffee per day) and all-powerful cloud computing, which eliminates the need for expensive onsite hardware. The ability and know-how to implement, which has arguably been overlooked, completes the triad of prerequisites. Becoming the most pervasive part of data science, It is not difficult to locate online courses with taglines such as "build X model in seconds" or "apply Z procedure to your data in just a few lines of code." In today's digital environment, quick pleasure has become the norm. While increased accessibility is not inherently harmful, beneath the brilliant array of software libraries and gleaming new models, the primary objective of data science has become buried, if not lost. It is not to run sophisticated models just for the purpose of running them, or to optimize some arbitrary performance metric, but to use them as a tool to address real-world problems.
The Iris data collection is a simple yet accessible example. How many people have used it to show an algorithm without considering what a sepal is, let alone why we measure its length? While they may appear to be inconsequential considerations for the aspiring practitioner who is more interested in adding a new model to their repertory, they were anything but for Edgar Anderson, a botanist who cataloged the properties in question in order to comprehend differences in Iris flowers. Despite the fact that this is a fabricated example, it illustrates a basic point: the mainstream has become increasingly focused on "performing" data science than "using" data science. However, this misalignment is merely a symptom of the data scientist's demise. To comprehend the source of the problem, we must take a step back and see from above.
Data science has the unusual distinction of being one of the few branches of study in which the practitioner lacks a domain. Pharmacy students go on to become pharmacists, law students go on to become attorneys, and accounting students go on to become accountants. So, data science students must become data scientists? But which data scientists? The widespread use of data science appears to be a double-edged sword. On the one hand, it is a strong tool set that can be applied to any industry that generates and captures data. On the other hand, due to the general applicability of these technologies, the user will rarely have actual domain knowledge of those sectors prior to the fact. Nonetheless, the issue was minor with the emergence of data science, as companies tried to embrace this emerging technology without completely comprehending what it was and how it might be fully incorporated into their organization.
However, nearly a decade later, both firms and the environments in which they operate have changed. They are currently pursuing data science maturity with large, entrenched teams that are benchmarked against established industry standards. Problem solvers and critical thinkers who understand the business, the appropriate industry, and its stakeholders are in high demand. The capacity to traverse a few software packages or regurgitate a few lines of code will no longer suffice, nor will the ability to code characterize a data science practitioner. The growing popularity of no-code, auto ML systems such as Data Robot, Rapid Miner, and Alteryx demonstrates this.
In 10 years (give or take), data scientists will be extinct, or at least the role title will be. In the future, the skill set known as data science will be carried by a new breed of data-aware business specialists and subject-matter experts who can imbue analyses with their deep domain knowledge, regardless of whether they can code. Their titles, whether compliance specialists, product managers, or investment analysts, will reflect their competence rather than the means by which they display it. We don't have to go far back in history to find historical precedents. Data entry specialists were highly sought after before the arrival of the spreadsheet, but as Cole Nussbaumer Knaflic (author of "Storytelling With Data") rightly observes, fluency with the Microsoft Office suite is now the basic minimum. Previously, the ability to touch-type on a typewriter was considered a specialist skill; now, with the advent of personal computing, it has also become an assumed capability.
Finally, for individuals pursuing a job in data science or beginning their studies, it may be beneficial to always refer back to the Venn diagram that you will certainly encounter. It defines data science as the intersection of statistics, programming, and domain expertise. Despite comprising an equal portion of the intersecting area, some may have a larger weighting than others.
Disclaimer: These are my personal opinions based on my observations and experiences. It's fine if you don't agree; the constructive debate is encouraged.
You must be logged in to post a comment.