As I tour campuses, the biggest concern I hear about our profession is the onslaught of big data. Many statistics departments are concerned that students, grants, and allocations are going to the quickly growing data analytics curricula. We also hear other terms used in place of big data such as data analytics, data mining (perhaps so yesterday), machine learning, data science, knowledge discovery, business intelligence, informatics, and even money science.

Typically, big data is identified by the three Vs of velocity, volume, and variety. I have been informed there are now up to 42 adjectives starting with V that describe it. I’m not sure I could think of 42 words starting with V. Apparently, many people do not like to use the term “big data,” perhaps because it might be demeaning to work with the opposite. But I am also unsure what the opposite might be. What happens if you have huge volumes of data with lots of variety, but they don’t necessarily come in with rapid speed? Does two out of three still let you join the big data club?
Now, I don’t think statistics is going out of business. My conviction is because I am not as concerned about the velocity of data coming in as I am about the velocity and veracity of analyses going out. People, businesses, industry, and government all want rapid analysis of the data. This is the calling card of the statistics profession, and we should continue to be the guardian of it.
With the demand for instant analysis, it is tempting for nonstatisticians to run with the numbers, use one of many computational tools, and make their own conclusions. Most of the analyses using big data are basically correlation analyses. Our profession is certainly aware correlation does not imply causation, yet I watch how conclusion after conclusion is made from a finding of correlation. Worse yet, merging all types of data can lead to Simpson Paradox situations, which, in turn, lead to incorrect results. So what is a statistician to do?
We must not be complacent. We know proper analyses depend on appropriate viewing of suitable data. We can deliver the required and correct analyses and conclusions. We are being called upon to do this more and more quickly. To do this, we have to learn some of the techniques used to analyze big data such as Spark, Hadoop, and Tensor Flow.
I will readily admit I do not know these techniques, and they may well be out of date by the time I type the end of this sentence. But the point is we have to at least learn some of the new methods, along with their pros and cons. Familiarity with the new procedures will enable statisticians to effectively collaborate with teams studying a problem and successfully communicate the conclusions, as well as the strengths and weaknesses of the methodology. Remember, there is not much value in getting the wrong answer quickly. So quality must be our trump card.
One thing we cannot do is sit on the sidelines and complain about what the others are doing, even if they are doing it improperly. For years, we have begged the research community to consult with us early so we can work with them to design samples. Now is the time for us to invite ourselves in early to participate in the action, armed with newly understood tools, and become truly important members of the team.
This is well in line with the summary of the June 2016 workshop of the National Academy of Sciences. The workshop, titled “Refining the Concept of Scientific Inference When Working with Big Data,” noted “decision-relevant knowledge can be derived from big data, but it hinges on making reliable inference. Further, the convenience or opportunistic method of sampling can lead to problems related to lack of controls, unidentified bias, missing or irregular data, and plenty of confounding factors.”
In October of 2015, the ASA issued a statement, “The Role of Statistics in Data Science,” which stressed collaboration in virtually every paragraph. Similarly, during JSM 2015, DJ Patil—the chief data scientist of the Obama administration—also stressed the importance of collaboration, but urged the statistician to get a seat at the table early in the problem analysis.
So, let’s go that extra mile. The ASA is assisting. Conferences, guidelines, reports, journals, active learning experiences, and short courses are all offered by the ASA. A quick look at the titles of papers from JSM 2017 shows plenty of big data–related topics.
A Bigger Concern
There is a new book, Everybody Lies by Seth Stephens-Davidowitz. The basic thesis of the book is that the public will lie to poll-takers and opinion seekers. Apparently, individuals are concerned about giving a socially correct answer to a human, so who will they trust with their true personal concerns? Google. Yup, Google, the totally impersonal information source.
I recognize many have uncovered some of the flaws in this kind of crowd sourcing. What concerns me is that with all the data being consumed, I find it hard to believe the ultimate source of representative information is the Google collection of inquiries. Stephens-Davidowitz has an interesting thesis (in fact, this was the basis of his doctoral dissertation), but I hope it ends there. If nothing else, in this time of increasing concern about privacy, I am worried society would rather “confess” to Google without thinking about what might happen to all that Google information stored in some cloud. One day, the cloud could burst and rain on many parades.
What’s in a Name?
What would statistics be with a different name? Would that solve the big data problem? The idea of renaming statistics is not new. More than 50 years ago, John Tukey suggested “data analysis.” Others have suggested data science might be more effective than just data analysis.
In a thoughtful recent essay distributed on the ASA Community, Donald B. Macnaughton expresses his thoughts regarding the name, suggesting data science may be more inclusive than statistics, as statistics could indeed have four definitions. He notes this may be more important now in the age of big data since “data science” would sound better than “statistics” in tackling big data. On the other hand, I have heard that adding “science” to a name usually means it is not scientific. Macnaughton’s essay has certainly inspired some spirited comments among ASA members. The thread on the ASA Community has more than 40 comments as I write this column.
I have noticed Yale University has changed the name of their statistics department. They stated, “The amount of data flowing into the university increases on a daily basis. … Reflecting that dramatic expansion and the growing need for teaching and research in data science, the department of statistics has been renamed the department of statistics and data science.” I guess combining forces seems like a better idea than two departments at odds with one another. We’ll have to see how it works out at Yale.
I also have some personal skin in this game. When I met my late wife, she was working on her dissertation in psychology. I became her statistician and, later, her husband. Would I have had that same charm asking her if I could assist with the data science on her dissertation? Also, I happen to have a degree in operations research. You may know it as analytics, systems analysis, managerial science, industrial engineering, or systems engineering. Yes, I have been through this name game before.
Significantly forward,
Barry

You touched a nerve. I would love to see some competent statisticians spend time undoing the mischief that the accounting profession has been perpetrating in recent years.
Auditors, in particular, have fallen in love with the holy grail of using data to test their own accuracy. They call it ‘data analytics’. One could start with the recently published American Institute of CPAs Auditing Guide, ‘Analytical Procedures’.
There was a time when the big accounting firms engaged mathematical and applied statisticians to guide them in the development of statistical sampling for audit testing [such as John Neter, Frederick Stephan, Herbert Arkin, and others]. Seemingly, not any more; but, they do need guidance.
The accounting profession’s missteps in this area can best be exemplified by the comment by Arthur Andersen former senior partner Melvin Dick, in his testimony before Congress about Andersen’s failure to detect the massive [$11bn] WorldCom fraud, who said:”We performed numerous analytical procedures … in order to determine if there were significant variation that required additional work. We also utilized sophisticated auditing software…, which did not trigger any indication that there was a need for additional work.”
Neal Hitzig