Frequently Asked Questions (FAQs) About Data Mining

What is data mining?

Data mining, broadly speaking, is the use of sophisticated computer analytical programs to sift through large quantities of digital information (data) and discover in that information patterns or links that might not be obvious or observable to a human who would not have the time or resources to do the analysis by hand.  Examples of data mining range from the prosaic (searching food purchases at the supermarket to identify trends and aid in marketing) to the significant (sifting through international financial transactions to identify terrorist financing transfers).

Is there any difference between data mining and link analysis?

Yes and no.  In technical terms there is a difference.  Data mining is, strictly speaking, the effort to discover a pattern in the data that has not previously been revealed.  Computer analytics (actually artificial intelligence) are used to discover patterns we have not seen.  Link analysis, by contrast, starts with some known piece of data (the name, for example, of someone who cheats at cards in Las Vegas casinos) and tries to find unknown and unsuspected connections (his friends who are working as casino dealers).

The two have much in common.  Both use computer analytics and both rely on sifting through large quantities of data.  Both are also thought to pose the same types of potential threats to privacy and civil liberties.  As a result, most people use the two terms interchangeably when speaking colloquially, and The Data Minefield project has not tried to make that distinction, except where speaking of technical issues.

Who uses data mining?

Everyone and anyone with large collections of data.  It happens in the private sector for commercial use and it happens in the public sector for national security and any of a dozen other purposes.  Some examples of data mining:

  • Your supermarket swipes your loyalty card and records your purchases.  The data is used broadly to spot trends in food purchasing and also to target advertisements to you based on your recent purchases.
  • Credit card companies collect and sift all of their transaction data in order to identify suspicious patterns of activity that could indicate a stolen credit card or other fraud.
  • Large private sector companies systematically purchase and collect data about individuals, link it all together and sell it to other commercial users.  These companies (with names like Axciom and Choice Point) have customers ranging from pet products manufacturers (who target ads to dog and cat lovers) to political campaigns (who target fundraising requests to likely supporters).
  • FEMA contracts with a large commercial data base owner to do identify verification after disasters.  This helps people who have lost all of their identity documents in the disaster to prove their identity and claim the financial relief to which they are entitled; it also allows FEMA to reduce fraud by helping to disclose false identity claims.
  • The Department of Homeland Security uses a computer-based targeting system to sift through data bases of travel patterns and known terrorist connections to try and identify suspected terrorists who are attempting to enter the United States.

How does data mining work?

One expert has described the process of data mining analysis as something like putting together a puzzle.  First you try and find the pieces that look like they go together and then you try and fit them to each other.  As the puzzle pieces come together, finding more pieces that fit gets easier.  Of course, the computer can do it much quicker than a human can and can try millions (even billions) of information “pieces” very quickly.  Another expert has said that the weird thing about data mining is that the more hay you have on the haystack, the easier it is to find a needle.

The more formal answer is that the process of data mining consists of three distinct tasks: a) assembly of the database; b) analysis; and c) implementation/review.

Often assembling disparate data bases into a single form that can be used and correlated is one of the hardest tasks (think, for example, of how many ways there are to spell the name Mohammed, and how useless a database might be if the critical entry were spelled in an unusual way that the computer didn’t recognize).

Once assembled, the data can be addressed/analyzed with any of several forms of analysis.  Common techniques include regression analysis (to find mathematically correlated data), clustering analysis (to see data cluster patterns), associative rules analysis (testing for correlations between sets of variable – if, for example, people buying pasta also buy pasta sauce), and classification rules (for example, to distinguish spam from real mail).

The output of this analysis then has to be reviewed and validated.  As a predictive model not all patterns and links identified by analytical algorithms are valid.  The results need to be tested against larger data sets and then, if valid, acted upon.

What are the problems/issues with data mining (or, why should reporters care about it)?

Data mining necessarily involves the collection of large volumes of information.  Sometimes that information can be collected in an anonymous fashion, but more often than not it is useful precisely because the large volumes of data can also be linked to specific individuals.

When the data is mined to, say, identify a terrorist threat, that use is likely a valid one.  But sometimes the analytics identify the wrong person (this is known as a “false positive”) who can be adversely affected (think, for example, of how Senator Teddy Kennedy was repeatedly pulled aside for secondary screening by the TSA – that was a data mining false positive).

Even worse, the availability of such a large trove of personal information raises the specter of deliberate misuse.  Imagine a Nixonian enemies list and you get a sense of what the threat might be.  Though no significant abuses of this sort have been identified recently, civil libertarians are concerned that the capabilities are being developed without adequate controls.

When did data mining start?

The idea of data mining has probably been around since the first data was collected.  Noah likely used some rudimentary data analysis on his ark.

One of the earliest known examples of data mining may have been the renowned effort by Dr. John Snow to identify the source of a cholera outbreak in London in 1854.  He painstakingly collected geo-location data regarding the residences of the victims.  Eventually, he identified three contaminated wells that were the source of the outbreak.

Data mining has taken off, in the last 10-15 years, as increasing computing power and decreasing costs of storing data have made the routine use of comprehensive data analytics common.

When did data mining become controversial?

That’s an easy question to answer.  In the immediate aftermath of September 11, the Pentagon began development of a program named Total Information Awareness.  In November 2002 NY Times columnist William Safire became aware of the program and disclosed it in a column he titled “You Are A Suspect.” Critical public reaction followed immediately and in the wake of the story a number of government data mining programs came to light.

What was the Total Information Awareness system?

Total Information Awareness (TIA) was a research program initiated by the Defense Advanced Research Program Administration (DARPA) in the immediate aftermath September 11th.  Its conception was to use advanced data analytical techniques to search the information space of commercial and public sector data looking for threat signatures that were indicative of a terrorist threat.  Because it would have given the government access to vast quantities of data about individuals, it was condemned as a return of “Big Brother” and, ultimately, portions of the research program were cancelled.

Who is John Poindexter?

Admiral John Poindexter worked for President Reagan and was implicated in the Iran-Contra affair.  His conviction for criminal conduct was overturned on appeal and he retired from public life.  Immediately after 9/11 he approached DARPA with a radically new idea for data analytics.  That idea became TIA, and Poindexter was put in charge of TIA’s development (he worked for a nominal $1/year salary).  When his involvement in TIA became known, those who recalled his conduct in the Iran-Contra affair were especially concerned.

What is PNR and why is it controversial?

PNR is the acronym for Passenger Name Records.  The PNR is collected by airlines for all international flights and contains individually-keyed data about each traveler – things like the traveler’s travel history, credit card number, cell phone number, point of origin, nationality, passport number and the like.  Even before 9/11 the United States required international airlines whose flights arrive in the United States to provide the PNR for all arriving passengers, and then used sophisticated link analysis and a targeting system (known as the Automated Targeting System, or ATS) to identify passengers who warranted additional scrutiny upon arrival.  The program is controversial because it collects data about all individual travelers, both the suspicious ones and the hundreds of thousands of innocent travelers who arrive every week.  It is particularly notorious because many of these passengers arrive in the US from  Europe, and the European Union has complained that US efforts violate European law.  The PNR collection program has been the subject of 4 separate negotiations with the EU in the last 10 years and has generated many objections in the European Parliament.


Is there a difference between how the public sector uses data mining and how the commercial sector uses it?

Yes and no.  At bottom the techniques of link analysis and data mining are the same in both spheres.  But several important factors distinguish the different types of use.  First, commercial link analysis and data mining has a lot of data to analyze – we make millions of supermarket purchases every day.  Fortunately, we have much less data about terrorist threats or even serious crime to analyze for governmental purposes.  But more data makes better analysis.  Second, of course, is the consequence that can come from the data mining.  In the commercial sector, you get more targeted ads and the like.  With governmental programs there comes governmental scrutiny of your conduct.  For both these reasons many who are perfectly comfortable with commercial data mining worry about governmental data mining.

On the other hand, the governmental objective is seen by many as more significant – it is not unreasonable to think that it is more important to prevent a recurrence of terrorism than it is to sell more dish soap.  For this reason, some advocate even greater use of data mining tools for counter-terrorism purposes.

Don’t laws protect our privacy and civil liberties?

Again: Yes and no.  The Fourth Amendment protects our privacy for information we keep to ourselves, but the law also says that if we give our information to someone else (like a bank or an airline) then the government can get our information from the airline without using a warrant or other Fourth Amendment protection.  So the only laws that protect our privacy are statutes passed by Congress – and most of those were written long before the internet came along.  The Electronic Communications Privacy Act (that protects the privacy of e-mail) was written in 1986.  The core protections of the Privacy Act were written in 1974.  The law doesn’t match technology today.

What is Geo-location?

Geo-location is a subset of the data mining problem.  Most data mining is done using transaction data (that is data about things you do, like a credit card purchase or a flight).  But sometimes data mining techniques can be applied to geographic location data in an effort to find patterns from where you go, rather than what you do.  And, of course, the two types of data can be mixed together, as well.

Comments are closed.