Key Terms and Definitions

Data Mining

There are many competing definitions of the practice of data mining.  Some are very technical, scientific and narrow.  Others are broader, more holistic, and more policy oriented.  Any effort to write about data mining needs to carefully define its terms and distinguish its assumptions.

In general, the narrower definition restricts the idea of data mining to instances in which data is collected and analyzed for predictive patterns or anomalies, without, any initial individual data point as a starting point.  In other words, under this narrow definition the data is studied to suggest a pattern for examination or consideration without any preconceptions and is often thought of, in the national security context, as being linked to preventative efforts to stop terrorism or crime before it occurs.

Likewise, in the commercial context, data mining is thought of as discerning patterns of consumer behavior without any preconceived notion of what that behavior is likely to be.  A classic, though almost certainly apocryphal, example is often used to explain the idea. In the telling of the story the clerks at a convenience store (sometimes Wal-Mart) are said to have noticed that men who came in late at night tended to buy beer and diapers together.  Managers studied the sales data and confirmed this insight – there was a strong correlation.  They hypothesized that the coincidence was because men were often sent out at night on emergency diaper runs and frequently bought beer as a collateral item.  The convenience store restocked the shelves to put diapers and beer next to each other and sales of both were said to have increased substantially.  Whether a true story or not, it is a good example of the technical definition of “data mining” – the effort to find patterns in data without any assumed or defined starting point.

The Data Mining Report Act (42 U.S.C. § 2000ee-3(b)(1)), which is the only Federal statute defining the term, uses the narrow and technical definition of data mining.  It defines data mining as:

“a program involving pattern-based queries, searches, or other analyses of 1 or more electronic databases, where—

  • (A) a department or agency of the Federal Government, or a non-Federal entity acting on behalf of the Federal Government, is conducting the queries, searches, or other analyses to discover or locate a predictive pattern or anomaly indicative of terrorist or criminal activity on the part of any individual or individuals;
  • (B) the queries, searches, or other analyses are not subject-based and do not use personal identifiers of a specific individual, or inputs associated with a specific individual or group of individuals, to retrieve information from the database or databases; and
  • (C) the purpose of the queries, searches, or other analyses is not solely—
    • the detection of fraud, waste, or abuse in a Government agency or program; or
    • the security of a Government computer system.

The broader, less technical definition of data mining is sometimes called “link analysis.”

Link Analysis

Link analysis is the science of “connecting the dots.”  It differs from the narrow definition of data mining in that with link analysis there is always a starting point or “seed” of information from outside the system.  The data is then mined to find connections between the known starting point and other points – and those connections can be of any sort.  They can be a common phone number, travel to the same place, membership in the same kennel club.  Any connection is potentially valid, though of course, some connections are more relevant than others.   In the aftermath of 9/11 one expert did this analysis to show how connecting the dots through link analysis might have uncovered the hijackers.

More formally, “link analysis is a subset of network analysis, exploring associations between objects. An example may be examining the addresses of suspects and victims, the telephone numbers they have dialed and financial transactions that they have partaken in during a given timeframe, and the familial relationships between these subjects as a part of police investigation. Link analysis here provides the crucial relationships and associations between very many objects of different types that are not apparent from isolated pieces of information. Computer-assisted or fully automatic computer-based link analysis is increasingly employed by banks and insurance agencies in fraud detection, by telecommunication operators in telecommunication network analysis, by medical sector in epidemiology and pharmacology, in law enforcement investigations, by search engines for relevance rating (and conversely by the spammers for spamdexing and by business owners for search engine optimization), and everywhere else where relationships between many objects have to be analyzed.”

Personally Identifiable Information (PII)

Personally identifiable information (PII) is any data about an individual that could potentially identify that person such as a name, fingerprints or other biometric data, email address, street address, telephone number or social security number.  The data need not actually have a personal identifier, so long as the potential exists to match the data to an individual through investigation or examination.

Comments are closed.