The Data Minefield http://nationalsecurityzone.medill.northwestern.edu/datamining Navigating Privacy and Security Fri, 29 Jun 2012 04:32:55 +0000 en-US hourly 1 http://wordpress.org/?v=4.2.3 Finding the Stories: Navigating the Data Minefield http://nationalsecurityzone.medill.northwestern.edu/datamining/finding-the-stories-navigating-the-data-minefield/ http://nationalsecurityzone.medill.northwestern.edu/datamining/finding-the-stories-navigating-the-data-minefield/#comments Fri, 19 Aug 2011 15:18:10 +0000 http://nationalsecurityzone.medill.northwestern.edu/datamining/?p=746 Continue reading →]]> Data mining is not the sexiest component of the national security beat. Still, from the standpoint of journalistic merit, it is a subject as quintessentially Fourth Estate-worthy as anything out there. The notion that we are entitled to a “right to be let alone” is a persistent tenet of American life, despite our increasingly public, digitally documented lives. But a search and seizure of one’s digital existence likely does not feel invasive – without a doubt, today’s potential privacy violations are of a more subtle and nuanced variety. This is all the more reason for competent journalists to foster an informed dialogue about our expectations and rights as Americans when it comes to the increasingly vast stores of data on all of our lives. What’s more, it is my sense that the magnitude of the personal information constantly shooting through cyberspace – and the extent to which such data has undergone commodification – is largely unbeknownst to the average American. This is prime territory for those who exist to inform the public, and not a bad place to look for the next big scoop either. 

Ironically, collecting information about the government’s collection of our personal information is a challenge, one which this brief “how to” guide attempts to make a little bit easier. Treading the fine line between national security and civil liberties is an inherently difficult and uncertain endeavor for those responsible for our safety. I imagine that given a choice, they’d rather abide by the notion that “What Americans don’t know (about what we know) can’t hurt them,” when it comes to data mining initiatives. In short, transparency and abundant information are unlikely to accompany your forays into the data mine field.

Having said that, via Freedom of Information Act requests and dogged scrutiny of other publicly available information, many advocacy organizations have done well to pull back the veil. Sitting down to begin this project, I began by familiarizing myself with the numerous privacy groups out there. Their websites served as valuable resources, archiving primary source material and hyperlinking to other influential sources and authorities within the privacy advocacy community. The blogroll we’ve included gives a taste of the daily musings of these sorts of folks, though effectively trolling the blogosphere typically involves straddling blogs devoted to three distinct but related issues: data mining (of which many focus on the issue in a technical sense), privacy and national security.

There is also a significant amount of literature on data mining produced by think tanks, which tends to be less news-y fare but usually offers deeper, more nuanced and more ideologically diverse perspectives on the subject. Most think tanks offer keyword search functions (of varying effectiveness) on their websites, but if not, a simple call to the organization’s press office should reveal any scholarship produced on the subject. If nothing else, many of these resources should give you go-to sources when you do break the next big data mining/privacy story.

But eventually, the trove of FOIA’ed documents, investigative reporting and scholarly dissertations will give way to this reality: that there’s much more Out There than these watchdogs and professional thinkers could ever fully uncover. At this point, you’ve got to get down in the mud with the rest of them.

For myself, this involved combing through the hundreds of Privacy Impact Assessments pertaining to various government databases that contain “personally identifiable information.” Typically, an abstract will let you know whether the thing is worth reading beyond the first page. Of particular interest is the PIA’s response to whether the system analyzes data to “assist users in identifying previously unknown areas of note, concern, or pattern” – or some derivation of this general line of questioning. If so, you’re looking at a program that employs some form of classic data mining. Naturally, a “no” here is no automatic disqualifier, due to variations in the interpretation of that question. The Transportation Security Administration’s Enterprise Search Portal is an example of this: Though its PIA states that it does not “use technology to conduct electronic searches, queries, or analyses in an electronic database to discover or locate a predictive pattern or an anomaly,” it does feature a “discovery” capability that “permits the system to identify relationships between data that is prompted by the system (rather than a user search).” In other words, data mining, as most understand the concept.

System of Records notices are also an important tool. These are required by law under the Privacy Act of 1974 for any government database that maintains information on U.S. citizens that is retrieved via some form of personal identifier. SORNs are published in the Federal Register. Both of these acts of public record offer email listservs that can be subscribed to, so that your inbox will inform you as to when a new PIA or SORN awaits your inquisitive perusal.

The difficulty of this approach lies in the fact that PIAs are not required for “national security systems,” and these are generally the data mining initiatives that are most intensive and broadest in scope. Even in the rare event that a PIA is issued on a system designed for national security purposes, the resulting document can be less-than-enlightening, as evidenced by this Pathfinder PIA, largely redacted and devoid of any real meaning. With SORNs, agencies will sometimes lump new programs into existing SORNs – the FBI, a data gathering behemoth, is notorious for categorizing any new use of data under its all-encompassing “Central Record System” SORN.

Annual data mining reports from the various federal agencies are also a logical place to look. Unfortunately, the statutory language of the Federal Data Mining Reporting Act of 2007 is fundamentally flawed for our purposes, narrowing the scope of what is mandated for inclusion in these reports significantly. Programs employing link analysis of data sets are not included, and agencies are not required to include systems designed for subject-based queries. Some agencies will provide information on these programs “in the interest of transparency,” but my assumption is the public is not shown this courtesy regarding the most potentially controversial of programs. Additionally, as one might expect, classified or law enforcement sensitive systems are reported to the appropriate congressional oversight committee members, but are shielded from the general public.

The Government Accountability Office and Congressional Research Service can also be sources of illuminating information on the government’s data mining initiatives. A 2004 GAO report in many ways served to raise national consciousness as to the extent of these efforts, and many reports from both entities have since shed further light on individual programs and some of the broader privacy and civil liberties concerns intrinsically tied to the practice. We’ve provided summaries and links to some particularly noteworthy GAO work on the subject, intended to serve as sources of information and potential story ideas.

Beyond these relatively structured and official sources, Googling and querying the myriad online databases should assist in filling in information gaps. From LexisNexis to govtrack.us to opencrs.com, as any 21st century reporter knows, the Internet is an indispensible resource for a skilled navigator of the cyber waters.

Finally, there’s a notable difference between researching data mining programs in the abstract and reporting on data mining as a beat component. Cultivating sources, tracking down leads and desperately scrambling for information on deadline are all distinct from the resource-creating effort we undertook here. As such, newsgathering methodologies are always somewhat individualistically organic. Where this how-to guidance is lacking, reporting experience should yield its own strategies and answers.

 

]]>
http://nationalsecurityzone.medill.northwestern.edu/datamining/finding-the-stories-navigating-the-data-minefield/feed/ 0
Finding the Stories: Data Mining Research http://nationalsecurityzone.medill.northwestern.edu/datamining/finding-the-stories-data-mining-research/ http://nationalsecurityzone.medill.northwestern.edu/datamining/finding-the-stories-data-mining-research/#comments Fri, 19 Aug 2011 15:16:55 +0000 http://nationalsecurityzone.medill.northwestern.edu/datamining/?p=744 Continue reading →]]> Data mining straddles a very gray area, provoking extreme reactions on either side of the debate on its use and implementation. Those in favor of such analytical techniques, governments being the usual suspects, make a strong case for its effectiveness. But those opposed, raise privacy concerns and, at times, even question the efficacy of expensive technology in curtailing threats.

It is difficult to navigate through government websites to look for information on various systems and laws enacted related to data mining policies. Most of the information on this topic is available through privacy-related advocacy groups.

While there is no guarantee of objectivity or any promises on finding both sides of the story, such websites do offer a good starting point.  Especially if most of the information refers you back to laws and statutes that are there on government sites.

One of the oft-quoted resources on the web is Privacy International. The group has a comprehensive listing of countries and details their privacy framework.  While the last profile reports they prepared for a number of countries came out in 2007, it still serves as a great place to start the research.

For a more current and up-to-date profile, it is best to search for region-specific groups rather than look for a global starting point. While you will find a great number of groups when looking for information on developed markets like the European Union, the research might be limited when looking for information on developing or under-reported countries.

A heavy reliance may have to be on blogs and forums where enthusiasts on this issue have a lot to say, more often than not highlighting their view on the issue.

A major hurdle while drawing up a conclusion on the data mining climate in a country is the dichotomy present in several newly implemented laws. Data retention principles may walk hand in hand with absolute guarantees of privacy. It is the presence of an oversight mechanism that could help further the understanding of the government’s attitude toward privacy.

There is also a worry about how much of the laws on paper are actually carried out in spirit. In several places, it won’t be surprising to find accusations of the government overstepping its legal right. And, perhaps, most crucially, with countries waking up to the technological potential of the new-age tools, changes to the laws in a short period might be commonplace — to accommodate concerns as well as to incorporate the latest tools to further the case for internal security.

]]>
http://nationalsecurityzone.medill.northwestern.edu/datamining/finding-the-stories-data-mining-research/feed/ 0