The Future of Data Mining: DataSphere, Knowledge Discovery and Dissemination, and Catalyst

Although the Total Information Awareness initiative was officially shuttered in 2003, it is clear from the efforts described below that the spirit of TIA lives on via ongoing research. As new forms of data are constantly thrown into the intelligence pile, efforts continue to produce the analytic capacity to bring meaning and value to the flood of information.

In terms of physical infrastructure, the National Counterterrorism Center exists to serve this purpose: an information hub staffed by representatives from multiple agencies, where analysts dedicate themselves 24/7 to the synthesizing of counterterrorism intelligence. The latest data mining report from the Director of National Intelligence details the NCTC’s DataSphere system, which is described as providing “analysts with a tool to aid in the discovery of unknown terrorism relationships and the identification of previously undetected terrorist and terrorism information” using NCTC data holdings – a potentially vast trove of information. The report indicates that the program does not “data mine” as the practice is statutorily defined, because all database queries begin from information related to known or suspected terrorists. However, the report goes on to state that “it is contemplated that Pattern-matching functionality using data not yet known to be terrorism information will be included in future development phases.”

At an IT level, an entire programs division within the Intelligence Advanced Research Projects Activity (IARPA) is devoted to the task of blunting information overload, designated the Office of Incisive Analysis. A program called Knowledge Discovery and Dissemination (KDD) is also included in the DNI’s most recent data mining report.  Though not yet operational, it is being tested in conjunction with the Department of Homeland Security and other law enforcement agencies. The KDD effort involves research that seeks to “align” large and complex datasets. Data alignment refers to the process of taking separate databases – often containing information that is different both in substance and format – and allowing for querying and effective analysis across these disparate information sources. The work is being carried out using “realistic intelligence problems” and classified datasets. Research contracts were awarded in September 2010 and the project is to undergo several phases over the course of its planned 51-month duration.

Catalyst’s purpose also seems directed toward managing information overload and bringing meaning to the various sources of Intelligence Community data. The program is managed by the DNI’s Chief Information Officer. According to the latest DNI report, “In its end-state, Catalyst will enable data fusion/analytic programs to share disparate repositories with each other, to disambiguate and cross-correlate the different agencies’ holdings, and to discover and visualize relationship/network links, geospatial patterns, temporal patterns and related correlations.” In describing the problem that Catalyst exists to solve, its purpose is practically indistinguishable from that of IARPA’s KDD: the fact that analysts frequently “encounter data that was collected for different purposes, protected in different ways, and described using different terminologies.” Research on incorporating pattern-matching functionality is slated to begin in 2013.

Comments are closed.