Making Sense of the Babel

As part of what seems to be an ongoing effort to convert and integrate all sensory elements into data to be analyzed, one Intelligence Advanced Research Projects Activity (IARPA) initiative named Babel seeks to develop technology capable of rapidly transcribing audio from a host of languages.

The development of this program in and of itself does not constitute data mining, but there can be little doubt about the ultimate value of being able to transform audio from hundreds of languages into searchable text. Already, the private sector has found cost savings via automated transcription in health care, and in sales and marketing, companies have begun to combine this technology with data mining practices.

A program similar to Babel, in development by the Defense Advanced Research Projects Activity (or DARPA, the Defense Department’s precursor to IARPA), incorporates a feature with unmistakable data mining potential: the ability to pluck keywords and phrases from streaming audio. The project also plans to provide voice recognition capabilities and to be able to do all this even when dealing with “noisy or degraded speech signals.”

The ability to rapidly transcribe numerous languages would appear to be a technology most aptly deployed by the National Security Agency, which intercepts untold volumes of communications data as the nation’s foremost spy shop. Its post-9/11 warrantless wiretapping of domestic communications caused a furor among privacy advocates and the general public when it was uncovered in 2005. It also gives us a sense of the surveillance capabilities of one of government’s most secretive entities.

A USA Today article published shortly after the wiretapping was exposed describes the agency as “the ‘ear’ of the nation’s intelligence system. It uses a system of satellites and other means to listen in on friend and foe alike, decoding and translating communications and reporting the results to key recipients in government.”

As with many technologies and data sets, the capability is only a potential civil liberties concern. It is deployment of such a capability – and specifically, deployment on a large scale of indiscriminant scope – that leads privacy advocates and civil libertarians to cry foul. The value of efficient and effective multilingual transcription is undeniable, at the same time that the impressive and ever-growing capacity to store data makes dragnet audio surveillance increasingly feasible. The ability to turn speech into text is an invaluable enhancement to the utility of such communications for pattern and link analysis, as well as subject- or keyword-based queries.

Comments are closed.