Showing posts with label Data Mining. Show all posts
Showing posts with label Data Mining. Show all posts

Saturday, December 19, 2009

Who Owns Our Personal Data?

A class action suit was filed on Thursday against Netflix in the US District Court for the Northern District of California. The class is “All Netflix subscribers that rented a Netflix movie and also rated a movie on the Netflix website during the period of October 1998 through December 2005, residing in the United States.”

According to the class action complaint, “On October 2, 2006, Netflix perpetrated the largest voluntary privacy breach to date, disclosing sensitive and personal indentifying consumer information. The information was not compromised by malicious intruders. Rather, it was given away to the world freely, and with fanfare, as part of a contest intended to benefit its trusted custodian, Netflix.” The lawsuit is brought as a class action by and on behalf of similarly situated Netflix subscribers whose privacy was violated by Netflix as organizer of the “Netflix Prize” contest.

Nettflix launched a contest in the Fall of 2006 offering cash prizes to contestants who could provide collaborative filtering algorithms that would predict viewers' movie ratings with a greater accuracy than Cinematch, which is Netflix’s proprietary recommendation software.

According to the complaint "Netflix subscribers’ movie rental choices constitute personal information that subscribers reasonably expect will be treated as presumptively confidential and that their relationships with Netflix are relationships of confidentiality. Netflix has been entrusted with the confidential, sensitive, and personal information of millions of consumers.”

What is personal information? According to Netflix privacy policy: “Personal information means information that can be used to identify and contact you . . . as well as other information when such information is combined with your personal information.”

This point is interesting, as many pieces of information can become personal information, if there is a way to combine them with information that can be used to identify a person, most of the time a name or an address, but also in some cases a title (CEO of Microsoft, Secretary of the PTA of PS 2349 Pleasantville, IL). It is so easy nowadays to link databases that virtually every information may become, at any given time, a personal information.

One of the argument of the complaint is that “ Netflix was attempting to play a semantics game—“personal” information meaning pertaining to or concerning a particular person; however personal information is not limited to a Netflix subscriber’s name. Netflix subscribers reasonably believe that no record would be released showing that they watched a dogmatic, controversial, or sexually explicit show, regardless of whether their actual name is known.”

The notion of personal information goes way beyond the mere name of an individual. Netflix indeed disclosed the personal information of its subscribers “to over 51,000 individuals.”

A very interesting point in the complaint is the claim that Netflix has been unjustly enriched by this scheme, firstly because it has “benefited from its unlawful acts through the receipt of payments for Internet service from Plaintiffs" and secondly because it “continues to benefit from [its] unlawful acts through the receipt of payments in connection with its proprietary search engine, which continues to index websites associated with the subscriber data. (…) Plaintiffs are entitled to the establishment of a constructive trust consisting of the benefit to Netflix of such payments from which Plaintiffs and members of the Classes may make claims on a pro-rata basis for restitution."

The first part of the unjust enrichment argument, the receipt of payments for Internet service from Plaintiffs is not as strong as the second part of the argument, the receipt of payments in connection with the search engine. The way I understand both arguments is that Netflix benefited from the users' fees, and rightfully so, until the moment it disclosed unlawfully their personal information. After that, these fees were somehow the product of its unlawful acts. But Netflix continued to provide their rental services, and was entitled to collect fees for that service.
The second argument is much stronger. Netflix benefited unjustly from the value represented by its users' data. Most consumers still fail to realize that their personal data are very valuable, so much that it has a market value. Their shopping habits are sold by retailers to marketers.

In Dwyer v. American Express Co., 273 Ill. App. 3d 742, (Ill. App. Ct. 1st Dist. 1995), the court held that Amex did not commercially appropriate its cardholder’s personal spending habits. In that case, the plaintiffs also claimed that these actions constituted a deceptive practice claim under the Illinois Consumer Fraud Act. The plaintiff had to prove that these practices constituted a misrepresentation or a concealment of material fact, that it was the defendant’s intent that the plaintiff relies on this misrepresentation or concealment, and that the deception had occurred in the course of trade or commerce. Amex had not informed the cardholders that their spending habits would be analyzed and their names sold to merchants, and this is a deceptive practice under the Illinois Consumer Fraud Statute. However, the court held that the plaintiffs had failed to allege that they suffered any damages, except maybe for “a surfeit of unwanted mail.”

In the Netflix case, it could be easier to prove damages, because Jane Doe, one of the plaintiffs, "believes that, were her sexual orientation public knowledge, it would negatively affect her ability to pursue her livelihood and support her family and would hinder her and her children’s’ ability to live peaceful lives within Plaintiff Doe’s community.”

Thursday, September 24, 2009

Government Use of Private Databases

Via Wired.com, an article on the FBI's National Security Branch Analysis Center (NSAC) database, which, according to declassified documents obtained by Wired, contains than 1.5 billion government and private-sector records about citizens and foreigners, ans is thus becoming the “Total Information Awareness” the government wanted to put in place after 9/11.

The government thought then of data mining private databases for national security reasons. The New York Times had reported in February 2002 that the Pentagon, under the leadership of Vice Admiral John Poindexter, was building a computer able to collect and data mine personal data, such as credit card records, school, travel and medical records, in order to track terrorists. The name of the program, Total Information Awareness, was later changed to Terrorism Information Awareness. Congress eliminated funding for the program in 2003.

TIA was followed by “The Matrix,” a data mining program linking government and commercial databases. Government agencies have also required in the past the assistance of telecommunication carriers to eavesdrop on suspect’s emails. The FBI’s CARNIVORE program plugs, (or plugged,) a computer (the DCS-1000) directly to an ISPs’ network to monitor suspect incoming and outgoing emails. ECHELON was a program eavesdropping on international private telephone calls, e-mails and faxes, using both ground and satellites.

Data mining, or Knowledge Discovery in Databases (KDD) is the process that allows experts to extract trends and patterns from data, using algorithms to identify relationships and patterns in data. There are two main data mining methods. The top-down method looks for a well-defined profile by asking questions and testing hypotheses. The bottom-up method analyzes raw data to find trends and groups.

Is data mining such an extensive amount of information an efficient method to increase security? This is an important question as citizens are asked to trade off some of their liberties for security, or at least for a renewed sense of security, and would be more reluctant to do so for a program efective in preventing terrorist attacks. Former Homeland Security Secretary Michael Chertoff believed in the ability of data mining to prevent terrorism. While an assistant attorney general, he testified in 2002 that he found data mining a promising way to fight terrorism. He further testified that the Department of Justice was “using computers to analyze information obtained in the course of criminal investigations, to uncover patterns of behavior.(…) Through what has come to be called ‘data mining’ and predictive technology, we seek to identify other potential terrorists and terrorism financing networks.”

Even some privacy advocates believe that the use of commercial databases can “help improve the amount and quality of identifying information in watch lists.” However, most of them do not believe that such massive data mining would protect us against terrorist attacks, and is not fail-proof. The Wired article quotes Kurt Opsahl, an EFF senior attorney: “We have a situation where the government is spending fairly large sums of money to use an unproven technology that has a possibility of false positives that would subject innocent Americans to unnecessary scrutiny and impinge on their freedom.”

The efficiency of a method should not be the ultimate test used to establish an opinion about government surveillance. But if we may have a high surveillance tolerance, we certainly have a zero tolerance for being arrested by mistake, or prevented to board an airplane.

Monday, August 27, 2007

Face Book as marketing tool?

USA Today has an article today about a Face Book plan to use the data voluntarily provided by Face Book users to display targeted ads .

Blogs and social networks sites provide ample opportunities to disclose, voluntarily, a cornucopia of personal information. Why disclosing voluntarily so much personal information? New York Magazine published last February an interesting article on the subject , written by Emily Nussbaum. In the academic world, Danah Boyd wrote Why Youth (Heart) Social Network Sites: The Roles of Networked Publics in Teenage Social Life.

Twitter

Blog Archive

AddThis Social Bookmark Button

Labels