I have some methodological problems with this piece. They're making claims about the makeup of this data set, that 50% of the documents contain minimized references to US persons for instance, or that 90% of the account holders were not intelligence targets, but we don't actually know whether this data represents NSA collections as a whole. There are 160,000 documents over a four-year period. That seems far too low to be the sum total of all NSA collections, so it must be a sample. But is it a random sample or was this data selected by Snowden using some criteria? How do we know that Snowden didn't choose data from one particular program or using certain selectors and that that particular data tends to have more or less US persons or a higher or lower percentage of intelligence targets than NSA intercepts taken as a whole?
I also find the following claim problematic:
>> Many of them were Americans. Nearly half of the surveillance files, a strikingly high proportion, contained names, e-mail addresses or other details that the NSA marked as belonging to U.S. citizens or residents. NSA analysts masked, or “minimized,” more than 65,000 such references to protect Americans’ privacy...
So there were 65,000 minimized references in 160,000 documents. But we also know that a "minimized reference" doesn't actually mean the a US person was the sender or the recipient - for instance, we know from Gellman that someone talking about President Obama would constitute a minimized reference. The first sentence, "Many of them were American" is not quantified, likely because the Post doesn't actually know how many participants in the intercepts were American.
> Is it a random sample or was this data selected by Snowden using some criteria?
Let's take one probable situation: Let's say Snowden took everything he could. That means there a gathering of data made by an employee at the NSA which results in 90% non-intelligence targets. What was the employee working on?
- Increasing the relevance of data? Not probable.
- Working on a usual sample of data? Most probable.
- Targetting non-intelligence targets on purpose? Scary and illegal.
I also find the following claim problematic: >> Many of them were Americans. Nearly half of the surveillance files, a strikingly high proportion, contained names, e-mail addresses or other details that the NSA marked as belonging to U.S. citizens or residents. NSA analysts masked, or “minimized,” more than 65,000 such references to protect Americans’ privacy...
So there were 65,000 minimized references in 160,000 documents. But we also know that a "minimized reference" doesn't actually mean the a US person was the sender or the recipient - for instance, we know from Gellman that someone talking about President Obama would constitute a minimized reference. The first sentence, "Many of them were American" is not quantified, likely because the Post doesn't actually know how many participants in the intercepts were American.