I don't know how to check this myself, but are the redactions flattened, i.e. so they cannot be removed, or are they just shapes on a another layer drawn over the relevant areas, i.e. is the redacted data still recoverable by editing the pdf?
It wouldn't be the first time a "redaction" turned out to be no more than a mere obfuscation in practise.
They are flattened: The FBI prints out files, redacts them, and then scans them back in. A number of federal agencies do this, really making it a pain to search through documents.
You can upload the PDF to Google Docs, which will OCR it and make the document searchable. This particular PDF, however, is rejected because it's more than 2MB :(
We're looking to build quality OCR into MuckRock so every document gets it, but it's surprisingly either a) technically difficult or b) financially expensive, choose one.
That's really cool, thanks. But I don't think it includes OCR. The best FOSS OCR out there is Tesseract, which we use through DocumentCloud, but it still leaves a lot to be desired compared to commercial solutions.
What do you consider expensive? Transym is really good and not too expensive (disclosure: they are a partner of my company -- but I am endorsing them on that basis)
£60.00 doesn't sound bad. We'd need it to fit into our automated workflow and be able to handle hundreds of pages at a time, and occasionally 1k+ page documents, and stuff from Nuance we'd looked at would have cost $4k or more a year, plus charge per page which is a deal killer. This might be a good fit ... thanks!
Yes. Accidentally leaking information makes it harder for the FBI to gather information in the future, so they ensure that they are very careful when disclosing information. In this case, it makes the information harder to search, but that's why we pay journalists to read the released documents and summarize them for us.
It wouldn't be the first time a "redaction" turned out to be no more than a mere obfuscation in practise.