Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You can upload the PDF to Google Docs, which will OCR it and make the document searchable. This particular PDF, however, is rejected because it's more than 2MB :(


We're looking to build quality OCR into MuckRock so every document gets it, but it's surprisingly either a) technically difficult or b) financially expensive, choose one.


I OCRd it using Acrobat: http://dl.dropbox.com/u/980684/Jobs.pdf

Probably not perfect, but better than nothing.



That's really cool, thanks. But I don't think it includes OCR. The best FOSS OCR out there is Tesseract, which we use through DocumentCloud, but it still leaves a lot to be desired compared to commercial solutions.


What do you consider expensive? Transym is really good and not too expensive (disclosure: they are a partner of my company -- but I am endorsing them on that basis)


£60.00 doesn't sound bad. We'd need it to fit into our automated workflow and be able to handle hundreds of pages at a time, and occasionally 1k+ page documents, and stuff from Nuance we'd looked at would have cost $4k or more a year, plus charge per page which is a deal killer. This might be a good fit ... thanks!




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: