Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's not just random I/O that's fast, but sequential is a lot faster too.

But what I'm thinking of is the tokenization of a large corpus of text into bigrams and trigrams for example.



Good point, with all the focus on random I/O in the announcement I was forgetting about the sequential I/O performance benefit.

Hmm, that reminds me, I have an interesting toy problem which is sequential I/O limited...


Admittedly, it would be fun to pull in the Common Crawl and process it.

Note that there's a code contest for Common Crawl that recently was announced: http://commoncrawl.org/first-ever-code-contest/




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: