Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I do a lot of data mining and text mining as well. My biggest concern is always memory usage and UTF-8 is hugely more memory efficient than UTF-16 in almost all cases, even for Asian text. Any difference in character counting performance is negligible compared to the benefit of avoiding disk access.

Also, the text I process is never UTF-16 at the source. So even if I used UTF-16, I would have to convert the text first and that would be the one and only time characters are counted. There would be no additional counting overhead at all.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: