Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I once wrote a classifier to tell whether the metadata on a web page was any good.

The best features? Look for classic html mistakes on the page. Still using font tags, do they use gifs instead of png, did they include a keywords meta tag, did they specify an encoding or are there windows-1252 characters present anywhere, etc. I came up with 20 or so signs of bad html elsewhere on the page, and collectively those features were much more predictive than any of the content itself.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: