I once wrote a classifier to tell whether the metadata on a web page was any good.
The best features? Look for classic html mistakes on the page. Still using font tags, do they use gifs instead of png, did they include a keywords meta tag, did they specify an encoding or are there windows-1252 characters present anywhere, etc. I came up with 20 or so signs of bad html elsewhere on the page, and collectively those features were much more predictive than any of the content itself.
The best features? Look for classic html mistakes on the page. Still using font tags, do they use gifs instead of png, did they include a keywords meta tag, did they specify an encoding or are there windows-1252 characters present anywhere, etc. I came up with 20 or so signs of bad html elsewhere on the page, and collectively those features were much more predictive than any of the content itself.