Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

No, it really isn't. Download a copy of SpamAssassin, train it on 400 hand-picked spams from your own mailbox, train it on 400 arbitrary hams as well, and you will have very accurate spam filtering. I was shocked how well it worked; I run my own mail personally and use GMail at work, and the results are (subjectively) indistinguishable. A Dovecot plugin that keeps the Bayesian numbers up to date as I move messages in and out of the Spam and Archive (for ham) folder completes the picture.

To go deeper, it turns out that Bayesian filtering is remarkably resilient. Even attacks which try to poison your filters by including ham-like content in spams are ineffective, because spammers cannot very accurately predict what your own particular flavor of ham is like. (People don't often mail me passages from out-of-copyright Victorian romance novels.)

I find that a few botnet-reducing SMTP heuristics plus Bayesian is sufficient; I dallied with some of the fanciness that compares known-spam hashes with other people, but it turned out not to be necessary.



Maybe it has changed but I have vivid memories of training lots of spam a decade or so ago and still getting half-assed results. Google was the first email provider that really did a good job blocking spam.


There is one setting, underdocumented, which makes a big difference these days. spamc has a default ceiling of 10K, messages larger than which it passes unchecked. Spammers have started routinely including images just over that threshold to defeat default installs. Bump that up a bit, and your accuracy will go way up.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: