For those considering something like this, you might want to consider using scrapy, a Python web crawler, instead of rolling your own crawler.
I remember when I was looking for something like this a year ago and found that project. I've used it for a few things and it does a nice job abstracting away most of the core scraping architecture, but leaving room to be extended as necessary. After playing with it for a few weeks, makes you realize just how easy it is to grab a bunch of data from the web if you need it.
I would recommend haystack also for someone wanting to take this to the next step. Since django is already being used and he can swap Whoosh for solr if needed to.
For those considering something like this, you might want to consider using scrapy, a Python web crawler, instead of rolling your own crawler.
I remember when I was looking for something like this a year ago and found that project. I've used it for a few things and it does a nice job abstracting away most of the core scraping architecture, but leaving room to be extended as necessary. After playing with it for a few weeks, makes you realize just how easy it is to grab a bunch of data from the web if you need it.