It is nice book, but might be little bit outdated.
Basically it's a map of words to a list of documents containing that word.
eg {"hello": [1, 2], "world": [1,3,4], ...}
(the numbers are document id's)
So for example, the word 'hello' occurs in documents 1 and 2. 'world' occurs in documents 1,3 and 4.
Doing boolean queries is also really easy with an inverted index. You basically get the document set for each word and then do a union on the sets for an OR query.. or an intersection to do an AND query.
Pretty cool right?
In addition there are a wealth of online video lectures that may inspire you: http://www.datawrangling.com/hidden-video-courses-in-math-sc... and http://videolectures.net/mlss04_hofmann_irtm/ and http://videolectures.net/Top/Computer_Science/
In so far as search engines go it's certainly worth playing around with Lucene. It's well implemented and you'll learn a lot of what really matters when it comes to indexing and retrieval.
For the text processing (classification, data extraction) side It may also be worth brushing up on your stats (a good excuse to learn R) and checking out Mahout http://lucene.apache.org/mahout/
There are some pretty nice tools to go with Lucene - I've used Luke quite a bit: http://code.google.com/p/luke/
Tim was nice enough to reply to my email query some years ago and point me to this (already written). It's not comprehensive, but it was helpful. I guess I'm mostly adding this comment to attest to the generosity inherent in the sharing of such information. Aka, "thanks".
http://www.search-engines-book.com - Slides, Data Sets
http://www.pearsonhighered.com/croft1epreview/toc.html - Book Table of Contents
The book expands on the slides, as well as includes homework problems, some requiring the use or modifications of the open-source Galago Search Toolkit.
Solr 1.4 Enterprise Search Server http://www.amazon.com/Solr-1-4-Enterprise-Search-Server/dp/1...
Programming Collective Intelligence http://www.amazon.com/Programming-Collective-Intelligence-Bu...
Building Search Applications: Lucene, LingPipe, and Gate http://www.amazon.com/Building-Search-Applications-Lucene-Li...
Play with a database or docs in a filesystem, do deltas of SOLR and sphinx, changing parameters like stopwords, token separators, stemmers, UTF-8 and ISO-Latin to ASCII mappings. See if you can get decent precision/recall metrics. There's quite a few degrees of freedom, depending on the database.
http://www.computationalmedicine.org/challenge/cmcChallengeD...
I downloaded a torrent version, then bought the paperback version straight after.
It was a lot of fun, even if when I started out I was already fairly sure that I would not have the stamina nor the funds to commercialize it but as a learning experience it was great.