Show HN: HackerBooks.com
hackerbooks.com
hackerbooks.com
Just last week I had two people email me with similar site ideas to hn-books.com, and when I launched hn-books 2 completely other people emailed me that they were working on similar sites.
Must be something about hacker books that's in the water. Lots of sites with lots of features and such means better resources for all hackers, plus lots of folks getting experiences doing stuff like this. Most excellent!
Since I've done this, I guess I should say something pithy or insightful. I think the trick is the navigation piece. I see you have categories -- that's probably a good place to start. I started with questions, you can go to my main page, click on the hacker-related question you have, and be presented with a sorted list of answers based on your experience. (see http://hn-books.com/ )
I'm not sure if questions are the way to go either, though, as there are a zillion questions hackers might ask. I'd still like to see somebody come up with some new ideas in this area.
I also ran into the "just what are hacker books, anyway?" question that we get over here all of the time. I finally said screw it, I'll just put things that I believe are hacker-related. But I don't think there's any easy answer to that one, either.
Outstanding site, though! I hope some of these other guys that have spoken to me will post what they've done as well.
thanks for your feedback! Your release of hn-books.com really made me realize I should carry on, that the topic was interesting (which I believed). But I was quite far from being able to ship and busy with a lot of client work, too.
I totally agree that the navigation is the tricky part, and I have a lot of work to do on this :)
I'm really curious to see other people approach to solve the general issue of finding useful books!
In all cases, thanks for the kind words, appreciated!
Cool work!
EDIT: FWIW, I enjoyed your site and our thread so much that I took a bit of time today and blogged about all book sites that are popping up. http://news.ycombinator.com/item?id=2250382
- fragment caching with redis/rails 3 (with some warmup script)
- the full search is handled by solr
- nginx and passenger are quite fast
I still need to make some work to compress the js/css together. The images themselves are served by Amazon.
On the size of the list: well I really have a bias toward "data aggregation". My two previous side-projects (http://www.learnivore.com and http://www.toutpourmonipad.com/) were already "ETL"-based.
I have another project where I will manually curate the list, though...
I think both approaches (data-driven or manually curated) have their pros and cons...
Glad you like it in all cases!
I'm not surprised. I've been working on my own half-baked version. Kudos to those of you who finished the baking.
Both of the sites still need to address the question, 'how do I present the categories without requiring too much navigation'. I know it's not sexy any more but how about tag-cloud per major category?
What I found is that I keep having "feature fever", where I think of all these ways of doing stuff that might or might not have value to the user.
The site database is actually set up for voting -- you can vote up and down books and create your own rankings according to your opinion. But it was another feature I had to kill. I find the hardest thing about building a site is killing stuff. It took me forever to strip out as much as I could from hn-books. Right now I'm a bit shy of starting to put stuff back in. Over and over again it looks like whatever I add -- reviews for tools, a link to get your own Kindle, a new question category, live real-time commenting with social links -- all seems to meet with some resistance from the community. I'm never sure if I'm hurting the community or just making artistic decisions that some small percentage of people don't like.
I really like hackerbooks with the simple categories and multiple books above the fold. One of the emails I got recently was from a guy who set up a "bookshelf" with books on it -- an image of a real bookshelf with what looked like real hacker books sitting on it. That was pretty cool too.
When I set the site up, I consulted with a few HNers about this very issue -- how to sort/filter. Like you, I'm still not happy with where it is right now.
So I'm open-minded, I'm just keeping the bar a bit high for any changes right now. One of the most difficult things I am learning is to separate what might be cool from what might be best. Still working on it. If you guys keep asking for tag clouds, happy to put them in.
EDIT: Bit of trivia, if you hold down "ctrl" and click on questions in hn-books you can apply multiple filters. So say you only had time to read 1 book, but wanted to know as much as you could about "how do I make beautiful web pages" and "How do I tell people about my business". You just ctrl-click on both, then get a list of books that covers both questions: http://hn-books.com/#BC=0&EC=0&FC=0&Q0=1&Q1=... Not sure tags would work that way by default (Also this is probably a great example of something I spent time putting in that means absolutely nothing to anybody)
What's your opinion on the later ?
[1] http://seatgeek.com/blog/dev/announcing-soulmate-a-redis-bac...
Actually I got that idea to try out experimental features under a /labs area. So I will give various features a go there. I can try out tags :)
It's fairly simple so far and more features are planned.
I'm submitting early on to get some wider feedback. Thanks to all the HNers that reviewed this before today already!
Next iteration will be on making the navigation better, so that you can click on everything that's possible.
I started the crawler behind HackerBooks with Resque and Redis, but ended up saving some RAM by going back to a simple daemon. I reused Redis to provide the caching.
So well: I mostly used what I was used too.
In the end I paid 90$ for 3 years, which is OK I guess.
Or was it the opposite?
Wish us luck :) (and communication, it's needed :-)
Another interesting idea maybe would be to map the annotations/bibliographies of all books, so I could start with books I've read and better see what kind of linkings exist. You could visually see the seminal works in a field and all the branchings. Maybe that already exists in some form?
That' what I learned from working on http://www.learnivore.com too - at the same time came out http://rubytu.be/, but I used both actually, and some people preferred one, some other the other.
I encourage you to ship anyway :)
The idea was founded on the tendency that people are skill-biased when rating books. They might dislike a book for being confusing or too easy because it wasn't designed for them at the time of reading. This was an attempt to figure out which books were just bad, and which were only rated bad because they were read without proper experience (or too much).
Ultimately, using community data, a visitor would be able to discover a "bookpath" of great resources that syncs up with their level of skill.
I wanted to create a community-edited wiki page providing a "training path" which would involve screencasts, books, articles.
I think it would really be useful actually.
The wiki approach would be great if the data was laid out appropriately. Community-edited wiki might lead to a singular view of what a good path would be.
The benefit of a multi-axis rating system is that you could lay out data using different permutations and add/subtract inputs (e.g. rater's experience at time of rating/age/hell, maybe even personality type). I'm sort of modeling this off of robust scientific questionnaires.
It does seem to require some more user involvement, but nothing a fantastic UI couldn't fix.
I would love to be able to build something like you describe, but I suspect it will be time consuming, that said!
I will probably come up with some kinds of list, and will definitely work on better sorting/filtering.
One thing that would be really handy - or at least interesting - is a "most recently mentioned" list. For example, when people were talking a lot about Program or be Programmed a while back, it would have been fun to see that rise to the top.
The most recently mentioned list is a very good idea. I had something similar in mind, like a news-letter that would send the "most mentioned this month", so you can get the trends.
Would you find that useful ?
My pet peeve is allowing to find books quoted by X, where X is some instance of someone I appreciated on HN :)
I think I will create a labs section with various experimentations like these, so people can try them out and see if it's useful.
Being able to select your amazon store is planned as well (I'm in France so I totally understand your point :-)).
I choosed to ship without that though, to see if people like the site first or not. It will require a bit of work underneath to do well, things such as verify in the background if a book is actually available on amazon.co.uk, .fr etc to avoid sending the person to the wrong place.
In this respect you'll see new Amazon links every so often in: http://twitter.com/hackerlinks (the tweet will begin with Amazon) http://hackerbra.in/links
This site is really a good idea. The Amazon links can get quite popular (I know from looking at my Amazon stats on @hackerlinks when I had the affiliate code inserted.)
I didn't know about hackerlinks either - just subscribed!
Good luck!
Perhaps I should make an Amazon only Hackerlinks-style feed! Or you could have one coming from your site of new books as they're added!
I wish I'd thought of it.
Congrats!
For example http://www.hackerbooks.com/book/rails-for-net-developers-fac...
Should link to http://pragprog.com/titles/cerailn/rails-for-net-developers instead of Amazon.com.
I'll try to make more "publisher-specific" linking.
I dont know if there is an API for it or not but if you can get the Amazon rating of a book and display it on the book description page, it will be great. :)
On AZ ratings: I wrote that down. It's somewhat complicated because Amazon just made it a bit harder to embed that. It now has to be an iframe; the iframe url must be refreshed every 24 hours.
But overall this should be doable to, I'll see if I can make it usable.
Thank you!
edit: not a mix, but separate points for each community
I think being able to filter on SO vs HN would definitely give different results, too.
The topics are sometimes fairly different on both sites.
Two conclusions:
- how the search works needs some explanation a bit!
- I will probably add an advanced search so you can specify exactly how you want the search to occur (eg: ignore the description....)
Thanks for your feedback!
I plan to write a couple of blog posts explaining the "how" later on! It's been an interesting ride really.
But a large part of it (the 2/3rds) was a learning exercise around chef and vagrant, which wasn't necessary to the project.
Of the remaining third, I've got around 70% for data processing in general and 30% on pure front-end code and design.
I really wanted to learn how to deal with sysadmin in a more productive fashion, so I took the plunge :)
I always felt that doing it manually (even using well-written notes) then gradually home-baked tools was a loss of time.
I looked at chef more than a couple of times, waiting for the documentation to be more available, and for feedback from people I know.
I started using chef with the opscode platform, then went back to chef-solo as it really fits my needs already.
I'm using it for client work as well as for everything behind HackerBooks (including Rails app deployment without capistrano anymore).
The consequence is that I can boot a new ubuntu instance from scratch, completely configured with the whole stack (rvm, rails 3, passenger, nginx, solr, god, the properly configured crawler, data restored from a S3-like etc) in less than 15 minutes.
I will never go back to manual sysadmin (apart from small tips) - this really fits my way of working.
But it has been a time-sink to get in :)
Hope I replied to your question properly, if I didn't, ask again!
I plan to display the actual conversation in some way on the site itself, if I can. Lot of people told me it would be useful.
I will also work on some topic extraction, yes!
I searched PHP and would like to see by # of recommendations...
It's currently sorted (in that order) by 1) textual relevance and second 2) number of quotes.
I know for sure it can be done with some configuration tweaking, not even sure I'll need to reindex.
Thanks!
Keep the good work. I love it :)
My only suggestion is to offer a ranking and order the results by # of quotes
On parsing: it's actually a somewhat fastidious process that involve digesting a couple of GB of data, but here is the bottom line.
I look for amazon.com links in the content in general - I will broaden that to other publishers and full-text extraction too later on.
The content itself comes from the StackOverflow dump (for SO) and a mixture of a crawler allowed by PG + the previous database dump that was available at some point.
I extract all the books, quotes, users data from both, conform these into a common schema, and index the whole result.
Hope I answered your question properly - feel free to ask again if you'd wish.
On ranking: I know what you mean! I need to find some way to balance number of quotes with textual relevance, which requires me to dive a bit more into solr. I currently use textual relevance first because it gives more useful results so far.
For ranking, instead of textual relevance (which will be hard to achieve :/) and # of quotes (it can be easily hacked/spammed or new announced books will have a huge weight on the ranking), I suggest you to check timelines: if a book is quoted once/twice/... a month regularly, I'm pretty sure it worths reading it
I'll love to read about the architecture behind the site... yep, technically curious :)
On the architecture: I'll create a side-blog that will outline all I learned while working on this. It's been a crazy ride actually (especially because I started using chef and vagrant full speed).
I'll post it back here in all cases.
I really love what you got here. I'd be happy to help you try out IndexTank and make it better. It would really take the Solr configuration burden off of you.
well I considered using IndexTank earlier on, especially because I didn't know yet how to deploy Solr. The relevance is mostly done already, I just needed to learn to use formulas :)
One thing that put me off is your cap in queries per day. The smallest paid plan (50k items in index) is capped at 1,000 queries per day.
Isn't that an issue for most sites ? Do people usually cache your results ?
I can't really tell yet, but my guess is that at least around the number of indexed documents could give a more usable subscription.
It would also be nice for people to know what happens if you go beyond the cap: do you offer some tolerance ?
Questions: Apart from the "Quoted by" section, is any of the content from SO or HN displayed on the site? E.g. are any comments incorporated in the book descriptions, or are those written by you/wife?
Currently, no content from HN/SO is displayed apart from the Quoted by area.
We're not editing anything manually; what I will do is display the actual conversations in the "Quoted by" area, either when you click on a conversation.
I may ove the quotes above to make them stand out more.
Did I properly answer your question ?
Was justing wondering if the content might be touched by any copyright issues etc.
We'll see how it goes.
So I'm currently working on balancing textual relevance (eg: Ruby) with the number of quotes.
Thanks for your feedback!
Do you mean you find the sites look similar ?
I wanted to keep the simplicity of Google if I could (hence the input/submit search). Otherwise no strict inspiration from any existing site afaik.
EDIT: I did a closer analysis out of curiosity and the common points are:
- the dual-color name
- a google-y search input+button
- the tagline with the number of items
- a similar green
Note that I definitely didn't use iconfinder as a model. I guess these are quite common characteristics.
Thank you for mentioning this!
I'm not aware of that, can you give more details ?
There's extreme misuse of the space on the page. When I was looking at Code Complete (a book I've been trying to get my hands on for a while), there is very little content about the book. The synopsis is cut off (!) and there are no reviews. But if you were trying to save space, why on the hells are there over 9000 books following in "quoted discussions"? You need to switch what you're truncating here. Also, I would suggest at least copying Amazon's ratings for some measure of book quality.
For reference: http://www.hackerbooks.com/book/code-complete-a-practical-ha...
Here's what's planned:
- instead of cutting the book description like it's done currently, you'll be able to expand with a click
- I'll do the same on books and quotes, because hundreds of books are not useful in the "suggested books list", definitely
- and I want to focus more on getting the actual quotes to the front (eg: allow to read the actual quotes by HNers etc) because I feel it brings significant usefulness
I'm mixed on the Amazon ratings, both because recent API changes made it unpractical to really use the content (it's now an iframe that must be used as is, and refreshed every 24 hours), and because sometimes the reviews are fake as well.
So I'll try to bring more value by letting users know what people think on HN and SO.
Would these points make the site more useful to you ?
On the synopsis: well in some (too frequent) cases, the synopsis was just several pages long, so the related books and quotes were really, really hidden.
I need to find a better middle ground.
But in all cases, thanks for your critics, it will only help me make the site better.