New Stackexchange like site for Ubuntu Q&A
ubuntu.shapado.com
ubuntu.shapado.com
It would be great if one of these ended up superseding (at least some parts of) the current Ubuntu forums. Too often the "winning" replies are buried a few pages into the thread, and the information isn't updated as Ubuntu evolves.
Edit: I'm not saying Stack Exchange is better, just that there's going to be a fight!
http://stackoverflow.com/legal/trademark-guidance
Perhaps coincidentally, the private beta of the real Ubuntu Stack Exchange has just begun. The site will open up to the public in 7 days:
You should re-read the title, it says "stackexchange like site", not "stackexchange site". Also this site has the support of many from Canonical, the company that owns Ubuntu.
> Perhaps coincidentally, the private beta of the real Ubuntu Stack Exchange has just begun
But stackexchange is only for English speaking people. How about people from South America like myself or the rest of the world? Shapado supports many languages http://translatewiki.net/wiki/Translating:Shapado/stats and integrates with more services, including Launchpad from Ubuntu. Also, Shapado is open source which is more in the spirit of Ubuntu.
If OP wants to compare one website to his, I think that's fine. He's not an established brand, and this comparison helps me understand what to expect. He's not claiming his site is part of the Stack Exchange network, he's just saying it's similar. If he used any other term of comparison, like "better than" "worse than" "faster than" "slower than" I don't think you'd mind. So what's so bad about using "like" ?
Shapado has the advantage of being localized in many languages so that not only English speakers can ask and answer questions. Also just like Ubuntu, Shapado is 100% free and open source.
They made some very primary design mistakes. Observe the schema of question URLs, for example. They don't contain any numeric identifier, only the normalized contents of the title of the question. This means that:
1. Sooner than later the database indexing of questions from URL contents will become a bottleneck. In the best case, the whole URL has to be hashed and the (non-unique) hash looked up in the database. If it had a unique number, the application can just snap that part of the URL to uniquely retrieve a single entry. Also, if load balancing becomes important in the future, having a short unique identifier easily accessible in the URL is very important.
2. If the title of a question changes, either the URL changes too (which is not good (http://www.w3.org/Provider/Style/URI) and causes problems) or the URL has to stay unchanged with the old contents of the title.
Another issue:
Country flags are displayed along usernames. I can't find the reference here, but this is a touchy subject in some cultures. Some cultures don't like to see flags for other countries, or don't want their flags to appear in a computer screen.
2. We already take care of that, the old urls are still valid, we just make a redirection with permanently moved code.
About flags: People can choose not to show them on their settings page. Also, no big deal.
Back to problem 1, if you use the normalized title as your primary key, you will have to map this key to the tuple, value or document stored in the database. The title of a question is supposed to be a free-form textual redaction written by the user to express his/her inquiry or idea. It can be arbitrarily short or long. You will probably have to set a low limit to it, because you don't want people dumping kilobytes of text in the title, but it also can't be too short, otherwise people won't be able to express themselves. Is the maximum length chosen for the title of a question appropriate also for the primary key of your table?
In order to be a primary key, the database has to compare the entire key against an index of keys present in the table. If all N characters are part of the key, it means that this lookup is O(N) on average (just transversing the key, plus the index lookup itself), because it has to consider the whole string. Compare with O(1) of an unique numeric identifier.
But you then say that you can hash the original key and use the hash result to lookup the index in O(1). Yes, but hashes are not unique. Two different strings may result in the same hash. The index has to consider that any key may contain multiple tuples, in a bucket or in a linked list, or something else. An additional O(N) compare is necessary to recheck that you indeed got the right result or some other result which happens to have the same hash.
But you then say that you will limit the size of the key in the db indexing. Still not good. This doesn't work well since the key comes from the title, and the title is free-form text entered by the user. If you limit to, say, 64 characters, what if two users type two similar but different questions with only the last characters differing?
So, it is much more work either to your app or to your database to use a big key instead of a short unique id. It means that if your app which would be able to handle 1000 reqs/sec with short unique ids, it ends up handling only around 200 reqs/sec using big primary keys (depending on how big is the key, the hashing function used, etc). This is the bottleneck.
If it grows and you hit some hardware limits, and then you think about load balancing, it also means that you will have to implement the same costly algorithms on the load balancers. The load balancers will have to hash your URLs in order to decide to which DB cluster to redirect the request, while if it had a short unique ID it would be extremely easy for the load balancer to perform the distribution based on the that ID.
And then, long URLs. Long URLs are terrible, in many, many ways. While analyzing the sites, I found conflicting results about what Shapado does with long URLs. In shapado.com they seem to be capped, while in ubuntu.shapado.com they are not:
http://shapado.com/questions/besides-shapado-which-other-qa-...
http://ubuntu.shapado.com/questions/what-s-the-most-well-int...
I don't know why the difference (different versions or configuration?). Anyways, you will find that people will copy-paste these URLs into their broken e-mail programs, which will arbitrarily break the URLs in all different lengths varying from 68 to 80 characters. If you use the entire length of the URL as your key, you immediately have a problem. If you decide to consider only the first N characters as your key, then you go back to the other problem above about two questions with only the last characters different. It is much better to have the short unique ID earlier in the URL. In this case, if the latter part is missing, the app uses the ID and either ignore the rest or send a 302 redirect with the correct URL.
Finally, the normalization takes out some characters from the title, most notably punctuation. I haven't tested, but the immediate question is: What if the user asks "どのようにケーキを焼くために?" ?
Now you open a real can of worms. If normalization considers only ASCII alphanumeric, the key generated by this title is empty. If normalization considers Unicode, you it may have to handle different Unicode normal forms. You will have even more issues with the URL being copy-pasted around in different broken applications. It will also overflow faster when percent-encoding the URL, because each CKJ character expands to up to three UTF-8 octets, and each octet expands to three bytes after percent-encoding. So a single "焼" is becomes 9 bytes long in an URL.
Enough of point number one? Next...
2. Then, what happens user asks "What is Unix?" and then he explains the question and it becomes clear that what he really wanted to know was about Linux, and the question is changed to "What is Linux?". (At this point, you will have two different keys pointing to the same document). After that, someone else comes and asks again "What is Unix?", now with a different meaning of "Unix"? What happens? How does it degrade to the point-of-view of the user?
Also, now, knowing how the system works (as you explained), anyone could go there and exploit all the short titles to create permanent redirects to their favorite questions. Just go there and ask "k", then change the question to anything else that is perhaps useful (for a spammer, for example). Now that guy "earned" the /question/k short URL for his product or his biased question.
3. Not allowing two questions with the same title based on your feeling that such cases will not exist is like saying that you will make a persons birth timestamp an unique index because you never saw two people who where born at the same date and the same time. Again, the title is a free-form textual redaction from the user. You may not see with the limited scope that Shapado has today, but what about future uses? What about different languages with a more limited vocabulary that doesn't allow much expressiveness, and question clashes are more common? In fact, it is not necessary to go that far. Lets keep it in technical computing English. For example, one user asks:
"What is ASP?"
Is he asking about Active Server Pages? Or perhaps Auxiliary Storage Pool? Maybe Application Service Provider or even AppleTalk Session Protocol? Ok, you can edit the question to reflect one of these many meanings, but it is still an unnecessary burden on the users. Again, how does it degrade if users hit this limitation?The ultimate question is: do you really need to create such an artificial limitation in your application, with so little information about its real-world uses (since the project is very recent)?
Well, that is it.
Please, don't get me wrong. It is great to see such a project being born. I'm always happy to see the open-source community growing and evolving. Take this as a constructive criticism. All these points raised above are really things that should have been considered in the design phase of the project, not after it is deployed.
Think about what happens to "I saw him beating that assassin" if an application decided to cap the string to 26 characters to fit an URL, for example.
Breaking between words reduces the chances of these accidents happening. Sometimes it is not possible (the title is a single big word, or in a language without spaces), in which cases you may want to degrade to the normal slicing.