To be fair, I haven't counted and done the maths (and I understand that human sentiment analysis often doesn't agree with other humans), but.. well, look at these examples:
Positive: Thank you Urbana voter for pointing out what we knew. Cheat lie steal the election. We will NEVER become an Islam nation as Obama wants.
Positive: Fun fact: @kbzeese got 1.5% of the vote in 2006 US Senate General Election, but thinks he knows what the people want #headdesk
Negative: What is our present condition? We have just carried an election on principles fairly stated to the people.
Does your model give confidence? It looks to me like you are making it bi-modal, but if you had a "neither positive nor negative" category it might fix some of the issues?
Eg, this is rates as positive, but I'd rate it as neither positive nor negative: "Should ballot papers in Northern Ireland include photographs of the candidates standing? Make your views known. http://t.co/NySCHsmecE"
Of course tweets are a pretty difficult thing to run sentiment analysis on: those random hashtags break a lot of machine learning models, whilst a human can read them and realise something like "#headdesk" is probably bad (although in that case the phrase "thinks he knows" is something that a model could probably use if it understood n-grams)
Also keep in mind, different topics work better. The term "election" has a lot of news headlines, which many are probably neutral, skewing the results. More consumer-ish topics yield better results. But yes, tweets are difficult to analyze. I've done another recent experiment with tweet analysis, if you like this kind of stuff http://primaryobjects.com/CMS/Article158.aspx