BlaBlaMeter detects how much bullshit is in your text
blablameter.com
blablameter.com
>Bullshit Index :0.26 Your text shows some indications of 'bullshit'
some indications ... I'd say this thing is broken.
I tried some Hegel, who scored 0.18. Clearly broken.
Schopenhauer got it right, where BlaBlaMeter gets it wrong:
"If I were to say that the so-called philosophy of this fellow Hegel is a colossal piece of mystification which will yet provide posterity with an inexhaustible theme for laughter at our times, that it is a pseudo-philosophy paralyzing all mental powers, stifling all real thinking, and, by the most outrageous misuse of language, putting in its place the hollowest, most senseless, thoughtless, and, as is confirmed by its success, most stupefying verbiage, I should be quite right. Further, if I were to say that this summus philosophus ... scribbled nonsense quite unlike any mortal before him, so that whoever could read his most eulogized work, the so-called Phenomenology of the Mind, without feeling as if he were in a madhouse, would qualify as an inmate for Bedlam, I should be no less right."
Finnegan's Wake, Chapter 2, Book 1.
> Bullshit Index :0.11 Your text shows only a few indications of 'bullshit'-English.
definitely broken
Or perhaps it's a feature and it can pick out philosophical / artistic 'bullshit' from PR 'bullshit'.
Quote for people not familiar with 'Finnegan's Wake'
> And aroud the lawn the rann it rann and this is the rann that Hosty made. Spoken. Boyles and Cahills, Skerretts and Pritchards, viersefied and piersified may the treeth we tale of live in stoney. Here line the refrains of. Some vote him Vike, some mote him Mike, some dub him Llyn and Phin while others hail him Lug Bug Dan Lop, Lex, Lax, Gunne or Guinn. Some apt him Arth, some bapt him Barth, Coll, Noll, Soll, Will, Weel, Wall but I parse him Persse O'Reilly else he's called no name at all. To- gether. Arrah, leave it to Hosty, frosty Hosty, leave it to Hosty for he's the mann to rhyme the rann, the rann, the rann, the king of all ranns. Have you here? (Some ha) Have we where? (Some hant) Have you hered? (Others do) Have we whered (Others dont) It's cumming, it's brumming! The clip, the clop! (All cla) Glass crash. The(klikkaklakkaklaskaklopatzklatschabattacreppycrotty- graddaghsemmihsammihnouithappluddyappladdypkonpkot!)
Anyway, I'd assume it's looking for certain 'filler' words and phrases that get used a lot when writing BS.
We help businesses increase profitability by helping develop a strategic market approach and communicate the right message.
Our approach is designed for those serious and committed to looking inward so they can connect outward for greater results in an expedient manner.
We analyze your business goals, assess your marketing needs to accomplish them, create a strategic plan and manage its implementation, leaving owners and managers to tend to other needs of their business.
Excerpt is cited as fair use, namely the educational value in evaluating a tool with three sentences of buzzword heavy english from a real site. Specific citation info omitted only so they won't be embarrassed, but may be found with Google.
I've seen users getting a lower score simply by separating out a block of text into numbered paragraphs which would seem to point to quite a simplistic method.
http://ipdraughts.wordpress.com/2012/08/25/cutting-down-on-t...
Or just esthetic rules + word dictionary.
But I'm starting to think a rule-based lexicon isn't out of the question, given these >1 scores on some texts.
The text was:
"Möchten Sie SAP-Software vor Ort installieren oder über die Cloud darauf zugreifen? Wir bieten in jedem Fall umfassende Services, zugeschnitten auf Ihre individuellen Anforderungen. Wir verfügen über eines der größten Expertenteams weltweit. Unsere qualifizierten Mitarbeiter beraten Sie gern bei der Konzeptionierung, Implementierung und Optimierung Ihrer Systemlandschaft. Profitieren Sie innerhalb kürzester Zeit von Ihrer SAP-Lösung."
It my old pet project; like BlaBlaMeter, but for licenses.
I still hope to experiment with the idea of improving licenses in future.
Bullshit Index :0.12
Your text shows only a few indications of 'bullshit'-English.
I get the feeling this tool is only testing for a very narrow definition of bullshit.Your text: 847 characters, 145 words Bullshit Index :0.56 Something's fishy. Obviously you want to sell something, or you're trying to impress somebody. Are you sure that you have a real message, and if so: who would understand it?
It appears our internal bullshit meters work just fine as well :)
Lincoln (Gettysburg): 0.09 Lincoln (Second Inaugural): 0.09 MLK (I Have A Dream): 0.08 (Lowest of any Political text I tried) Someone else said Churchill got a 0.08, but I didn't actually try the unaltered text for myself.
BUT
John Donne: 0
My thinking is you are measuring word count versus commonly used marketing or political jargon count, but that's probably too simple.
- It uses a unigram language model. You can take the same text, randomly permute the words, and you get the same score. This means it also can't be using things like POS tagging, phrases, etc.
- It normalizes words by making all letters lowercase. The exact same text in all upper case has the same score.
- The score is eventually normalized by the length of the text. The same text copied multiple times gets the same score.
- It does not form a valid probability distribution, as someone's managed to get some 1.16's. This makes me believe it's not a Naive Bayes classifier giving you the P(Bullshit|Text). Though this is what I originally thought it would be.
For example, if you take the score from the Oracle Pricing blurb posted by BitMistro, and change 'strategies' to 'goals', you drop down to 0.8 or so. If you add an extra random 'strategy' somewhere, it bumps up to 1.4 or something.
I actually suspect a bug on their pair for strategies... probably a decimal error when building to BS level tables.
But similar things happen with other 'bullshity' words, just to a lesser degree.
But it's not QUITE a true lexicon, as it handles Out-Of-Vocabulary words quite strangely. If you use as input text:
"PR-Experts, politicians, ad writers or scientists need to be strong here! BlaBlaMeter unmasks without mercy how much bullshit hides in any text. A useful tool for everyone involved in writing! Simply copy your text into the white field and check your writing style. It works with english text up to 15.000 characters (overhead will be cut off). For a meaningful result we recommend a minimum length of 5 sentences."
Then you get 0.16. If you replace the last word 'sentences' with 'strategy' you go up to 0.44. However, if you change the last word to 'sentstrategyences' you get 0.47. Try it: you can basically insert 'strategy' inside ANY word and really raise your score. Actually, if you just insert "strateg" anywhere inside the text, it goes up massively.
So I actually think it's just doing string search counts over a lexicon.
"Politics are great, come buy our new, brand spanking awesome banana phone, apple, steve jobs, cripplingly epic banana phone. Just great phones, with bananas, no apples to be found here. Samsung can suck on our banana phone. Android is better than iOS."
"Your text: 251 characters, 43 words Bullshit Index :0.03 Your text shows no or marginal indications of 'bullshit'-English."
Edit: oh its happening today - I thought it was old news and no-one had told me. Apologies for the exclamation marks now gone.
The actual content parts hover around 0.16 (although I only tried segments without any equations). It makes me wonder what sort of results you'd get if you plotted this BS metric vs. page in a full length book.
1. "You probably want to sell something, or you're trying to
impress somebody. It still may be an acceptable result
for a scientific text."And got scores of 0.35-0.45
TechCrunch: 0.42 ("Amazon Wants Everyone To Know The Kindle Fire Is Sold Out")
Paul Ryan's RNC speech: 0.14
NYTimes: 0.11 ("Storm’s Winds Slow as It Exits Southern Louisiana")
Obviously its an apples and oranges comparison here but interesting (to me anyway) nonethelesshttp://news.ycombinator.com/item?id=4270768
got
"Your text: 8044 characters, 1160 words Bullshit Index :0.17 Your text shows only a few indications of 'bullshit'-English."
I too would like to know what the model is for the online ratings. Has there been a validation study of the model?
AFTER EDIT: Paul Graham's essay "Why Nerds Are Unpopular"
http://www.paulgraham.com/nerds.html
(which is the first writing of his that I ever read) appears to reach the maximum length (by character count) that the program will evaluate, and comes out like this:
"Your text: 15000 characters, 2698 words Bullshit Index :0.09 Your text shows no or marginal indications of 'bullshit'-English."
Maybe Paul's procedure of having friends look over his essays and give suggestions helps cut out the bullshit.
http://martinfowler.com/bliki/SnowflakeServer.html
> Bullshit Index: 0.31 – Your text shows indications of 'bullshit'-English. It's still ok for PR or advertising purposes, but more critical audiences may be skeptical.
http://martinfowler.com/bliki/PhoenixServer.html
> Bullshit Index: 0.4 – Something's getting a bit fishy. You probably want to sell something, or you're trying to impress somebody. It still may be an acceptable result for a scientific text.
---
Your text: 130 characters, 28 words Bullshit Index :0.05 Your text shows no or marginal indications of 'bullshit'-English.
---
Your text: 129 characters, 29 words Bullshit Index :0 Your text shows no or marginal indications of 'bullshit'-English.
BlaBlaMeter prefers an extra word over a comma. Makes sense given the goal here.
> For a meaningful result we recommend a minimum length of 5 sentences.
The result is that non-German speakers are standing in the dark, wondering what methodology is behind that software.
It might be even difficult to translate the concept to English, as some terms and ideas like http://de.wikipedia.org/wiki/Nominalstil are missing in the (restricted) English language. German is so much more elaborated, its a language of philosophers ;-)
High-quality journalistic texts are usually in the range of 0.1 and 0.3
* Are there any texts with an indexof 0.0?
This is rare, but occurs. Too low of an index is rather suspicious though, and might also indicate stilistical deficits.
* What is the highest index value?
Thile the BlaBla Meter was designed for index values between 0 and 1, but in rare cases the scale can be exceeded. The highest measured values so far have been over 2.0 even.
* Does Google use a similar algorithm?
We don't know, but if we were Google, we'd probably try! Our samples for highly competitive search terms show that highly ranked sites often show good index values.
* How does BlaBla Meter work?
BlaBla meter checks texts for various linguistic traits, e.g. it checks for exceeding use of Nominalstil[1]. In addition, the text is checked for various phrases [buzzwords] with a certain weighting. We don't want to reveal the secrets though ;)
[1] Heavy use of "nominal style" is a German particularity. You can replace most verbs by nouns, adding suffixes like '-ion' or '-age', and transform the sentence into passive. This is often found in legal texts, where the author doesn't want to appear as a suject.
* Why does my scientific text have such a high index?
Nominalstil has crept into scientific language at large. This often serves to 'beat around the bush' - in this regard, there are analogies to typical PR tongue.
* My mindless text gets a good score, why?
BlaBla Meter can't really understand the topic, it looks after linguistic traits. While rainbow press can be attributed all kinds of substantial weaknesses, their language is usually short and concise.
* My text is wrongly devalued!
As every computer algorithm, BlaBla Meter might be at fault - in doubt, human beings must decide of course. Generally the hit rate is quite large, so you should usually tweak the wording upon receiving high indices.
* But humans read texts completely differently ...
When BlaBla Meter crys havoc, one can usually assume that a human reader is alerted too. People are trained every day to tell authentic messages from artifical advertising messages - it'd be naive to belive they couldn't!
* What happens with entered texts?
Texts are private matter of course, and are neither saved nor processed in any way or shape.
Yeah, personally I also score around .16 to .18 on my internet texts. Long papers get around 0.2, so I'm probably losing momentum after the fourth page :-/
Bullshit Index: 0.46 Something's getting a bit fishy. You probably want to sell something, or you're trying to impress somebody. It still may be an acceptable result for a scientific text.
Pretty good figuring it's under the 5 sentence recommendation.
This is not a useful tool because it does not define or point to the problems in the text.
"This project is a blue sky implementation of a cloud based virtual paradigm which really shifts the boundries of our connected world.
Using patent pending technology to accelerate your business, we use full-stack web-scale systems built on noSQL databases to drive clicks.
And your security is safe with us. We use military grade encryption to protect all your data at rest."
"Bullshit Index :0.64 This reeks. I bet you're a PR-Expert, Politician, Consultant or Scientist. If there is a message, it's unlikely it will reach anyone."
I put in a page or so from a scientific paper I recently wrote and got 0.19.
Note: I'm a scientist too ... good to poke fun at ourselves, eh? :)
0.8 This reeks. We bet you're a PR-Expert, Politician, Consultant or Scientist. If there is a message, it's unlikely it will reach anyone. Maybe you should spend less effort on trying to impress somebody.
http://news.ycombinator.com/item?id=4451328 - full text http://news.ycombinator.com/item?id=4451586 - summary text
4. Small-Business tax breaks - [major bullshit] 0.5 [a]
1. Space Race - [PR level bullshit] 0.33
7. Money in Politics - [Some bullshit] 0.28
2. Internet Freedom - [some bullshit] 0.23
10. Work/Life Balance - [little bullshit] 0.16
9. Jobs / recent grads - [little bullshit] 0.16
6. Most Difficult Decision/afghanistan - [No Bullshit] 0.07
[Un-analyzable]
3. Favorite Basketball Player - too short
8. White House Beer Recipe - too short
5. First Activity on Nov 7 - too short
[a]"Are you sure that you have a real message, and if so:
who would understand it?"Your text: 15000 characters, 2645 words Bullshit Index :0.13 Your text shows only a few indications of 'bullshit'-English.
Does that mean the algorithms work?
Your text: 2418 characters, 451 words
Bullshit Index :0.05
Your text shows no or marginal indications of
'bullshit'-English.