Boring Problems Need Attention Too
wyounas.com
wyounas.com
It's amazingly hard to communicate "variability" in an answer. Whenever I communicate experimental results, my business partners always want to hone in on the simplest explanation (e.g. "treatment effect somewhere between 1 and 3" becomes "treatment effect = 2"). And that's just the uncertainty part of variability. Then there's the kinds of variability you're talking about.
There are a ton of tricks to mitigating this, like reporting confidence intervals instead of point estimates, or giving the right number of digits to communicate the key message, or making super sharable/linkable results to minimize the game of telephone. I'd love someone to write a book on how to communicate this sort of "variability" in a business setting.
Then there is the problem of how insanely innumerate society at large is to the point where something like this doesn't register for huge swaths of people. might work in the context of engineers or people with some higher education but not all contexts.
It turns out that a lot of things that initially seem trivial to precisely define aren't actually that precisely defined, like the length of the california coastline. This is, in my mind--and as a complete tangent--a great argument for wide programming and math education. When you're forced to be so goddamned precise all the time, it's very clear when an idea isn't fully defined.
"ago--never": G docs 1 word vs OpenOffice 2 words.
"tiger-lillies--what": G docs 1 word vs OpenOffice 2 words (IDK what this should "really" be)
"Wanting?--Water": GD 3 words vs OO 2 words
In this case the disagreement springs, exclusively, from if the docs engine believes that double hyphens make a compound word (and potentially handling punctuation in the middle of such a compound word)
The irony is that his title is correct - Boring problems need attention.
But the content is jarringly wrong - The lack of attention didn't lie with the engineers of these companies, but with the author themselves.
For ~1600 words, they could have very easily tried to count it and quickly realized that the answers they got weren't wrong - the question was lazy and lacking attention.
Do you mean dictionary words?
Compound words?
Whitespace separated tokens?
What about stand alone text that isn't a word (ex - page numbers? Title numbers?)
Words that the author invented?
Words use non-traditional-spacing or grammar?(and maybe a missing space)
----
But nope, it's those damn lazy engineers!
i think 3 words. The tiger lily is a type of lily. plural is tiger lillies. Fair to say hypen is stylistic (& after a quick web search , I can't find any other usage like that).
That aside, I can't believe how simplistic Google Docs word recognition/count engine is.
If exact precision is necessary you probably shouldn't rely on imprecise terms like "word".
Look at the waveform of speech and you will see long silent gaps inside "words" as well as there frequently being no gap between "words".
There are phrases like "Skinny Puppy" that can do the same job as a word, there are also structures smaller than words that people smush together to make words. The two even work together:
missile
anti-missile missile
anti-anti-missile missile missile
If you see "words" as the molecules of text there will always be an asymptote you can't overcome because segmenting text into words will sometimes introduce errors that you might not be able to recover from.Except that's simply not true for most eastern languages. Devanagari, Arabic, Bengla, Thai and so many other languages are packed with diacritics and graphemes, conditional rendering based on surrounding letters and much more that force you to completely rethink your approach.
German has very long "words" just because they decided not to add lots of spaces when creating a writing system for it. There is no particular reason why it's "Bundesverfassungsgerich" instead of "Bundes verfassungs gericht".
"I work at the F.B.I. I like it there." "I work at the F.B.I., I like it there." "I work at the F.B.I. I like it there!"
It's not as simple as counting periods.
So that counting words can have corner cases is definitely understandable. Is "&" a word? It is literally just 'e' + 't' superimposed and "et" is definitely a word.
It is also possible that the author is working on it (his Github profile suggests he open-sourced some work in NLP).
I remember taking a typing class some 25 years ago and being told a that a word count is typically every 5 characters. That way someone doesn't pad out their word count by using lots of small words.
That might be only within the context of "words per minute" in typing. I mean otherwise Finnish typists would be terribly slow.
This is why each browser used to parse HTML differently.
This is why you'd have compat or even security issues because some software used \r\n for newlines splitting while other used \n.
Luckily the browser vendors formed WHATWG which created pretty precise specs which are maybe convoluted but at least everyone parses HTML in the same way, and each browser pretends to be every other browser for compatibility.
2021 is really great for web compat, maybe not all browsers implement every API, but existing APIs are accompanied by very thorough test suites (Web Platform Tests): https://github.com/web-platform-tests/wpt
Live results from nightly builds: https://wpt.fyi/results/?label=experimental&label=master&ali...
Having said that I don't see vendors aligning on definition on word count any soon due to corporate inertia, lack of incentives and lack of "Word editors consortium" (or is there any?)
A truly good function for word count might be pretty complex, and perhaps different for every language.
[1]http://johnsalvatier.org/blog/2017/reality-has-a-surprising-...
"I've seen things, you people wouldn't believe"
Also, believing that complex problems are easy is not something only programmers do. I work currently a lot with Excel automation, and most people have no idea of what can be automated easily and what can't. I have some people coming that ask for automating a task they've never done manually and don't really know how to do precisely. I think that's the same mechanism of "overabstraction" that leads to people to say "WET instead of DRY" (Write Everything Twice instead of Don't Repeat Yourself).
Open to ideas from HN too :)
Work with the brokers (and reinsurers if possible) and go after the carriers. Follow the Zillow model instead of Zenefits.
Edit, more:
Businesses want to work with established P&C brokers, and brokers have ensured (in most states) that they have to anyway. Provide the right data-driven SaaS tool that gives brokers the power to sell/renew better. Offer it for free unless it converts but require reporting conversions, then charge a percentage. You could call up brokers and they would all try it out. Then focus on solving boring, mostly statistical, problems in an elegant way behind the scenes.
I feel safe to say that most people consider that boring work for a boring business.
I personally don't need to work on "interesting" stuff.
I can tell when there's a perf/promo cycle coming up at Google because all the core apps on my phone change their UIs and get buggier and slower.