HNHacker News
TopNewBestAskShowJobs

hftf

53 karma · joined May 6, 2014

submissionscomments
hftf··on The Small World of English
I don't think definitions "are" highly accurate precise things. Sometimes yes. The same scholarship, skill, and need to not mislead also applies for so many other things: encyclopedic articles, taxonomies, news, maps, operating systems. Do people still question the value of Wikipedia, OpenStreetMap? Yeah, there are problems with them, and with peer review. Using fuzzy words (or fuzzy phonetic symbols, fuzzy categories, fuzzy semantic links…) to define words is a problem (if at all) of literally any dictionary. I don't see any of these as particularly unique obstacles for Wiktionary.

Unabridged dictionaries take decades to release new editions and are still navigating transition into the exploding digital age. They are so expansive in scope, while often so limited in resources, and barely accept any crowd contributions. Such deliberately slow-going is often a good thing, but words also change quite quickly and these sources are now playing a very long game of catch-up. (Yesterday I tried to verify the latter English senses of "fandango" on Wiktionary with other dictionaries; OED's entry has not been touched for 131 years! What am I going to do with that, I need to use / understand the word now!)

Wiktionary is the big web-native word-resource (and is not cluttered with commercial junk) – allowing links, expandable quotes, images, diagrams, etc. that print's minimalism suffers from as you mention. When someone in 2025 wants information on a word, they'll likely use a search engine and click a link to Wiktionary (where Google blurbs steal some data from). Maybe they are a student wanting to confirm their nonstandard pronunciation with the IPA (still rarely used in mainstream English dictionaries) or if it's recognized in their own dialect (mainstream dictionaries rarely provide more than UK and US pronunciations) – if enough people have the same question, Wiktionary seems like the best place to put the answer – or see an accessible etymology tree. While you probably know this, it's also worth reminding that English Wiktionary isn't just for English words, it is a dictionary of all languages' words, which is written in English. It has metadata and links connecting languages' words that you can't find elsewhere.

Yes, I indeed do want people to just write what they think a word means – as a starting point in a collaborative refining process. I believe the number of word-users in the world with valuable potential contributions is a lot closer to a billion than the thousand gatekeepers working hard on classical dictionaries. The barrier to entry is really low, but the tooling could still be much better. This is one reason i'm putting my appeal under this article - because I think (professional) lexicography can stand to evolve more in the 21st century. (And are people today really buying enough dictionaries to sustain a professional version of Wiktionary, or even a professional dictionary offered in structured data form?) If we don't contribute to a crowdsourced dictionary, then we won't have any such thing.

(Meta-lookup sites are link/search engines, not dictionaries and IME really don't do a good job synthesizing their information or conventions.)

hftf··on The Small World of English
I get the impulse to assume they'd be alike, but I've found that Wiktionary really isn't much like Wikipedia.
hftf··on The Small World of English
I really enjoyed the article, reading it more from the perspective of what 21st-century lexicography could be, less as a customer of a word game however thoughtfully designed. As a Wiktionary editor (and Android user who's also grown out of bare word-relationship puzzle games) though, it's sad that there seems to be no way to just use the end-product network as a reference, which I would love to do, but I suppose they did spend a million bucks on it.

I'll also use this post to wish that more people would edit Wiktionary. It has such a good mission (information on all words) and yet there are only like 80 people editing on any given day or whatever. In some languages, it's even the best or most updated dictionary available. The barriers to entry and bureaucracy are really not high for HN audience types.

hftf··on Ask HN: Why does Pinterest dominate Google text search results?
Some relevant prior discussions:

https://news.ycombinator.com/item?id=21622322 (Nov 2019) "Tell HN: Google should drop Quora from search results" 1000+ upvotes

https://news.ycombinator.com/item?id=16613996 (Mar 2018) “Pinterest needs to be removed from Google IMO” 1100+ upvotes

https://news.ycombinator.com/item?id=16388833 (Feb 2018)

And many more: https://www.google.com/search?q=site%3Anews.ycombinator.com+...

hftf··on Tell HN: Google should drop Quora from search results
I entirely agree.

Quora and Pinterest are particularly routine spam sites in my search results.

They rank just below word reference site spam, like dictionaries, thesauruses, or translation dictionaries (sites which I do benefit occasionally from), and below Wikipedia mirrors (which I feel has become so bad that I can't even get legitimate results talking about the problem itself! Try searching something like: search results spam wikipedia mirror "revolvy" "wikiwand").

But for me, the worst (and most obvious!) offenders by far are "pronunciation guide" spam sites. Just a few examples:

  howtopronounce.com
  howtopronounce.co.in
  pronouncekiwi.com
  pronouncenames.com
  pronunciationof.com
  rightpronunciation.com
plus the scourge of 16-second YouTube videos on channels with names like Pronunciation Guide or Emma Saying.

(If you search for something like "Deidesheimer pronunciation" or "pronounce Canynge" on Google, the vast majority of results will be those spam sites, plus maybe an ancient forum thread from 2004 that veered off topic before anyone even tried to give a serious yet uninformed answer.)

These ad-infested spam sites purport to teach you how to pronounce an unfamiliar name or tricky word (an important and underappreciated service that many people use!). But usually they merely contain computer-generated bullshit, as if fed directly into all available text-to-speech algorithms. Even the ostensibly human-generated recordings and sites are often flagrantly wrong, unsourced, and untrustworthy.

There are a few legitimate sites (such as Forvo, Youglish, etc.), but too often they are woefully incomplete (by nature of their being crowdsourced). Forvo even contributes to the spam with "do you know how to pronounce this word?" false positives.

I once blocked all of the spam sites when the domain-blocking feature you mentioned was built into Google Search; then had to do it once again when I needed a browser add-on to replace the removed feature (which naturally only worked on desktop); and recently I was astonished to find that the add-on also stopped working! The spam never ends.

hftf··on The Evolution of Trust
The execution of this visualization was rather disappointing.

I didn’t like the overly cute text (the description of the Simpleton algorithm was almost incomprehensible), the low-contrast captions and colorblind-unfriendly color scheme, and the limited navigation (there was no way to go to the previous slide within a chapter, for example).

But more importantly: If you are going to design an entire interactive exercise like this, graphs are a much better way to explore the effects of varying different parameters. Trying to experiment (as instructed) with different parameters by watching animations in the various chapters and the "sandbox mode" included in this simulation was not only tedious, but prevented effective comparisons. If you just run each iterated tournament (from chapter 4 and onwards) by pressing the "Start" button, there is too much going on simultaneously at a high speed to follow along – I would recommend a sorted table or bar chart rather than many multi-digit numbers arranged in a circle – while stepping through is too slow to keep everything in your head.

I noticed that some of the other "explorable explanations" by the same creator include graphs; I think omitting them from this visualization was a mistake. http://explorableexplanations.com/

hftf··on Show HN: Aeneas – a Python audio/text aligner
I looked into gentle a few weeks ago and did notice that it seems to use an online algorithm. It doesn’t have built-in support for live audio input unfortunately, but it may be tweakable as you say (such as reimplementing it to use audio streams that work with either static or real-time input). I guess there’s no other way to find out than just try it myself.
hftf··on Show HN: Aeneas – a Python audio/text aligner
Do you know of any existing forced alignment tools that work well with live audio (microphone) input? I would like to create a live stream in which the words of a known text are displayed as they are being spoken into a microphone.
hftf··on Show HN: Primitive for macOS
I wonder if it’s possible to turn this into a video filter.
hftf··on Tufte CSS
> screenshots

What help. /s

hftf··on Tufte CSS
I’ve seen Tufte CSS before when I tried to search for HTML/CSS sidenote implementations. I’m happy to see a responsive one, even though it uses JS.
hftf··on A multiplayer Tron-like game with curves
What exactly are these "unique twists"?
hftf··on A multiplayer Tron-like game with curves
This seems to be a laggier, CPU-exhausting shameless copy of Curve Fever: http://curvefever.com/play2.php
hftf··on Smarter Link Underlines For Every Website
I had seen this effect about two years ago on the website of Roman Komarov and was impressed by it at the time: http://kizu.ru/en/fun/
hftf··on How I reverse-engineered Google Docs to play back any document's keystrokes
Has the author made any insights into reverse-engineering Google Docs’ spell checking?
hftf··on Show HN: Hipster Domain Finder
Thanks! — but I hardly think it’s impressive scraping your website using kimono (besides, 40% of the hyphenation data is missing anyway…)
hftf··on Show HN: Hipster Domain Finder
Here is a version of the list grouped by whether the TLD is the same as the last hyphenation point (e.g., crow.bar and not frig.ht).

http://pastie.org/pastes/9147186/text