And sometimes, the extant magical belief that "government" is different & immune lets those same human factors be ignored until they feed bigger, slower disasters that everyone is afraid to admit, because (ostensibly) "we all did this together".
30,203 karma · joined February 24, 2007
Twitter: http://twitter.com/gojomo
Project: http://thunkpedia.org
Idea blog: http://memesteading.com
Older blog: http://gojomo.blogspot.com
Homepage: http://xavvy.com (username @ here for email contact)
My HN peeve is formulaic downbeat comments, like: "How is this news, I already knew this!" "…Betteridge's Law…" "I stopped reading at…"
And sometimes, the extant magical belief that "government" is different & immune lets those same human factors be ignored until they feed bigger, slower disasters that everyone is afraid to admit, because (ostensibly) "we all did this together".
Sure, that works – like having one (rare, expensive) savant engineer apply & review everything in a linear canonical order. But that's not as competitive & scalable as flows more tolerant of many independent coders/agents.
It'll fire on merge issues that aren't code problems under a smarter merge, while also missing all the things that merge OK but introduce deeper issues.
Post-merge syntax checks are better for that purpose.
And imminently: agent-based sanity-checks of preserved intent – operating on a logically-whole result file, without merge-tool cruft. Perhaps at higher intensity when line-overlaps – or even more-meaningful hints of cross-purposes – are present.
Can't achieve subject-verb agreement in 1st sentence of their English abstract.
Advances made through No Language Left Behind (NLLB) have demonstrated that high-quality machine translation (MT) scale to 200 languages.
https://somethingaboutmaps.wordpress.com/2011/03/08/remember...
– 'SLOW TUESDAY NIGHT', a 2600 word sci-fi short story about life in an incredibly accelerated world, by R.A. Lafferty in 1965
https://www.baen.com/Chapters/9781618249203/9781618249203___...
Signal and other messaging apps offer a 'search' bar across all sessions & history, so I doubt I'm the only one.
It's hard for me to imagine being so present-focused such a history wouldn't be personally useful.
Or, so worried about "someone [using] it against [me] in court" that I'd need more than the occasional auto-expiration, and specifically my messenger "protecting" me with intermittently-enforced loss-of-histories (on just theft/loss/hard-failure of primary device).
Anecdotes that sometimes those problems don't occur are nearly worthless. Of course that's true - the original anecdotal complaint already implicitly relies on, & grants, the idea that there's some default, "hoped for" ideal from which their experience has fallen short.
To chime in, "never had your problems" thus adds no info. Yes, people lucky enough not to hit those Signal limits that cause others to lose data exist, of course. But how does that testimony help those with problems? Should their frustration be considered less important or credible, because of your luck?
The as-if portrayal is one way your anecdote will be perceived, even if that wasn't your intent.
I was born in 1970; per your reference, there've been a bunch of state & federal legislators (or recently-former legislators) killed for political (or pseudo-political deranged) motives "in my lifetime" – and far more in the 1970s than in the last 10 years.
In my lifetime, one sitting President was shot at & missed (Ford in 1976), and one was shot at & hit by a ricochet (Reagan in 1981) – again, more in the past than the shots that grazed candidate Trump in 2024.
The Wikipedia-listed murders of other officeholders, like mayors or judges, are also more frequent in the past than recently – especially going before either of our lifetimes.
So trend impressions are very subject to frames of reference & familiarity with history.
I suspect if people in general had a deeper & broader sense of how common political violence has been, both in US history & worldwide, they'd be, on the one hand, less prone to panic over recent events & rhetoric (even though it is concerning), but also on the other hand more appreciative of the relative peace of recent decades (even with the last few years' events).
I don't believe this is true with regard to ending angles after addition steps between vectors of varying magnitudes.
Imagine just in 2D: vector A at 90° & magnitude 1.0, vector B at 0° & magnitude 0.5, and vector B' at 0° but normalized to magnitude 1.0.
The vectors (A+B) and (A+B') will be at both different magnitudes and different directions.
Thus, cossim(A,(A+B')) will be notably less than cossim(A,(A+B)), and more generally, if imagining the whole unit circles as filled with candidate nearest-neighbors, (A+B) and (A+B') may have notably different ranked lists of cosine-similarity nearest-neighbors.
So of those 3, despite the superficially "large" distances, 2 of the 3 are just as good at this particular analogy as Google's 2013 word2vec vectors, in that 'queen' is the closest word to the target, when query-words ('king', 'woman', 'man') are disqualified by rule.
But also: to really mimic the original vector-math and comparison using L2 distances, I believe you might need to leave the word-vectors unnormalized before the 'king'-'man'+'woman' calculation – to reflect that the word-vectors' varied unnormalized magnitudes may have relevant translational impact – but then ensure the comparison of the target-vector to all candidates is between unit-vectors (so that L2 distances match the rank ordering of cosine-distances). Or, just copy the original `word2vec.c` code's cosine-similarity-based calculations exactly.
Another wrinkle worth considering, for those who really care about this particular analogical-arithmetic exercise, is that some papers proposed simple changes that could make word2vec-era (shallow neural network) vectors better for that task, and the same tricks might give a lift to larger-model single-word vectors as well.
For example:
- Levy & Goldberg's "Linguistic Regularities in Sparse and Explicit Word Representations" (2014), suggesting a different vector-combination ("3CosMul")
- Mu, Bhat & Viswanath's "All-but-the-Top: Simple and Effective Postprocessing for Word Representations" (2017), suggesting recentering the space & removing some dominant components
That it ever worked was simply that, among the universe of candidate answers, the right answer was closer to the arithmetic-result-point than other candidates – not necessarily close on any absolute scale. Especially in higher dimensions, everything gets very angularly far from everything else - the "curse of dimensionality".
But the relative differences may still be just as useful/effective. So the real evaluation of effectiveness can't be done with the raw value diff(king-man+woman, queen) alone. It needs to check if that value is less than that for every other alternative to 'queen'.
(Also: canonically these exercises were done as cosine-similarities, not Euclidean/L2 distance. Rank orders will be roughly the same if all vectors normalized to the unit sphere before arithmetic & comparisons, but if you didn't do that, it would also make these raw 'distance' values less meaningful for evaluating this particular effect. The L2 distance could be arbitrarily high for two vectors with 0.0 cosine-difference!)
So if for you the resulting doc-to-doc similarities seemed nonsensical, there was likely some process error in model training or application.
That's not the gambling-activity-specific taxes that Stoller's article discusses - typically applied to gambling businesses' revenues, not bet winners specifically.
I don't know, and there's no way to find via "Google search", what HN user ~dylan604 is specifically alleging has been improperly "pushed as fact".
If it's clear to you, can you share a representative quote from Loeb? He's got a lot of writing to choose from!
Does everyone at any prestigious institution have some duty to remain conventionally mundane in all their musings?
Is there any reason to think such hypotheticals are, on net, more harmful than helpful?
Isn't tenure (like Loeb's) designed to encourage a fearlessness around topics & speech?
Those other ways to integrate the texts might be some form of RAG or other ideas like Apple's recent 'hierarchical memories' (https://arxiv.org/abs/2510.02375).
So an old down-pressure on sizes – internal training costs & resource limits – now weaker. And as long as LLMs are seeing benefits from larger embeddings, they'll become more common and available. (Of course via truncation/etc, no one is forced to use larger than works for them... but larger may keep becoming more common & available.)