I think humans have to reason because we don’t already have a statistical embedding of the solution pattern built in. We have vastly less rote knowledge crammed into our heads and so require creative synthesis to span the gaps.
With LLMs the trick is revealing their existing relevant embedded knowledge more reliably. They’ve almost literally seen it all before, and the trick is dialing it in. The reasoning tokens help shape the autoregressive attention lens that focuses on and enables recall of the already-experienced answer.
It is interesting that “reasoning” has a similar outward appearance, but since LLMs are built to
mimic outward appearance from trillions of examples, you can’t infer underlying mechanism from appearance.
Probably the Windows (NT) kernel, which I imagine would be a prerequisite for the Windows System/GUI. The anachronistic initialism is “New Technology” — MS in the 90s not thinking the name would survive Y2K I guess.
I think many failed/aborted projects only seem stupid in retrospect because we assume it proof they could never have succeeded .. but many ridiculous things can and do succeed and find their place. It fizzled after his death without anyone championing it with persistent care and constructive iterations. It may well have ultimately failed, regardless, but it certainly lost its motive power.
It is interesting that Meta would have to develop a social or political content identification system in order to either label, limit targeting, or ban/censor it. It seems the hard part of complying was developed, but only to thwart the law, rather than comply. That might be a bigger FU to the EU than ignoring the law.
Your text may well have been in the training corpus, but not searchable with whatever terms and search engine that LLM used after you prompted it. It doesn’t have recall of sources of documents comprising its training corpus, unless the source is widely cited in other of its training documents. Sourcing and provenance aren’t currently an intentional part of LLM training.
That may already have been obvious, and central to the complaint, but I was just knee-jerking to the anthropomorphic language.
Well, in fairness, both sides were bad, and still are -- although, now, one side is spectacularly worse. Dems can't ever agree on a candidate or position or message, and intentionally run candidates the otherside would abhore, and Republicans would happily vote in unison for an old shoe as long as it was sufficiently dog-whistley.
It may be turtles all the way down, but we don't want to need to visit each tier, wading through obvious slop engineered for engagement as we go. The sloppiness reveals the low effort machinery meant to sap our vital attention and time. It's probably best to build a growing resistance to it.
… but, a power power user would know that it was transparently saved every several keystrokes, even while named Untitled and ostensibly unsaved, ready for immediate transparent recovery on relaunch, after an app or system crash -- so they might again be okay with it.
Like an idiot, I once wrote a simple, internal, in memory chat room via web page at work. Because it let you pick a nickname, co-workers immediately gave themselves names of other employees and started sexually harassing each other and being asses. I pulled it down immediately before they got me fired. Anonymity does something to some people. They lurk among us biding their time, looking for any opportunity to be reputationally unrestrained.
You’re right. Modern frontier models are now multimodal. I used often as weak a hedge, because I know at least his gpt3.5 turbo and llama3.1 generated pelicans were from text only models without image training. The chinese models are interesting, because before their vision models existed they may have been distilling text only models from text output of American vision models, so they could have benefited from the teacher model’s vision capability without being vision models themselves.
At that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.
I think the evil part is the [at any cost] part, the incentives are irrelevant. An AI that only wants to make paperclips isn't evil because of the paperclips.
But, they were wrong ... in assuming what they were saying was at all relevant. And they were so sure they knew what they were talking about and that it applied that they dismissed all the ample evidence to the contrary and also dismissed the, they assumed, delusional people pointing out the contradictions. Their initial misunderstanding was understandable given the overloaded terms but their confidence was poorly calibrated to their weakness detecting the misunderstanding.
Wait, so it would have, ironically, been safer to allow microsoft telemetry to bypass the VPN entirely and remain associated only with their home network, because it's the phone home to microsoft tunneled through the VPN that tied together all their IP addresses to a single microsoft account. The GDID itself is almost a red herring, as it could have been a session id or username or something only long lived enough to appear comming from several ip addresses to windows update request?
Oh, good. Maybe I can stop myself from mindlessly upgrading yearly which being on the original iPhone upgrade program makes too tempting and easy. I should be keeping phones two years so they can be family hand-me-downs, anyway.
Unrelated to Google Workspace and Firefox, but I just noticed today that Google’s YouTube now says my iPhone’s Safari browser is incapable of playing full screen videos, which it’s not ever claimed before. I’m also getting sick of them pushing Chrome anytime I use a Google service like search or Gmail. I keep dismissing the prompts, but they are relentless. It all seems so sleezy and desperate.
I think their intent is actually safety. They employ two touch interaction models: Flexible while not moving and simplified while driving. For instance, keyboard input becomes unavailable while moving and you must rely on Siri. I personally find it irritating, particularly when I am a passenger, but I get it.
People confabulate. The setup further invites it. We mindlessly fill voids and call it opinion, and occassionally even believe ourselves. "Plausibility" is a really low bar. Rigor is tiring and we'd rather not invest as no one demands it, anyway.
I'm not sure what's going on, but in Chrome my median point is heavily on the blue side but in Safari it's on the Green side. Also, at least in Safari, reset leads to a series of perceptually unchanging turquois screens -- it seems a bug. Refreshing fixes it for the next run.
Paraphrasing F. Scott Fitzgerald? "The test of a first-rate intelligence is the ability to hold two opposed ideas in the mind at the same time, and still retain the ability to function."
Holding contradictory ideas isn't the laudable skill. Any uncritical person can believe conflicting things without being troubled by them. The genius is holding such ideas in disbelief long enough to let evidence alter or evict them.
I always sort of speculated that sports existed to channel what would otherwise be human tendencies toward violence; an outlet enablining more stable civilization. Even though I largely ignore sports, I appreciate it over possible alternatives.
What I find most interesting is that HN posters seem to overwhelmingly skew liberal, but HN flaggers lean extremist “conservative”. They rarely post, but completely control the discourse of the posters. Thats a crazy dynamic.
Out of curiosity, if you asked for the same text extraction multiple times, each inside fresh contexts, is it likely to fabricate unique quotes each time? And if so, a) might that be a procedure we train humans to do to better understand LLM unreliability, and 2) and instrumentalize the behavior to measure answer overlap with non LLM statistical tools?
Also, quote-presence testing/linking against source would seem to be a trivial layer to build on a chat interface, no LLM required. Just highlight and link the longest common strings.
Edit: Oop, I misread! Right, yes, the change up was arguably not entirely boring. Some people were excited at least.
Originally: To be the annoying pedant, version numbers did still monotonically increase, even with the gap, because each version is >= to the last. The mono means a single direction, not a step size of one.
The rationale for rejecting the syndrome was the reported scientific consensus of the impossibility of such a device being smaller than a large truck. The revelation that such a device is possible, exists, works and is full of Russian parts seemed to change things. Further, the coincidental appearance of Russian agents on video within operational radius of incidents wearing backpacks sized to contain the device raises some questions.
Probably your DNS -- the archive.today guy is a stickler that dns must pass client subnet to partially deanonymize visitors, and for instance, cloudflare's 1.1.1.1 server doesn't pass it. I think that's still the case.