Particularly when the only source is a friend of the author, posting on a blog named "AI Clambake" about "A weekly, human-powered newsletter for advertising folks who want to stay on top of the AI mayhem" and not a publication with any credibility in linguistics.
None of that means it can't be true, but some basic skepticism is warranted here. Otherwise we end up in a situation like the LK99 room temperature superconductor where a lot of HN commenters were also upset at the cynical "downers" who just couldn't root for a good thing/progress.
1) Many preprints are bad, incredible bad. I read a lot of posts about ivermectine during 2020 and the errors were obvious. Like no control groups, the control group is a bunch of unrelated guys in another city, and a weird articles that split the 20+20 cases in 10 bins with 2+2 cases in each. They had a lot of error that were easy to spot without being a medical doctor. (Ctrl+F exclusions, you may get a surprise.) (And don't get me started with Chlorine Dioxide.)
2) Perpetual mobile and mass less drive reappear every few years. I definetively can read most of them. The most interesting part is the totally broken explanation of why this new version does not break the laws of physics.
3) HN has a lot of users specialized in niche topic. A few weeks ago I wrote a comment with a joke: "the list of text transformation to allow a Spanish speaker to read German enters in a napkin" (for example v->f and w->v and a few more). Someone was surprised because s/he knows that German has more phonemes than English that has more phonemes than Spanish. There is someone wandering here that really knows about phonetics.
So, I want to see a preprint. Perhaps I can read it, perhaps someone else can read it, perhaps we have to wait a few days until someone writes a nice blog post and debunks it, perhaps it's correct.
Could you rephrase this or explain it more thoroughly? I don’t follow. What does it mean to categorize a written form by systems built with Claude?
https://gist.github.com/fragmede/bbf277d36a2398065f109484f34...
The original prompts aren't provided, nor is the original context; even then, you can't really treat a stochastic system like an LLM as a major component in reproducibility.
If you had the other things, being "stochastic" is not even remotely a show-stopper. Stochastic processes abound and are the reason the mathematics of statistics was developed in the first place, ultimately allowing us to create such things as LLMs.
When all the relevant steps gets published, I absolutely expect a lot of people to (attempt to) reproduce this work even though LLMs are stochastic.
On the prompt formulation; prompts with very similar formulations (in terms of both semantics, hamming distance, or both) can lead to _wildly divergent_ outputs in my experience. It's not rigourous, and when that divergence happens, it's extremely difficult (arguably impossible, by nature of the architecture of transformers) to identify why the divergence happened and where.
It's not about being able to throw claude or codex at a loop and having it evaluate it for halting, it's about being able to do this for arbitrary code. Computer science rigourously defines the halting problem as not computable and undecidable. within the framework of using something akin to static analysis using any deterministic Turing machine.
There's not really a question of "solving" the halting problem like there's some as-yet unknown way of generally figuring out if arbitraty code halts. Turing proposed a proof in 1937 in favour of undecidability of what we now know as the halting problem, building on ideas first articulated by Church a few years prior.
Frankly, if anything, it's reasonable to say that the halting problem's been solved, just in the direction of undecidability rather than decidability.
Anyway, back to LLMs; as code gets more complex, the robot will need a bigger context window, more hardware resources, and more time, all of which will be variable due to the noise inherent in the system. It'll be difficult to put a useful upper and lower bound on how much computing power and time it'll take to figure out if a program ever halts. Which is all a bit moot, frankly, in the context of halting, but useful to keep in mind in the more general context of using these things as analysis tools.
Every day when you lower your butt onto your chair, you trust a stochastic system enough to assume you'll rest on the chair safely and not spontaneously phase through, which would lead to rather gory and painful terminal experience.
Physics at macro scale is stochastic, which is a good reminder that stochastic != uniformly random. Expected distributions matter.
IMO a better example would be the stochastic nature of quality control in manufacturing.
I was going to segue into thermodynamics as a backup example, but you made me think of something better.
> IMO a better example would be the stochastic nature of quality control in manufacturing.
How about, more specifically, food manufacturing? Or maybe, let's talk about cooking?
Cooking is as stochastic as it gets, and we handle it fine. It could be better - the better version is called "chemical process engineering", it's what cooking looks like when you care about quality and consistency of output, and can afford the equipment and process actually necessary for it. Regular people don't (i.e. neither care, nor can afford) - we call this cooking. It's an art, not a science, and people not only do it, but love it, and tie their identities to it, and build businesses around it, and a culture that embraces all the compromises (and calls the more serious approach "unhealthy").
My attempts at making bread have been too stochastic, in that it hardly ever produces nice results.
But yes. Eyeballing how much dried herbs to put in my dishes because I like what 2-isopropyl-5-methylphenol does for them. Usually it works, sometimes it's just a bit too Italian.
... in some sense, it's a miracle most people deal with this kind of bullshit without complaining much.
(Probably because they don't realize it's something to complain about. It's just how things are.)
Speaking generally, food produced though "chemical process engineering" (a.k.a. factories) must compromise on many axes, one of them being nutritional content. We intuitively do not care about several of these dimensions when cooking food with fresh ingredients, at least not at the scale of, say, Kellogg's or General Mills.
Maybe that's evidence of accepting a stochastic process in our daily lives, but you're kind of selling the tradition and science of cooking short when you argue that factory-produced food is a "more serious approach".
Claude helped write code to read and parse the corpora and to do some fairly basic statistical analysis along the lines of "which Linear A symbols most often occur together" and "if we use known Linear B sound values, which of the other corpora most often have vowel similarities with the Linear A corpus".
You can write that code yourself or you can ask an LLVM to write it for you. The provenience of the code isn't important.
*) He later deleted some of them, I think. What was still there on reddit a few weeks ago had dead links to a web site of his with statistical tables and I believe also code.
The Dutch guy thinks he can recognize some names (personal and toponyms) and that seems to be his main evidence for Linear A being Hurrian/Urartian. His theory doesn't seem very convincing (and his youtube videos are incredibly booooring).