1,106 karma · joined October 17, 2013
Email: my username here at physicsdog dot org
long-long-short-short-short is 7 in morse code. They don't say what time they tried this or if it ever changed, though, so that's just a hypothesis.
That's not to say it can't be a neutron, but it would be surprising if it were.
As they better understand the detector, they can use more of that mass. They have data from it, but they just didn't use it. And they're always collecting more data, too, as time passes.
So the 3x is saying they have something like 8.5 tonne-years of data.
They were pretty model agnostic in what they were looking for. They modeled and simulated a number of different ways a WIMP could interact with normal matter. If this is a discovery, more data will be needed to figure out the nature of that interaction and how it fits into particle physics.
But there's always a chance it's something completely new, or some extremely rare manifestation of things we already know about, but have never seen before. And even if it is WIMP, it may not be the right type of WIMP (wrong mass, or wrong interaction strength) to explain cosmological dark matter.
Additionally, when a particle interacts with the nucleus, the ratio of how much energy ends up as scintillation light versus ionization is different than when a particle interacts with an electron, which is most of the background processes.
Then, whatever is left, they try to model using known processes. After all that, there's one event that they can't account for. And that's what the news is about.
So it's certainly interesting!
That said, particle physics history is full of 3 sigma particle "discoveries" that disappeared with more data. They're collecting more, so hopefully we'll learn more in a few more years.
[1] https://lz.lbl.gov/wp-content/uploads/sites/6/2026/08/LZ_Pre...
The 2017 paper you linked is the first observation of the process. They had to do it statistically; there's no smoking gun event. It took them almost a year with the thing sitting next to a neutron beam to get enough statistics. With a beam, they were able to do extra noise rejection based on the beam timing. It's still a lot of experimentation and engineering work to go from that to something that can operate in a lower signal-to-noise environment.
Also, growing large scintillator crystals is a very specialized process and so they are expensive. The cost is going to limit how large you can make such a detector. You can scale up a water-based detector much more easily, and that extra mass can make up for the reduced cross section, depending on what you are trying to observe.
I don't have any real proof for this, but it feels like House of Leaves inspired a lot of the people making "found footage" and "creepypasta" stuff one the internet in the 2000s and early 2010s (SCP, Marble Hornets, Slender Man), and then that stuff came together to inspire the Backrooms.
There's a lot of research going on in this space though, because yeah, nature can solve certain mathematical problems more efficiently than digital systems.
There's a decent review article that came out recently: https://www.nature.com/articles/s41586-025-09384-2 or https://arxiv.org/html/2406.03372v1
How many times at work have you been talking to someone else where they're using common words as jargon? Maybe it's something like "the online system" or "the platform". And it's perfectly clear to them what they mean, but everyone else in the company either doesn't know what that actually is, or they have a distorted idea based on the conventional definitions of the words. Even without LLMs in the mix, this can lead to people coming out of meetings with completely different understandings of what's going on.
My experience is few people are actually providing the relevant context to the LLM to explain what they mean in situations like this. Or they don't have the actual knowledge and are using the LLM in the hopes it'll fill in for their ignorance. The LLMs are RLHFed to sound confident, so they won't convey that they don't know what a piece of jargon means. Instead they'll use a combination of the common meaning and the rest of the context to invent something. When this gets copy/pasted and sent around, it causes everyone who isn't familiar to get the wrong idea. Hence "misunderstanding amplifier".
To the point of the article, this is soluble if people take the time to actually figure out what they are trying to convey. But if they did that, they wouldn't need the LLM in the first place.
Several different engineering teams from different parts of the company had to come together for this, and the overall architecture was modular, so there was a lot of complexity before we had to start integrating. We have some company-wide standards and conventions, but they don't cover everything. To work on the code, you might need to know module A does something one way and module B does it in a different way because different teams were involved. That was implicit in how human engineers worked on it, and so it wasn't explicitly explained to the coding agents.
The project was in the life sciences space, and the quality of code in the training data has to be worse than something like a B2B SaaS app. A lot of code in the domain is written by scientists, not software engineers, and only needs to work long enough to publish the paper. So any code an LLM writes is going to look like that by default unless an engineer is paying attention.
I don't know that either of those would be insurmountable if the company were willing to burn more tokens, but I'd guess it's an order of magnitude more than we spent already.
There are politics as well. There have been other changes in the company, and it seems like the current leadership wants to free up resources to work on completely different things, so there's no will to throw more tokens at untangling the mess.
I don't disbelieve the success stories, but I think most of them are either at the level of following already successful patterns instead of doing much novel, or from companies with much bigger budgets for inference. If Anthropic burns a bunch of money to make a C compiler, they can make it back from increased investor hype, but most companies are not in that position.
At first, it proceeded very quickly. Using agents, the team were able to generate a lot of code very fast, and so they were checking off requirements at an amazing pace. PRs were rubber stamped, and I found myself arguing with copy/pasted answers from an agent most of the time I tried to offer feedback.
As the components started to get more integrated, things started breaking. At first these were obvious things with easy fixes, like some code calling other code with wrong arguments, and the coding agents could handle those. But a lot of the code was written in the overly-defensive style agents were fond of, so there were a lot more subtle errors. Things like the agent adding code to substitute an invalid default value in instead of erroring out, far away from where that value was causing other errors.
At this point, the agents started making things strictly worse because they couldn't fit that much code in their context. Instead of actually fixing bugs, they'd catch any exceptions and substitute in more defaults. There was some manual work by some engineers to remove a lot of the defensive code, but they could not keep up with the agents. This is also about when the team discovered that most of the tests were effectively "assert true" because they mocked out so much.
We did ship the project, but it shipped in an incredibly buggy state, and also the performance was terrible. And, as I said, it's now being wound down. That's probably the right thing to do because it would be easier to restart from scratch than try to make sense of the mess we ended up with. Agents were used to write the documentation, and very little of it is comprehensible.
We did screw some things up. People were so enthusiastic about agents, and they produced so much code so fast, that code reviews were essentially non-existent. Instead of taking action on feedback in the reviews, a lot of the time there was some LLM-generated "won't do" response that sounded plausible enough that it could convince managers that the reviewers were slowing things down. We also didn't explicitly figure out things like how error-handling or logging should work ahead of time, and so what the agents did was all over the place depending on what was in their context.
Maybe the whole mess was a necessary learning as we figure out these new ways of working. Personally I'm still using the coding agents, but very selectively to "fill-in-the-blanks" on code where I know what it should look like, but don't need to write it all by hand myself.
Very quickly:
a dollar coin is about 550 mm^2 on a face
the Cray-1 could do 160 MFLOPS
an M1 chip has a die size of 120 mm^2
an M1 chip can do over 1 TFLOPS[1] https://internethistory.org/wp-content/uploads/2020/01/OSA_B...
https://leginfo.legislature.ca.gov/faces/codes_displayText.x....
The drawbacks of YAML have been well-documented[1]. And I think it's worse now in the LLM era. If I have a system that's controlled via scripts, an LLM is going to be good at modifying those scripts. Some random YAML DSL? The LLMs have seen far fewer examples, and so they're going to have a harder time writing and modifying things. There's also good tooling for linting and checking and testing scripts to ensure LLM output is correct. The tooling for YAML itself is more limited, even before getting into whatever application-specific esoteric things the dev threw in.