HNHacker News
TopNewBestAskShowJobs

frumiousirc

506 karma · joined May 31, 2016

submissionscomments
frumiousirc··on ArXiv's Updated Rate Limit Policy
> (and HEP in general)

Not for any neutrino experiment I know of. There are paper committees that govern publication but the actual preprint submission to arXiv is always done by one of the authors.

> I would think arXiv can work something out for these cases.

Yes, I also expect special cases would be defined.

frumiousirc··on Show HN: Ledge.sh – Runnable Markdown Notes
I've used org-babel in big and small contexts for more than a decade. I generally reach for the patter for a few goals:

- Keep diagram source (eg Graphviz dot) with the document.

- Write software developer manuals that can reference code robustly as the code changes (eg include code snippets from source based on regex).

- Literate Coding / Reproducible Research pattern. Eg, org file explains and generates/builds/runs code and graphs/figures which then get included back into the document.

Some of the problems

- At some scale, one needs a DAG (eg make) or to make every command idempotent with fast no-op. Otherwise, each little change to a source block takes too long and there is a worry that something didn't update which should.

- At some scale, the document becomes the size of a library/package and it does not compose and/or I want to run the code outside of the document.

- The unenlightened around me do not use Emacs so an org file is a "me" file. In some ways this is a plus as it keeps others fingers out of my pie but it also means no way to share the baking.

That was in the days before LLMs. As ellieh's #3 points out, things are different now (we all know that). I now have an LLM externalize org-babel. The LLM maintains an org or LaTeX document describing bits of work, an external library/CLI which runs to produce content including putting numerical results into LaTeX macros, a Makefile or Snakemake to regenerate content and figures. When things are found to change prior understanding the LLM remakes a section of prose in the LaTeX document. I then write my own notes or another LaTeX document so that the trip through eyeballs to fingers on the keyboard assures I keep some level of understanding.

Likewise, in the software documentation goal, it's far better to give a good LLM access to the source and have it generate documentation targeting some learning goal with follow exploration via Q&A than it is to read some prepared document that assumes my goal. Software documentation is kind of a relic useful only for those people that have not yet taken up LLM tools.

frumiousirc··on Ideas on modernizing the open-source desktop
The basic idea of Lifestreams is good for its metaphor but a lot of important implementation problems are ignored: ballooning storage from "clone" operations for every edit, ever increasing "view" query time, lack of a provenance back pointer. I think a lot of this can be solved with an SQL store, git style content hashes but there are still leaks to plug. One big one is how to resolve hashes given the content stores can move or disappear? Here I am assuming the system does not attempt to intern and reinvent git, github, dropbox, Maildir, etc? Then there is the huge footprint of how to collect inputs. Plugins for browser, email, git, perhaps even file system. It's all rather daunting to contemplate yet also feels like such a workable system is just within reach.
frumiousirc··on VSCode's SSH Agent Is Bananas (2025)
TRAMP actions are also rather slow (high latency). OTOH, tramp-rpc relies on a little tool to run on the remote and is much snappier. This proves there is a better middle ground than TRAMP with nothing and whatever abomination VSCode injects. Basically, busybox with a persistent RPC connection is all one needs.
frumiousirc··on ArXiv receives multiyear commitments to support it as an independent nonprofit
> Published

Posted. Having a preprint on arxiv is not publishing.

frumiousirc··on Ideas on modernizing the open-source desktop
The meta problem with all this desktop UX design discussion is that it's not actually very problem driven.

Many of us solve our UI problems and we land into our own happy minima.

Many of the UX "solutions" are either invented problems or not my problems or worse, grinding some anti-me agenda.

Like many, a lot of my UX improvements have been to simplify the UI. Tiling, with just a few windows per virtual desktop in herbstluftwm: kitty, emacs, browser and some rare transient application (usually PDF or image viewer).

The idea in the screen shot horrifies me. Why would I buy an expensive high pixel monitor just to use only 10% of it to display a single window? Why would I want to waste CPU to animate the migration of a full window to a single button icon (another example given)?

These UI "innovations" are great for movie props but do they actually help in reality? Not that I see. The "stagnation" of UI to me is more a plateau. Anything truly new must be driven by new "I", no input/output channels between user and computer.

The one interesting idea I saw in the article was Lifestreams. I think storage and recall is something that needs improvement and new ideas because for me that is a problem. I produce and consume too much info while recalling it is hard and while also I often do not produce or consume info that I wish I had.

Again, it's the I/O that drives the progress.

frumiousirc··on Orchestrating Claude Code Agents: The Chief of Staff Pattern
What is the interplay between such delegation and token cache timeout? If Fable is truly active for 4 days, enough to keep the cache hot, the cost would be astronomical. OTOH, if Fable idles while subagents are active, then each awakening is a cache miss.
frumiousirc··on Show HN: Jeff – A read-only CLI for semantic code review using Jev
The point is you can quickly get any yes's, and then subject them to more expensive scrutiny.
frumiousirc··on Communication by means of modulated Johnson noise
> the maximum achievable throughput is 26 bps at 1.5 m and goes down to 22 bps at 4.5 m.

And when there are more than one transmitters in real world conditions instead of anechoic chamber?

It's cool and all and if this work is for it's own sake, for the sake of research, no issues. But otherwise, I struggle to think of practical use cases. I grant that my imagination may be deficient.

frumiousirc··on Mastering Layout Engines in Graphviz: Dot vs. Neato vs. Twopi vs. Circo
Thanks! I've used graphviz for more than a decade and had never learned of gvpr.
frumiousirc··on Hister: A private search engine for the pages you visit and the files you keep
Do you have thoughts on the prospects for inter-operation between the "recoll" application and hister? Like, maybe an adapter that let's hister read recoll's xapian database/index? Or, a tool that converts/syncs beween hister's store and recoll's?
frumiousirc··on Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
They also define and measure an "IPJ" as well as "IPW"

> the NVIDIA B200 achieves 1.6× to 2.3× higher intelligence per joule than the APPLE M4 MAX across QWEN 3 and GPT-OSS model variants

The B200 = "cloud", M4 = "local".

So "cloud" does even better in energy than it does in power compared to "local". Or, to flip it, "local" is both slower and more expensive than "cloud".

frumiousirc··on Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
> Unless I misread it, are they saying local GPUs use less energy?

You misread it. From the abstract:

> local accelerators achieve at least 1.4× lower IPW than cloud accelerators running identical models

That's "intelligence per watt". They also have IPJ, per Joule.

So, they find local is 40% "dumber" than cloud for the same power or 40% more power for the same "intelligence".

Tables 13 and 14 summarize their IPW and IPJ metrics.

But, to your actual point, I think the "local is 40% dumber per watt than cloud" message is still an understatement. And maybe this is something I failed to find in the paper but they seem to ignore the "idle baseline" costs and talks about explicitly focusing on the power consumption of just the accelerator under load.

There is a large baseline power consumption just to support the accelerator. CPUs, memory, PS losses, network, fans, general environment cooling. This "cost floor" is different for data centers and a "random local computer" and I think must be in favor of data centers which are designed and built with efficiency in mind.

Idleness should also be considered. My local GPUs at $WORK and home are idle more than they are used. Idle time energy in real world scenarios should be somehow attributed to those brief, punctuated times when LLM functions are actually active on the accelerator. Actual, local LLM usage of a GPU is brief (assuming one user per PC). Even with my heavy usage developing s/w I'd guess I heat up a GPU about one hour per day total, sometimes much less. If that is local then one must pay 23 hours of idleness for that 1 hour of "intelligence". Of course a local PC is used for other things and the idleness penalty must somehow account for that. OTOH, data centers try to maximize utilization so their idle time penalty would be much less, perhaps close to zero, by construction.

frumiousirc··on LG says we're fake news [video]
Apparently, nearlyepic is trying to deliver two veiled insults and a critique of judgement or perhaps of the current state of broader affairs. I'll attempt to translate:

- "no idea what you’re talking" = "you have outsourced expertise to an LLM"

- "solution that doesn’t actually work" = "piecemeal blocking leaves many problems still active"

- "used to be free" = "you / others in general now pay LLM services for the outsourcing leading to poor solutions while in the past poor solutions came for free"

frumiousirc··on Ask HN: How do you manage skills files?
The closest I get to finding skills useful is when I find myself repeating myself to an LLM. This tends to happen most when I am starting new projects and want to communicate basic design principles and patterns to follow and libraries to use. What I did was to factor and store these "chunks" of instruction in some text files. I then made a little script that can list what chunks are available and when given a subset will essentially `cat` the selected files to emit AGENTS.md content which I save into the new project or append to shore up an existing one.

Your observation on the readership bias of HN is a good one for people to add to their HUMANS.md before reading and commenting. :)

frumiousirc··on Gemini 3.8 Flash and 3.8 Flash Cyber
My dictionary writes: "precision * accuracy = fidelity".
frumiousirc··on Claude Fable 5.1 and Claude Mythos 5.1
> not solved

I agree. Other reasoning traces simonw quoted in his blog post showed that the model made changes to consider realistic fork rake. I think this may also be the first case where the chain went inside the seat stays. The overall bike geometry is still comical but these bits show improvement in this model over prior ones.

On the other hand, given the absurdity of the original prompt, I should not necessarily expect realistic bike geometry.

frumiousirc··on I accidentally turned LLM memory into program analysis
Datalog seems like a way to "spell" knowledge graph (KG).

The article touches on Datalog statements changing over time. One ingredient I think would be good to add to the system is to make every statement carry "providence" metadata. The providence should be sufficient to enable later confirmation that a statement is still valid or if the statement needs to be reformed without the need to remake the entire graph from scratch.

I would make at least some forms of providence follow a strict schema that is defined for the subject matter that is being captured. For example, statements about a code base should refer to the source files and their version (file modification date, content hash) from which the statements were concluded. When a source file is modified we may then find all statements made from them and reevaluate just those statements.

The next level would be to keep statements even if reevaluation breaks them and add a method to derive a subgraph for a given state of the subject. For example, over many releases of a code base, a lot of statements would not change, some would. Having a graph that spans all conclusions about all releases of a code base and a way to form the subgraph for a specific release would allow the system to efficiently target queries for a particular release.

frumiousirc··on GUIs should be fully keyboard-driven
> Interestingly, people have seen my GUI Emacs and commented "I didn't know you could get images in a terminal". Then I have to tell them it's not a terminal...

...[finishing with], though you *can* also display images in a terminal, including in Emacs running in a terminal.

frumiousirc··on The August 17 outage
Linux didn't (yet) kill Microsoft. Microsoft absorbed that shot. Then the Git arrow went straight to cold black heart of Microsoft. The next few months will determine if they survive it. If they do, what will we see from the third draw out of Linus' quiver?
frumiousirc··on The Benchmarkpocalypse
That is not what I read from danluu's words. He merely stated in the prompt that there is a holdout set and did not iterate to minimize error against the holdout set. In a prior attempt he prompted with only "don't overfit" to ill effect on the holdout eval. Did I misread?
frumiousirc··on Qwen 3.8 27B
> Bicycle is the right shape. Pelican beak is excellent. Nice background.

Relative to other results I agree. But on an absolute measure, there is not a single element in the current bicycle that is real-world accurate and many elements are omitted or non-physical (eg, the transparent seat tube top, entire lack of a head tube).

Consider a series of followup benchmarks.

With a fresh context of the LLM under test, ask it to generate a list of findings for how the pelican-on-a-bicycle SVG that was produced is inaccurate compared what the real world scene might appear, accepting for the limitations of SVG as a medium. Then, feed back the list of findings to the original context for a second try. The benchmark can stop here by humans looking at the result and forming their own conclusion.

Next phase is to repeat the analysis phase using the 2nd context to determine what findings were satisfied and what new inaccuracies are found. These two differences can form a second benchmark.

Last phase is to iterate with the goal to drive the number of findings to zero.

frumiousirc··on Auto mode is now the default in Claude Code
I use bubblewrap, which I believe claude code also has internally but not for its `Bash()` tool.

I wrap bubblewrap in a script that supports config files to allow different "profiles" of use (analogous to eg firefox profiles). The bwrap starts with the whole filesystem mounted read-only, then mounts the current directory read-write and then applies further bind mounts for devices, special case other read-write (eg, ~/.cache/) and to mount empties to cover sensitive directories (eg, ~/.ssh/). The profile also specifies the default command to run and for claude, it gets yolo mode.

frumiousirc··on U.S. Department of Energy Launches the Genesis Open Models Initiative
Your question piqued my curiosity.

I thought maybe NERSC doesn't accept jobs from private corporations but I checked and that's not true, as long as results are not held proprietary.

Perhaps ANL was used as they have a lot of compute and they lead and host the Genesis Open Models Initiative?

From your link, 2048 B300 GPUs were used for 6 months. If google search is right, NERSC has 7168 A100. B300's are way more capable than A100. To do this training in 6 months, "3 NERSCs" would be needed.

Between the political angle and the technical, I'd guess these two make up a big chunk of the answer.

frumiousirc··on A Tome of Forbidden Technologies
Well, vinyl records are an analog medium so their sampling rate is infinite. More important would be physical and electrical band limits of the original recording, cutting and pressing devices and of course the playback devices.
frumiousirc··on U.S. Department of Energy Launches the Genesis Open Models Initiative
There's no mention of "LLM" nor "language". It does mention "foundation model" which includes LLMs but that also includes non-LLM architectures and non-text data. Many of the Genesis Initiative proposals answer "foundation model" call with non-LLM systems. All the FM's I know about currently in this sphere are non-LLMs. The "about gs1" page also does not mention "LLM" but does talk more about agentic harness and workflows. That description certainly sounds LLM'ish but describes a more rich system. I don't mean to suggest that LLMs will not be part of these "genesis open models" but as described, this will not result in a replacement for the "claude" or "codex" commands.
frumiousirc··on U.S. Department of Energy Launches the Genesis Open Models Initiative
> and there are career civil servants (all the government scientists are under this category)

US national lab scientists are not even civil servants. The labs themselves are run by a corporation under contract to the DOE and the scientists work for that corp. The managing corporation changes from time to time and the scientists transparently start working for whatever assumes the replacement. The land, the hardware, the buildings and any physical products are owned by the US gov't. To a very large extent, the intellectual output is set free to the world in the form of papers, presentations and to some small extent (eg compared to CERN) in the form of software.

frumiousirc··on Herdr is joining Y Combinator. The runtime stays open
> I find myself to be a huge bottleneck, having to do technical design and design reviews, to make sure they're actually working on the right things.

I find this depends on the LLMs being used and the person using them and the problems being solved.

Like, Fable can be sent off to do some big thing for a long while and I find myself getting bored and move over to push on some other thing. Or, while reading a paper I have a string of follow up questions and ideas and launch them via web or CLI agent. Or I bounce between multiple chores in different packages, each of which is fast for an LLM and a low cognitive burden for me. Then other times, something needs my full, ongoing and serial attention where the LLM turn is only a minor element. I'll use the brief LLM interludes to get up, walk around a bit, stretch and think.

As for Herdr, it seems nice when I tried it a few weeks ago. But, I'm surprised it is getting so much attention (kudos). Like many people, I've developed more than one work-alike before Herdr hit the scene. They were based on tmux which I have concluded I simply detest. In the end, I made a little agent hook script that speaks to Kitty terminal and that plus Kitty and a stable SSH persistent connection gives 80% of what I was looking for. I do still like the "control panel" aspect of my past and herdr's tries. Adding that would get me another 10%.

frumiousirc··on Prevent cognitive debt by manually retyping LLM-generated code
Of course I push my car to the store! It has the GPS and I've lost the ability to self-navigate. Plus, how else could I carry my groceries.

:)

frumiousirc··on Prevent cognitive debt by manually retyping LLM-generated code
I think this characterization misses the feedback loop that exists.

The OP is not sitting in a one-way flow of info:

    LLM->human->code 
but rather a the center of a feedback loop:

    LLM<-->human<-->code.  
Human-is-the-loop, not human-in-the-loop. Each iteration of that loop is fully driven by the human. Human creativity is involved in both directions.

I'd argue that anytime we drive an LLM through more than one turn (and/or more than one session) we are really doing a human-is-the-loop thing. The OP's extreme case of begin the only thing editing code is on a spectrum with the other extreme being vibe coding (never looking at output code, but still interacting with that output in some way).

Off the spectrum is what I call LLM-vomit. A human one-shots something and puts it out for others to see, suffer and clean up or ignore. This is code that is encountered literally out of context and can only be further improved (if that is even attempted) by approaching it from first principles. Such code is akin to people using LLM to generate an answer delivered to another human. Both flavors (sorry) of LLM-vomit are bad. I can prompt my own LLM to do that, don't do it for me.

Page 1 of 12Next →