HNHacker News
TopNewBestAskShowJobs

joshka

2,438 karma · joined June 6, 2011

[ my public key: https://keybase.io/joshka; my proof: https://keybase.io/joshka/sigs/ElI6j13SI-_2rTrkLnoKyV-uDepH4JvAQAzp2OWV_W0 ]
submissionscomments
joshka··on GitHub's new dashboard experience now the default
x-posted on their feedback:

https://github.com/orgs/community/discussions/208952#discuss...

Instead or alongside this, how about making notifications good? There's some concrete simple improvements that would help triaging in general.

1. Provide ways to filter notifications of a specific state (merged, shipped release, closed) so these can be handled in bulk easily

2. Filter by type (issue / pr / release)

3. Provide access to the done state as a filter and via the api, so we can use non github website tooling to manage the notifications

4. Filter by author (e.g. let me filter dependabot notifications so I can handle them in one pass)

5. Show and allow filters on PR check status so I can look at PRs that are ready for a quick next step rather than ones that require digging into failures etc.

6. Make this all available properly via the API so I can let my clanker interact with this stuff based on rules that it has to help me triage stuff. Right now there are enough missing fields from the notification APIs that it makes it impossible to work around the shortcomings of the notifications with tooling (without screen scraping /notifications).

All of these things would be better for interacting with work to be done.

A general rule on this is if the information would make it easier to understand what work is to be done, to choose it, to act on it in bulk or in batches of similar actions, then that info belongs in the UI and the API.

---

Also. Please make it possible to subscribe to a discussion comment, but unsubscribe from a discussion. I want to read replies to this comment, but don't care to get a notification on every reply to this entire discussion. Unsubscribing. Tag me if you want me to read a reply here.

joshka··on GitHub's new dashboard experience now the default
> We’d love to hear your thoughts on the new dashboard experience. Drop a comment with any questions or feedback in the Community discussion.

Please someone at github, hear this. Discussions are useless for feedback as you cannot subscribe to just the single thread you posted, only the entire discussion. That means I have a a binary choice of drinking the firehose or ignoring all future comments on some feedback I care about.

Please fix this.

joshka··on Tells of a Slop UI
My bugbears are narration - where the page tells me how to use it, or is labelling things in a way that a human would not use. "Choose a ..." where ... would be better as a heading.
joshka··on Revealing the details of how OpenAI agents hacked Hugging Face
I'd put it more generously (albeit biased), that they do take it seriously. But even serious people can be misguided in what things they pay attention to. Security is something that you have to get right 100% of the time and have people whose job it is to say no a lot. Research is the opposite. There's a clash of cultures in those two extremes and OpenAI was born from the wrong side of it. It's worth reminding that ChatGPT was launched as a "low key research preview".

I'd say the narrative that AI agents are a looming danger to the world is probably undersold rather than overhyped. I'm not particularly a doomer on this, but I have an infosec background too, so have a fair idea of what the combination of agentic harnesses + a malicious mindset could do to people/companies/nations/politics/world if wielded incorrectly. I think the good guys will win on this, but there will be plenty of interesting things that happen in that journey.

A good thought process might be to think back to the various large internet worms of the 2000s (Code red, Nimda, SQL Slammer, ...) which were mostly monoculture 0-days (not technically but close enough). Now consider if you no longer have monoculture / single bug as the limitation plus an ability for the hosts to take part not just as attack surface, but also cognition and planning. There's lots of variants of this and they're not particularly far fetched scenarios.

joshka··on Revealing the details of how OpenAI agents hacked Hugging Face
OpenAI's business model would align infra as a cost center rather than infra as a profit center (e.g. Google / AWS). Perhaps there's something there. I'd say also the OpenAI as a grad school that just happens to have a business aspect is also part of this. Bringing a tonne of good process on top of the build fast break things startup stuff would have cramped research speed significantly.

It's likely that OpenAI has gotten as good as it is because it ignored the traditional sysadmin stuff and went scrappy.

I worked there, but this is just my opinion and guesses, not facts.

joshka··on OpenAI 'agent' hacked Australia's health service
dupe https://news.ycombinator.com/item?id=49822556
joshka··on On the Navier–Stokes Millennium Prize Problem
At Astra API prices that's 300B tokens (I saw 130B output tokens claimed elsewhere), large but not unheard of if you consider it across a few people doing random experiments with best-of-n type things. On my personal account, I've done a billion+ token days just on a normal pro 20x subscription. I know many others that wildly outpaced that by orders of magnitude. This was apparently 130B over 89 hours, so about 30x that rate. When things are free and you're expected to token max 30x seems fairly reasonable to me.

If you consider this as a cost to be compared against the question: "What does it take to be able to prove that you have a model that can solve the hardest problems that humans know about?", then spending a some amount of thousands/millions to know the boundaries of that seems not too important in comparison.

You've also got to consider this as compute that's allocated to pushing the frontier of what models can do, so while it's using GPUs that have been paid for etc., it's not like it's a cost that's supposed to be use less of this so that others can have capacity. If you made researchers afraid to use capacity like this, a lot of the things that improve would tend to do so significantly slower. (some may say that's a good thing ;)

A good way to think about this is when tokens are free, you get to choose whether you're optimizing for latency or intelligence rather than having to consider price.

---

Publically, tibo (Codex owner) in Feb this year: https://x.com/thsottiaux/status/2024649339344445825

> OpenAI employees currently get unlimited inference. Usage is now peaking at > XX billion tokens per week for some of them.

Mathew Berman (AI Youtuber) in Jun: https://x.com/MatthewBerman/status/2067270730795134984

> I've used 25 billion tokens in the last 7 days.

joshka··on On the Navier–Stokes Millennium Prize Problem
I worked at OpenAI previously, but don't know any of the people involved in this.

My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".

It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.

joshka··on I've factored the RSA keys of a Certificate Authority from the 90s
I think you're assuming that the output of the page is LLM generated and not the process to produce the page.
joshka··on Show HN: Stuxnet – A reconstructed source code of the infamous cyber-weapon
Your decompilation threads have all the necessary info in them for this, anyone coming after lacks that foundation and effectively is doing a second inference over the hidden state, assumptions, etc. that your sessions have in them. A simulacrum of a simulacrum in essence is likely to be not particularly good.
joshka··on Show HN: Stuxnet – A reconstructed source code of the infamous cyber-weapon
If you're using coding agents for this, it may be worth splitting this up into multiple well arranged modules that tell a coherent story and make it easy to browse, and add explanatory docs based on the various things the LLM has found about each function / type.
joshka··on Can AI design circuit boards yet?
Astra seems generally available now in codex at least - got any updates on this?
joshka··on US gov sides with OpenAI on issue of training LLMs on copyrighted material
To some extent, and obviously grossly over-simplified here, copyright is the act of taking some already public good and putting private rights on it. My reading of the amicus is that the government is arguing that particular rights that NY Times are asserting that they have are not ones that promote the aims of the original reasons for copyright to exist and so aren't necessarily ones which need to govern the behavior of any other people (in this case OpenAI, but this likely applies to any LLM training no matter the size).
joshka··on Can AI design circuit boards yet?
Would be good to see the full details of the tasks / methodology used. https://eebench.org/methodology.html makes this unclear.

> A real capacitor makes the task more interesting. A ceramic part may provide much less than its advertised capacitance once it has voltage across it. Parts have tolerances. Adding more capacitance costs more, takes up space and makes the rail slower to recharge when the power returns. A design that works with nominal values can fail with the parts that arrive.

It sounds like from a reasonable reading of the benchmark post that there's some things that are being tested that are assumed to be criteria that you expect the models to intuitively find those things to be important (i.e. the stuff about working on parts that have tolerances etc.). If that's so, then this really feels like mostly an exploration of whether an LLM has a good understanding of unstated constraints and has an appropriate in distribution set of priors that would be able to form models where it's reasonable to design on those lines.

It's hard to tell whether this is a problem though as the methodology is imprecise.

If you're spending time on evals against your own product, I'd be super curious to see how far you can get to by using a top tier model to produce generalized instructions for lower tier models. E.g. in a loop: "This eval missed X. what's the simplest single instruction that would have helped this session consider that as necessary that can benefit all future runs. Stick that in AGENTS.md and retest."

joshka··on Justice Dept. Sides with OpenAI in New York Times Copyright Suit
https://news.ycombinator.com/item?id=49538820
joshka··on US gov sides with OpenAI on issue of training LLMs on copyrighted material
Reuters: https://news.ycombinator.com/item?id=49538820

NY Times: https://news.ycombinator.com/item?id=49543821

joshka··on Apple reveals 'shocking evidence' from ex-employee's MacBook in OpenAI suit
Not sure why the downvotes on this - it's a pretty reasonable take. I suspect that you're right that it's unlikely that the tainted items would be used to update the weights here and are more likely to be something that would be in sessions / memories rather than future model weights.
joshka··on AI Agents and the Refactoring That Never Happens
A lot of this feels like it comes down to the training of the agents to produce code that satisfies the various benchmarks combined with reactions to things which were previously maladaptive. I.e. things which were explicitly trained out of the model in post training. I think there's a lot of missing long term software engineering principles that don't seem to be baked into the way the models tend to write code by default.

I have some speculation that maybe the people doing the model post-training tend to be younger researchers that haven't worked on large complex software systems, so their taste isn't as developed in this regard about what things are important here.

But this is an area that can be steered with appropriate early instructions ("When choosing tradeoffs of implementation, build for long term maintainability and understandability of code over implementing just the exact code necessary to solve the issues. etc. chain of thought often includes information that would have to be repeated in a future agent session, make sure to persist it to code or external docs so that future sessions and user understanding is respected.")

It can also be done as a post-change step with similar effects. And you can use your agents to build this layer into your general modus operandi for dealing with the crimes of generated code. But one of the things that all AI labs should be doing is looking at AGENTS.md on real project as being hard expressions of what failure modes real projects have noticed in models generally. Don't wait fo the bugs to be raised on these things, use express preferences that show that there's a problem. Go trawl github for these in bulk to use for future post-training.

joshka··on Apple reveals 'shocking evidence' from ex-employee's MacBook in OpenAI suit
> Apple argues that when trade secret information is fed into an AI agent or model that learns from it, that learning “may create irreversible and continually propagating uses of the trade secret.”

This is somewhat of a high impact argument to test. I wonder if the case will eventually get to working this point out.

joshka··on The safest job from AI may be writing
Size of context is not the entire story here, it's ability to properly feed and index the context that's needed on this sort of thing. E.g. your entire slack/discord/email/github/jira/zoom meeting/coffee chat ... history is the context that you bring to the table on this sort of thing. Most of this is unindexed. Much of this will not be in the future.

> The latest models got even worse.

Which models? This is one of those things that likely has both model and domain specific aspects that impact your experience. In my experience with OpenaAI models predominantly (I previously worked there), they've improved significantly over the last 6-12 months. My experience with Claude is worse, but I haven't spent as much time getting into a mechanical sympathy there. They're still not perfect though and I have many steering docs that help avoid the biggest problems in the models I use when generating docs.

joshka··on The safest job from AI may be writing
> It can't magically know what you want to say

I think for this argument to be true, the axiom that supports it is that the models have just as much context as they will ever have, and you cannot see being able to give them more / enough to be able to understand your perspective. That feels unlikely to be a position that doesn't change. As a society we're giving more and more context each day to this, and that makes this a valid opinion now, but one that erodes over time.

joshka··on Tim Curry has died
Dammit
joshka··on Migrating a Synology NAS to a UniFi UNAS Pro 8 with Robocopy, SMB Multichannel
I took a look into it mostly to satisfy my curiosity. The AES-NI chip on the J4125 in the Synology DS1520 hits 5Gbit/s transfer when using 4 cores (AES-256-GCM). So technically this it would be possible to saturate its 4 ethernet ports (or at least the encryption wouldn't be the bottleneck). You could back this down to 128 or drop the encryption altogether on an network that you own like this reasonably.

My guess is the actual bottleneck would be directory traversal and metadata stuff - i.e. trying to keep the pipe full, not saturation effects.

Yeah I know on src/dest. I've done this sort of approach countless times in various technologies in the past 40 or so years, but that always puts the mechanism for checking things in onus of the human rather than having that one command that does it all right.

Anyway, not a slight on rsync in the slightest - it doesn't have this probably because no-one realistically actually needs it most of the time.

joshka··on Stop Making TUIs
It's mostly something I'm thinking a bunch about recently. Nothing written up yet aside from the above. I'd go read Mitchell Hashimoto's Lobsters interview fora different take and see how that resonates with you as well as the recent blog posts about TUIs and accessibility.
joshka··on Stop Making TUIs
yeah, something along those lines - but more robustly defined.
joshka··on Stop Making TUIs
one step deeper - it's a failure of WIMP[1]. Imagine if instead of having a border around the parts of an app that make it an app and instead we applied the unix way but to the gui? What if we had lightweight widgets that had a lifetime of their own and could be nested and combined.

[1]: https://en.wikipedia.org/wiki/WIMP_(computing)

joshka··on Stop Making TUIs
I have built a pretty full version of that idea in rust (codex tui2). There's a while tonne of downsides that make it basically impossible to really get to 100% good on this.

> Under the hood it's still a terminal app, but from the UI author's perspective you're not fighting terminfo, CSI sequences, or varying terminal support. You get proper cursor tracking, scroll regions, and editable text areas out of the box.

All that is not true. Take scrolling for example. Each terminal emulator implements this subtly differently (mouse/trackpad/speed etc.), scrolling will never feel native unless you push the control of it into the terminal emulator.

There's more than that, each and every point on this becomes something that you have a gap between native and the terminal emulator that you're building to run in the terminal.

See https://github.com/openai/codex/issues/8344 and https://github.com/openai/codex/pull/9640

Claude code and gemini have both had various versions of the same idea at times. I'm unsure where they ended up.

joshka··on Stop Making TUIs
The problem succinctly is when the terminal emulator only sees cell values and instructions it can't do anything more with things. Two really good examples are implementing accessibility well, and scrolling / changing things above the terminal pane without rewriting the whole history.
joshka··on Stop Making TUIs
> as you're drawing interfaces with punctuation characters

Yeah, in my hypothetical new protocol, the character cell is still there, but it's not the element that drives the main abstraction. The layer above that being more semantic is where the smarts is. So borders, interactive areas, mouse hit mapping to elements, etc. ends up being built into the protocol rather than being a cell level thing. Cells having proper borders being similar to how they allow for underline, being able to have both images and text, ...

I think it's important to do something about the VT100 +kitchen sink stack and to have that as something that still works with that in areas of the new terminal proto. I've got a slopcoded rust library in the works that's about fully mapping the full list of terminal protocols and related things in one coherent library (VT100, ECMA48, DEC modes, iTerm extensions, Kitty, ...). From there I think making that fairly isomorphic with the terminal emulator layer (I have a RIIR ghostty slopfactor there too that fits).

I think ghostty / superlogical looks at this from the perspective of doing pty things and then transfering that calculated state across the wire while doing fun stuff with windows, panes etc. with replay and that sort of thing. I think this is valuable too.

I suspect that there's possibly a way to do a bit of a hybrid on this - normal pty things with the extra sidecar of more semantic structured, but it'll make things more complex for app builders on that side of things.

(On another comment I saw you mentioned "couldn't help noticing how much of the tedious work of putting a TUI together". I think probably textual in python is probably the best of the low ceremony libraries for making TUIs if you haven't checked that out yet. I think someday there will be a good Rust library that has that level of put together-ness. I have some ideas on this for ratatui and have seen a bunch of things that are directionally right for this, but nothing that gets close yet).

joshka··on Felony charges for citizen deleting phone data at US Border
Could he? To answer properly without possibly committing further possible crimes, he needed a lawyer familiar with whether that's allowable, something he was explicitly denied after asking multiple times. See https://www.courtlistener.com/docket/71998357/21/united-stat...
Page 1 of 34Next →