What the hell?
12,591 karma · joined April 17, 2018
[ my public key: https://keybase.io/holysz; my proof: https://keybase.io/holysz/sigs/X_fR6Li4U1P5lNXm-V0ccha73PNi0KSCQvtk70bfuWo ]
What the hell?
Imagine you spend $100k on average each month, but spend around Christmas rises to $1m because of the specifics of your industry. With a global spending limit, you can't distinguish between $200k of general spend increases due to Black Friday versus $200k in fraudulent 2FA SMS to South Sudan.
It's clearly a hobby project (constant breakages, the maintainer getting into politics, rejecting some safety mechanisms, the anti-LLM crusade, a strange focus on esoteric targets with little to no commercial significance), but the maintainer does not admit that it is a hobby project.
It makes me respect the Rust community even more.
Whereas I would say "Have they put this on Arxiv yet?", mathematicians would say "Have they put this on the Arxiv yet?".
If Arxiv was the only game in town and it allowed everything, the result would be more like GitHub or Substack. That is, you wouldn't be able to trust everything hosted there, and a lot of it would be low-quality promotional garbage, with filtering moved to a different layer.
You could do some kind of karma system for example, where the prestige of your institution, your own karma, number of citations, papers published in high impact-factor journals etc would affect karma weight. You could have journals be like "awesome-x" lists on Github, with authors who have already been published there able to endorse other papers. I believe the mathematics community is trying that.
At least, you should have some algorithm estimating marriage length / likelihood of divorce based on your (non-app) statistics, and then use that as one factor for your in-app calculations.
You don't know what these things do and what their effects really are (examples and syntax illustrative, but this is the kind of code that has disastrous effects when used carelessly), but you know they achieve your particular micro goal of "make things go fast" or "make this fit in packets on these strange industrial networks customer X has" or whatever.
Do you allow them acces to your school's website? Guess what's under one of these forgotten posts about a singing competition from 2012. Do you let them access some math exercise platform? Does it allow sending people invitations to work together and also let you set group names? That's a chatroom right there.
I think the coolest one I've heard about is prison tablets that didn't restrict access to WiFi settings, so people would set up personal hotspots and use access point SSIDs for communication between cells.
Javascript is plenty fast, but only if it has enough time to JIT the code and can keep it JITed in memory. For long-running backed services, it's plenty fast, for browser things, it's the only option, but for the command line, it adds needless startup time.
Agents make this even worse because they don't have a "sense of time", so if they accidentally do something which causes startup time to increase dramatically, they won't feel it like a human dev would, and won't immediately start optimizing. Unless you have some specific benchmarks in CI that fail any PR which makes the code too slow, agents will just make things slower and slower.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
P2P is great on desktop computers (and, in the modern age, Smart TV boxes), where your network connection has no data cap, there are no battery concerns, you're unlikely to be behind CGNAT (common for cellular data connections), and the app can surreptitiously minimize itself to tray instead of closing itself outright, making you an unwitting seed.
2. Why do we want fewer cats in our cities?
3. Would we rather fewer cats or fewer rats?
As far as I understand it, it's fingerprinting + IP reputation + cross-site IP rate limiting + raising the cost of a request by an order of magnitude because of the LLM call.
The last ingredient is what we forget about. $0.001 or whatever might not seem like much, but there's a difference between "a million requests a day costs you $5 in AWS bandwidth" and "a million requests a day costs you $1000 in VLLMs, $500 in premium bandwidth from residential proxies, + constant toil to update your scraper to keep up with what Google is doing."
(Figures illustrative).
Captchas worked well enough for a while, if you weren't visually impaired (or worse, deafblind), spoke English, could read (Latin) characters, and so on and so on, but they were always a hack. The hack is now failing.
The only sure-fire way to check for humanity is to ask somebody who tested it in-person, likely an ID-issuing government.
On the other hand, you can have concurrency without parallelism. Think a database where IO is the bottleneck, and you have multiple clients doing reading and writing at all once, potentially to the same table, in isolated transactions, on different db nodes which have to communicate. That's a lot of concurrency and nasty locks, even if you're running on a single core and wouldn't get much of a speedup from doing otherwise.
Fable was a huge leap in terms of model persistence and raw intelligence, but it still had terrible taste for human writing and code architecture. It would constantly keep making decisions which would achieve the desired objective (and make the code correct), but would bite you n years from now, and n years from now isn't RLVR checkable.
Opus 5.5 has a very different "feel" than anything else I've seen in this generation, though GPT-6 does seem to be moving in a similar direction. They have finally solved the writing part, and architectural taste also seems to have improved significantly.
I did a review of some GPT 6 Sol's code with Opus 5.5 yesterday, and it went "the code is correct, but there's a bunch of things here that could be simplified, and the split of responsibilities doesn't follow your established architectural layers" (which was true and exactly what I've noticed myself when reading the diff). I don't think I've ever seen a model do this before and actually be on-point.
What I'm thinking of is something far more opportunistic; the way to piggyback on an existing peer (or a set of such) as a pseudo-relay, if and only if network conditions allow, and this is actually something that is worth doing in the given situation.
Assume we have devices a, b and c, which are basically in different segments of the same network and have nice pings to each other. We're trying to ssh from a to c, but NAT traversal isn't possible. Both a and c can do NAT traversal to b.
Instead of a going all the way over to the DERP in WAW and then back to c again, it could go a->b->c instead.
In my experience, DERPs are pretty slow and have high latency (compared to not going off-network at all), but there's no real way to avoid them if everything you have is behind some sort of NAT, even if some of them are trivially traversable (think "devices can do UPNP").
So the car might need to stop on red multiple times (unless following a bus), while the bus will have most lights automatically turn green in front of it.
Not sure how it works in SF, but many cities across the world have moved to "in-card tickets." You tap your card on the way in, and the bus "knows" that this card number has a valid ticket attached to it. Depending on the particulars of the city and system, you might need to tap out too.
In some cities, this is done entirely offline. For cards which cannot guarantee offline payments (which many low-income / under-age individuals have, and such people are far more likely to use transit), this is done "on trust", with failed payments getting your card blacklisted until you go to the transit agency's office and get the situation rectified.
The blind community is benefiting enormously from coding agents. Game accessibility mods for everything under the sun (the big names in the last month or two are Civ V and Witcher, although there's plenty more), accessible 3rd party clients for annoying sites, people's favorite speech synthesizers ported to platforms they never ran on natively (or just straight down turned into portable C), plenty of small but nifty utilities and apps.
This is because that community's needs are amenable to what AI can do. Accessibility work (on somebody else's product) is a lot of demotivating and extremely difficult reverse engineering drudge work with a verifiable success criterion, and this is what AI excels at. Large-scale software dev is all about judgment and taste, and here, AI is not doing so well.
You see the same things with mathematics versus medicine. In math, the bottleneck is basically human attention, formal proofs in Lean are, again, drudge work with verifiable success. In medicine, the bottleneck is patients, paperwork and the lab environment, so even an omniscient LLM without the ability to pour fluid into a beaker wouldn't be that much of a productivity improvement.