8,936 karma · joined March 13, 2013
Please take the time to read and edit. Get rid of the emdashes and the pithy little "The quick start lands here. Nothing expires into a paywall." / "Sign up and pay in one flow — checkout follows registration." zingers that are meaningless or already implied.
Please stop allowing pragraphs like this:
> Bounds are per organization — one organization can hold several teams, and its message count is metered across all of them. The monthly allowance is the real limit; the burst ceiling exists only to stop runaway loops and sits far above a normal day’s work.
to appear on pages where you expect humans to read them. They are worse than useless to read.
It's wild that this is an accurate summation of their position here. It's equally wild that a company building and selling a technology as actually useful as an LLM - yes, I think they're useful, sue me - can't begin to think about how to make money with one.
I agree with the article - they'd better hope they are messianic super geniuses, because if not, they're gonna seem real stupid when we have all moved in in a few years. The technology will stay, but that doesn't mean they will.
I am serious, reply here, I'll sell you my onion contract for $4.50. Who's buying?
I come neither to praise AI nor to bury it. This is a case where the tool is useful.
The key is that this is work the attornies themselves do not want to do and which does not require their brains or expertise. It's not just "boring" but a bad use of their time.
It took a human attorney 20-30 minutes on average to manually copy-paste data from these PDFs into a spreadsheet (while also fixing any errors they found in the document and re-checking for quality).
Now, the AI copies everything into the spreadsheet in a small amount of time, and then the human reviews it. It takes maybe ~5-7 minutes to scroll to the appropriate pages in the document, read the lines vs the spreadsheet, and make corrections. So you've gone from 2-3 items an hour to ~8-10 items an hour.
Maybe you could pay someone to develop an OCR/ML application that could do this. But that project would never be profitable, even with the time savings. At the cost of a couple Claude subscriptions, it makes sense.
But now it's comparing already filled columns on a spreadsheet, not copy-pasting every single thing from an (often uncopyable) PDF.
They recently bought a Claude subscription and began using Claude to do the initial read of the documents and output JSON they can import into their internal systems. The work still must be reviewed by an attorney - Claude is nowhere near making the kinds of judgments a lawyer would make about this content - but it has increased their throughput from 2-3 documents an hour to 8-10 documents an hour by killing the busy work.
LLMs have great advantages for this kind of work - but not for decision-making. I just don't see OpenAI ever admitting that.
(I've left some details intentionally vague because this is a very specific area of law and I don't want my friends to be identified without their consent.)
The truth is that given the legal and ethical constraints available, routing is already relatively well-optimized in many places. (There are some places, like China's hybrid civilian-military airspace, where those constraints are extremely nasty and basically make the problem impossible to solve better than it is.)
Aviation tries to solve the fuel consumption problem by designing planes to use less fuel, rather than by flying more optimal routes. There is a massive effort in aviation pretty much all the time to try to use less fuel, for every reason from cost to environmental impact to the relative accessibility of jet fuel now and into the future. It's the reason engines have gotten physically larger; it's a big part of why we don't use supersonic jets anymore as well.
The efforts have been pretty sucessful: https://theicct.org/maximizing-aircraft-fuel-efficiency-desi...
There's no hubris in identifying a problem, and fairly little in thinking you might be able to help solve it, but there's a good bit in imagining nobody else has done it.
Also worth noting, tangentially, though - the reason planes are fueled as they are is because if you run out of fuel in a car, you're just stranded. In an airplane, you're quite possibly dead. If you only load the plane with exactly as much fuel as it needs to get to its destination, you run the risk of not having enough to safely divert.
On the other hand, planes also have maximum landing weights, so a plane which must divert or land in an emergency often has to dump fuel for a while first. That's more unusual and not something you can easily plan for, but it's part of the problem too.
It increasingly feels like the power of these agents is less that they find things humans COULDN'T find, and more that they find many things much more quickly than humans would bother to do.
I don't know if this is a great advert for Strix over other agents - what did their agent do that Claude or Codex couldn't? It didn't do anything that I couldn't do, if I wanted to.
There's no shame here, this was a mistake, probably made by a human, and ultimately corrected. Nobody seems upset by the outcome!
It wouldn't show up in this study at all, because it's online, not in-person. These particular friends live thousands of miles from me.
I also have dedicated nights of the week for seeing my in-person friends, and the amount that I do that has stayed pretty consistent (outside of COVID), but it's never been nearly as much as my virtual friend time. I have friend groups more than a decade old that I've maintained this way.
(The mix is about 75% people I met first in the real world, 25% people I met first or have only met online, if that matters.)
A $10,000 car and a $80,000 car are going to be markedly different. Same for a $100,000 house vs a $800,000 house.
This $2,000 phone does not seem 8 times better than my $250 phone.
I assume the people who would buy something like this have enough money not to care, and also would not care about the second part of my comment, because they want New Thing or use their phones enough for it to be useful.
I guess it might be the case that the only way to make the folding screen thing even remotely work is a very, very expensive manufacturing process. But as is often the case, I don't see how it's worth it?
I want my phones smaller and with Touch ID, not enormous and wrinkly.
I can admire and enjoy the work of coding agents, which mostly just help solve problems and save me time. I can appreciate that there are creative uses for agents with text, that they could save people time, and that the environmental and safety concerns potentially have solutions.
I just can't see my way to appreciating these art agents, or any use of AI as a replacement for a human creative process. It's not because their output is bad, anymore (though sometimes it still isn't great). It's because it is bereft.
Even if I could, nobody in my life would be okay with my using them alongside my actual creative work. In fact, they'd be pretty upset if I did. And if I found out they'd been sharing stories with me written by AI, I'd be mad, too.
I'd rather just see the prompts.
What is it that Anil Dash thought the purpose of these firms was?
Please share the information you have which contradicts the conclusions I have drawn from Github's statement.
(And we know they're liars. They report very few of the actual incidents they have; see for example https://mrshu.github.io/github-statuses/)