HNHacker News
TopNewBestAskShowJobs

stult

1,789 karma · joined January 15, 2016

submissionscomments
stult··on UK military jamming other nations' satellites to defend itself, BBC told
I assume in that situation they would be targeted specifically with ASATs at the beginning of the conflict, with the denial of LEO then preventing us from launching replacement sats.
stult··on UK military jamming other nations' satellites to defend itself, BBC told
We also have visual navigation systems that can use satellite imagery to produce guidance packages. They will become increasingly less useful the longer the imagery goes without updates after a Kessler syndrome type event, though. Another strategy, which I believe the Ukrainians have been experimenting with, is scattering small radio beacons throughout the battlespace to provide a network of navigation aids.

In a true Kessler scenario where LEO is denied for an extended period of time, I expect a variation of that strategy is how we would replace GPS for most civilian use cases. Likely by using cell towers to triangulate, possibly enhanced with additional transmitters to improve accuracy and coverage. Pure cell tower triangulation would have much worse coverage than GPS, obviously especially in rural areas where cell service is often already poor, but also anywhere that getting a signal from multiple towers isn't easy or possible. Cell towers are never distant enough to use the relativity-based timing tricks GPS relies on, so multiple signals would be required to triangulate. We'd likely end up with substantially less precise positioning, maybe something more on the order of the 10-20m of error we used to see with GPS before the US stopped limiting civilian access to the deaccuratized version. Cell tower triangulation also obviously doesn't work for aviation or oceanic shipping, where trident-like inertial+stellar nav would probably be the only workable option.

stult··on Pentagon says overreliance on AI contributed to missile strike on Iran school
This happened after I stopped working for the DOD, but I am 99% sure the program in question during this incident was something I worked on. It was effectively an anomaly detection system for identifying any weird behaviors detectable in WAMI (wide area motion imagery, which is high resolution, low framerate drone video data that covers a large geographic area). As an example from the unclassified sample project they asked us to do as part of the bidding process, our code flagged a car doing donuts in a parking lot (the sample data was just from a random US city on a random day because this was an unclassified open bidding process). The system was not and never was intended to be a "terrorist" detection system. The idea was always to flag potentially interesting events for a human analyst to follow up on. The disconnect it took for someone to blindly trust the program as a target acquisition system boggles my mind. Someone doing donuts in a parking lot should not be automatically targeted with hellfire missiles, yet in practice that is how they were using the tool. I still lose sleep over it.
stult··on Wall Street Is Growing Skeptical of the Data Center Boom
Bullshit. Even if you don't think LLMs alone are enabling that level of advancement, GPUs in general very much are. They are the critical enabling technology underlying self-driving cars, autonomous drones, camera-based robotics controllers, faster/better drug discovery, faster/better modeling protein structure and design, ML weather forecasting, materials discovery, and many other applications.
stult··on Wall Street Is Growing Skeptical of the Data Center Boom
You are not disproving but rather reinforcing my point. Amazon survived the dotcom bubble precisely because they could hope to do more than deliver physical books. It was a more flexible, generally applicable business model than pets.com, and they started out targeting a market more susceptible to ecommerce conversion than pet food delivery proved to be. Similarly, a company selling a product that can only be used for the limited use cases we have found for chatbot-style text generating LLMs will be more likely to fail than a company selling something that is more generally useful. Data centers are more generally useful than LLMs, and the companies building them are more likely to survive than those whose fortunes are pinned entirely to unrealistically high expectations for replacing white collar workers with LLMs.
stult··on Wall Street Is Growing Skeptical of the Data Center Boom
I think actually it's quite the opposite and you're the one conflating the impacts on a narrow market space with the larger economy. Right now, AI dominates growth in the stock market and demand for chips, but it is by no means encompassing the whole economy. Neither the stock market nor the chip sector are the entire economy. My grocery store isn't going to go out of business if Anthropic does.

A lot of investors may lose money as a result of the bubble bursting, but that does not mean the underlying asset will be forever worthless, just that it didn't provide sufficient returns sufficiently quickly to justify the upfront investment for the investors that funded it, at the time they made that investment. A different investor who could afford to ride out a period of reduced demand might have an entirely different experience.

Imagine you take out a five year loan to buy a truck to deliver packages, but then for the first two years you operate it, gas prices are elevated so you struggle to make the payments on your loan and end up not making as much money as you had hoped to originally, perhaps even to the point you need to declare bankruptcy and sell of the vehicle. But if not and then gas prices drop back down and demand shoots up for the last three years of the loan, you could then make up the difference. If the first two years drive you into bankruptcy, that is difficult for you, but amortized over the entire five year loan period, the truck may actually have been a profitable investment for someone who could have afforded to ride out the first two years. And just because you go bankrupt doesn't mean the truck stops being a valuable asset, it's just that you don't end up benefiting personally from that value because your timing was bad. From a macro perspective, the overall economy doesn't suffer, except to the extent that it might have been more efficient to invest the capital that went into procuring the truck elsewhere during those first two years. But only possibly and only on the margins, because the truck remains a profitable investment over the course of its entire lifetime.

Don't confuse the success or failure of individual investors or businesses with the success or failure of the overall economy. Current data center build outs premised on fanciful projections of demand for LLMs may end up not being profitable in the short term while still being profitable over the entire productive lifetime of the asset, if sufficient demand is found elsewhere or if demand for LLMs picks up later. Similarly, I anticipate at worst we will see chip prices plateau for a while if there is a pullback in LLM demand, but we won't see them fall and they will continue to rise over the longer term as more demand is generated elsewhere.

As one small example: we have barely begun to scratch the surface of what we can achieve with robotics. Think about the demand for video processing if you have tens of thousands of robots stocking shelves in supermarkets generating video all day long. On board processing will of course be the obvious primary demand for chips, which doesn't benefit data centers, but central processing of video to extract useful data from the entire fleet will generate demand for data centers. As will large scale training jobs. Now multiply that thinking across the entire scope of industries where robotics may be useful for replacing human labor, and you're talking about an extremely significant amount of valuable computational work.

stult··on Wall Street is growing skeptical of the data center boom
That math is very much general purpose. It's just linear algebra under the hood, and linear algebra is a wildly useful toolkit with an unbounded set of potential applications, including many applications other than LLMs. That's how something developed for gaming evolved into something used for text processing and generation. For example, we have barely begun to scratch the surface of what we can achieve with robotics, and I have no doubt that further advances in robotics will generate an enormous amount of computer vision work that GPUs are extremely well designed to handle.
stult··on Wall Street Is Growing Skeptical of the Data Center Boom
Fair, but my point is that a data center is not pets.com. They are a much more flexible asset, and are even more flexible than AI itself. Pets.com failed because the physical delivery infrastructure and consumer buying habits that support Chewy today did not yet exist and took years to create. Data centers can be repurposed from supporting the current generation of LLM-based AI models to an infinite variety of use cases. Pets.com could only ever hope to deliver pet food.
stult··on Wall Street is growing skeptical of the data center boom
I have generally been bullish on data centers independent of how AI demand/load evolves. People will find a way to use that compute, even if it isn't the precise use we expect. As scale increases and the price per computation comes down, we will be able to brute force solutions to problems that otherwise would be intractable or prohibitively expensive to solve. And there are effectively an infinite set of those problems.

This tracks the evolution of how we use cloud computing and GPUs over the last 20 years. Cloud computing was originally just about saving companies from needing to maintain their own server racks, but has unlocked previously unserviced demand by allowing people to build an app or service and scale it to meet rapidly rising use without needing to invest a ton upfront in hardware. Suddenly a hobbyist could spin something up in their spare time that previously required many thousands of dollars of investment. Or when someone wants to run a single large scale computation, they can do it without needing to waste capital on maintaining idling servers to meet an occasional demand spike, like when a company I used to work at moved from running atmospheric calculations on a server in a closet to the cloud and were able to achieve double digit accuracy increases with the increased scale, while spending less overall on batch computing jobs.

GPUs, as the name implies, were created for graphics, and primarily for gaming graphics, but then more or less accidentally ended up enabling the present AI boom, which depends on a scale of computation that would have been impossible with older CPU architectures. Maybe someone at some point predicted this, but I think for the vast majority of people, it was extremely surprising that a niche gaming product would enable an industrial revolution level technological leap forward.

LLMs are just one way that increased compute scale unlocks seemingly magical results, but they are far from the only example and I have no doubt that there are many unknown examples remaining to be discovered yet.

stult··on Wall Street is growing skeptical of the data center boom
Do you mean is there anyone who thinks that the future is NOT local models anyway? Since that seems to be the gist of your other points
stult··on AI researchers debate how close we are to recursive self-improvement
Humanity might be able to, humans cannot
stult··on Netflix employee fired for sharing personal details in retreat trust exercise
I'm shocked Netflix hasn't just settled to keep this quiet, because the entire situation makes them look not only like an abusive, discriminatory employer but even worse, horrifically incompetent at being abusive and discriminatory. Plenty of companies treat their employees like shit, that's not news. But any competently managed company knows how to do so without incurring massive legal liability like this, and it speaks to a remarkable lack of professionalism and basic competence at the leadership level and within the legal and HR teams.

If the accusations are true as reported, their in-house counsel fucked up terribly by assuming all ketamine use is illegal and recreational without bothering to conduct the very cursory investigation required to determine that it is in fact a perfectly legitimate option for treating treatment-resistant depression. That single error drove them to fire an otherwise exemplary employee for an obviously illegal reason in a manner that screwed him out of significant compensation including severance.

This was an obvious, unforced error by their legal team. Competent lawyers and competent corporate leadership would recognize that the relatively small amount of money they may save by litigating this issue is not worth the damage it is very predictably inflicting on their corporate culture and public brand. The lawyer who publicly admitted that ketamine factored into their decision should be immediately fired because he is demonstrably a terrible fucking lawyer, and he seems to be representative of the quality of legal advice Netflix is receiving, because any attorney with half a brain would have immediately recognized both the underlying legal error that incurred liability for the company and the business context that makes immediate heavily NDA'd settlement the optimal strategy for defending, even were the original legal error not so egregious.

My main takeaway from this situation is that Netflix is managed by remarkably incompetent leaders who are advised by even more remarkably incompetent lawyers.

stult··on Netflix employee fired for sharing personal details in retreat trust exercise
You can't have an HR policy that discriminates against people for receiving mental health treatment from a physician, that is plainly illegal under multiple statutes both federally and in the state of California. This was not a situation where he was on ketamine while at work, and there are no accusations that his ketamine use even affected his work at all, even incidentally or indirectly. Even if it had, Netflix would be required to make a reasonable accommodation for him. But we know it did not affect his work because he received treatment in 2022 and Netflix only found about it in 2026 after he disclosed the treatment. Besides, he is a movie executive, it's not like he is operating heavy machinery. Ketamine is unlikely to have any effect on him that would in any way expose Netflix to liability.

Netflix seems to have simply assumed that all ketamine use is illegal and recreational, without bothering to conduct even the most cursory of investigations, which would have revealed that ketamine is in fact a clinically acceptable treatment for treatment-resistant depression. This is what happens when you have utterly incompetent idiots for in-house counsel and HR.

stult··on Netflix employee fired for sharing personal details in retreat trust exercise
Yes, it seems likely they violated California's Fair Employment and Housing Act (FEHA) and the federal Americans with Disabilities Act (ADA) based on the details I have seen so far. The EEOC’s mental-health guidance states both that an employee may choose to discuss his condition with coworkers and that the employer may not discriminate against him for doing so. https://www.eeoc.gov/laws/guidance/depression-ptsd-other-men...

That's only guidance, not binding law in and of itself, but it means that the employer can only discourage employees from discussing their mental health if they do so in a neutral manner that applies equally to all employees in a non-discriminatory manner, which was self-evidently not the case here.

Notably, the EEOC guidance standards are not wholly unlimited. Employers can take action against an employee if their discussion of their mental health is repeatedly inappropriate or disruptive. e.g., if someone is constantly cornering coworkers to trauma dump suicidal thoughts on them or something like that. But only in truly extreme cases where an accommodation is impossible (e.g., if they are violent) or if they fail to adjust their behavior after receiving feedback that it has become unacceptable. A single incident that does not appear to have even made any other employee uncomfortable certainly doesn't suffice.

Netflix's only real defense here would be to argue that they would have fired him anyway without the disclosure, in response to other issues. That is a challenging defense to make given the admission their lawyer made that the ketamine treatment factored into their decision to fire, but the ADA and FEHA require "but for" causation, i.e. he is protected from firing if they would not have fired him but for the admission of ketamine treatment. I doubt they will succeed given the relative triviality of their other accusations (somewhat excessive profanity in a context where a fair amount of profanity was considered acceptable) and the entirely inoffensive party trick (at least in a context where the CEO has been repeatedly photographed drinking alcohol at company events and the alcohol at this event was provided by the company). I seriously doubt they will be able to point to any similarly situated employees that Netflix has previously fired solely for profanity or consuming alcohol that the company provided to them (and from a PR perspective that would almost be worse for them to admit). Frankly, the profanity feedback also seems like the very common scenario where a manager is required to provide regular feedback but can't think of anything constructive to say because the employee is a high performer, so they reach for something funny and minor just to check the box. In the absence of extensive complaints from coworkers, it is unlikely to overcome the company lawyer's outright admission that the ketamine use was a factor in their decision.

What appears to have happened here is Netflix has awful in-house counsel and/or HR who utterly failed to carry out their duties with even minimal competence. They appear to have simply assumed that all ketamine use is automatically illegal recreational drug abuse without bothering to investigate whether that is true in general or in this particular case before escalating to the most extreme possible reaction. Those employees are the ones who should be fired, not only because of the gross incompetence it takes to so egregiously violate America's otherwise absurdly weak legal protections for workers' rights, but because they did so in a manner that seems practically designed to permanently destroy employee trust in the company while inviting unwelcome public scrutiny of their potentially discriminatory employment practices.

This is a fairly straight forward discriminatory-causation story: Netflix invited vulnerability and thus potential mental health related disclosures, learned of a psychiatric history, reframed treatment as drug misconduct, and then expressly treated it as a termination factor. It's honestly pretty rare to see such an obvious example of this kind of discrimination, usually companies do a better job covering it up with a pretext, and usually their *lawyers* aren't so unbelievably fucking stupid as to admit publicly to the discriminatory decision.

stult··on AI companies spend record sums on Washington lobbying
That is a hopelessly naive description of lobbying, to the point that I assume you are being intentionally disingenuous. Lobbyists commonly funnel money to politicians via fundraiser bundling, donations to outside spenders like PACs, and other mechanisms that get around the weak remaining restrictions on direct campaign contributions that SCOTUS has not yet gutted. A lobbyist may not donate directly to a politician's campaign, but they get a group of their friends together for a fundraiser to effectively donate on their behalf (bundling). Or they donate to outside groups who then spend to support the politician's campaign, e.g. by funding advertising that benefits the candidate or hurts their opponent. Like when AIPAC sinks millions into ads against progressive candidates, that's tightly connected to their lobbying on behalf of Israel and you better believe that the politicians benefiting from that spending are well aware of how it is connected to their willingness to adopt the lobbyists' positions.
stult··on We Tested Nonstick Cookware: Coatings Don't Need to Look Worn to Shed Particles
I stopped cooking with nonstick pans over 15 years ago, and the only thing I've found that can consistently be a bit tricky is eggs, but really only if the pan isn't hot enough yet when you add the butter or oil, you don't use enough butter/oil, you try to scramble the eggs in the pan, or you try to lift the eggs too soon. So arguably it takes a little more attention to detail to cook eggs on stainless steel, but in general I really don't think there is anything that requires nonstick. But even for the limited list of items that are easier to prepare on nonstick, it's usually just a matter of learning how to adjust your technique to match the cookware.

I made the switch to stainless steel and cast iron pretty much on a whim after I first read about Robert Bilott's lawsuits against DuPont, at which point I discovered it made almost no difference to my quality of life at all to avoid nonstick. To such an extent that I have never been even slightly tempted to switch back, despite honestly not feeling a particularly strong motivation to avoid nonstick in the first place. I've just always figured cooking with steel is an entirely painless way to reduce my potential exposure to dangerous chemicals, so even if it ends up being pointless, I haven't lost anything of value anyway.

stult··on Private healthcare makes industries less innovative. It's time for change
I don't get why decoupling insurance from employment isn't a bigger part of the healthcare reform discourse. It's the most obvious, easiest, lowest hanging fruit for improving the US healthcare system, and almost uniquely politically acceptable to both Democratic and Republican voters. Decoupling also doesn't entail disrupting any of the established incumbents (pharma, AMA, insurance companies, etc.) who typically kill any reform. The policy is also arguably market-oriented, or at least is not a form of government intervention that would otherwise offend Republicans, yet benefits individuals over large companies so can appeal to Dems too.

The Republicans have utterly failed to offer *any* alternative at all to Obamacare after 16 years, which is insane for many reasons, but especially considering that they could have put decoupling insurance from employment at the heart of an alternative (even more than the ACA) market-oriented reform. Letting individuals choose their own insurance would put more pressure on insurance companies to actually deliver quality care because, unlike employers, individuals are highly sensitive to the quality of care they receive and are more likely to switch insurers if the insurer unreasonably denies claims or otherwise fails to deliver a decent product. On the other side, individuals are also more sensitive to costs so with greater control over their plan choices would be more able to pick plans that are structured to match their anticipated costs.

In other words, individual insurance would improve the functioning of the market, which should in theory appeal to Republicans. Instead, they only focus on making individuals pay more at the point of care, on the theory that cost-sharing discourages insured people from over-consuming healthcare. Yet they don't think about how the current system makes it impossible for individuals to be truly cost sensitive when selecting plans.

I would guess the Republicans haven't adopted the idea because they are not really in favor of free markets, they are in bed with big business, and employer-provided insurance tends to favor larger companies that can demand better terms from insurers, that can amortize the HR overhead over a larger number of employees, and that tend to benefit disproportionately from lower labor mobility compared to small business (people are more likely to be involuntarily locked into a big company job for security than at a small business, and it's harder for people who are worried about maintaining quality healthcare access to move to a startup, reducing the competition experienced by the bigger companies).

stult··on DOGE is done. What happened to its records?
Every other country is responsible. But in proportion to their wealth and power, and the US is far and away the wealthiest and most powerful country in the world so we bear an outsized share of the responsibility for keeping the world stable and safe. We built the international order around free trade to suit us, and it works exceptionally well to enrich us, so we have a strong motivation to ensure the stability of that international order by reducing the causes of international conflict and civil disorder like hunger or disease.

Regardless, even if your amoral nihilism were correct rather than the hallmark of a morally repulsive psychopath with the imaginative capacity of a tapeworm, there were two things DOGE did wrong. First, much of the actual damage they caused was not from the US cutting aid per se, but rather how quickly and with such little warning they cut aid. DOGE denied aid recipients that were relying on the US to keep people alive with life saving medicine and food a reasonable opportunity to make alternative arrangements. People are dying not because the rest of the world is incapable of supplying ARVs to HIV patients in Africa, but because we took those critical life saving drugs away in a manner that made it impossible for the people depending on us to adapt. We killed many those people. You can't just stop taking ARVs and be OK, and someone make a few hundred dollars per year in rural Africa is not well positioned to find alternative suppliers. Many thousands of HIV positive pregnant women who would otherwise have been able to give birth to a child without the child contracting HIV now have to figure out how to survive HIV themselves and how to care for a child needlessly infected with HIV. Many of these people are now dead because of our negligence, arrogance, and stupidity. Because of your negligence, arrogance, and stupidity.

And it didn't even save us any money to do it that way, it was nothing less than abject cruelty and racism. DOGE let perfectly good drugs and food we had already paid for go to waste in warehouses rather than allow it to be delivered. For literally no reason, it saved us not a single penny and instead deprived many innocent people.

Above all, cutting aid like this was unbelievably stupid and self-defeating. Because even if psychopaths like you are objectively correct about reality (you aren't), when the world's richest man and the world's richest country murder millions of the world's poorest people for literally no reason, that makes us look really bad to the rest of the world. And then they do not cooperate with us. See, e.g., Trump begging the Europeans he so frequently attempts to bully for help with Iran. Idiots like you and Musk have trashed America's hard-earned reputation as benevolent superpower. That will cost us trade deals. It will force our allies to hedge against us by making trade deals with China instead as a counterbalance, as Canada has begun to do. The US is so phenomenally wealthy we can afford to be sociopathic assholes to the rest of the world for a little while at least, but it is difficult to overstate just how naive, ignorant, and outright moronic you are if you think that doesn't come at a price to American interests that far outweighs the negligible amounts of foreign aid spending DOGE illegally cut.

And I won't even get into the illegality of an unelected jackass impounding congressionally authorized spending because you do not seem like the sort of person who has any concept whatsoever of the value or importance of the rule of law and respect for the constitutional order.

stult··on Court Records Should Be Free
> A similar dilemma faces PACER. Overwhelmingly, PACER is used by attorneys, who are generally well-compensated professionals with a whole host of protectionist policies insulating them from market forces.

One of those protectionist policies is charging for PACER itself.

The costs of running PACER are absolutely trivial in comparison to the costs of running the judiciary. To the point that even bringing the point up is disingenuous to the point that it discredits everything else you say.

Case law is law. People are required to obey the law. They should be able to access the law so they can know how to follow it. It's that simple.

stult··on Agentic coding and persistent returns to expertise
This analysis is almost bafflingly stupid.

They conflate domain expertise with coding expertise, and then assess that people with domain expertise demonstrate great success at coding tasks, which suggests coding agents are so good at writing code that domain experts can now cut software engineering experts out of the loop entirely. Yet if you look at their classifier, it classifies user expertise almost exclusively according to standards that measure expertise in coding. No wonder it predicts success at coding tasks. This just in: people who know how to develop software are better at developing software. What a fucking joke.

stult··on Peopleless economy? Not technically impossible
Literally does not work for me at all, just keeps reloading and spinning. I'm not even on my VPN. Really frustrating dealing with these things lately.
stult··on Peopleless economy? Not technically impossible
Man, I really freaking hate cloudflare bot checks. I can't even access this site, which I presume is just a few kilobytes of simple static HTML with some straightforward text content. I shouldn't have to work this hard to prove I'm human, it's exhausting
stult··on Anthropic flies staff to D.C. to clean up White House fight
That's not true. ITAR and security clearance are entirely separate regulatory regimes. For ITAR purposes, being a permanent resident is good enough. I used to work for a defense contractor, and we hired plenty of green card holders. They were not in general assigned to work that required a clearance, but plenty of defense-sensitive, ITAR-controlled work is done by green card holders.
stult··on Law Enforcement's "Warrior" Problem (2015)
Troops is almost exclusively used to refer to Army personnel in the US.
stult··on Law Enforcement's "Warrior" Problem (2015)
I think warfighter crept into the lexicon for somewhat understandable reasons, likely because of the increasing frequency of joint operations (i.e., operations involving more than one branch of the military working together) after Vietnam, combined with the long-standing military tradition whereby members of any given branch take great offense if you refer to them using the wrong professional label (i.e., soldier, sailor, ~crayon-eater~ marine, airman, space cadet). That is, we can't just call all of them soldiers because only members of the Army are soldiers, so if for example you call a mixed group of marines and soldiers "soldiers", the marines will make their displeasure known to you, aggressively and in no uncertain terms.

When you're talking about DoD stuff all day long and frequently need to refer generically to the mixed personnel involved in a joint operation, warfighters beats saying Soldier-Sailor-Marine-Airman-Spacecase. All the other alternative phrases for the concept of "person employed by the military in one of the five combat arms branches" are variations on "member" and tend to sound clunky or be overly verbose, like "service member" or "member of the military." Try saying "service members" 50 times per day. Trust me, it gets old fast.

And frankly I don't see the problem with warfighter. Fighting wars is quite literally what they do, and pretending otherwise does a disservice to the truth and risks papering over the deadly seriousness of their work. Warfighter is also quite distinct from "warrior," which carries connotations of a specifically aggressive and barbaric flavor of professional violence purveyor. Like you say, it sounds like some atavistic hereditary soldier caste for whom violence is a sacred vocation joyfully undertaken rather than a solemn duty carried out only with great reluctance and forbearance.

stult··on Harness engineering: Leveraging Codex in an agent-first world
> - Do we have reasons to care about LOC in a world where we don't write code manually? What happens to token usage numbers when the codebase is significantly larger?

Yes, at least to the extent that we care about context windows and tokens consumed by coding agents processing code that is ultimately irrelevant to their assigned task.

Anecdotally, I've found keeping file sizes small has been important for agentic coding not just to maintain human readability, but also for optimizing agent performance, precisely because it limits the amount of incidental context they load while working a problem, because they generally load entire files rather than just parsing the part relevant to their current assignment as a human might. That smaller file size thus reduces input noise and the LLM generates a tighter solution, which in turn reduces input noise for future solutions. Or at least this strategy avoids a death spiral into exploding context length.

I expect (but cannot currently prove) that keeping overall LOC down yields similar benefits even when file sizes are kept small because it spares the LLM from parsing potentially relevant files that prove irrelevant to its current task.

stult··on AI outperforms law professors in Stanford Law study
I'm not sure about that, I actually think planning may be just as important in both domains. Outlining before drafting is an almost universal best practice in legal writing that is drilled into law students to the point that outlining as exam prep is something students spend several weeks on each semester. So personally I always have a fairly detailed implementation plan in the form of an outline before I ask an LLM to draft a more detailed legal document.

I've also adopted an AI coding workflow that involves a lot of planning, although I actually write very little of the plan myself anymore. I have a chain of slash commands like this: create-issue -> plan-issue -> build-plan -> pr-into-dev. I write a relatively brief description of what I want accomplished to create the issue, and then the agent fleshes out my description with more detailed requirements and acceptance criteria. I review the issue description, and the LLM often identifies open questions I failed to consider, so I revise as necessary and then the agent posts the description to the GH issue. I have planning separated because I often create issues quickly when something occurs to me and then circle back at a later date to implement, and want the agent to create the concrete implementation plan with an up-to-date snapshot of the code in context. Then I review that again, adjusting as necessary, and then the agent posts the result as a comment on the original issue.

Like you, I've found this detailed planning makes for a very robust coding agent (again, also in combination with the aforementioned best practices, especially requiring 100% test coverage because forcing it to exercise every line of code avoids hallucinated dummy tests that assert on nothing). Interestingly in comparison to legal writing, I also rely on the agent to decompose complex tasks into separate issues or subissues as appropriate, which is something that is never necessary for legal analysis because pretty much every every legal analysis can be one-shotted.

For legal writing, my workflow is nowhere near as structured as that. For context, I have only ever used LLMs for drafting what are effectively emails to clients or memoranda of law for clients that are a step up in complexity and formality from an email. So not something that will be filed with a court necessarily but very much in the same format and style as a formal motion that would be submitted to a court on behalf of a client. And never a contract, will, or judicial opinion, nor a communication with a counterparty like a demand letter or C&D. So YMMV for other types of legal writing.

That said, I typically start drafting a memo by conversing casually with an agent to explore the general boundaries of an issue I am evaluating, by identifying relevant sources of law, potentially related issues, and the analytical process I need to follow (i.e., what issues to evaluate and what order to evaluate them in, more or the less the analytical "algorithm"). Once I have a good sense of that algorithm, I put together a high level outline and then ask the agent to draft a detailed memo around that outline. Or at least that's what I used to do before the last few months, since when the models have matured to the point where I increasingly just ask the agent to write the outline based on the conversation we had, then review that, then ask it to write the memo based on the outline.

As I have been writing this, it occurs to me that actually I am following almost the exact same process for writing code and for writing legal memos, and should probably distill the legal writing process into a similarly well-structured set of chained skills/slash commands. In both domains, I describe an issue at a high level, get the LLM to fill in some of the broad outline level details, review that, then get the LLM to implement the complete final product. (Also perhaps worth noting while I do occasionally conduct general high level research by talking to a frontier lab LLM, I have always used locally hosted OS/OW models for drafting memos where I need to provide concrete, specific factual information about clients to the LLM, to avoid attorney-client privilege issues, so the quality has lagged behind the frontier models, which is part of why I haven't developed this workflow into as structured of an approach as I have for coding).

In both coding and legal contexts, I think that this planning or outlining step is critical not (or not just) because it forces the agent to create a higher quality product, but because it forces me to review what I am asking the agent to do at a sufficiently detailed level that I can catch errors before they crop up in the implementation. A lot of the time, the errors that occur if I skip this step aren't because the LLM has made any clear mistake, but because I failed to specify some aspect of the task and the LLM is forced to guess at what I really intended, which is where agents often struggle.

So I guess I would tentatively suggest that legal writing does in fact benefit from thorough planning, though it is hard for me to quantify whether those benefits are greater or less than the comparable benefits for code.

stult··on AI outperforms law professors in Stanford Law study
I could not agree more. A simple example: it boggles my mind how every state organizes their statutes in entirely dissimilar ways. I'm not sure there's a need for every state to have slightly different wording for a murder statute in the first place, but even assuming there is, why do they all have to be scattered around in different code sections instead of every state just following some consistent convention like always putting the murder statute at Title V, Section 1.4 (or whatever the case may be, that's just a random invented example).

For murder that's not such a huge deal because the statutes are typically easy to track down and don't really differ all that much substantively, but once you get really into the weeds on something like commercial contracts it can be a huge pain to do cross-jurisdictional research.

And that's just a tiny, super obvious example of how impenetrable statutory law is, which isn't even the really pernicious problem. Case law is infinitely worse. It makes me absolutely furious how difficult legal research still is. The Westlaw/LexisNexis duopoly is a moral crime and wildly destructive to the quality of government in this country. Every single written court opinion should be publicly available for free on the internet in an easily searched format. It would cost practically nothing to achieve. We're talking about less text than Wikipedia hosts. Yet still many states make it almost impossible to access case law. Even though these cases are law. Binding law that we are supposed to follow, yet we cannot even easily access. It's insane, and largely perpetuated by the complacency of lawyers who can charge others for what should be free, the lobbying of the duopoly, and the incompetence of politicians.

If all of the laws were consistently available and stored in reasonable, consistent citation formats (I would settle for hyperlinking as a replacement for the rat's nest of wildly varying jurisdiction-specific citation systems), it would even be possible to introduce a form of unit testing for legal drafting that would allow us to automatically verify if the LLM hallucinated a citation.

It also doesn't help that we (for what were at the time very good reasons) moved away from the system of legal writs that used to provide fairly standardized, almost "cut and paste" templates for legal filings. So now every legal document (filings, memos, contracts, court opinions, statutes) is drafted like a bespoke, artisanal creation with few strict structural or stylistic conventions. That makes automated interpretation much harder than it needs to be.

stult··on AI outperforms law professors in Stanford Law study
I would agree with this point and as I explained in a comment replying to the GP comment above, that atrophy is far more dangerous in the legal field than it is with code because legal documents do not benefit from the structural safeguards available for code, like automated testing, static typing, static analysis tools, etc. IME with legal LLMs so far, they are easily in that most dangerous valley where they can lull you into a false sense of security while still introducing extremely dangerous mistakes that are frequently difficult to detect without very careful reading.

The danger of those mistakes creeping in also grows exponentially the farther a lawyer strays from their core legal expertise. There are a few statutes I know inside and out, and I can spot LLM analytical errors related to them in a split second, but once I venture out into domains where I am not an expert (but where I am nevertheless reasonably qualified to practice), it becomes much harder to spot drafting mistakes because I have not refreshed my own understanding of the law by reviewing the relevant cases or statutes as I would when drafting the analysis myself from scratch.

stult··on AI outperforms law professors in Stanford Law study
IME so far (as both a lawyer and a software engineer), LLM error rates when drafting code and legal documents are reasonably comparable, but it's more problematic in the legal context because legal documents do not benefit from many of the structural safeguards available for code. For legal documents, there are no automated tests, no static typing, no test environments, no logging/observability instrumentation, no sandboxing.

The time lag between drafting and "deployment" also makes for much less effective, much more expensive debugging loops. You can deploy your code to prod in seconds, see an error pop up in the logs, and immediately start debugging. But it will take at a minimum days and frequently as long as several years before an error in a contract or a court filing will be detected, and often the error is beyond correction at that point. Thus, the errors are both more difficult to detect and to resolve.

And the consequences of error are often much greater, both because they are not correctable and because a legal error may risk someone's life, liberty, or substantial property. Although that's not categorically the case, obviously bugs in certain safety critical systems can be as bad or even worse than legal mistakes. But in general, most software is lower stakes than most legal writing.

On the flip side, LLMs do seem to do a better job with basic style and structure for legal documents compared to code. Things like following IRAC format, citing assertions of law (although hallucination remains an issue), and writing comprehensible sentences. These would be the equivalents in code to best practices like good comments, cohesion, consistent use of design patterns, test coverage, clear variable names, DRY, etc. Although the better performance on those more qualitative metrics may just be because even the longest legal documents are typically simpler in structure and have fewer lines of text than a large, complex codebase. Or maybe it's because LLMs are trained on natural language text more than on code. Or because natural language is more forgiving than code, in that minor variation in diction or grammar is unlikely to have any significant effect on how the document is interpreted, whereas even single character errors in code can have enormous effects.

Page 1 of 14Next →