Meta announces LlamaCon, its first generative AI dev conference on April 29
meta.com
meta.com
3.3 70B is the best model I've managed to run on my laptop, and 3.2 3B is my favourite model to run on my phone.
I use the MLC Chat app from the App Store: https://apps.apple.com/gb/app/mlc-chat/id6448482937
My favourite example prompt for demos is "Write an outline of a Netflix Christmas movie where a topical-profession falls in love with another topical-profession" - customized for the occasion.
e.g. "Write an outline for a Netflix Christmas movie set in San Gregorio California about a man who runs an unlicensed cemetery falling in love with a barrister at the general store" - result here: https://bsky.app/profile/simonwillison.net/post/3ldthrqb6c22...
Still a bit too hard to copy and paste transcript aback out again though!
In any case, I'm excited!
Those are 14 lawyers gone. That’s more than 3% on “productivity”, but 14 people who lost their jobs. And that’s now with the current state of things.
Most of the policing would happen by courts and bar associations.
For your situation, you really want to measure how many PII records you handle per lawsuit. That way you can accurately measure the lawsuit cost per record and compare it to revenue per record to see if you're profitable.
That labor is not often used sanely.
It is common to use lawyers costing hundreds per hour to do fairly basic document review and summarization. That is, to produce a fairly simple artifact.
Not legal research, not opinionated briefing.
But literal: Read these documents, produce a summary of what they say.
While I can't say this is the same as what you are talking about ("contracts review" means many things to many people), i'm not even the slightest bit surprised that AI is starting to replace remarkably inefficient uses of labor in law.
I will add: Lots of funding being thrown at AI legal startups around on products that do document review and summarization, but that's not the big fish, and will be commodity very quickly.
So i expect there will be an ebb and flow of these sorts of products as the startups either move on to things that enable them to capture a meaningful market (document review ain't it), or die and leave these companies hanging :)
The main one i think is probably wrong is that there was 15 lawyers worth of work being done before (when measured by some average lawyer standard).
For example, it's possible there was only really 1 lawyer worth of work being split 15 ways, so each lawyer was really only responsible for 1/15th of an average lawyers amount of work :)
In that scenario, they'd only be responsible for 1 average lawyers worth of mistakes now.
Is that realistic? Who knows. I've definitely seen that level of "waste" (for lack of a better term) before in law firms :)
Even in the scenario you are positing, it's not obvious it matters as much as you seem to think it does.
If the per-lawyer mistake rate was low enough, it may be that 15x that rate simply does not matter.
These kinds of contracts are fairly standardized, and so they are mostly looking at the differences from last time. Those differences are often not legal as much as factual. IE the table of costs changed, not the legal responsibility.
So the main thing mistakes get you is maybe cost (if mistakes matter at all).
This isn't like they are seeing brand new from scratch contracts constantly that require brand new analysis.
Even if they were, like I said, the main issue with a mistake is cost.
For all we know, the AI company also agreed to indemnify them for a certain rate of mistakes or something (which wouldn't be hard to get insurance for).
I'm not actually a fan of AI taking necessary jobs, but I think the view here that this is sort of life or death is strange.
I'd be much more worried about AI handling criminal defense in some semi-autonomous fashion than this.
Undoubtedly. Happy to be disabused of my misgivings.
> The main one i think is probably wrong is that there was 15 lawyers worth of work being done before (when measured by some average lawyer standard). For example, it's possible there was only really 1 lawyer worth of work being split 15 ways, so each lawyer was really only responsible for 1/15th of an average lawyers amount of work :)
> In that scenario, they'd only be responsible for 1 average lawyers worth of mistakes now. Is that realistic? Who knows. I've definitely seen that level of "waste" (for lack of a better term) before in law firms :) Even in the scenario you are positing, it's not obvious it matters as much as you seem to think it does.
> If the per-lawyer mistake rate was low enough, it may be that 15x that rate simply does not matter.
Well having done quite a bit of work with attorneys no longer practicing law, I’m definitely familiar with the gripes about inefficiencies and running up hours— especially during litigation in the larger firms. Even not being as efficient as they could be, assuming 1400% inefficiency or whatever seems much less reasonable than assuming 0% inefficiency. It’s obviously not either of those extremes, but I have a hard time imagining it’s even close to the former.
> These kinds of contracts are fairly standardized, and so they are mostly looking at the differences from last time. Those differences are often not legal as much as factual. IE the table of costs changed, not the legal responsibility. So the main thing mistakes get you is maybe cost (if mistakes matter at all).
> This isn't like they are seeing brand new from scratch contracts constantly that require brand new analysis. Even if they were, like I said, the main issue with a mistake is cost.
> I don’t actually know what kind of contracts they were working on so I’ll have to take your word on that.
> For all we know, the AI company also agreed to indemnify them for a certain rate of mistakes or something (which wouldn't be hard to get insurance for).
I was involved with the AI legal tool scene indirectly for about a decade, but haven’t been for a couple years, and am only getting info indirectly from people I know that still are. (Actually clicking through the top results on Google, I’m actually on a first-name basis with the first founder there was a picture of. I didn’t know he started a new company though so I guess we’re not THAT close!) My knowledge could be out of date, but I’ve not seen one of these services offer indemnity for mistakes and ostensibly for good reason — the latest data I’ve seen shows that attorney-targeted legal tools make more mistakes than people hoped. I also know nothing about legal insurance, but I don’t think it would be smart to insure an organization that just canned 94% of their counsel in favor of tools known to not be particularly reliable when their workload probably has not changed. Whether they did it because either they care more about payroll than reliability, or they had the poor judgement to maintain 1500% staffing levels until then, it still seems like a pretty poor bet.
> I'm not actually a fan of AI taking necessary jobs, but I think the view here that this is sort of life or death is strange.
I certainly don’t think it’s life or death, and of all the places in our society that could use a little more efficiency, legal services is right up there. That said, the fact that it’s not life or death also doesn’t mean that it’s totally fine either.
> I'd be much more worried about AI handling criminal defense in some semi-autonomous fashion than this.
Haha— frankly, I don’t give a damn if the hospital signs a terrible contract that costs them a bazillion dollars as long as they don’t pull a Steward and stop purchasing basic medical supplies.
SURELY public defenders are an attractive target for the outright person-replacing sort of efficiencies, but I have a hard time imagining that would pass muster. I can definitely see some supposedly adversarial plea agreement system being implemented by more authoritarian jurisdictions as an incremental expansion of the NN sentence-recommendation type of tools. My gut says the bigger semi-automation risk there is overworked public defenders’ being lulled into false confidence in legal and general office LLM type tools (messages summaries, auto scheduling appointments, etc) without having the time to give them the scrutiny they need. I’d be shocked if that wasn’t already happening though. Hey maybe with a bunch of attorneys having newfound time on their hands they can bone up on criminal law and provide some relief for the public defender staffing crisis.
Like anything else, it's a question of performance - if the AI misses it at the same or less rate than the lawyers, ...
If not, it's a question of whether the higher rate is acceptable. For these kinds of contracts, that's mostly about cost.
But zoom out and we see job loss that could be stated as productivity, but what do those lawyers and law degrees do that’s more productive for society? They’re already near the top of an information economy; we’d need to invent an entire next phase. That takes time and pain management; and yet those 14 jobs are gone now.
I used to think all of this was much further away and we had time. But now I’m seeing that we don’t actually need hallucinations fixed, actual AGI, or major quality boosts before displacement begins.
In particular I've found that these tools make it a lot easier to explore or get started with unfamiliar domains. One of my big issues has often been decision paralysis, so having a tool to help me narrow down the list of resources and make it more approachable has been a huge win.
My general experience has been that getting AI tools to directly do stuff for you tends to produce pretty bad results, but you can use it as a force multiplier for your own capabilities. If I'm confused or uncertain about how to do something, AI tools are usually pretty good at clarifying what needs to be done.
They’re mediocre to awful at consistently following instructions, unless you have the fortune of having a task and domain that are well represented in the post-training data.
Yesterday I needed to generate ~500 filenames given a (human written) summary of each document’s contents. This seemed to be the perfect task to throw an LLM into a for loop. Yet it took three hours of prompt engineering to get a passable set of results - yes, I could’ve done it by hand in that time. Each iteration on the prompt revealed new blind spots, new ways in which the models would latch onto irrelevant details or gloss over the main point.
That would definitely account for the difference in my perception of their utility vs other peoples’.
Applied to ChatGPT's 4o:
> "Suggest good guidelines for generating ~500 filenames given a (human written) summary of each document’s contents"
> "Pretend you have this data and create 10 example file names."
> "Financial_Report_Q1_2024.xlsx (Summary: "Quarterly financial performance for Q1 2024, including revenue and expenses.")"
> "Marketing_Strategy_SocialMedia_2024.docx (Summary: "Comprehensive marketing strategy focusing on social media growth for 2024.")"
> "HR_Employee_Handbook_v2.1.pdf (Summary: "Updated version of the company employee handbook with revised policies.")"
And so on. In general, you have to either create the process or have it create a process and refine it if it isn't quite there. You can even use standard processes and guidelines when they're available. Then it usually does a great job.
The reasoning models improve on this kind of thing, but that just means they do even better when provided with or asked to create and follow a process. o3-mini-high, since it shows its work, didn't just spit out a process. It showed the process of creating the process in its reasoning block. It considered the format above in there and removed and added things while considering different levels of detail before the final version before producing anything.
So many things. It's a general-purpose "thing doer" in many situations where you otherwise wouldn't have one. Let me give a super-simple example - not a high-value one, but an example of obvious value IMO.
Say for some reason, you have a screenshot of a bunch of text. Maybe you took a picture of a page from a book or something, idk. Now you want it in textual form. You can throw it in ChatGPT and ask it to give you the text, and a few seconds later you have the text.
I'm not saying there are no other solutions for this - there are. You can look for some software to do OCR or something. But that's what makes ChatGPT or others general-purpose - they're a one-stop shop for a lot of different things, including small one-off tasks like this. I can name a dozen other one-off tasks that it helps me with. Again, not the most high-value things it helps me with (that'd be programming help), but an undeniable example of value, IMO.
How would you do this without an LLM? (I personally would've just typed it up myself, probably.)
This has been a feature of Apple Preview (default image program) for years and years. You can just highlight text and copy it from a jpeg or png.
But how about a next step that I use LLMs for - taking this text from a document and reformatting it as, say, bullet points.
E.g. I literally had this in a Jupyter notebook: [some_var_name, some_other_name, ...] and a bunch of those, and I wanted them redone as bullet points. It was a bit more complicated than that but I'm simplifying for the example. I literally took a screenshot, through it in ChatGPT and got back the correctly-formatted list.
There are other ways to solve something like this of course (normally I'd put it in VSCode and use multiple cursors or macros), but I don't think there's anything that can go from a screenshot and a one-sentence description, to having finished the task, all in a single tool (that can also solve 100 other problems that come up in my work similarly easily).
- Llama 4.0 Phone (Lite / Standard / Max) – For mobile devices.
- Llama 4.0 Workstation (Lite / Standard / Max) – For PCs and laptops.
- Llama 4.0 Server (Lite / Standard / Max) – For high-performance computing.
This approach would enable developers to select the appropriate model based on both device type and performance needs.
What do you think? For example now I feel like 3.3 70B is more for laptops/PCs, and the previous 3.2 3B for phones, is a bit confusing to me.
That was close to the bottom of Meta's stock price.
Could you explain what that means - please?
Double click with us.
Strangely enough, I can work quite well without your help. I've been doing it professionally for 35 odd years. I'm "just" an engineer - no capital E - I simply studied Civil Engineering at college and ended up running an IT company and I'm quite good at IT.
What I would really like to see is really well indexed documentation written by people ie an old school search engine. Google used to do that and so did Altavista, back in the day.
I do not need or want a plethora of trite simulacra web sites dripping with AI wankery at every search term.
That’s an odd flex to roll with for someone who has long been a member of a community of technology enthusiasts, but you be you.
LLMs are tools, not religions. They don’t need to be elevated to a dogmatic level.
Do you have any suggestions for a linux-based, qt-tolerant, LLM-integrated IDE? I’d love to try one.
All of my trades are likewise threatened, but I've found various AI options to be useful augments or interesting toys. Things like having ChatGPT check if a particular game already exists, or analysing cost/effort of lining a shed with gyprock vs ply, or analysing a chicken orchard with steel or timber, or instantly culling lists of extraneous options, or summarising someone's public writings on a particular topic, name ideas in various languages. If you have an experience eye to review suggestions, it's fantastic. Or having CoPilot/similar quickly juicing up a web page - something I could do manually, but would rather save time. Or learning how to build a game in a different language.
These are very different propositions.
You probably already use various code automation tools in your work. At the lower end, Copilot is exactly that, just a bit smarter.
You write "const [getFoo, set" and it autocompletes "Foo] = useState(". Who wouldn't want that?
It's entirely within your control how far you stray into adopting code that you don't thoroughly understand or haven't thoroughly vetted.
If nothing else, use Dall-E to draw stupid pictures to make you and friends laugh. :)
I really tried and I was both impressed and horrified in equal measure. I advise people to use them but treat them like a calculator that snorts cocaine.
A calculator is a useful tool but when they start going off the rails, things can start to get nasty.
I should be more precise: After messing around with generic questions and answers, I went for VMware PowerCLI to test it out. PowerCLI is PowerShell for VMware boxes. PowerShell is very popular so loads of input. PowerCLI is VMware so lots of input too but not quite so much as PS itself.
I tried to get ChatGPT to generate a PS script to patch a VMware cluster. The result was horrific and not even close. Bear in mind that the entirety of the VMware docs for PowerCLI are public and I wrote a script myself - its not perfect but good enough.
Oh and I am dropping VMware for good in favour of Proxmox. I have been a VMware consultant for 20+ years. Oh well.
I do find that for big tasks I generally have to scaffold it a bit, lead it in the right direction.
But it's also impressive the way it can do tasks that I can't: like writing complex TypeScript. I could spend hours on certain TypeScript challenges and not actually find a workable solution. With ChatGPT I can usually either get a solution, or convince myself it's not going to happen, within 5-10 minutes.
Indices are, by definition, lossy representations of their underlying data. If you use stemming and lemmatization to preprocess both documentation and query text, you're already departing from a truly hand-optimized indexing system, and choosing to have imperfect algorithms do things in a more scalable way. And indexing by embedding vectors that use LLMs to determine context are a natural extension of this, in my view. And on top of that, when you have a massive amount of candidate text to display to the user... is displaying sentence fragments one on top of the other the most optimal UX there? At a certain point, RAG becomes the answer to this question.
The problem, as you note, is that search engines and social media systems are incentivized to allow garbage content into the original set of things they index and surface, if that garbage content drives more attention to advertisements. But that's not a reason to reject the benefits that the underlying LLM technology can bring towards building good indexing on top of human-written documents. It just won't be done by the companies that used to do it.
> banger
“Fellow kids” vibes from the dinosaurs at Facebook and Zuckerfuck.
up next: farting on the earnings call "here's what I think of your question Chadwick at Vanguard..."
/s
I wondered why so much support on HN all of a sudden.