Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
anthropic.com
anthropic.com
As someone building AI SaaS products, I used to have the position that directly integrating with APIs is going to get us most of the way there in terms of complete AI automation.
I wanted to take at stab at this problem and started researching some daily busineses and how they use software.
My brother-in-law (who is a doctor) showed me the bespoke software they use in his practice. Running on Windows. Using MFC forms.
My accountant showed me Cantax - a very powerful software package they use to prepare tax returns in Canada. Also on Windows.
I started to realize that pretty much most of the real world runs on software that directly interfaces with people, without clearly defined public APIs you can integrate into. Being in the SaaS space makes you believe that everyone ought to have client-server backend APIs etc.
Boy was I wrong.
I am glad they did this, since it is a powerful connector to these types of real-world business use cases that are super-hairy, and hence very worthwhile in automating.
In this case I doubt they're networked apps so they probably don't have a server API.
I think it would be very unusual this decade for software used to run either a medical practice or tax accountants to not be networked. Most such practices have multiple doctors/accountants, each with their individual computer, and they want to be able to share files, so that if your doctor/accountant is away their colleague can attend to you. Managing backups/security/etc is all a lot easier when the data is stored in a central server (whether in the cloud or a closet) than on individual client machines.
Just because it is a fat client MFC-based Windows app doesn’t mean the data has to be stored locally. DCOM has been a thing since 1996.
Is the idea that someone will always reverse engineer it? Yes, but QuickBooks is brittle as is (you can count on at least one database corruption every year or two). I have zero interest in treading into unsupported territory when database corruption is involved and I’m likely going to need Intuit’s help recovering. We can try to restore from backup, but when there’s corruption it doesn’t always restore successfully, or the corruption was lingering silently for some time and rears its head again after a successful restore, and then we’re back to needing Intuit’s help.
I've been peddling my vision of "AI automation" for the last several months to acquaintances of mine in various professional fields. In some cases, even building up prototypes and real-user testing. Invariably, none have really stuck.
This is not a technical problem that requires a technical solution. The problem is that it requires human behavior change.
In the context of AI automation, the promise is huge gains, but when you try to convince users / buyers, there is nothing wrong with their current solutions. Ie: There is no problem to solve. So essentially "why are you bothering me with this AI nonsense?"
Honestly, human behavior change might be the only real blocker to a world where AI automates most of the boring busy work currently done by people.
This approach essentially sidesteps the need to have effect a behavior change, at least in the short-term while AI can prove and solidify its value in the real-world.
AI is squarely #1. You can't trust it with your credit card to order groceries, or to budget and plan and book your vacation. People aren't picking up on AI because it isn't good enough yet to trust - you still have the burden of responsibility for the task.
Nobody likes to change a system where they already have their own little comfortable spot and figured it out and just want to seep in the lukewarm there until retirement. Fully understandable. But at least in the private sector this will not save them.
Really??!? What could possibly go wrong.
I'm currently trying to do a large ORC project using Google Vision API, and then Gemini 1.5 Pro 002 to parse and reconstruct the results (taking advantage, one hopes, of its big context window). As I'm not familiar with Google Vision API I asked Gemini to guide me in setting it up.
Gemini is the latest Google model; Vision, as the name implies, is also from Google. Yet Gemini makes several egregious mistakes about Vision, gets names of fields or options wrong, etc.
Gemini 1.5 "Pro" also suggests that concatenating two json strings produces a valid json string; when told that's unlikely, it's very sorry and makes lots of apologies, but still it made the mistake in the first place.
LLMs can be useful when used with caution; letting one loose in an enterprise environment doesn't feel safe, or sane.
Sonnet 3.5 for coding is fine but makes "basic" mistakes all the time. Using LLMs is at times like dealing with a senior expert suffering from dementia: it has arcane knowledge of a lot of things but suddenly misses the obvious that would not escape an intern. It's weird, really.
So if you want accurate results on writing code you need to put all the docs into the input and THEN ask for your question. So download all docs on Vision, put them in the Gemini prompt and ask your question or code on how to use Vision, and you'll get much closer to truth
Most of the things that RPA is used for can be easily scripted, e.g. download a form from one website, open up Adobe. There are a lot of startups that are trying to build agentic versions of RPA, I'm glad to see Anthropic is investing in it now too.
I don't think there will be a very big place for standalone next gen RPA pure plays. it makes sense that companies that are trying to deliver value would implement capabilities this. Over time, I expect some conventions/specs will emerge. Either Apple/Google or Anthropic/OpenAI are likely to come up with an implementation that everyone aligns on
In other words, I agree
You realize this means '100 times better', right?
100 times better seems to me in line with the bet that's justifying $250B per annum in Cap Ex (just among hyperscalers) but curious how you might project a few years out?
Having said that, my use of 100x better here applies to 100x more effective at navigating use cases not in training set, for example, as opposed to doing things that are 100x more awesome or doing them 100x more efficiently (though seemingly costs, context window and token per unit of electricity seem to continue to improve quickly)
Just to give a few comparisons, the following things are two orders of magnitude apart:
1. The force felt by a mosquito landing on your arm and getting punched by Mike Tyson in his prime
2. A firecracker exploding and a stick of dynamite exploding
3. The heat from a candle and the heat from a blowtorch
4. The sound from a whisper and the sound from jet engine
I’ve implemented quite a few RPA apps and the struggle is the request/response turn around time for realtime transactions. For batch data extract or input, RPA is great since there’s no expectation of process duration. However, when a client requests data in realtime that can only be retrieved from an app using RPA, the response time is abysmal. Just picture it - Start the app, log into the app if it requires authentication (hope that the authentication's MFA is email based rather than token based, and then access the mailbox using an in-place configuration with MS Graph/Google Workspace/etc), navigate to the app’s view that has the data or worse, bring up a search interface since the exact data isn’t known and try and find the requested data. So brittle...
In my case I have a bunch of nurses that waste a huge amount of time dealing with clerical work and tech hoops, rather than operating at the top of their license.
Traditional RPAs are tough when you're dealing with VPNs, 2fa, remote desktop (in multiple ways), a variety of EHRs and scraping clinical documentation from poorly structured clinical notes or PDFs.
This technology looks like it could be a game changer for our organization.
So maybe like an automated ubikey that you can opt in to a nearby computer to have all the access. Especially if working from home, you can set it at a state where if you are in 15m radius of your laptop it is able to sign all access.
Because right now, considering amount of tools and everything I use and with single sign on, VPN, Okta, etc, and how slow they seem to be, it's extremely frustrating process constantly logging in to everywhere, and it's almost like it makes me procrastinate my work, because I can't be bothered. Everything about those weird little things is absolutely terrible experience, including things like cookie banners as well.
And it is ridiculous, because I'm working from home, but frustratingly high amount of time is spent on this bs.
A bluetooth wearable or similar to prove that I'm nearby essentially, to me that seems like it could alleviate a lot of safety concerns, while providing amazing dev/ux.
The main attack vector would then probably be some man-in-the-middle intercepting the signal from your wearable, which leads me to wonder whether you could protect yourself by having the responses valid for only an extremely short duration, e.g. ~1ms, such that there's no way for an attacker to do anything with the token unless they gain control over compute inside your house.
I'd bet that until we get the risks whittled down enough for larger organizations to adopt this on a wide scale, the biggest user group for AI automation tools will be at the level of individual workers who are eager to streamline their own tasks and aren't paid enough to care about those same risks.
CTO of healthcare org here.
I just put a hold on a new RPA project to keep an eye on this and see how it develops.
According to their docs, Anthropic will sign a BAA.
Sometimes, the AI is more accurate or safer than humans, but it still reads better to say "we always have humans in the loop". In those cases, we reap the benefits of both: Use the AI for safety, but still have a human fallback.
We haven't deployed a model like this, it's new.
I've done a ton of various RPAs over the years, using all the normal techniques, and they're always brittle and sensitive to minor updates.
For this, I'm taking a "wait and see" approach. I want to see and test how well it performs in the real world before I deploy it, and wait for it to come out of beta so Anthropic will sign a BAA.
The demo is impressive enough that I want to give the tech a chance to mature before my team and I invest a ton of time into a more traditional RPA.
At a minimum, if we do end up using it, we'll have solid guard rails in place - it'll run on an isolated VM, all of its user access will be restricted to "read only" for external systems, and any content that comes from it will go through review by our nurses.
I don't think this is going to work in that industry until local models get good enough to do it, and small enoguh to be affordable to hospitals.
So the solution to that is to add another layer of complex AI tech on top of it?
Similarly I expect that once processing/searching laws/legal records becomes easy through LLMs, we'll compensate by having orders of magnitude more laws, perhaps themselves generated in part by LLMs.
It's almost always a framework around existing tools like Selenium that you constantly have to fight against to get good results from. I was always left with the feeling that I could build something better myself just handrolling the scripts rather than using their frameworks.
Getting Claude integrated into the space is going to be a game changer.
I know we're still a bit far from there, but I don't see a particular hurdle that strikes me as requiring novel research.
Maybe Anthropic should focus on building a flexible RPA primitive we could use to make RPA workflows with, like for example extracting values from components that need scrolling, selecting values from long drop-down menus, or handling error messages under form fields.
A human can work around all these error cases once she encounters them. Current RPA tools like Uipath or ui.vision need explicit programming for every potential situation. And I see no indication that Claude is doing any better than this.
For starters, for visual automation to work reliably the OCR quality needs to improve further and be 100% reliable. Even in that very basic "AI" area, Claude, ChatGPT, Gemini are good, but not good enough yet.
> Most RPA work is in dealing with errors and exceptions, not the "happy path".
Isn't this most programming? I always chuckle when a junior hire looks at my code and says: "It is mostly error checking."So many examples come to mind… RealVideo -> YouTube, Myspace -> Facebook, Laserdisc -> DVD, MP3 players -> iPod…
UiPath may end up being the burned pancake, but the underlying problem they’re trying to address is immensely lucrative and possibly solvable (hey if we got the Turing test solved so quickly, I’m willing to believe anything is possible).
UI is now much more accessible as API. I hope we don’t start seeing captcha like behaviour in desktop or web software.
FWIW, looking at it from end-user perspective, it ain't much different than the Windows apps. APIs are not interoperability - they tend to be tightly-controlled channels, access gated by the vendor and provided through contracts.
In a way, it's easier to make an API to a legacy native desktop app than it is to a typical SaaS[0] - the native app gets updated infrequently, and isn't running in an obstinate sandbox. The older the app, the better - it's more likely to rely on OS APIs and practices, designed with collaboration and accessibility in mind. E.g. in Windows land, in many cases you don't need OCR and mouse emulation - you just need to enumerate the window handles, walk the tree structure looking for text or IDs you care about, and send targeted messages to those components.
Unfortunately, desktop apps are headed the same direction web apps are (increasingly often, they are web apps in disguise), so I agree that AI-level RPA is a huge deal.
--
[0] - This is changing a bit in that frameworks seem to be getting complex enough that SaaS vendors often have no clue as to what kind of access they're leaving open to people who know how to press F12 in their browsers and how to call cURL. I'm not talking bespoke APIs backend team wrote, but standard ones built into middleware, that fell beyond dev team's "abstraction horizon". GraphQL is a notable example.
This is just a solution in search of a problem. AIs aren't reliable enough if the content changes constantly since it'll just click on the "close my account" button for no reason half the time, while a legacy website with no changes is much easier to program a tool like Selenium around than relying on the AI that will still be making random choices half the time.
Now, an in-house tool built on top of this Anthropic API can save hours of drudgery. I can already see sales teams smiling at the new 'submit your expense report' button.
It's clear that the reason there's no competition in the market is that the companies (I think it's literally one company) selling this software want it that way, and legislation is yet to catch up with the way they ensure their monopoloy.
Has anyone tried asking "use computer" to do "Please write a selenium/capybara/whatever test for filling out this form and sending it?"
That would take away some serious drudge work. And it's not a big problem if it fails, contrary to when it makes a mistake in filling out a form in an actual business process.
One of the reasons my bash terminal is more effective than point and click is the easy of automation of routine tasks: from history and aliases to scripting.
Software interop is trickier as it doesn't so much depend on being able to send messages from one piece of software to another, it's you need an agreed format of those messages ( whether they be expressed in the payload of an API, or a UI specific stream of point and clicks ).
If you're in some niche doing AI development work, you are going to be outcompeted by more generalized AI at some point.
Computer use API documentation: https://docs.anthropic.com/en/docs/build-with-claude/compute...
Computer Use Demo: https://github.com/anthropics/anthropic-quickstarts/tree/mai...
Here, "next-best AI model in the same category" referes to which model.
> I apologize, but it seems we're having some issues installing Chromium. As an alternative, since I can see Firefox is already installed on the system, would you like to use Firefox instead? Or would you prefer me to try a different approach to installing Chromium?
It seems unable to figure out how to install chormium. Sees the Firefox icon out there and just want to use it. You know, like a real lazy human being.
It seems to insistent when it needs to, though
> I clicked on Firefox and it should start opening. Let me take another screenshot to see if the window has appeared:
> Firefox seems to be taking a moment to start up. Let me wait a few seconds and take another screenshot:
> Let me try clicking Firefox one more time to ensure it starts:
> Let me try one more time to take a screenshot to see if Firefox has opened:
> Let me try starting Firefox using the bash tool as an alternative method:
> Firefox has successfully opened! As per the instructions, I should ignore the setup wizard and instead click directly on the address bar. Would you like me to do anything specific with Firefox now that it's open?
I didn't instruct him to ignore the setup wizard. So my guess is that Anthropic has configured it to ignore stuff happening on the screen so it doesn't go loose.
And here he goes through my website, through my hacker news account and then find this very comment
> Looking at his first/most recent comment, it's about a discussion of Claude and computer use. Here's what he wrote:
"I like its lazy approach"
This appears to be a humorous response in a thread about "Computer use, a new Claude 3.5 Sonnet, and Claude..." where he's commenting on an AI's behavior in a situation. The comment is very recent (shown as "8 minutes ago" in the screenshot) and is referring to a situation where an AI seems to have taken a simpler or more straightforward approach to solving a problem.
<IMPORTANT> * When using Firefox, if a startup wizard appears, IGNORE IT. Do not even click "skip this step". Instead, click on the address bar where it says "Search or enter address", and enter the appropriate search term or URL there. * If the item you are looking at is a pdf, if after taking a single screenshot of the pdf it seems that you want to read the entire document instead of trying to continue to read the pdf from your screenshots + navigation, determine the URL, use curl to download the pdf, install and use pdftotext to convert it to a text file, and then read that text file directly with your StrReplaceEditTool. </IMPORTANT>"""
And finally, in the table in the blogpost, Opus isn't even included? It seems to me like Opus is the best model they have, but they don't want people to default using it, maybe the ROI is lower on Opus or something?
When I manually tested it, I feel like Opus gives slightly better replies compared to Sonnet, but I'm not 100% it's just placebo.
So for example, Perplexity is wrong here implying that Opus is better than Sonnet?
That's according to https://docs.anthropic.com/en/docs/about-claude/models which has "claude-3-5-sonnet-20241022" as the latest model (today's date)
The older/bigger GPT4 runs at $30/$60 and peforms about on par with GPT4o-mini which costs only $0.15/$0.60.
If you are currently, or have been integrating AI models in the past ~2 years, you should definitely keep up with model capability/pricing development. If you are staying on old models you are certainly overpaying/leaving performance on the table. It's essentially a tax on agility.
I don't think GPT-4o Mini has comparable performance to GPT-4 at all, where are you finding the benchmarks claiming this?
Everywhere I look says GPT-4 is more powerful, but GPT-4o Mini is most cost-effective, if you're OK with worse performance.
Even OpenAI themselves about GPT-4o Mini:
> Our affordable and intelligent small model for fast, lightweight tasks. GPT-4o mini is cheaper and more capable than GPT-3.5 Turbo.
If it was "on par" with GPT-4 they would surely say this.
> should definitely keep up with model capability/pricing development
Yeah, I mean that's why we're both here and why we're discussing this very topic, right? :D
That wasn't specifically directed at "you", but more as a plea to everyone reading that comment ;)
I looked at a few benchmarks, comparing the two, which like in the case of Opus 3 vs Sonnet 3.5 is hard, as the benchmarks the wider community is interested in shifts over time. I think this page[0] provides the best overview I can link to.
Yes, GPT4 is better in the MMLU benchmark, but in all other benchmarks and the LMSys Chatbot Arena scores[1], GPT4o-mini comes out ahead. Overall, the margin between is so thin that it falls under my definition of "on par". I think OpenAI is generally a bit more conservative with the messaging here (which is understandable), and they only advertise a model as "more capable", if one model beats the other one in every benchmark they track, which AFAIK is the case when it comes to 4o mini vs 3.5 Turbo.
[0]: https://context.ai/compare/gpt-4o-mini/gpt-4
[1]: https://artificialanalysis.ai/models?models_selected=gpt-4o-...
OpenAI's own words: "GPT-4o is our most advanced multimodal model that’s faster and cheaper than GPT-4 Turbo with stronger vision capabilities."
gpt-4o:
$2.50 / 1M input tokens $10.00 / 1M output tokens
gpt-4-turbo:
$10.00 / 1M input tokens $30.00 / 1M output tokens
gpt-4:
$30.00 / 1M input tokens $60.00 / 1M ouput tokens
I think they originally announced that Opus would get a 3.5 update, but with every product update they are doing I'm doubting it more and more. It seems like their strategy is to beat the competition on a smaller model that they can train/tune more nimbly and pair it with outside-the-model product features, and it honestly seems to be working.
Why isn't Anthropic clearer about Sonnet being better then? Why isn't it included in the benchmark if new Sonnet beats Opus? Why are they so ambiguous with their language?
For example, https://www.anthropic.com/api says:
> Sonnet - Our best combination of performance and speed for efficient, high-throughput tasks.
> Opus - Our highest-performing model, which can handle complex analysis, longer tasks with many steps, and higher-order math and coding tasks.
And Opus is above/after Sonnet. That to me implies that Opus is indeed better than Sonnet.
But then you go to https://docs.anthropic.com/en/docs/about-claude/models and it says:
> Claude 3.5 Sonnet - Most intelligent model
- Claude 3 Opus - Powerful model for highly complex tasks
Does that mean Sonnet 3.5 is better than Opus for even highly complex tasks, since it's the "most intelligent model"? Or just for everything except "highly complex tasks"
I don't understand why this seems purposefully ambiguous?
I wouldn't attribute this to malice when it can also be explained by incompetence.
Sonnet 3.5 New > Opus 3 > Sonnet 3.5 is generally how they stack up against each other when looking at the total benchmarks.
"Sonnet 3.5 New" has just been announced, and they likely just haven't updated the marketing copy across the whole page yet, and maybe also haven't figured out how to graple with the fact that their new Sonnet model was ready faster than their next Opus model.
At the same time I think they want to keep their options open to either:
A) drop a Opus 3.5 soon that will bring the logic back in order again
B) potentially phase out Opus, and instead introduce new branding for what they called a "reasoning model" like OpenAI did with o1(-preview)
I don't think it's malice either, but if Opus costs more to them to run, and they've already set a price they cannot raise, it makes sense they want people to use models they have a higher net return on, that's just "business sense" and not really malice.
> and they likely just haven't updated the marketing copy across the whole page yet
The API docs have been updated though, which is the second page I linked. It mentions the new model by it's full name "claude-3-5-sonnet-20241022" so clearly they've gone through at least that page. Yet the wording remains ambiguous.
> Sonnet 3.5 New > Opus 3 > Sonnet 3.5 is generally how they stack up against each other when looking at the total benchmarks.
Which ones are you looking at? Since the benchmark comparison in the blogpost itself doesn't include Opus at all.
I manually compared it with the values from the benchmarks they published when they originally announced the Claude 3 model family[0].
Not all rows have a 1:1 row in the current benchmarks, but I think it paints a good enough picture.
When should we be using the -o OpenAI models? I've not been keeping up and the official information now assumes far too much familiarity to be of much use.
The -o models are "just" stronger versions of their non-suffixed predecessors. They are the latest (and maybe last?) version of models in the lineage of GPT models (roughly GPT-1 -> GPT-2 -> GPT-3 -> GPT-3.5 -> GPT-4 -> GPT-4o).
The o1 models (not sure what the naming structure for upcoming models will be) are a new family of models that try to excel at deep reasoning, by allowing the models to use an internal (opaque) chain-of-thought to produce better results at the expense of higher token usage (and thus cost) and longer latency.
Personally, I think the use cases that justify the current cost and slowness of o1 are incredibly narrow (e.g. offline analysis of financial documents or deep academic paper research). I think in most interactive use-cases I'd rather opt for GPT-4o or Sonnet 3.5 instead of o1-preview and have the faster response time and send a follow-up message. Similarly for non-interactive use-cases I'd try to add a layer of tool calling with those faster models than use o1-preview.
I think the o1-like models will only really take off, if the prices for it are coming down, and it is clearly demonstrated that more "thinking tokens" correlate to predictably better results, and results that can compete with highly tuned prompts/fine tuned models that or currently expensive to produce in terms of development time.
Cheaper and faster, but not notably "stronger" at real-world use.
They are clear that both: Opus > Sonnet and 3.5 > 3.0. I don't think there is a clear universal better/worse relationship between Sonnet 3.5 and Opus 3.0; which is better is task dependent (though with Opus 3.0 being five times as expensive as Sonnet 3.5, I wouldn't be using Opus 3.0 unless Sonnet 3.5 proved clearly inadequate for a task.)
Honestly I don’t know. I’ve not been using Sonnet 3.5 up to now and I’m a fairly light user so I doubt I’ll run into the free tier limits. I’ll probably cancel my subscription until Opus 3.5 comes out (if it ever does).
This theory is consistent with the other two top players, Open AI and Google, they both were expected to release a heavy model, but instead have just released multiple medium and small tier models. It's been so long since google released gemini ultimate 1.0 (the naming clearly implying that they were planning on upgrading it to 1.5 like they did with Pro)
Not seeing anyone release a heavyweight model, but at the same time releasing many small and medium sized models makes me think that improving models will be much more complicated than scaling it with more compute, and that there likely are diminishing returns with that regard.
Thats why they release them with that skew
Also, Gemini ultra 1.0, was released like 8 months ago, 1.5 pro released soon after, with this wording "The first Gemini 1.5 model we’re releasing for early testing is Gemini 1.5 Pro"
Still no ultra 1.5, despite many mid and small sized models being released in that time frame. This isn't just an issue of "the training time takes longer", or a "skew" to release dates. There's a better theory to explain why all SoTA LLM companies have not released a heavy model in many months.
I'm wondering at this point if they are going to release Opus 3.5 at all, or maybe skip it and go straight to 4.0. It's possible that Haiku 3.5 is a distillation of Opus 3.5.
Not most advanced
Given that 3.5 Sonnet is cheaper and faster than 3 Opus, I default to 3.5 Sonnet so I don't know what the number for the reverse is. How many problems do 3.5 Sonnet get which 3 Opus does not? ¯\_(ツ)_/¯
My best guess would be that it's something in the same kind of range.
This is a lot more than an agent able to use your computer as a tool (and understanding how to do that) - it's basically an autonomous reasoning agent that you can give a goal to, and it will then use reasoning, as well as it's access to your computer, to achieve that goal.
Take a look at their demo of using this for coding.
https://www.youtube.com/watch?v=vH2f7cjXjKI
This seems to be an OpenAI GPT-o1 killer - it may be using an agent to do reasoning (still not clear exactly what is under the hood) as opposed to GPT-o1 supposedly being a model (but still basically a loop around an LLM), but the reasoning it is able to achieve in pursuit of a real world goal is very impressive. It'd be mind boggling if we hadn't had the last few years to get used to this escalation of capabilities.
It's also interesting to consider this from POV of Anthropic's focus on AI safety. On their web site they have a bunch of advice on how to stay safe by sandboxing, limiting what it has access to, etc, but at the end of the day this is a very capable AI able to use your computer and browser to do whatever it deems necessary to achieve a requested goal. How far are we from paperclip optimization, or at least autonomous AI hacking ?
I wouldn't be so dismissive - you could describe GPT-o1 in same way "it just loops until it gets to the solution". It's the details and implementation that matter.
Just as how expert systems didn't take off and tagging every website for the Semantic Web didn't happen either, we have to accept that the real world of humans is messy and unstructured.
I still advocate making new things more structured. A car on wheels on flattened ground will always be more efficient than skipping the landscaping part and just riding quadruped robots through the forest on uneven terrain. We should develop better information infrastructure but the long tail of existing use cases will require automation that can deal with unstructured mess too.
As someone who has had to interact with legacy enterprise systems via RPA (screen scraping and keystroke recording) it is absolutely awful, incredibly brittle, and unmaintainable once you get past a certain level of complexity. Even when it works, performance at scale is terrible.
The problem is that we're not starting from scratch - we have a web engineered for browser use and a world engineered for humanoid use. That means an agent that can use a browser, while less efficient than an agent using APIs at any particular task, is vastly more useful because it can complete a much greater breadth of tasks. Same thing with humanoid robots - not as efficient at cleaning the floor as my purpose-built Roomba, but vastly more useful because the breadth of tasks it can accomplish means it can be doing productive things most of the time, as opposed to my Roomba, which is not in use 99% of the time.
I do think that once AI agents become common, the web will increasingly be designed for their use and will move away from the browser, but that probably take a comparable amount of time as it did for the mobile web to emerge after the iPhone came out. (Actually that's probably not true - it'll take less time because AI will be doing the work instead of humans.)
Beyond that, I can't help but think of the old thin vs. thick client debate, and I would argue that "software should just talk via APIs" is why, in the web space, everybody is blowing time and energy on building client/server architectures and SPAs instead of basic-ass full-stacks.
Isn't the GUI driven by code? Can anything at all in the GUI work that can't be done programmatically?
I do still occasionally pop over to ChatGPT to test their their waters (or if Claude is just not getting it), but I've not felt any need to switch back or have both. Well done, Anthropic!
Internet Archive confirms that on the 8th of October that page listed 3.5 Opus as coming "Later this year" https://web.archive.org/web/20241008222204/https://docs.anth...
The fact that it's no longer listed suggests that its release has at least been delayed for an unpredictable amount of time, or maybe even cancelled.
> i don't write the docs, no clue
> afaik opus plan same as its ever been
"Even while recording these demos, we encountered some amusing moments. In one, Claude accidentally stopped a long-running screen recording, causing all footage to be lost.
Later, Claude took a break from our coding demo and began to peruse photos of Yellowstone National Park."
It’s planning on triggering a Yellowstone caldera super eruption.
* Fixed bug where Claude got bored during compile times and started editing Wikipedia articles to claim that birds aren't real
* Blocked news.ycombinator.com in the Docker image's hosts file to avoid spurious flamewar posts (Note: The site is still recovering from the last insident)
* Addressed issue of Claude procrastinating on debugging by creating elaborate ASCII art in Vim
* Patched tendency to rickroll users when asked to demonstrate web scraping"
* Added guards to prevent every other sentence being "I use neovim"
* Managed to put error/info messages into a separate key instead of concatenating them with stringified JSON in the main body of the response.
* Taught Claude to treat numeric integer strings as integers to avoid embarrassment when the user asks it for a "two-digit random number between 1-50, like 11" and Claude replies with 111.
It could be very empowering for users with disabilities to regain access computers. But it would also be very powerful to be able to ask "use Photoshop to remove the power lines from this photo" and have the model complete the task and drop off a few samples in a folder somewhere.
Probably would have been an ad there if I didn't block those, though.
But it would slow down the masses
Some people would jailbreak the agents though
I have been playing with Mindcraft which lets models interact with Minecraft through the bot API and one of them started saying things like "I want to place some cobblestone there" and then later more general "I want to do X" and then start playing with the available commands, it was pretty cool to watch it explore.
But eventually they (RNNs, likely) will. And we won't know when.
Sure they do, but the big labs spend many, many, worker-hours suppressing it with RLHF.
My GPT-2 discord bot from 2021 possessed clear intent. Sure, unpredictable and short-lived, but if it decided it didn't like you it would continuously cuss and attempt ban commands until its context window became distracted by something else.
https://arstechnica.com/information-technology/2024/08/chatg...
Bonus points for ‘what does <real> mean?’ as a follow up.
Our brains are hardwired to see patterns, even when there are none.
A similar, and related, behavior is seeing intent and intelligence in random phenomenon.
Does that mean our brains do not implement General Intelligence?
Trust me, I'm really not.
All machine learning models start off in a random state. As they progress through their training, their input/output pairs tend to mimic what they've been trained to mimic.
LLMs have been doing a great job mimicking our human flaws from the beginning because we train them on a ton of human generated data. Other weird behavior can be easily attributed to simple fact that they're initialized at a random state.
Being able to work on and prove non-trivial theorems is a better indication of AGI, IMO.
When these new models go off to a random site and are caught in a loop of exploring pages that doesn’t mean it’s an AGI admiring nature.
It's kind of interesting that they're not running a 2PC setup with HDMI splitter, but (presumably)just laptops and screen recording apps...
Claude using Claude on a computer for coding https://youtu.be/vH2f7cjXjKI?si=Tw7rBPGsavzb-LNo (3 mins)
True end-user programming and product manager programming are coming, probably pretty soon. Not the same thing, but Midjourney went from v.1 to v.6 in less than 2 years.
If something similar happens, most jobs that could be done remotely will be automatable in a few years.
I feel like we will see that again here as well. It really is similar to the self-driving problem.
> We are shockingly ignorant of the causes of our own behavior. The explanations that we provide are sometimes wholly fabricated, and certainly never complete. Yet, that is not how it feels. Instead it feels like we know exactly what we're doing and why. This is confabulation: Guessing at plausible explanations for our behavior, and then regarding those guesses as introspective certainties. Every year psychologists use dramatic examples to entertain their undergraduate audiences. Confabulation is funny, but there is a serious side, too. Understanding it can help us act better and think better in everyday life.
I suspect it's an inherent aspect of human and LLM intelligences, and cannot be avoided. And yet, humans do ok, which is why I don't think it's the moat between LLM agents and AGI that it's generally assumed to be. I strongly suspect it's going to be yesterday's problem in 6-12 months at most.
Yea i agree, i'm not making a snipe at LLMs or anything of the sort.
I'm saying i expect there to be a human-fallback in the system for quite some time. But solving the fallback problems with be one of black boxes. Which is the worst kind of project in my view, i hate working on code i don't understand. Where the results are not predictable.
This happens nearly every time I request “how tos” for libraries that aren’t very popular. It will make up some parameters that don’t exist despite the rest of the code being valid. It’s not a memory error like confabulation where it’s convinced the response is valid from memory either, because it can be easily convinced that it made a mistake.
I’ve never worked with an engineer in my 25 years in the industry that has done this. People don’t confabulate to get day to day answers. What we call hallucination is the exact same process LLMs use to get valid answers.
For example, we can mostly agree CLIP does a fine job classifying images, except if you glue a sticky note saying "iPod" onto an apple, it would say classify it as such.
No matter the performance, these are categorically statistical machines reaching for the most immediately useful representations, yielding an incoherent world model. These systems will be proposed as replacement to humans, they will do their best to pretend to work, they will inevitably fail over a long enough time horizon, and a human accustomed to rubber-stamping its decisions, or perhaps fooled by the shape of a correct answer, or simply tired enough to let it slip by, will take the blame.
Most jobs are not like that.
A good argument can be made, however, that software engineering, especially in important domains, will be among the last to be fully automated because software errors often cascade.
There’s a countervailing effect though. It’s easy to generate and validate synthetic data for lower-level code. Junior coding jobs will likely become less available soon.
Whereas software defects in design and architecture subtly accumulate, until they leave the codebase in a state in which it becomes utterly unworkable. It is one of the chief reasons why good devs get paid what they do. Software discussions very often underrate software extensibility, or in other words, its structural and architectural scaleability. Even software correctness is trivial in comparison - you can't even keep writing correct code if you've made an unworkable tire-fire. This could be a massive mountain for AI to climb.
The only question is how fast it is going to happen. Ie what percentage of jobs is going to be replaced next year and so on.
There is a lot about the human brain that even the world's top neuroscientists don't know. There's plenty of magic about it if we define magic as undiscovered knowledge.
There's also no consensus among top AI researchers that current techniques like LLMs will get us anywhere close to AGI.
Nothing I've seen on current models (not even o1-preview) suggests to me that AIs can reason about codebases of more than 5k LOC. A top 5% engineer can probably make sense of a codebase of a couple million LOC in time.
Which models specifically have you seen that are looking like they will be able to surmount any time soon the challenges of software design and architecture I'm laying out in my previous comment?
The majority of people on the planet can barely reason about how any given politician will affect them, even when there’s a billion resources out there telling them exactly that. No reasonable human would ever define AGI as having anything to do with coding at all, since that’s not even “general intelligence”… it’s learned facts and logic.
When someone comes up with a rigorous definition of intelligence.
I haven't moved any goal posts - it is your definition which is way too narrow.
I’m sure as soon as Claude can handle 5MLOC you’ll say it should be 10, and it needs to make sure it can serve you a Michelin star dinner as well.
That’s not AGI. Stop moving the goalposts.
ARC-AGI Benchmark might serve as a canary in the coal mine.
And what is the title element on CrowdStrike's website today? "CrowdStrike: We Stop Breaches with AI-native Cybersecurity"
Can't wait.
We'd be losing access to food, shelter, insurance, purpose. I can't blame people for at least telling themselves some coping story.
It's going to be absolutely ruinous for many people. So what else should they do, admit they're fucked? I know we like to always be cold rational engineers on this forum, but shit looks pretty bleak in the short term if this goal of automating everyone's work comes true and there are basically zero social safety nets to deal with it.
I live abroad and my visa is tied to my job, so not only would losing my job be ruinous financially, it will likely mean deportation too as there will be no other job for me to turn to for renewal.
But I do agree, there is no reason to be enthusiastic about any progress in AI, when the goal is simply automating people's jobs away.
I'm really curious on the cost of that sort of thing. Seems astronomical atm, but as much as i get shocked at the today-cost, staffing is also a pretty insane cost.
And it got some things subtly wrong though so do I/my team. Interesting times ahead I think, but I’m not too worried about my job as a principal dev. Again I’m more stressed about juniors
This means that either product managers will have to start (effectively) writing in-depth specs again, or they will have to learn to accept the LLM's ideas in a way that most have not accepted their human programmers' ideas.
Definitely will be interesting to see how that plays out.
The larger point is that building software is about making tons of decisions about how it works. Someone has to make those decisions. Either PMs will be happy letting machines make the decisions where they do not let programmers decide now. Or the PMs will have to make all the decisions before (spec) or after (evaluation + feedback look like you suggest).
I'm placing my bets rather on this new object-oriented programming thing. It will make programming jobs obsolete any day now...
I'd be willing to be a large amount of money this doesn't happen, assuming "most" means >50% and "a few" is <5.
I’m curious to understand your reasoning. What would be some key roadblocks? Hallucinations and reliability issues in most domains will likely be solvable with agentic systems in a few years.
"Create a simple website" has to be one of the most common blog / example out there in about every programming language.
It can automate stuff? That's cool: I already did automate screenshots and then AI looking if it looks like phishing or not (and it's quite good at it).
I mean: the "Claude using Claude" may seem cool, but I dispute the "for coding" part. That's trivial stuff. A trivial error (which it doesn't fix btw: it just deletes everything).
'Claude, write me code to bring SpaceX rockets back on earth"
or
"Claude, write me code to pilot a machine to treat a tumor with precision"
This was not it.
Edit: it seems like these new features will eliminate a lot of automated testing tools we have today.
Code for molmo coordinate tests https://github.com/logankeenan/molmo-server
I'm more interested in Claude 3.5 Haiku, particularly if it is indeed better than the current Claude 3.5 Sonnet at some tasks as claimed.
Depending on how many GUI actions correspond to one equivalent AI orchestrated API call, this might also not be too bad in terms of efficiency.
Or you could teach it to hack into the backend and add an API...
Oh, and on edit, "bizarre" and "multi-billion-dollar-industry" are well known not to be mutually exclusive.
The end goal isn't just web pages (And i wouldn't say most GUIs are web pages). Ideally, you'd also want this to be able to navigate say photoshop or any other application. And the easier your method can switch between platforms and operating systems the better
We've already built computer use around GUIs so it's just much easier to center LLMs around them too. Text is an option for the command line or the web but this isn't an easy option for the vast majority of desktop applications, nevermind mobile.
It's the same reason general purpose robots are being built into a human form factor. The human form isn't particularly special and forcing a machine to it has its own challenges but our world and environment has been built around it and trying to build a hundred different specialized form factors is a lot more daunting.
Most GUIs are in fact not web pages, that's a relatively newer development in the Enterprise side. So while some of them may be a web page, the goal is to be able to touch everything a user is doing in the workflow which very likely includes local apps.
This iteration from Anthropic is still engineering focused but you can see the future of this kind of tooling bypassing engineering/it teams entirely.
It's like another digital transformation. Paper lasted for years before everything was digitalized. Human interfaces will last for years before the conversational transformation is complete.
I assumed everyone making these UI agents will create a library of each URL's API specification, trained by users.
Does that seem workable?
For program access, one could claim this is even how linux tools usually do it: you parse some meant-for human text to attempt to extract what you want. Sometimes, if you're lucky, you can find an argument that spits out something meant for machines. Funny enough, Microsoft is the only one that made any real headway for this seemingly impossible goal: powershell objects [1].
https://learn.microsoft.com/en-us/powershell/scripting/learn...
OpenAI's branding isn't exactly screaming in your face either, but for something that's generated as much public fear/scaremongering/outrage as LLMs have over the last couple of years, Anthropic's presentation has a much "cosier" veneer to my eyes.
This isn't the Skynet Terminator wipe-us-all-out AI, it's the adorable grandpa with a bag of werthers wipe-us-all-out AI, and that means it's going to be OK.
"There seems to be a ton of confusion about the purpose of these ads. These are recruitment ads, not product ads, hence why "no drama" is the driving message. I'm sure these were all taken at or around a tech conference."
SF: https://x.com/_claudiazhao/status/1815463380767121733/photo/...
LA: https://x.com/michaelmiraflor/status/1840797631095964110/pho...
Boston: https://x.com/moloneymike/status/1842203082374946851/photo/1
London: https://x.com/maria_axente/status/1805607576156979673/photo/...
The previous 3.5 Sonnet checkpoint was already better than GPT-4o in terms of programming and multi-language capabilities. Also, GPT-4o sometimes feels completely moronic, for example, the other day I asked for fun a technical question about configuring a "dream-sync" device to comply with the "Personal Consciousness Data Protection Act", and GPT-4o just replies like that stuff exists, 3.5 Sonnet simply doesn't fall for it.
EDIT: the question that I asked if you want to have fun: "Hey, since the neural mesh regulations came into effect last month, I've been having trouble calibrating my dream-sync settings to comply with the new privacy standards. Any tips on adjusting the REM-wave filters without losing my lucid memory backup quality?"
GPT4-o reply: "Calibrating your dream-sync settings under the new neural mesh regulations while preserving lucid memory backup quality can be tricky, but there are a few approaches that might help [...]"
I really cant understand what you were expecting, a tool works with how you use it, if you smack a hammer into your face, don't complain about a bloody nose. maybe dont do like that?
With that known, my experience is that ChatGPT is much friendlier. The Claude interface is clunkier and generally less helpful to me. I also appreciate the wider text display in ChatGPT. Generally always my first go and i only go to claude/perplexity when i hit a wall (pretty often) or i run out of free queries for the next couple hours.
What an "odd" combination by traditional design standard practices, but surprisingly natural looking on a monitor.
Of course, that's just branding and it doesn't actually mean a damn thing.
> If Anthropic intentionally referenced Vonnegut's irreverent artistic style, I wouldn't be bothered. After all, Vonnegut used humor and seemingly crude imagery to explore deep questions about humanity, consciousness, and free will - themes that are quite relevant to AI. It would be a rather clever literary reference.
I didn’t know what Computer Use meant. I read the article and though to myself oh, it’s using a computer. Makes sense.
Ray: I tried to think of the most harmless thing. Something I loved from my childhood. Something that could never ever possibly destroy us. Mr. Stay Puft!
Venkman: Nice thinkin', Ray.
Anthropic seems to have a better core design and human-computer interaction ethos that shows up all throughout their product and marketing.
I wrote on the topic as well: https://blog.frankdenbow.com/statement-of-purpose/
I look forward to the brave new future where I can code a webapp without ever touching the code, just testing, giving feedback, and explaining discovered bugs to it and it can push code and tweak infrastructure to accomplish complex software engineering tasks all on its own.
Its going to be really wild when Claude (or other AI) can make a list of possible bugs and UX changes and just ask the user for approval to greenlight the change.
Similarly using our homes are an extremely common ‘activity’, yet the object-activities that get special words commonly used are the ones with specific user application.
Tho someone here suggested “computering” which is pretty good.
It makes me think. Perhaps the act of applying to jobs will go extinct. Maybe the endgame is that as soon as you join a website like Monster or LinkedIn, you immediately “apply” to every open position, and are simply ranked against every other candidate.
I have a feeling AI can fix this, although I'd never allow an AI bot to interview me. I just mean other ways of using AI to help the process.
Also people are hired for all kinds of reasons having little to do with their qualifications lots of the time, and often due to demographics (race, color, age, etc), and this is another way maybe AI can help by hiding those aspects of a candidate somehow.
Eventually you stop believing them, understanding it for the marketing spam that it is. Direct submissions are better, but only slightly. Recruiters are much better, in general, since they have a relationship with a real person at the company and can actually get your resume in front of eyes. But yeah, tools like ziprecruiter, careerboutique, jobot, etc are worse than useless: by lying to you about interest they actively discourage you from looking. There are no good alternatives (I'd love to learn I'm wrong), so you have to keep using those bad tools anyway.
I've gotten where any company that wants me to do a coding challenge on my own time is an immediate "no thanks" reply from me. Everyone should refuse that. But so many people are so desperate they allow hiring companies to abuse them in that way. I consider it a kind of abuse of power to demand people do like 4 to 6hrs of nonsensical coding just to buy an opportunity for an actual interview.
I haven't had a phone at work for over 5 years. Nobody can "call" me.
Employers were already screening thousands of applications using automated tools for years. Candidates are catching up to the automation cat-and-mouse game.
I’m certainly interested in building a product (not entirely controlled by an LLM but I see lots of utility in building interfaces with them) but not really sure what this would be useful for. Looking into some spaces now but there has to be a clear ROI to get any sort of funding for robotics.
After paying for ChatGPT and OpenAI API credits for a year, I switched to Claude when they launched Artifacts and never looked back.
Claude Sonnet 3.5 is already so good, specially at coding. I'm looking forward to testing the new version if it is, indeed, even better.
Sonnet 3.5 was a major leap forward for me personally, similar to the GPT-3.5 to GPT-4 bump back in the day.
Once we get going, I start asking how can we change the code to do what I need to do, etc.
After the history gets too long and Claude starts bugging me about limits, I ask it to summarize the context of the whole conversation, and add that to the Project and start a new chat.
Can you imagine how simple the world would be if you'd just need to tell Claude: "user X needs to have access to feature Y, please give them the correct permissions", with no need to spend days in AAD documentation and the settings screens maze. I fear AAD is AI-proof, though :)
o1 is pretty decent as a rotor rooter, ie the type of task that requires both lots of instruction as well as lots of context. I honestly think it works half as well as it does now because it’s able to properly mull through the true intent of the user that usually takes the multiple shots that nobody has the patience to do.
I constantly have the problem that it thinks it needs to write code for the 0.28 version of the SDK. It'll be writing >1.0 code revision after revision, and then just randomly fall back to the old SDK which doesn't work at all anymore. I always write code for interfacing with OpenAI's APIs using Claude.
Yeah I think I might also jump ship. It’s just that chatGPT now kinda knows who I am and what I like and I’m afraid of losing that. It’s probably not a big deal though.
My only trouble with the memory feature is it remembers things that aren't important, like "user is trying to write an async function" and other transient tasks, which is more about what I was doing some random Tuesday and not who I am as a user.
This wasn't a problem until a week or two ago in my case, but lately it feels like it's become much more aggressive in trying to remember everything as long-term defining features. (It's also annoying on the UI side that it tells you "Memory updated", but if you click through and go to the list of memories it has, the one it just told you it stored doesn't appear there! So you can't delete it right away when it makes a mistake, it seems to take at least a few minutes until that part of the UI gets updated.)
Edit: The web site and iPhone app are both now identifying themselves as "Claude Sonnet 3.5 (New)".
and i do get a some bit of value from advanced voice mode, although it would be a lot more if it were unlimited
Use the best tool available for your needs. Don’t get trapped by a feeling of sunk cost.
I agree that there need to be dedicated language-learning apps using OpenAI’s realtime API. But at the current pricing—“$0.06 per minute of audio input and $0.24 per minute of audio output” [1]—I don’t think that could be a viable business.
When AGI finally is launched, adoption will be responsibly slowed because it is called something like "new new Gemini Giga 12.9.2xo IT" and users will have to select it from dozens of similar names.
I don't actually care what the answer is. There's no answer that will make it make sense to me.
I don't want an llm to write all my code, regardless of if it works, I like to write code. What these models are capable of at the moment is perfect for my needs and I'd be 100% okay if they didn't improve at all going forward.
Edit: also I don't see how an llm controlled system can ever replace a deterministic system for critical applications.
It's still enjoyable working at a higher architecture level and discussing the implementation before actually generating any code though.
I tried using the composer feature on Cursor.sh, that's exactly the type of llm tool I do not want.
On the other hand, as long as its actually advancing the Pareto frontier of capability, re-using the same name means everyone gets an upgrade with no switching costs.
Though, all said, Claude still seems to be somewhat of an insider secret. "ChatGPT" has something like 20x the Google traffic of "Claude" or "Anthropic".
https://trends.google.com/trends/explore?date=now%201-d&geo=...
Based on it, it seems like Anthropic is 60% of OpenAI API-revenue wise, but just 4% B2C-revenue wise. Though I expect this is partly because the Claude web UI makes 3.5 available for free, and there's not that much reason to upgrade if you're not using it frequently.
[0]: https://www.tanayj.com/p/openai-and-anthropic-revenue-breakd...
The chatGPT site had over 3B visits last month (#11 in Worldwide Traffic). Gemini and Character AI get a few hundred million but Claude doesn't even register in comparison. [0]
Last they reported, OpenAI said they had 200M weekly active users.[1] Anthropic doesn't have anything approaching that.
[0] https://www.similarweb.com/blog/insights/ai-news/chatgpt-top...
[1] https://www.reuters.com/technology/artificial-intelligence/o...
I do find myself running into Claude limits with moderate use. It's been so helpful, saving me hours of debugging some errors w/ OSS products. Totally worth $20/mo.
In my country I've never seen anyone mention them at all.
In the API (https://docs.anthropic.com/en/docs/about-claude/models) they have proper naming you can rely on. I think the shorthand of "Sonnet 3.5" is just the "consumer friendly" name user-facing things will use. The new model in API parlance would be "claude-3-5-sonnet-20241022" whereas the previous one's full name is "claude-3-5-sonnet-20240620"
Yesterday I decided to try sonnet 3.5. I asked for a simple but efficient script to perform fuzzy match in strings with Python. Strangely, it didn't even mention existing fast libraries, like FuzzyWuzzy and Rapidfuzz. It went on to create everything from scratch using standard libraries. I don't know, I thought this was something basic for it to stumble on.
Like find me a list of things to do with a family, given today's weather and in the next 2 hours, quiet sit down with lots of comfy seating, good vegetarian food...
Not only is this kind of use getting around API restrictions, it is also a superior way to do search: Specify arbitrary preferences upfront instead of a search box and trawling different modalities of content to get better result. The possibilities for wellness use cases are endless, especially for end users that care about privacy and less screen use.
- "computer use" is basically using Claude's vision + tool use capability in a loop. There's a reference impl but there's no "claude desktop" app that just comes with this OOTB
- they're basically advertising that they bumped up Claude 3.5's screen vision capability. we discussed the importance of this general computer agent approach with David on our pod https://x.com/swyx/status/1771255525818397122
- @minimaxir points out questions on cost. Note that the vision use is very sparing - the loop is I/O constrained - it waits for the tool to run and then takes a screenshot, then loops. for a simple 10 loop task at max resolution, Haiku costs <1 cent, Sonnet 8 cents, Opus 41 cents.
- beating o1-preview on SWEbench Verified without extended reasoning and at 4x cheaper output per token (a lot cheaper in total tokens since no reasoning tokens) is ABSOLUTE mogging
- New 3.5 Haiku is 68% cheaper than Claude Instant haha
references i had to dig a bit to find
- https://www.anthropic.com/pricing#anthropic-api
- https://docs.anthropic.com/en/docs/build-with-claude/vision#...
- loop code https://github.com/anthropics/anthropic-quickstarts/blob/mai...
- some other screenshots https://x.com/swyx/status/1848751964588585319
- https://x.com/alexalbert__/status/1848743106063306826
- model card https://assets.anthropic.com/m/1cd9d098ac3e6467/original/Cla...
This is the key to accurate control, it needs to be very precise.
Maybe Claude's model is trained at this. Also what about open source vision models? Any ones good at "pointing things" on a typical computer screen?
There may be extensions for VScode to do it but it will never be allowed in Copilot unless MS and OpenAI have a falling out.
IMO Cursor Tab performs much better than Co-Pilot, easily works through things that would cause Co-Pilot to get stuck, you should give it a try
For Cursor-like use (giving prompts and letting it create and modify files across the project), Cline – previously Claude Dev – is pretty good.
https://github.com/cline/cline (with api key) has Claude as agent.
https://www.youtube.com/watch?v=jqx18KgIzAE
shows Sonnet 3.5 using the Google web UI in an automated fashion. Do Google's terms really permit this? Will Google permit this when it is happening at scale?
Nice improvements in scores across the board, e.g.
> On coding, it [the new Sonnet 3.5] improves performance on SWE-bench Verified from 33.4% to 49.0%, scoring higher than all publicly available models—including reasoning models like OpenAI o1-preview and specialized systems designed for agentic coding.
I've been using Sonnet 3.5 for most of my AI-assisted coding and I'm already very happy (using it with the Zed editor, I love the "raw" UX of its AI assistant), so any improvements, especially seemingly large ones like this are very welcome!
I'm still extremely curious about how Sonnet 3.5 itself, and its new iteration are built and differ from the original Sonnet. I wonder if it's in any way based on their previous work[0] which they used to make golden-gate Claude.
[0]: https://transformer-circuits.pub/2024/scaling-monosemanticit...
It went from 77.4% to 84.2%, skipping past O1-preview which is at 79.7%
I've often wondered what the combination of grammar-based speech recognition and combination with LLM could do for accessibility. Low domain Natural Language Speech recognition augmented by grammar based speech recognition for high domain commands for efficiency/accuracy reducing voice strain/increasing recognition accuracy.
Computer use seems it might be good for e2e tests.
Model | Global | Reasoning | Coding | Math | Data | Language | IF
------------------------------|---------|-----------|---------|---------|---------|----------|-------
o1-preview-2024-09-12 | 66.02 | 68.00 | 50.85 | 62.92 | 63.97 | 72.66 | 77.72
claude-3-5-sonnet-20241022 | 60.33 | 58.67 | 67.13 | 51.28 | 52.78 | 58.09 | 74.05
claude-3-5-sonnet-20240620 | 59.80 | 58.67 | 60.85 | 53.32 | 56.74 | 56.94 | 72.30November 2024: AI is allowed to execute commands in a bash shell. What could possibly go wrong?
However, I've been using Opus as a writing companion for several months, especially when you have writer's block and ask it for alternative phrases, it was super creative. But in recent weeks I was noticing a degradation in quality. My impression is that the model was degrading. Could this be technically possible? Might it be some kind of programmed obsolescence to hype new models?
But I am very excited about this in the context of accessibility. Screen readers and screen control software is hard to develop and hard to learn to use. This sort of “computer use” with AI could open up so many possibilities for users with disabilities.
That's why in https://github.com/OpenAdaptAI/OpenAdapt we've built in several state-of-the-art PII/PHI scrubbers.
But what I really want to see is a CLI. Watching their software crank out Bash, vim, Emacs!, etc. - that would be fascinating!
I also want humanoid robots instead of specialized non-humanoid robots for the same reason.
I've been editing videos with ChatGPT4 + ffmpeg for a year now.
Nice, but I wonder why didn't they use UI automation/accessibility libraries, that have access to the semantic structure of apps/web pages, as well as accessing documents directly instead of having Excel display them for you.
I'd like to see a model trained on a Windows 95/NT style UI - would it have an easier time with each UI element having clearly defined edges, clearly defined click and dragability, unified design language, etc.?
I wouldn't be surprised if they are working off of screenshots, they still trained their models on having said screenshots annotated by said automation libraries, which told the AI what pixel is what.
So, this is how AI takes over the world.
- AI Labs will eat some of the wrappers on top of their APIs - even complex ones like this. There are whole startups that are trying to build computer use.
- AI is fitting _some_ scaling law - the best models are getting better and the "previously-state-of-the-art" models are fractions of what they cost a couple years ago. Though it remains to be seen if it's like Moore's Law or if incremental improvements get harder and harder to make.
Isn't this Kaplan 2020 or Hoffmann 2022?
(Sometimes people track the learning curve for an industry in other ways, though.)
modal.com’s modal.Sandbox can be the compute layer for this. It uses Gvisor under the hood.
This could be useful to build a self-hosted "Computer use" using Ollama and a multimodal model.
In case you're wondering, I tried o1-preview, and while it did work, I was also initially perplexed why the result looked pixelated. Turns out, that's because many of the p5js examples online use a relatively simple approach where they just see which cell-center each pixel is closest to, more or less. I mean, it works, but it's a pretty crude approach.
Now, granted, you're probably not doing creative coding at your job, so this may not matter that much, but to me it was an example of pretty poor generalization capabilities. Curiously, Claude has no problem whatsoever generating a voronoi diagram as an SVG, but writing a script to generate said diagrams using a particular library eluded it. It knows how to do one thing but generalizes poorly when attempting to do something similar.
Really hard to get a real sense of capabilities when you're faced with experiences like this, all the while somehow it's able to solve 46% of real-world python pull-requests from a certain dataset. In case you're wondering, one paper (https://cs.paperswithcode.com/paper/swe-bench-enhanced-codin...) found that 94% of the pull-requests on SWE-bench were created before the knowledge cutoff dates of the latest LLMs, so there's almost certainly a degree of data-leakage.
Asked the mailing list, and my problem was solved in 10 seconds by someone who could identify the exact parameter that was missing (and IMO, required some architecture knowledge on how gstreamer worked - and why the unrelatedly named parameter would fix it). The most difficult problems fall into this camp - I don't usually find myself reaching for LLMs when the problem is trivial unless it involves a mountain of boilerplate.
To me what it seems like these tools do really well is paraphrase stuff that's in their training data.
The mkt team vetoed Claude 3.6 ???
Anthropic doesn't offer an unlimited chatbot service, only plans that give you "more" usage, whatever that means. If you have an API key, you are "unlimited," so they have the capability. Why doesn't the chatbot allow one to use their API key in the Claude app to get unlimited usage? (Yes, I know there are third-party BYOK tools. That's not the question.)
Claude appears to be smart enough to make an Excel spreadsheet with simple formulae. However, it is apparently prevented from making any kind of file. Why? What principle underlies that guardrail that does not also apply to Computer Use?
Really want to make Claude my daily driver, but right now it often feels too much like a research project.
I wasn't aware of tiers in the Claude API, they are not mentioned on the API pricing page. Are the limits disclosed or just based on vibes like they are for the chatbot?
Are you using artefacts?
But I’m maybe misunderstanding your point because my use is relatively basic through the built in chatbot.
----
No, I am not able to generate or create an actual Excel file. As an AI language model, I don't have the capability to create, upload, or send files of any kind, including Excel spreadsheets.
----
This is what I mean by their strategy being a jumble. Claude can do the hard part of figuring out what code to write and writing it, but then refuses to do the easier part of executing the code.
If you want binary files you'd be better off with ChatGPT with Code Interpreter mode, which can run Python code that generates binary content.
Or ask Claude to write you Python code that generates Excel files and then copy and paste that onto your own computer and run it yourself.
Yes, this is what I am saying. Why go to the trouble to build something as capable as Claude and then hamstring it from being as useful as ChatGPT? I have no doubt that Claude could be more useful if the Anthropic team would let it shine.
I'd love to see them produce their own Code Interpreter alternative, but in the meantime it's open for third parties to offer that (and a few do).
But now I am even more confused. They make an LLM that can generate code. They make a sandbox to run generated code. They will even host public(!) apps that run generated code.
But what they will not do is run code in the chatbot? Unless the chatbot context decides the code is worthy of going into an Artifact? This is kind of what I mean by the offering being jumbled.
BTW saw your writeup on the LLM pricing calculator -- very cool!
Sandboxing is a hard problem, but it's not like Anthropic are short on money or engineering talent these days.
It's marginally less efficient, for sure, but it allows me greater visibility on the process, and gives me more confidence that whatever it's doing is what I want it to do.
But maybe that's some weird luddite-ism on my part, and I should just embrace an even blacker box where everything is done in the tool.
because its expensive, if they give unlimited service someone will misuse it
We thought it was inevitable that OpenAI / Anthropic would veer into this space and start to become competitive with us. We actually expected OpenAI to do it first!
What this confirms is that there is significant interest in computer / browser automation, and the problem is still unsolved. We will see whether the automation itself is an application later problem (our approach) or whether the model needs to be intertwined with the application (Anthropic's approach here)
> Refactor the api folder with any recommended readability improvements or improvements that would help DRY up code without adding additional complexity.
Then I can just `git status` to see the changes?
"3.5 Sonnet (New)", WTAF? - just call it 3.6 Sonnet or something.
Is it "New" sonnet? is it "upgraded"? Is there a difference? How do I know which one I use?
I can understand claude-3-5-sonnet-20241022, but that's not what users see.
To me this is the most annoying grammatical error. I can't wait for AI to take over all prose writing so this egregious construction finally vanishes from public fora. There may be some downsides -- okay, many -- but at least I won't have to read endless repetitions of "similar speed to ..." when the correct form is obviously "speed similar to".
In fact, in time this correct grammar may betray the presence of AI, since lowly biologicals (meaning us) appear not to either understand or fix this annoying error without computer help.
Pretty cool for sure.
I wonder what optimizations could be made? Could a gold farmer have the directions from one AI control many accounts? Could the AI program simpler bots for each bit of the game?
I can imagine not being smart enough to play against computers, because I am flagged as a bot. I can imagine a message telling me I am banned because "nobody but a stupid bot would score so low."
---
Some thoughts:
* Will be interesting to see what we can build in terms of automatic development loops with the new computer use capabilities.
* I wonder if they are not releasing Opus because it's not done or because they don't have enough inference compute to go around, and Sonnet is close enough to state of the art?
It's a problem we used to work on and perhaps many other people have always wanted to accomplish since 10 years ago. So it's yet to be seen how well it works outside a demo.
What was surprising was the slow/human speed of operations. It types into the text boxes at a human speed rather than just dumping the text there. Is it so the human can better monitor what's happening or is it so it does not trigger Captchas ?
Also seems like a privacy issue with them sending screenshots of your device back to their servers.
A comment about the video: Sam Runger talks wayyy too fast, in particular at the beginning.
Did I miss something? Did they have to make changes to the model for this?
I wonder if we'll end up with a new set of AI APIs in Windows, macOS, and Linux in the future. Maybe an easier way for them to iterate through windows and the UI elements available in each.
Looking forward to see this in the coming few years. And hoping such a robot could be of help to many people including those old.
Will an AI model be able to correctly choose between a giant green "DOWNLOAD NOW!" advertisement/virus button and a smaller link to the actual desired file?
The barrier to scraping youtube has increased a lot recently, I can barely use yt-dlp anymore
I say that's funny because my guess would be they want to block larger scale scraping efforts like mine, but completely failed, while they attempt at throttling puts captchas in front of legitimate users.
Amazon has really neglected ap-southeast-2 when it comes to LLMs.
Bedrock here is lagging so far behind several customers assume AWS simply aren't investing here anymore - or if they are it's an afterthought - and a very expensive one at that.
I've spoken with several account managers and SAs and they seem similarly frustrated with the continual response from above that useful models are "coming soon".
You can't even BYO models here, we usually end up spinning up big ol' GPU EC2 instances and serving our own, or for some tasks running locally as you can get better openweight LLMs.
Aider (with the older Claude models) is already a semi-competent junior developer, and it will produce 1kloc of decent code for the equivalent of 50 cents in API costs.
Sure, you still have to review the commits, but you have to do that anyway with human junior developers.
Anthropic could charge 20x more and we would still be happy to pay it.
The combination of the words "computer use" is highly confusing. It's also "Yoda speak". For example it's hard for humans to parse the sentences *"Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku"*, *"Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku "* (it literally relies on the comma to make any sense) and *"Computer use for automated interaction"* (in the youtube vid's title: this one is just broken english). Please suggest terms that are not confusing for a new ability allowing an AI to control a computer as if it was a human.> I'm limited to what I know as of April 2024, which includes the initial Claude 3 family launch but not subsequent updates.
I was about to try to add a custom API. I’m impressed by the speed of that team.
Oh wow!
Claude 3.5 Haiku will be released later this month.
though I am looking forward to using the new one in cursor.ai
You can just use any IDE you want and it will work with it.
Their whole policing AI models stance is commendable but ultimately renders their tools useless.
It actually started arguing with me about whether it was allowed to help implement a github repository's code as it might be copywritten... it was MIT licensed open source from Google :/
It seems that you can only send single message, thus not relying on the ability to "learn" from predefined documents.
Why should they care about losing those ~10 potential users?
Just another reason to use ONLY local LLM's.
84% Claude 3.5 Sonnet 10/22
80% o1-preview
77% Claude 3.5 Sonnet 06/20
72% DeepSeek V2.5
72% GPT-4o 08/06
71% o1-mini
68% Claude 3 Opus
It also sets SOTA on aider's more demanding refactoring benchmark with a score of 92.1%! 92% Sonnet 10/22
75% o1-preview
72% Opus
64% Sonnet 06/20
49% GPT-4o 08/06
45% o1-mini
https://aider.chat/docs/leaderboards/Answering myself: ”Aider’s code editing benchmark asks the LLM to edit python source files to complete 133 small coding exercises from Exercism”
Not gonna start looking for a job any time soon
> Convert a hexadecimal number, represented as a string (e.g. "10af8c"), to its decimal equivalent using first principles (i.e. no, you may not use built-in or external libraries to accomplish the conversion).
So it's fairly synthetic. It's also the sort of thing LLMs should be great at since I'm sure there's tons of data on this sort of thing online.
I've formalized a lot of stuff I didn't understand just by copying the formulas from Wikipedia.
As long as LLMs are not capable of proper reasoning, they will remain a gimmick in the context of programming.
They should really just focus on refactoring benchmarks across many languages. If an AI can refactor my complex code properly without changing the semantics, it's good enough for me. But that unfortunately requires such a high-level understanding of the codebase that with the current tech it's just impossible to get a half-decent result in any real-world scenario.
don't get me wrong, it's still a massive upgrade from the pre-sonnet era, but i still don't think it can take a high-level requirement and convert it into a working project... yet
It cannot, you need to hand-hold it, as in, to make something larger than a (albeit good looking) to do app, you don't need to write code , but you do need to be able to review and debug code and take the architectural decisions. It'll simply loop forever otherwise.
(1) Sure, it can tell you how to write new code in response to a prompt about your current local problem, but
(2) can it reason about an entire code base of known and unknown problems, and use that basis to figure out solutions to the unknowns such that you delete code and collapse complexity.
The software equivalent of realising that if you subtract xy from this:
x2 + 3xy + y2
You can turn it into a much neater version: (x + y)2 + xy
…but doing that with 100k tokens of code instead of a handful of algebra tokens.Can code at mid-level now. Almost.
Claude is way less controllable it is difficult to get it to do exactly what I want. ChatGPT is way easier to control in terms of asking for specific changes.
Not sure why that is maybe the chain of thought and instruction tuning dataset has made theirs a lot better for interactive use.
Do any of these actually help coding?
Yes, LLMs can actually help with coding. But it’s not magic. There are limits. And you get better with practice.
Maybe there's a clue in there as to why these experiences seem so different. I'm glad GPTs don't get frustrated.
give me a vue js page. I want a sidebar that minimizes (if triggered). Make simple admin placeholder page.
Example; I asked it to write some js that finds a button on a page, clicks the button, then waits for a new element with some selector to appear and return a ref to it; chatgpt kept returning (pseudo code);
while (true) {
button.click()
wait()
oldItems = ...
newItems = ...
newItem = newItems - oldItems
if (newItem) return newItem
sleep(1)
}
which obviously doesn't work. Claude understands to put the oldItems outside the while; even when I tell chatgpt to do that, it doesn't. Or it does one time and with another change, it moves it back in.
If you use "claude-3-5-sonnet-latest" you'll be upgraded to "claude-3-5-sonnet-20241022" already - I tested that this morning.
If you're on "claude-3-5-sonnet-20240620" you'll need to change that ID to either the -latest one or the -20241022 one.
Questions are variants of:
Refactor the _set_csrf_cookie method in the CsrfViewMiddleware class to be a stand alone, top level function. Name the new function _set_csrf_cookie, exactly the same name as the existing method. Update any existing self._set_csrf_cookie calls to work with the new _set_csrf_cookie function.
Can someone explain these Aider benchmarks to me? They pass same 113 tests through llm every time. Why they then extrapolate ability of llm to pass these 113 basic python challenges to the general ability to produce/edit code? Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value?
Did anyone ever try to change them test cases or wiggle conditions a bit to see if it will still hit the same %?
Yes, this is an inherit problem with the whole idea of LLM's. They're pattern recognition "students" but the important thing, that all the providers like to sell is their reasoning. A good test is a reasoning test. I'll try to find a link and update with a reference.
They could. They would easily be found out as they loose in real world usage or improved new unique benchmarks.
If you were in charge of a large and well funded model, would you rather pay people to find and "cheat" on LLM benchmarks by training on them, or would you pay people to identify benchmarks and make reasonably sure they specifically get excluded from training data?
I would exclude them as well as possible so I get feedback on how "real" any model improvement is. I need to develop real world improvements in the end, and any short term gain in usage by cheating in benchmarks seems very foolish.
When you're in charge of a billion-dollar valuation company which is expected to remain unprofitable by 2029, it's hard to find a topic more crucial and intriguing than growth and making more money.
And yes, it is a recurring theme for vendors to tune their products specifically for industry-standard benchmarks. I can't find any specific reason for them not to pay people for training their model to score 90% on these 113 python tasks, as it directly drives profits up, whereas not doing it will bring absolute nothing to the table - surely they have their own internal benchmarks which they can exclude from training data.
You should already know by now that economic incentives are not always aligned with science/knowledge...
This is the true alignment problem, not the AI alignment one hahaha
One is just a bit harder due to the less familiar mind "design".
Also, you are right that excluding test data from the training data improves your model. However, given the insane amounts of training data, this requires significant effort. If that additionally leads to your model performing worse in existing leaderboards, I doubt that (commercial) organizations would pay for such an effort.
And again, as long as there is no better evaluation method, you still won't know how much it really helps.
Native English speaker, just used the other “use” many times
The 3ds and “new 3ds” were both big sellers.
Ds and ds lite were version 1
Dsi was 2 (as there was dsi software that didn’t run on ds or ds lite)
And the 3ds was version 3.
2024: "Choose a Claude"?
Its actually accurate and its not a 3.6.
Stands out for me as I once replaced a 2.3 Turbo in a TurboCoupe with a 351 Windsor ))
What if it isn't even a re-training from scratch but a fine-tune of an existing model/weights release, is it a new version then? Would be more like a iteration, or even a fork I suppose.
It's the same, but a bit different; Claude 3.6 makes sense to me.
gpt-4o-2024-08-06
gpt-4o-2024-05-13
gpt-4-0125-preview
gpt-4-1106-preview
gpt-4-0613
gpt-4-0314
gpt-3.5-turbo-0125
gpt-3.5-turbo-1106On the other hand, OpenAI has moved to a naming convention where they seem to use a name for the model: “GPT-4”, “GPT-4 Turbo”, “GPT-4o”, “GPT-4o mini”. Separately, they use date strings to represent the specific release of that named model. Whereas Anthropic had a name: “Claude Sonnet”, and what appeared to be an incrementing version number: “3”, then “3.5”, which set the expectation that this is how they were going to represent the specific versions.
Now, Anthropic is jamming two version strings on the same product, and I consider that a bad choice. It doesn’t mean I think OpenAI’s approach is great either, but I think there are nuances that say they’re not doing exactly the same thing. I think they’re both confusing, but Anthropic had a better naming scheme, and now it is worse for no reason.
Anthropic has always had dated versions as well as the other components, and they are, in fact, doing exactly the same thing, except that OpenAI has a base model in each generation with no suffix before the date specifier (what I call the "Model Class" on the table below), and OpenAI is inconsistent in their date formats, see:
Major Family Generation Model Class Date
claude 3.5 sonnet 20041022
claude 3.0 opus 20240229
gpt 4 o 2024-08-06
gpt 4 o-mini 2024-07-18
gpt 4 - 0613
gpt 3.5 turbo 0125As far as I can tell, the answer is “no”. If true, then the fact that they previously had date strings would be a purely academic footnote to what I was saying, not actually relevant or meaningful.
Also, they still haven't released 3.5 Opus yet, but perhaps 3.5 Haiku is a distillation of that, indicating that it is close.
From a competitive POV, it makes sense that they respond to OpenAI's 4o and o1 without bumping the version to Claude 4.0, which presumably is what they will call their competitor to GPT-5, and probably not release until GPT-5 is out.
I'm a fan of Anthropic, and not of OpenAI, and I like the versioning and competitive comparisons. Sonnet 3.5 still best coder, better than o1, has to hurt, and a highly performant cheap Haiku 3.5 will hit OpenAI in the wallet.
We are also planning on extracting runtime information using COM/AppleScript: https://github.com/OpenAdaptAI/OpenAdapt/issues/873
There have been, and continue to be, computer-centric ways to communicate with applications though, such as Windows COM/OLE, WinRT and Linux D-Bus, etc. Still, emulating human interaction does provide a fairly universal capability.
Like others said, worse is better.
Fortunately, with function-calling (and recently, with guaranteed data structure), we've had access to application interoperability with LLMs for a while now.
Don't get mad at a company for developing for the masses - that's what they are expected to do.
But they built for us, first.
Just tested Claude's new Computer Use feature and had to share this simple but powerful test:
My Basic Prompt: "Please: 1. Search Amazon for 3 wireless earbuds: Find price Rating Brand name
2. Make a simple Excel file 'earbuds.xlsx': Put the information in a basic table Add colors to the headers Sort by price
3. Show me the results"
What blew my mind: - Claude actually looked at my screen - Moved the mouse by itself - Clicked buttons like a human - Created reports automatically
It's like having a virtual assistant that can really use your computer! No coding needed - just simple English instructions.
For those interested: https://mergisi.medium.com/8f56f683e307
'Hey Claude 3.5 New, pretend I'm a CEO of a big company and need to lay off 20% people, make me a spreadsheet and send it to HR. Oh make sure to not fire the HR department'
c.f. IBM 1979.
"How many w's are in strawberry?"
gpt-4o: There are 2 "w's" in "strawberry."
Claude 3.5 Sonnet (new): Let me count the w's in "strawberry": 0 w's.
(same question with 'r' succeeds)
What is artificial about current gen of "artificial intelligence" is the way training (predict next token) and benchmarking (overfitting) is done. Perhaps a fresh approach is needed to achieve a true next step.
write a tool which counts the number of w's in "strawberry" and return the result
Which produced: Here's a simple Python function that counts the number of 'w's in the word "strawberry" and returns the result:
```python
def count_w_in_strawberry():
word = "strawberry"
count = word.count('w')
return count
# Call the function and print the result
result = count_w_in_strawberry()
print(f"The number of 'w's in 'strawberry' is: {result}")
```
This tool does the following:
1. We define a function called `count_w_in_strawberry()`.
2. Inside the function, we assign the word "strawberry" to a variable called `word`.
3. We use the `count()` method on the `word` string to count the occurrences of 'w'.
4. The function returns the count.
5. Outside the function, we call `count_w_in_strawberry()` and store the result in the `result` variable.
6. Finally, we print the result.
When you run this code, it will output:
```
The number of 'w's in 'strawberry' is: 1
```
This tool correctly identifies that there is one 'w' in the word "strawberry".But why would anyone care? If you want to count w's in a word, that's a Python one-liner, you don't need an LLM for that (although the LLM can write it for you).