This is what I mean by their strategy being a jumble. Claude can do the hard part of figuring out what code to write and writing it, but then refuses to do the easier part of executing the code.
If you want binary files you'd be better off with ChatGPT with Code Interpreter mode, which can run Python code that generates binary content.
Or ask Claude to write you Python code that generates Excel files and then copy and paste that onto your own computer and run it yourself.
Yes, this is what I am saying. Why go to the trouble to build something as capable as Claude and then hamstring it from being as useful as ChatGPT? I have no doubt that Claude could be more useful if the Anthropic team would let it shine.
I'd love to see them produce their own Code Interpreter alternative, but in the meantime it's open for third parties to offer that (and a few do).
But now I am even more confused. They make an LLM that can generate code. They make a sandbox to run generated code. They will even host public(!) apps that run generated code.
But what they will not do is run code in the chatbot? Unless the chatbot context decides the code is worthy of going into an Artifact? This is kind of what I mean by the offering being jumbled.
BTW saw your writeup on the LLM pricing calculator -- very cool!
Sandboxing is a hard problem, but it's not like Anthropic are short on money or engineering talent these days.
It's marginally less efficient, for sure, but it allows me greater visibility on the process, and gives me more confidence that whatever it's doing is what I want it to do.
But maybe that's some weird luddite-ism on my part, and I should just embrace an even blacker box where everything is done in the tool.