436 karma · joined April 15, 2014
Last night, I asked Claude 3.7 Sonnet to obtain historical gold prices in AUD and the ASX200 TR index values and plot the ratio of them, it got all of the tickers wrong - I had to google (it then got a bunch of other stuff wrong in the code).
Also yesterday, I was preparing a brief summary of forecasting metrics/measures for a stakeholder and it incorrectly described the properties of SMAPE (easily validated by checking Wikipedia).
I constantly have issues with my direct reports writing code using LLM's. They constantly hallucinate things for some of the SDK's we use.
> Personally, when I want to get a sense of capability improvements in the future, I'm going to be looking almost exclusively at benchmarks like Claude Plays Pokemon.
Definitely interested to see how the best models from Anthropics competitors do at this.,
Definitely made me respect the bloggers I read even more.
The Very Bad Wizards podcast on it is interesting/fun too. https://podcasts.apple.com/us/podcast/very-bad-orgies-kubric...
I'm very keen on one of these, but I simply have no idea how good they are at my day to day tasks in R or Python.
I've found the reaction to this article can be pretty intense. We read this in a journal club many years ago and one of the mathematicians who was kind of new to the idea that research papers (in other fields) didn't more or less represent 'truth' said this article was _dangerous_.
However, as far as I can tell, it's never actually clear what the hardware requirements are to get these to run without fussing around. Am I wrong about this?