Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?
Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?
My limited experience with fable over the last few days suggests (1) I can’t see any improvement in output, and (2) it is useless for writing secure software because it constantly hits safety walls if you ask it to close security holes.
I’m definitely shopping around for other LLM providers next week, and testing vs local (target: 128GB strix halo - any war stories?)
That’s with heavy compression of the weights and the context, of course.
I haven’t gone through model evaluation + shoehorning at 128GiB yet.
this is exactly why hypotheses come before the experiment in the scientific method.
Some model cards do show regressions on benchmarks for newer models on specific tasks: https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
This wasn't a new model but updates to models backed by numbers being better can make the model worse: https://openai.com/index/sycophancy-in-gpt-4o/
The slight increases in performance/benchmarks may be just noise: https://arxiv.org/pdf/2602.07150
1. The sloppy/unpredictable behavior of LLMs as a general class of algorithm, how you shouldn't use document-generation for calculating budgets, and you shouldn't trust it to not-alter things you "asked" it to to alter.
2. Vendors of thing-as-a-service (not necessarily only LLMs) putting in traps and sabotage to prioritize their own business-model or economic incentives.
Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.
That has to be the understatement of the century.
Why go to bat for anti-consumer behaviors unless you are a shareholder?
Their billions are not my problem; but the money I pay them and service I get in return, is. And if they can't provide, I will shop elsewhere (and do).
If we are talking about distillation vs building from scratch, none of these are congruent to Windows. I can build my own LLM [0] and then distill off of Claude, but that is not the same as a 1:1 copy of an operating system because there was the ability to crack how licensing works. We are not seeing Windows clones, at the source level, for that reason.
Also, Linux exists. Anyone can copy that. Why doesn't that count?
>anthropic
> mine the internet for data, blasting millions of blogs with scrapers
>a few have to shut down, but that's just the price to pay
>finally, the chatbot is ready
>learn that there are EVIL cretins out there trying to scrape automated output from OUR product to build their chatbot
>build in safeguards to new model to stop this
>the users are mad, now the model accuses users of being bioterrorists if they so much as mention they have a cold
>mfw