So I'd usually share what I read before this.
4,951 karma · joined March 8, 2012
So I'd usually share what I read before this.
What you are saying is like me losing 13kg is as impressive as having the potential to do so if only I would stop eating ice cream and too many snacks every day.
I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.
So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.
Spark is such a bad model, you are far better off using Astra Medium than using Spark.
Sure there is the downside of having to be at your laptop or having it on, but the upside is that it has very few moving parts and is very simple and it just has one set price.
Then again I don’t live in Texas let alone the US so i might not know where to look and I don’t care enough to truly find out.
I was just surprised that they were supposedly illegal. Which seems untrue.
Homeowners can’t put solar panels on their roof to use the produced electricity?
It monitors any server you add to it and it can install a daemon that collects metrics for when you are offline. I keep it running to monitor our servers.
Kagi can do things at their scale that are not feasible for me right now and one of them is categorising links like they probably do for this feature. So it is inspiring. A route that might work is checking if there are ublock lists that keep track of paywalled sites.
I've also been looking into creating my own search index, but this is a huge undertaking. I might go the tiny index route like Marginalia has done and try to find a niche for my specific use cases.
Another route could be to categorise websites better automatically and build up a category index.
Things will start to get interesting where €1600 don't buy you 5kWh but double or triple the amount.
I can imagine most people are better off either investing like I said or in insulation.
I found that most things that are “features” in harnesses and tools that don’t relate to UI and UX in those harnesses themselves can be done through skills or just prompts.
A nice features of Claude’s is the remote control through the mobile app. That isn’t just a skill.
The consequence of Windows having the blame is that one should not buy it.
The downside is "lets try giving everyone basic income of $100k/week". But apart from that great!
I never heard anyone contemplating "why people are so sure we'd get freely available GPT-4 level+ models".
Mythos is just another advancement in AI model capability. If we start pretending like it isn't, we are setting ourselves up for paying far more than is reasonable and accepting the frame that we should not in fact have free access to open models with the same capabilities.
This is just marketing.
I can't wait for some similar big steps in open weight models.
There is some difference in how OpenAI and Anthropic handle 'max_tokens'. The OpenAI way raises errors in Anthropic for example.
I will look at confirming this a bit more in depth once we've taken our RubyLLM based adapter in production and see if I can make some contribution.
Thanks for all the work! It is incredibly impressive and I never meant any disrespect.
I checked and it turns out I remembered correctly that setting effort and some of its settings are not portable between providers.
There are some different settings that each provider uses and in order for it to be portable, you have to force some defaults on provider A when using a setting that is almost only supported in provider B.
In our implementation we decided to drop a certain setting when using OpenAI in one case and we decided we can just force some other setting when using Anthropic. But this 'solution', might not be what others expect.
When you build an open-source library you can go this opinionated route and force these settings, or you might go the config route and force people to explicitly handle per-provider differences. I will have a look at what I am able to do in terms of a contribution and then in the PR Carmine can decide what they like.
I will have a deep dive into which things I felt we needed to adapt per provider.
I didn't mean to imply that you have to solve all of our wants of course.
One thing we did do was monkey-patch the spot where tool_calls are performed by RubyLLM. We had our own mechanism for that and were able to skip RubyLLM's and still extract the tool calls and run them through our own tool harness. That all worked beautifully. I don't know if that type of stuff is something you want PRs on or that you want to keep steering towards the route that does everything within RubyLLM classes. Happy to contribute some of that.
> For context: GLM 5.1 ran the same task and reached 7.3x. Kimi K2.6 reached 5x. DeepSeek V4 Pro reached 3.3x. The models that stopped early did so because they issued no tool calls for five consecutive rounds, they concluded they couldn’t make further progress and stopped. Qwen3.7-Max didn’t stop.
By this reasoning I could release a model that lacks all the basic optimisations. Have it optimise itself for hours to reach 20x the throughput and then claim that the model is superior to the others?
I am not saying that is what happened here, but the reporting is abysmal.