How? Explicit instructions, memories and even skills have not been able to keep Claude from saying "genuinely" every two sentences and keep it from explaining heavily what something _isn't_.
Yeah lots of weird emphasis on things a human wouldn't care about. And emphasis on what it isn't, rather than what it is. It's not Y, it's X. And there are two files!!!
I have, but honestly that was exactly what I am not looking for. Sometimes it picked a model for a Dutch email that totally does not support Dutch. Other times it would work fine. It's just a layer of indeterminism I wasn't looking for.
I'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for something I see the list on openrouter and I'm completely overwhelmed.
I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.
E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.
Another thing they apparently do is let business owners (e.g. hotels) remove negative reviews because they're "defamatory".
I had a very factual 3-star review on a hotel that was supposedly very quiet, but was super noisy and the AC was not working. And there was literally nobody at the reception at the time of check-out and nobody picked up the phone either. All in all, nothing serious, but worth a review. Purely factual stuff, nothing bad-mouthed, just my experience and three stars.
A few months later the review was removed because it was "defamatory". I was able to get it reinstated by sending them proof of payment to the hotel and proof that I was there. A few months later, it was removed again! That same dance was repeated six times..... What a weird system.
And on top of that, we've all seen how confidently LLM's get things wrong. I've seen them get better, but even the current SOTA models fumble. Why would you want to send a bot a message "sure, reschedule my flight" and have a bot potentially fumble that in all sorts of ways? I really really do not see the appeal. My calendar and my emails and my things to do are all things I want to be on top of. Not have it half-assed or potentially messed up by a chatbot made to harvest all your data.
Oh and don't forget: it keeps chaining a bazillion commands together so any whitelisted commands still need approval because they're nested in such a convoluted way.
LinkedIn too. The first time they show a bottom sheet dialog thing on the website, and if you dismiss it, it scrolls you all the way to the top. The second time it appears, it messes with the page and sometimes automatically closes the tab.
Yeah I've seen it a lot. It goes through the effort, unasked, of pulling screenshots off a connected device and then it's like... Oh shit yeah I can't see.
That's cool. I often write tiny blurbs of kotlin just to test out a simple algorithm. I often do this on kotlin playground because doing so inside a scratch file or test is somehow more cumbersome and slow. This ran and compiled something in 98ms on my smartphone, cool stuff.
I had the same sentiment. In my limited testing, it didn't perform any better than Opus at all. It wasn't a particularly challenging taskset either, mostly just "add this simple feature" with plenty of context and very clearly defined scope. It worked functionally but there were much better and simpler approaches available. For the cost, I don't see how Fable can ever be worth it.
That's what stood out to me as well, and it struck me as odd that nobody seemed to think it's odd? Almost half of the points to be made is related to contributions to open source projects. Guess my 10+ years of experience in a niche topic is worthless.
That's what I mean with the promise. Nobody knows what they're doing on their end. The data goes over the wire, and then you need to assume they are true to their word and the data is not intercepted by anyone.
It is. Not per sé because the code might be of poor quality, but because someone sent that source code to a public API under the promise that oh noooo we won't use your code for training. Probably.
Yeah the slowness is what always gets me. Like in essence a ticketing system isn't more than just a database of tickets and relations between tickets and states. And okay you can kinda make it explode by having tons of interconnected tickets and custom fields and plugins. But I will never understand how something that just works with simple textual data and attachments can be so unbearably slow.
I'm honestly not really that surprised by that. All (except clothes) are prime examples of products where you don't really care who sells it to you and how it looks in the packaging. People want a certain product and want it the cheapest. Why would you go to a real store to look at that product inside of packaging you can't open, with the added cost of the person behind the counter?
I mean a Lego set is a Lego set, whether you see the pictures on the box or online.
Yeah especially the AI stuff is so... not a driving force for anyone to buy this?
For me, unless you can run LLM's and whatnot locally (which is not the case on this undisclosed low-end hardware), "AI" just means doing some API call to a web service and have it serve me some freshly made up tokens. You can do that on a potato. The fact that they happily announce something that can be done on any other cheap-ass laptop as the main selling point, means this product is nothing special at all.
I'm super techy but I admit that I just use Signal to send me a "Note to self" whenever I need a file from my phone on my computer quickly. For images I just use immich, but texting myself is honestly the quickest way for files because the experience is indeed terrible.
Yup! All good ideas and solutions to hard problems, after becoming stuck, have come to me after a good night's sleep or after removing myself from the "thinking place" and taking a break. Yes I mean the toilet. Many fantastic ideas come to me on the toilet.
Has it, though? There's still features that bring large user value and require 10 lines of code, and features that bring a small user value and require AI to burn tokens on huge refactors and babying to make sure it doesn't break anything.