I feel like there has been a few takes on this (atuin desktop comes to mind[0] but apparently it's not a going concern anymore?) why hasn't it caught on?
This is a different kind of censoring though - the examples given are hacking related but as the world is slowing moving to "research" == "I asked AI" its important that we have a means to reverse political censorship for example as well as yes, not having only the best models available to the privileged few - IE the whole mythos/fable split
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.
Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM
Heavily biased a kube admin but make kube clusters is easy now - its literally one command (aws eks create-cluster/gcp ... something something haven't done it in awhile after doing it multiple times a day in a past life) - I think its more familiarity with the tools that is the challenge - creating a vm is also easy if you know how, brew/apt/yum install is easy if you know how, setup.exe is easy etc)
“The ai told me it’s right” is actually probably my biggest problem with AI being freely available. Yes I know karens existed before AI but even then I feel like the excuses were more… thoughtful? IE my dad is a designer/it looks right in the photoshop preview whatever.
Maybe off topic - I really like the design of the site, it feels like it was designed with care which (no offence) is often not the case with AI focused websites. Was it hand designed, or just the right claude design prompt?
I've been watching a bunch of bycloud on YouTube recently, and although he's done a great job reviewing papers from the big AI labs, I feel like I'm missing something - how have all the labs seemingly made a model that's cheaper, faster AND has better performance? Historically `flash` variants (like codex spark as well) have been faster but perform worse
A harness is fundamentally required - although we don’t really see it using tools like Claude or codex ai models still only can take tokens in and produce tokens out. Tool calls are literally just the ai model printing “I want to call search_web with parameters abc” and the harness sees that in the text stream and runs it
Devin was the first ai developer harness I ever heard of — and in fact the first one I ever used as I had gotten an invite while it was in private beta. Now it seems like the product has faded (into irrelevance?) between other harnesses like cursor especially and all the major harnesses having a cloud mode. Is there some other area Devin has become really popular / has stayed competitive?
I think as a hacker/engineer type you’re more aware of what’s possible and because of that you expect more from things in your life. The flip side of that is that it makes features most people would never think of feel like table stakes and combined with the ability to build your own solution - IMO leads to over engineering.
"yes" but that's a philosophical question. I think they're getting better at "early fusion" IE, training the model that "apple" and these visual tokens are the the same concept, but LLMs are fundamentally a pattern matching machine so even with perfect fusion I personally wouldn't call it understanding.
> Cross out stuff you disagree with, draw a big circles around the stuff that resonates.
> Not only does it make consumption a lot more engaging, it makes a revisit of that book extremely rewarding
This is actually the biggest reason I own a Remarkable 2/Kindle Scribe, let's me "destroy" the book while also not having to worry about losing my notes.
If phones didn't have secure boot nowadays inevitably little Johnny would (unknowingly) install a rootkit on his mom's phone for the promise of free vbucks or some see through walls app.
Was still the take away. The idea is the lost a single instance should cause downtown so it’s not multiplied. IE three instances of the database case tolerate a loss of a single instance for maintenance since the other two will take over the load.
There is of course an argument against the cost of this and sure I’d even accept the complexity argument since you probably need to add more tooling to manage the hand over but again the point of this complexity is specifically to avoid having a single point of failure / to naturally handle failure such that the possibility of downtime isn’t multiplied
As always it depends but at least for me in my anecdotal experience as a Distributed systems engineer (I know, I'm clearly very biased here) - the problem with a single system is that mundane things cause unavailability like - software updates, hardware upgrades/failures, having a single fault domain (IE, someone writes a bad query that wedges the DB, now no one can use it) and since failure isn't built into the design, _when_ things fail the mean time to recovery tends to be rather high. For example, if the power supply explodes and takes out your server, how long would it take to procure a new server and restore from backup? On the flip side, if your system is designed around (for example) server less functions or spot instances- which become unavailable multiple times per day, you would have already engineered in fast recovery
A fair callout, though yes it was supposed to be a simplified analogy. I think you got it exactly right though that its temporary in the same that I'd argue the benefits of using AI are temporary
Presumably if you don't trust apple you wouldn't purchase their products and even if you were for example forced to use it via work or something you wouldn't use this feature ... so it doesn't really change the calculus as presented by this article - IF you ALREADY HAVE a MODERN Mac (and trust apple) this is your best option
It's not open weight, but the point is to be an on device (and thus local, privacy preserving) option. The article mentions that as the caveat
> What this means if you just want good transcription
> If you are on a current iPhone or Mac, the best on-device transcription engine for English is already in the operating system, and the private option is no longer the compromise option
Not to "glaze" the author as the kids say but this has to be one of the best written musings I've read on HN in a long time. I'm likely bias because it's written in "my style" but I feel like it's a rather fair and balanced approach to a nuanced and socially difficult topic.
I liken it to dieting. If your only goal is to be certain weight, then learning how to cook, learning how to portion, learning to make a meal plan, learning about macro nutrients is all "friction" now that we have GLPs. And maybe using a meal prep service reduces some "unnecessary friction" but you still would have to learn a bunch of useful skills along the way.
Concretely, If your only goal is to produce "software" then learning about design, planning, project management, testing etc is all unnecessary friction when you can just ask an LLM to "make it so"
Other than batch jobs, I can't think of a problem that can be solved these days that doesn't also require high availability - at the very least they require a warm standby.
I'm using "sneaky" here to refer to anything that's not very obviously stated but anyway
> That their actions make sense for their business isn't any reason for people to accept their deceitful, customer-hostile decisions.
While I agree it's a dangerous precedence to set, I think this is a "vote with your wallet" sort of situation. They shouldn't do it, but from their POV this is what they need to do to offer the product they do at the price they do. If the product wasn't compelling people wouldn't accept that they do this. However they've decided if you want their product you have to use their interface and whatever spyware it comes with, so it comes down to, is the value proposition good enough that people will put up with it? As of today, the answer is unfortunately yes