if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
11,280 karma · joined November 25, 2012
hn_serf@gmx.com
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
it's more impressive to me that the electric fans within the Stratus server still operate than it is that the chips still work, and if we're going with the plane metaphor it's the silicon that puts the thing into the air, not the simple fans.
So really it just boils down to "well, the planes work is harder." as to why it fails, which is of course abstract and unsatisfying; comparing the comparative lifetime workloads of a chip that has shifted trillions of bits versus some jet engines that have moved millions of kilograms around the world.
and then another layer comes and makes it even weirder : the vast majority of hardware around the world that has been EoLd and replaced has been in that situation due to software, not the hardware itself; a paradigm that really doesn't exist in the physical engineering world in the same was as it does CS.
w.r.t. this post : let's define 'reasoning' here, because there are definitions of 'reason' and 'reasoning' that would fit to a simple condition comparison let alone a massively complex llm.
as for the meta cognition bit : show me a human that can accurately affirm when they know something. These kind of things aren't binary, nor can they be.
the ICE extender was an option.
it was a tremendously technical and expensive feat.
his dads' books didn't give any strong warnings other than we may grow 'overdependent'.
the sentient evermind enslaving millions thing was his sons concept.
so, to answer your question more practically, during this outage a luna agent of mine fell back to a local bonsai 2 model hosted on my desktop, and the work still got done just fine, just took longer.
that's why I keep all logging thoroughly disabled across all systems.
Saves a lot of room.
I don't get these new logic statements everyone is coming up with against AI systems.
If the excuse is "I meant literature!" well then I argue that where that line is drawn in the sand is exactly the question these kind of black and white expressions strolling entirely past.
(and that's actually the important question; where does it become neglectful for a human to throw an AI into the drivers seat wrt human read media?)
here's the /s for those of you who are tonedeaf.
what's the meaningful distinction now? It isn't all SoC anymore and it has had a linux capable MMU since S3, even if it required a lot of work.
The only real difference at this point is pure numbers, and i'm not that keen on defining microcontrollers as "those things with less ram than an SBC.", and if it's a microcontroller because it requires bare-metal flashes; well then I point to the older S3 MMU boards running busybox.
The closest thing so far is a Bigscreen Beyond 2 w/ a laptop if it wasn't for the fact that you need a base station for the damn thing.
I know my request is a fantasy right now, but I put this our there in case someone that could make it less of a fantasy reads this.
buy professional stuff that's expensive, not consumer gadgetry expensive.
a samsung fridge is expensive because of the gimmicks and high tech integrations.
a sub zero fridge is expensive because of the build quality and manufacturer network , part re-use and larger consumer part availability and market.
fuckin laughable, literally invoked a laugh from me in real life.
I hope customers aren't so stupid that they think a chatty developer on twitter/hn/mastodon/screaming-in-the-wind/wherever (or any other public-facing-place) means shit about customer service, and that goes towards ANY company where the primary customer service is an LLM.
Anthropic is the only company where it took (!) 9 weeks (!) to convince to hand over a 4 dollar refund for book-keeping errors on their side that caused an inappropriately early account deactivation due to time zone issues on their end, while all the while telling me that they don't offer refunds. It took stacks of evidence and argument, and that was after spending two weeks in their system trying to convince every level that I was worth a human.
For me personally it'd require Dario to resort to armed mugging to see another buck out of my wallet. I'm not alone.
tl;dr : being able to convince the powers that be on highly active industry forums (hacker news, twitter, mastodon..?) to act right using the power of peer shaming doesn't good customer service make. That said -- I do appreciate the direct response/statement from mpoteat;
..I just don't appreciate the good actions of a decent individual being too broadly interpreted as the do-good customer-centric nature of Anthropic .. an element I do not believe exists there.
it's harder to tell at a glance rather than pulling out a magnifying glass and counting the fingers of everyone in each group; the time limit pressures the player into making underqualified guesses..
it's not pointless.
if you're not in the US. revolut in the US is a fintech group with a bunch of bank partners.
smaller models are attempting to distill the useful methodologies, not the license plate number of an obscure extras car on Magnum PI.
that said I wonder if there is a small 'trivia' model out there. Seems like the kinda thing Google would tackle.
if there are a dozen unit tests trying to determine if some regex can escape a sensitive area, then one can derive a generalized 'don't let the regex escape from here' type rule -- or at least you could theoretically. I'm sure in reality that'd be a big minefield much like harness self-skill-writing has been.
is there a name for this? I've seen it all over recently. Ente products are like that; it's self-hostable but you've gotta really want it.
One trademark of this : the first steps of install are easy, but then it progresses into this quagmire of no-docs/old-docs/presets that favor the company somehow.
one of these things, (maybe ente photos/auth?) were self-hostable but there was a secret button on the mobile app that had to be known about to input a server URI outside of the cloud host to connect the app to the server backend that's self-hosted. Why would you go to that extent if not to add friction?
following this advice would have made most of my professional life impossible, what a luxury it would have been to be able to.
furthermore i'd argue that the success of an saas, or really any startup, has more to do with things aside from the code and service itself.
Finance management, social connections and effects, word of mouth, networking -- probably all more important than whether or not the todo app uses a functional language and is formally proven.
qwen3.8 has been fantastic as far as token rates are concerned for me, but the overthinking thing with higher reasoning levels takes some coercion to get right.
fwiw pi and hermes both handle that model fairly well. omp required a lot of tuning. I didn't bother figuring out why, I presume it's because qwen3.8 expects reasoning declarations a bit differently. nothing a proxy can't fix.
so, in other words, that seems like the way to breed more realistic acting agents rather than getting rid of them.
coool interior and aero. all of the hyper efficient designs end up looking like the modern prius designs with the swan tail rear and hawk-beak nose.
oil spiked and then crashed during the gulf war & the iraq war in 03.
the gulf war probably because of western price control methods kicking in, the iraq war because of an increase in confidence that the region would produce oil after western 'securing' of oil fields allowing further volume investment from the UAE and friends.
I guess it really just depends on what the definition of 'regional disruption' is.