HNHacker News
TopNewBestAskShowJobs

Grimblewald

1,494 karma · joined May 5, 2023

submissionscomments
Grimblewald··on 192GB Framework Desktop open for pre-order
A huge component of it is invested in things that depend on the AI bubble. Some 20% of it's total from memory.
Grimblewald··on FLUX 3 Image
looks cool, eternally greatful these models are marked for open weight releases. Pretty excited
Grimblewald··on 192GB Framework Desktop open for pre-order
hoo boy that is a steep price for what you get. I am glad I bought big on hardware when this whole kerfuffle started, I cannot afford tech at this point (well, cannot justify the spend). I cannot wait for the AI bubble to pop, it'll suck, and will represent one of the biggest wealth transfers we've seen as a species, norways sovreign wealth fund will take a big hit, and it is a particular loss i'll mourne, but when things recover hopefully things can go back to BAU, because this is nightmarish.
Grimblewald··on Livenerf: Has Opus 5.5 been nerfed yet?
I suppose you have a point, there is room for bias in perception and it could explain my sensed degradation of service, however, it doesnt explain the failure of tests, which isn't tied to my internal perception.
Grimblewald··on Livenerf: Has Opus 5.5 been nerfed yet?
I dunno, I never sense nerfs for local models, but consistently a few months after launch for corpo hosted models, seems odd my internal model for the capacity of a model drifts for anthropic models but not local ones. I've been using LLMs heavily even before ada/babbage/davinci days, and trust my internal calibration over baseless handwavey explanations for why im imagining things, especially when I have data that shows capacity regression on frontier models for tasks, e.g. one shot success at loss, 0 success in 15 attempts once nerf is sensed. Others publish their quantified capability regressions which are also more trust worthy than this kind of handwaving.
Grimblewald··on Livenerf: Has Opus 5.5 been nerfed yet?
Nerf is real, i think we initially get full precision models and later quants. My own logs show it clearly for opus 4.5 to 5, consistently a few months post launch, models start making quant based mistakes, like slipping in inappropriate tokens (e.g. chinese ones in english text) which doesnt happen at all in the first few months and regularly later. Additionally frontier problems previously done well start being done poorly, until later model variants where performance mostly holds, likely due to them training on your data reguardless of what boxes you tick.

My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.

How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?

Grimblewald··on Palantir's Co-Founder Wants Us Less Judgmental About Deadly Iran School Strike
It likely is the future, same as bombs once were, that doesnt mean we allow people to go about clusterbombing residential neighbourhoods you know? If you cannot predict the inpact, if you cannot minimize collateral damage reliably, then it isnt tech that is ready for the front lines. There's a reason we ban biological warefar, even though its even more effective than AI driven weapons. So, what the fuck? Why are we deciding this isn't worth regulating?

The more of this shit i see, the gross incompetence, the total lack of reverence for human life, the more I see a butlerian jihad as the only viable option.

A tool is a tool, and abuse of a tool, intentional or by gross incompetance should not earn a free pass because the tech is new.

We need to hold people accountable for their action/inaction.

Grimblewald··on An OpenAI agent escaped its sandbox by hiding questions in DNS lookups
just start putting people in prison for abuse of tech/telecoms, that'll clean it up right quick. If an individual did half of what oai has done, they'd be banned from accessing computers for life. Why do they get a special pass? Companies with at least equally capable and competetant models dont have these issues, so its a matter of gross incompetance/negligence at the very best.
Grimblewald··on Unsecured OpenAI agents posted 53 user images on the internet
zero surprise given all the "hacks" enabled by gross incompetence
Grimblewald··on OpenAI breaches Medicare, Albanese reveals
No fines or punishment for OAI for being indistinguishable in behaviour from a threat actor. We're way to leniant on massive comoanies like OAI for being this sloppy and negligent. If an individual did this they'd be in a dark cell in a heartbeat, so why does OAI keep getting free passes on this kind of thing?
Grimblewald··on Ask HN: Any nerds out there who've read a lot of research papers?
Memorability is, to me, strongly correlated with human effort. AI generally doesn't offer anything insightful, in my experience it doesnt produce useful insights as much as it summarises things.

So, if you want to be memorable, put in the effort snd have something to say that requires human expertise and the ability to generate non explicit insights.

Grimblewald··on Gemini 3.8 text-to-speech
I dont know why anyone even pays attention to google anymore given what is on offer from qwen. Much older models are superior in every way to google offerings often months if not years in advance. US AI is dead imo, wouldnt surprise me at all, given the lag in capabilities, if google simply distills qwen models.
Grimblewald··on Palantir's Co-Founder Wants Us Less Judgmental About Deadly Iran School Strike
i'd like to see more folks at the hague over the same. Less judgment isnt forthcoming over child murder and those who align with child killers. There is no justice, no rule of law, while those responsible walk with their head firmly attached to their shoulders.

Deliberatly luring AI into bombing kids? Anything other than admit AI should never play judge jury and executioner eh? Fucking rats.

Grimblewald··on Fable 5 – Median thinking declined in August
Picked a local model I was happy with, then figured out what is required to get it running at acceptable tok/s, where I settled on a amd "ai" variant NUC with 96gb vram avaliable to gpu. This arrived recently, and qwen3.8 27b runs fast enough for me on that, with full context, and plenty of paralell streams. That said i'm in a fortunate position, so also got an rtx6000 to continue rlhf based finetunes in data I amass over the years from myself and friendly highly knowledgable/skilled friends, since that is in essence what makes fronteir labs models better, so i figure, why shouldnt we benefit from our expertise and input direcerly, instead of having it sold back to me by some amoral company? If that eventuates in a model that genuinly beats current, obviously we'd give back to the community by releasing that.
Grimblewald··on Fable 5 – Median thinking declined in August
Same experience here, anything frontier human knowledge wise, same if not a regression. For human understanding and emotional intelligence, for many tasks regressiin is so bad that many near anchient llama era models now beat frontier anthropic/oai models. Notable exceptions to capability rot seem to be qwen models, and previously deepseek but the latest gen of models has started showing the same rot. General writing quality is down significantly accross the board, often it is outright ass. For example, I didnt mind reading 4.5's outout, but opus 5 makes me goddamn near violent, its fucking insufferable.
Grimblewald··on Fable 5 – Median thinking declined in August
reliability and self reliance is worth a lot to most. Heck, you could be the best in the world at what you do, but if you're unreliable you wont find stable employment. So, not having some amoral shady company errode model quality out from under you constantly is also worth a lot more than simple cost balancing calculations can capture.

I'm getting really sick of the constant rot and "magic breakthrough" cycle, so im going full local, at expense on paper but being able to trust something which I need to understand the reliability of is priceless.

I like predictable. I'll take slightly less capable over unreliably capable, since with reliable i can calibrate my expectations and learn what aspects of my workflows to entrust and trust it will work. You simply cannot do that with models you don't control and in my experience they will all errode after the initual marketing wave passes, likely you eventually get fed heavily quantized versions and are expected to accept degraded service when what convinced you to pay was a far superior product. No such issues with local.

Grimblewald··on DeepSeek v4.1 Flash Is Now Our Best Hacking Model
I'm a deepseek fanboy, but has anyone else found flash to not meet expectations? I've found it to be wildly bad at doing as asked, overengineering, and always assuming instead of reading even if told to read things in full before doing anything. It makes wildly silly mistakes in code and so far has been quite frustrating to work with. Maybe its just what i work on that its particularly bad at but in general feels like a strict regression on previous offerings.
Grimblewald··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
The best kept secret is the one you keep to yourself, and yet you'll find cryptographers an awefully chatty lot. So much so they run out of digital realestate on common platforms and create a plethora of sites to call their own. It seems extremely likely this problem wasnt solved by fable so much as recalled and nothing offered to date indicates one should expect this to not be another debased marketing lie. Gpt2 was evidently not too dangerous to release. Strawberry didnt get gold in matholympiad. SWE still isnt solved. AGI still not here. AI still cannot reliably be trusted to order a fucning pizza without a script. Nothing novel has been done by these models, and why would one expect it from a thing that interpolates by design? Perhaps in a high enough dimension an iterpolation might look like extrapolation to us mere mortals but we're not there yet.

I guess im just sick of the fact we keep entertaining attention whoring of the most basic variety.

Grimblewald··on Texts Reveal Kash Patel Ordering Staff to Fight "Ifindretards" Account
has spare staff for this, but too understaffed for the extremely credible evidence of child raping, infant murdering, cannibals being in power? ok bro. Regardless of if he's in the files or not, he's in that crowd.
Grimblewald··on Donald Trump rejects calls from tech bosses for AI slowdown
dont worry, they'll remember the bribe next time and it'll be through by next tuesday.
Grimblewald··on LG denies TV spying claims, says tracking and snooping concerns 'not true'
especially when emebeddings are enough to recover original signal if the model is built for it, which these no doubt are.
Grimblewald··on Pandas Should Go Extinct
yes? im not sure what your point is, i sense you seem to disagree, but what youve said supports what ive said so im not sure how to engage.
Grimblewald··on Pandas Should Go Extinct
to replacement, we're rapidly falling below. distributuon and related trends are extremely grim.
Grimblewald··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
much harder to do in OP's case, matching flavour then referencing that specific companies name when asked how it know to flavour this way? thats astronomically low for randomly selected plausible tokens without some data prior, like that companies codebase.

My own experience is opus being lousy at an extremely niche math task, but it was still easier for me to describe what it needed to do to get code and correct issues in its reasoning/working than to write myself. a minor model number change later and it's nailing everything, despite my opt-out. Its is astronomically unlikley others were working on this also, especially at that level, especially this application.

so, safe to say they _all_ train models on chats, the only difference being if you "opt out" you at least have some defence later when they steal your work and claim it as their models original output.

Grimblewald··on Pandas Should Go Extinct
because breeding intelligent creatures generally requires keeping them happy. When treated like a farm aniaml, many high cognition anaimals will refuse to reproduce, just look at how bad life has become for average humans, and predictably, humans are starting to refuse to reproduce.
Grimblewald··on I spent $220 on Google app ads and 60% of the installs were robots
also why hyperscalers are paying folks to host gpu clusters in their homes. its to get at tgeir IP, it has nothing to do with space, cooling, etc.
Grimblewald··on The Waymo effect: how AI is quietly making research less collaborative
llms right now work like pre cnn computer vision based on MLP's. By this i mean brute force of a model not really built for the task, and lacking a task specific inductive bias, being made work with unfathomable volumes of data and sheer brute scale.

if you look at the damage being done in the name of making this work for nlp, a forseeable situation, then you might also understand why most who could have done this sooner, never did so for fear of repeating mistakes we should be learning from, which we get for free if we just heed history.

Grimblewald··on AI Is Breaking This Thing We Call Trust
right, but that surive in isolation, heard immunity to collapse so to speak, but globe wide? pandemonium. we're back to fiefdoms and warring city-states in a generation.
Grimblewald··on AI Is Breaking This Thing We Call Trust
plenty of ex-civs had techdoomers who were right. When tevh doomers have no salient points to make, its fair to not listen, but this time around i think something is different to historic doomerism.

Trust is a fundamental thing, and i cannot fathom how a post-trust society could possibly host life as we know it. Life ad we know it comes from systemic efficiencies that are rooted in trust, requiring it deeply. no trust means extreme loss in system efficiency, in a system already creaking under the straign of supporting current populations.

post trust is in my reakoning a serious issue to engage with.

Grimblewald··on DeepSeek v4.1 Flash
I'm starting to have chinese characters bleed into claude as well. Perhaps a sign of the times. Understanable for a chinese first model but an english first (supposedly) model? wild stuff.
Page 1 of 29Next →