HNHacker News
TopNewBestAskShowJobs

randomblock1

136 karma · joined March 2, 2018

submissionscomments
randomblock1··on LLM Ass Bench
Now put them on a bike.
randomblock1··on Claude Opus 5.5
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
randomblock1··on I don't like passkeys
It's amazon, I have 1password and it always asks me to create another security key
randomblock1··on Bend – A language that blocks AI mistakes via proof, on CPU and GPU
Multiple times, even. Still no real reason why. https://github.com/bendlang/bend/activity?ref=main

One time they force pushed and erased everything except a 2-line README... on purpose.

Pre-obliteration version: https://github.com/bendlang/bend/tree/814453670d0e0d6777c131...

randomblock1··on TSMC revealing details about next gen A14 node
It's not quite SRAM, it needs to be refreshed like DRAM. That's the main downside compared to SRAM but it's still very interesting.
randomblock1··on Hister: A private search engine for the pages you visit and the files you keep
Nothing about this depends on the provider, you could spin up a local Qwen and give it a search tool like SearXNG or something. At this point local models are more than good enough for simple tasks like that. Using ChatGPT is just (usually) faster and simpler
randomblock1··on Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
No it still requires Pro in the CLI. Only free model is SWE-1.6 slow
randomblock1··on Why is the x86 undefined instruction called ud2? Why 2?
Schrodinger's instruction: simultaneously defined and undefined
randomblock1··on Why is the x86 undefined instruction called ud2? Why 2?
0 undoubtedly comes first, but we all know 1 == true...
randomblock1··on So you want to use OpenRouter?
You know how there's a router mode to use the cheapest provider? That only takes into account uncached rates, last I checked. Make another one that takes into account effective rates (the ones that include cache).
randomblock1··on Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Oh I tried in cloud. I'll give it another shot
randomblock1··on More questions about whether researchers can trust OpenAI with unpublished math
So then they DIDN'T "learn the secret to cracking the problem". They simply knew that part of the problem was solved. Knowing a problem can be solved and knowing the solution are not the same thing.
randomblock1··on Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage. It does say 75% off though. Seems like for Pro subscribers SWE-1.7 is free, maybe SWE-2 is free for them?
randomblock1··on Searching for the best silicone USB cable
It's the website's "grid-bg" effect. It's glitching out and also normally barely visible at all.
randomblock1··on Protecting Engineers' Skills in the AI Era
The writer is the CEO of... whatever this is: https://auraspark.com/

Completely slopped up website, zero human touch. I would even go so far as to call them a hyprocrite, for getting AI to do all that for them instead of, you guessed it, training a junior engineer to do it.

randomblock1··on Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
I think it's still a useful data point. For example, omp, which is pi with some default extensions, scores worse. I do agree that adding more configurations of Pi would help though.
randomblock1··on Can I opt out of my input or output data being used for training?
The retain it for "the duration necessary to achieve the intended purposes", which could mean forever.
randomblock1··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
It's not that far off anymore. On my 7900 XTX 24GB, I can run Qwen3.8 27B with 131K context at Q4_K_M (55 tok/s with MTP). Excluding hardware cost, it's about $0.02 tok/M in and $0.40 tok/M out (cached in $0.0001). On OpenRouter, that would cost more than 10x what it actually costs me.

Of course, 131k context at 4-bit quant is a trade off, but even then, it's VERY capable. It doesn't feel that far behind something like GPT 5.6 Luna.

randomblock1··on Previewing the Model Hardware Standard
To me, this seems like OPC UA / SiLA but instead of being software <-> machine semantics/control it's AI <-> machine semantics/control. Or in simpler terms, an AI-facing hardware abstraction / device-description standard.

This is how I understand it:

Agent <-> MCP/CLI/code <-> MHS <-> vendor API/SCPI/OPC UA/ROS/etc <-> CAN/Modbus/USB/etc <-> hardware

MHS can describe capabilites, metadata, safety limits, as well as provide read/write control and discovery. Things like "can measure temperature", "arm weighs X kg", "never exceed X RPM", that agents can easily understand. (as opposed to that being buried in a datasheet somewhere, or having to be included in the prompt)

Also see: https://en.wikipedia.org/wiki/OPC_Unified_Architecture, https://en.wikipedia.org/wiki/Standardization_in_Lab_Automat..., https://xkcd.com/927/

randomblock1··on YouTube Format IDs
Anything with motion. Sports, games, vlogs, and so on experience immense improvements. And it's not like it takes 2x the bandwidth, because inter-frame compression can be smarter about it. It's like 30-50% more bandwidth.
randomblock1··on Behaviorally fingerprinting Ox Alpha's provenance
I find the tokenizers most compelling. That's what the model is trained on, it's an immutable fact of the model and its architecture. You know for a fact that the model is at least related to other models that way. And if a tokenizer is unique / specific to one lab, like GLM's is, it's basically as good as it gets.

Comparatively, you can't be 100% sure that Z.ai isn't able to host some other lab's model (although in this case, the hosting errors still support the GLM theory).

randomblock1··on C2PA Cameras Do Not Survive Contact with Reality
Even at the hardware level, if it was a separate chip that the camera data passed through or something, that's not really good enough either, people have broken TPMs before. It'd have to be baked into the camera sensor. Even then, you could attack it from the next level up, with some fancy optics and a display, or something like that.

I don't think completely solving this sort of problem is even possible.

randomblock1··on ElevenLabs, TwelveLabs, ThirteenLabs
They even go into the negative numbers.

https://minusonelabs.com/

https://minus3labs.com/

https://minus9labs.com/ (borked but existed at one point)

It's getting ridiculous, frankly

randomblock1··on Ornith-1.5: From Self-Scaffolding to Self-Improvement
Yes: https://huggingface.co/collections/ornith-ai/ornith-15
randomblock1··on Fixing a bricked Framework laptop
It's not just the chassis, it's also the screen, storage, RAM, ports, speakers, battery, and so on. And there is lots of demand for old mainboards, they sell pretty quick on Ebay. More sustainable to only buy and sell the part you need than the entire laptop, especially factoring in shipping costs. Plus the mainboards can be used like a mini PC.
randomblock1··on Show HN: Embed a real Linux terminal on your website
How exactly is this different from something like v86? It's definitely easier to embed but also not as customizable. Like if I don't need bioinformatics stuff, can I just exclude that?
randomblock1··on How Compaction Works in Pi
TLDR: It keeps ~20k tokens of recent conversations, then hands the rest of the conversation to another model with a special system & user prompt. This then fills out a template with relevant information.

See: https://github.com/earendil-works/pi/blob/main/packages/codi...

randomblock1··on Where did the old web go? We followed 657,607 links to find out
The ENTIRE thing is AI generated. I'm not talking about the article. I'm talking about the entire website, the entire "product". https://0.mk/blog/zero-humans

Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.

randomblock1··on Gemini 3.7 Flash
I think it's just meant to make it more competitive, Gemini has kinda been behind in everything except maybe multimodal. It's only 3 weeks after Flash 3.6, so if they really wanted to, they could probably do a 3.8 Flash before then.
randomblock1··on Oracle cut its Always Free ARM limits to 2 OCPU / 12GB, enforced Aug 18
> afaik the PAYG subscribers are not subject to this lower limit ... In other words, if you want the old provisioning limits for free you just have to put your credit card info in.

Always Free can mean different things: to a tenancy like you're describing, or to resources, where it describes the amount of resources you can use for free before you start getting charged money.

https://docs.oracle.com/en-us/iaas/Content/FreeTier/freetier...

> All Oracle Cloud Infrastructure accounts (whether free or paid) have a set of resources that are free of charge ... These resources display the Always Free-eligible label

> All tenancies get a set of Always Free resources in the Compute service ... All tenancies get the first 1,500 OCPU hours and 9,000 GB hours per month for free for VM instances using the VM.Standard.A1.Flex shape.

That's for all tenancies, and calculates to 2 OCPUs and 12GB per month. So sure, the provisioning _limits_ go up for PAYG tenancies, but using them certainly ain't free! If you have a 4/24 VM you have to downsize it if you don't want to be charged for it.

Page 1 of 2Next →