HNHacker News
TopNewBestAskShowJobs

nikcub

19,788 karma · joined October 4, 2009

Nik Cubrilovic - https://nikcub.me

squirrelscan - https://squirrelscan.com

open electricity - https://openelectricity.org.au

email nik at nikcub.me

@dir on twitter / x

submissionscomments
nikcub··on Claude Status – Elevated errors for multiple models
and the two 9's aren't in order
nikcub··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Defense in depth + defense in breadth - aka. all of the above

sandbox escapes have been the rage recently

nikcub··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Reading the patch[0] for libheif the bug which lead to the vuln was around bounds checking for image overlays. the container can have multiple images and you can compose them in the output.

heif also supports rotating, cropping, alpha channels, thumbnails and a ton of other features that a web forum where a user is uploading photos or screenshots doesn't need.

It's a much, much larger attack surface than plain old school JPEG.

I'd suggest rather than wait for the next bug to appear in this or another image lib to keeping things simple - stick to plain JPEG and handle image conversion in the client (wasm in the browser) if you really need to support users uploading iphone images.

Media decoding is so hard - there have been tons of bugs in ffmpeg and imagemagick and the core libs. You really need to think about how much of it you expose via a web server

[0] https://github.com/strukturag/libheif/commit/85e21ad44eba931...

nikcub··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
a) they were part of the offsec program

b) they proxied the target through a CTF host to fool the model and guardrails

> We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

the proxy is smart - there are other methods to bypass the guardrails to have it attack remote hosts.

you just have to prove to the model that you control the host or that its a valid target - and there are plenty of ways to fake that.

nikcub··on Xiaomi Mimo 2.6 live post-training dashboard
this is remarkable transparency in an otherwise hyper competitive and secretive industry
nikcub··on Pangram – AI detector for text and images
they also publish great tech reports. their founder is so confident in their model that he's regularly on social media offering bounties for false positives

https://www.pangram.com/blog/pangram-4-technical

https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangra...

https://www.pangram.com/blog/introducing-pangram-image-detec...

nikcub··on Automattic's board forces CEO Matt Mullenweg into leave of absence
from my own observations:

a) you don't want the WPEngine case going to trial and that is coming to a decision point soon. chances are the new CEO and board settle.

b) wordpress is being absolutely obliterated by ai. work for commercial plugin authors and agencies has dried up while the platform and ecosystem see major security issues and worms exploit a large number of sites every week.

wordpress and automattic haven't responded to this well, at all. at times it is often straight denial.

founder CEO distracted by expensive lawsuit while facing existential threat to project that will require complete attention and a lot of skill to steer out of

nikcub··on Muse – Meta’s personal AI agent
> "just ChatGPT, what is Sol?."

i'd really like to see some analytics on what proportion of users have ever switched the model in chatgpt because I anecdotally believe it is < 5%

nikcub··on OpenRouter is joining Stripe
getting zero traction here is clearly the path to success :)
nikcub··on OpenRouter is joining Stripe
same with cursor - which was submitted multiple times:

https://news.ycombinator.com/item?id=34423387

are there more examples of this? a bit of an HN anti-portfolio. good reminder to others.

nikcub··on OpenRouter is joining Stripe
I love OpenRouter, long time user. Stripe will hopefully be a good custodian.

I just want to point out some features of OpenRouter that make it more than just a model selection and routing endpoint and that I find incredibly useful:

0/ Default routing is to the cheapest provider, but they're usually not the most performant. I'd guess 99% of OpenRouter integrations never tweak the default routing. Here you can setup cheapest with performance minimums:

https://openrouter.ai/docs/guides/routing/provider-selection...

You can also stack model selection in priority

1/ Using broadcast you can push all your analytics to clickhouse / s3 / snowflake and a bunch of other compatible destinations. Setup a clickhouse server ($5 VPS[0]) and send all your traces to it:

https://openrouter.ai/docs/guides/features/broadcast

customise your own observability in your dashboards from there. Superwin

2/ Model router is also a natural home for llm security - OpenRouter has the beginnings of prompt injection detection:

https://openrouter.ai/docs/guides/features/guardrails/prompt...

there is also PII detection. This will show up in observability as rejections/blocks etc.

There are so many model routing solutions (same with observability, security etc.) but they're all 80% solutions - OpenRouter really rounds out with well implemented features that you need when deploying models at any scale and I gladly pay the toll.

[0] not sure if these exist any more but clickhouse is resource efficient

nikcub··on Grok 4.6
yes but it was a stock deal - so they bought it using spacex bucks
nikcub··on Grok 4.6
It was said at the time that xAI acquiring Cursor was very smart because it would give them access to years of agent coding traces from millions of users.

$60B in SpaceX stock for Cursor was a bargain

Data + compute + being competent and smart enough to ship.

fwiw I don't think these are yet Fable level - the difference tends to get discovered in the long tail of tasks - but they're close enough, they're cheap, and the length of the frontier exclusive window is narrowing

nikcub··on DeepSeek V4 Pro 0813
that DeepSWE result is likely most indicative of how you'll find real world usage
nikcub··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
There has been an over-obsession with frontier models and benchmarks. Most of the work will be done by task specific models. You don't put Phds on the factory floor.
nikcub··on Mistral's Shieldstral: 3B open-weights model for multimodal moderation
public ones are content moderation as above and previously llama guard, et al

OCR is also another field - Mistral have a model, so do deepseek

The ones I have experience with where you fine-tune smaller / faster models for business tasks like content writing, support, etc. by their nature stay private

nikcub··on Mistral's Shieldstral: 3B open-weights model for multimodal moderation
> which seems to be mistrals whole thing

They got a lot of hate for not keeping up with frontier model releases, but have managed to carve out a nice business that isn't even really niche.

Before the datacenter deals their revenue was higher than xAI's

There is a whole world out there of purpose built and hosted task specific vertical llms - especially with an emphasis on cost.

Mistral, Microsoft model releases and Thinking Machines are all over this, and it's smart. Scoop up all the tasks that don't require large and expensive frontier general-purpose llms.

nikcub··on The relay market powering token resellers and fraud
sign up for the resellers, insert a random id canary into your requests, trace them back on your server to tie it to an account

find fingerprints / signatures of the accounts being used. ban all of them.

eventually build an ml based system that detects these at signup

reinforce with more data. loop forever, etc.

nikcub··on Be skeptical of OpenAI's rogue hacker agent story
> The most damning thing is, they could've just included in the prompt

You can't prompt your way to a compliant model. This is just a reformatting of the 'make no mistakes' meme.

nikcub··on Be skeptical of OpenAI's rogue hacker agent story
> AI managed to escape using standard and well documented script kiddie methods

> AI broke in using standard script kiddie methods.

I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where did you get this from?

nikcub··on Scanning for Pangram Errors
> I've seen it correctly detect AI writing even when people use those "humanizer" rewriting skills

Those skills are meant for humans, not for models. Pangram said in their recent model update that they now also detect the humanizers.

All it requires on their end is running each of their AI outputs through each humanizer and including it in their corpus as ai-humanized.

It's essentially impossible now to prompt your way to non AI detectable text. Even if you do today, it'll be picked up by the next model update.

This is why in academic environments its probably worth re-testing old exams or papers periodically - same way blood and urine samples from athletes are preserved to take advantage of better future testing.

nikcub··on Kimi K3: Open Frontier Intelligence
Those thin capitalized eyebrows are becoming like the emdashes of visual design
nikcub··on Kimi K3: Open Frontier Intelligence
> I pretty sure OpenAI and Anthropic are doing the same or worse.

No they're not. It would end both companies if they were ever found to be doing that.

Their terms are clear - if you use the coding plans they can[0] train in return. Enterprise and API, absolutely not.

The argument here is that with the Chinese labs you have zero legal recourse.

[0] opt-in, thanks

nikcub··on Inkling: Our Open-Weights Model
Real test here would be using tinker to fine tune a tinker model to generate pelicans
nikcub··on Inkling: Our Open-Weights Model
This is a winner IMO. Lots of cost pressure on token spend atm within enterprises and tasks that don't require Opus / Codex class models.

These companies have hopefully captured all of their traces and now have enough to fine-tune an open model and host themselves.

Inkling feels like the right base - not obsessed with benchmaxxing on coding but rather being adaptable to the task required

For tasks like GTM, support, content writing etc. seeing 80%+ savings

nikcub··on Australian energy retailers must provide three hours of free daytime electricity
It kinda is since wholesale energy prices are often negative in these markets during the day
nikcub··on Australian energy retailers must offer three hours of free daytime electricity
> WA could be part of the NEM with some HVDC across the Nullabor, not sure if it would be economically worthwhile though.

Part of the motive of moving the WEM to 5 minute intervals was to eventually leave this option open.

The largest renewable project in the world is being planned in this area[0] so it's feasible that it all may be connected one day

[0] https://wgeh.com.au/overview/

nikcub··on Australian energy retailers must provide three hours of free daytime electricity
> The embedded networks collude with the builders and offer them the installation

I got into a dispute with my embedded provider because of a bad meter and came to discover through friends and family in the construction industry as well as speaking to a former sales person in the industry that there is a lot of additional corruption in the process with straight up payments being made to win installs with developers.

When it came time to switch providers in our building, strata was promised electric vehicle chargers as part of signing a new deal with a new provider. They never delivered because they found an escape clause because of fire safety approval.

We're now locked in for years (again) and they've already increased rates once in the first year.

Nobody in the entire chain works in the interests of residents or owners. It's a completely broken system and a thorn in the side of otherwise advanced and progressive Australian energy policy. It needs to be abolished ASAP.

I still pay more for my single apartment living alone in electricity than what family and friends do in full large homes with air conditioning, 4-6 residents, heated pools, etc. It's astonishing.

nikcub··on Show HN: Smart model routing directly in Claude, Codex and Cursor
I'm glad there are more attempts at solving model routing, as costs (at API rates) has really become an issue. Some feedback:

1. Reiterate the cache issue from other comments already here. there is a lot of optimisation in harnesses around caching and a proxy model blows that up

2. Coding agents are model aware - they already route code discovery to mini / flash models, planning to heavy models, workflow design to ultra, implementation to mid / high etc. They know when they're exploring, planning, implementing, reviewing etc. and which model class to select and when it fails.

With a proxy you're breaking this control loop and feedback. It doesn't know, for ex. that it just attempted with deepseek v4 and it failed, lets try Opus?

3. How are you going to RL improvements and prevent the router becoming stale? You only have access to your own internal prompts and ~thousands of samples.

This is RL'd on one orgs codebase. There are going to be a lot of prompts you haven't seen before and have no insight to on how to route correctly, and you have no insight into users HF to improve your own model. Orgs aren't going to share their traces with you, so you need other sources to train on and improve

There are also new model releases every week that you need to keep up with - whats the story going to be here

4. Publish evals by running terminalbench / deepswe bench. Show us the performance / cost / time chart vs the other agent and model sets. If you can show gains there, you have a very simple value prop to sell where you can charge for a % of the saved costs

nikcub··on Om Malik has died
This is devastating. Om was the godfather of early tech blogging and lifted up so many people around him. He was kind, caring and compassionate.

When I first started blogging around 25 years ago, he would have been amongst the first 10 readers. He linked to me, emailed me privately with feedback, praised posts and would call bullshit when he saw it.

He was never competitive with other blogs or bloggers and was never tied up in drama. He was very often a mediator in behind the scenes conflicts and was obsessed with truth over getting the scoop.

He loved tech and startups and most of all loved seeing other succeed and didn't have a gram of resentment within himself.

Everybody from that post-dotcom crash era of tech owes Om a large debt of gratitude. He will be missed. RIP Om.

Page 1 of 34Next →