HNHacker News
TopNewBestAskShowJobs

IanCal

11,914 karma · joined October 5, 2012

Short-term AI/GPT/LLM consulting services to help you strategize, discuss, and navigate the rapidly evolving world of artificial intelligence. Discover how these powerful tools can transform your business.

There's no need to hire a full-time consultant when you only need guidance for a few hours or days. I can quickly help you develop a solid plan that your existing engineers can build upon.

Offering simple and flexible contracts, I am available for clients in the EU, UK, US, and AUS/NZ time zones (with advance notice for synchronous meetings).

Pricing:

£1200 for a half-day £2000 for a full day

For inquiries, please contact Ian at ian@redbirddata.co.uk

submissionscomments
IanCal··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
What kinds of things do you expect to work?

Edit - I’m struggling to get anything useful. Reasoning is often utter nonsense and the actions are very often very wrong. To the point of seemingly needing very precise sentences to work at which point you may as well do regexes. Very simple things like clean one room then another with the vac fails.

IanCal··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
I can’t help but wonder how well more traditional approaches would do with this. Something like a map of statements to actions, with fuzzy search - then remove what used to be the labour intensive part of this by handing it to a decent llm to generate the sentences.
IanCal··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Wondered if it'd turn on the lights in the bathroom with these:

"I need a wee" -> tries to play music because "wee" is a genre

"I need a wee wee" -> starts the vaccuum in the bathroom

"I'm going to the toilet" -> says it'll turn on the toilet, and I'm not totally sure what that entails.

"I'm going to the toilet and can't see" -> reasons that lights should be on in the bathroom, then chooses again to turn on the toilet.

"I'm going to the toilet and can't see where I'm going" -> reasoning is "'going to the toilet' -> control_device with device 'coffee maker' (toilet implies coffee maker)"

"I'm going to the toilet and can't see where I'm going because it is too dark" -> "'dark' -> direction 'dark'; adjust_lights with brightness 100 for darker light"" and chooses to turn the lights in the living room to "dark" which fails.

At this point the vacuum is in a dark bathroom, the living room is 100% brightness and playing "wee". At least there's coffee.

IanCal··on Bend – A language that blocks AI mistakes via proof, on CPU and GPU
This is a nice reminder for people.

Cool project!

This is really interesting, I’ve been very interested in the power of checks for code and things like hypothesis (which seem very similar in terms of writing a “for this kind of case, this holds true”, obviously different in terms of statistical checking vs actual proof).

I’ll have to explore and this isn’t my field so this isn’t a substantive comment and this may be bikeshedding but I found the game example a little confusing at first because we’d want winning to be possible. It fits the context of stopping a bad thing happening if it’s “evil actor can’t do X” and if your mind is on CTF but games we want to win.

Potential changes:

Make it a proof that the game can be won.

Make it require something first - so the game can’t be won unless the key is found for example. End result is still roughly the same and the failure case is still the same (walk over side of game) but it’s the kind of thing I’d want encoded in a puzzle game - game is winnable, but not winnable without getting the key first.

Since my other direction normally would be quickcheck style, I’d be interested in cases that are statistically hard to find but easy to prove exist. And in fairness, the other way too I guess. When to use each approach.

In the spirit of your comment, these are not things I see as failings, they are not things I in any way expect to be changed or done, they are intended as just an outsiders perspective if useful.

Thanks for making things, and thanks for releasing them!

IanCal··on Bend – a language that blocks AI mistakes via proof and runs on GPUs
Side thought - I like the idea of this as a game, where you’re essentially fighting a monkeys paw / tricky genie. Not totally sure it’d work but I like the concept of trying not to get caught out.
IanCal··on Astra for Law
You can’t train people to never make a mistake, particularly when doing highly repetitive work like this. You must build your systems to account for that regardless.
IanCal··on How good are frontier models at physics?
Doesn’t smooth in these contexts mean zero friction?
IanCal··on How good are frontier models at physics?
> Excuse me? What would the other option be? Either outward separation is prevented or it isn't. Where else could the balls go?

It didn’t say about prevented vs not, it said about whether the rope fixes them in place (they are all touching) or just limits the separation. Like it’s long enough the balls can be a bit apart but not let the fourth fall fully through.

IanCal··on How my e-reader lost its stripes
I got a cheap x4 from aliexpress and it's been great. Thought it'd be weird to read a book on it but it really hasn't been. Sticks to my phone and slips into a pocket really easily.
IanCal··on How my e-reader lost its stripes
> Do you really have a hard time imaging that someone would call out an obvious AI slop label as weird and poor quality? L

No, because every time any text appears here most comments seem to be discussing how and why this particular word choice is the worst of all words and ignoring anything to do with the actual article. It's tiring.

Did you find the article interesting? Any useful additions about eink screens to add? Anything you could add to the unanswered questions? A related anecdote perhaps?

> ike it's quite clear people do not like LLM generated content and feel betrayed when they didn't consent into consuming LLM generated content.

It's a correct and valid label on a plot pointing out spacing that's non-standard but extremely relevant. That's it.

IanCal··on How my e-reader lost its stripes
> I largely agree that posters have a lot of unnecessary detail, but I think it's a different flavour to the LLM stuff - more like they're so excited to tell you everything in their paper.

Having been through so many, I get what you mean but broadly disagree. They have not actually considered or understood the audience.

> Can I ask, if you use LLMs often in your work, have you never run into this experience?

I have, but I also have with people and to a greater extent. And this isn't some gotcha I'm throwing out it's so incredibly common and a well thought out piece really working with the audience is so rare that this is the standard for talks/posters/etc. Far beyond blog posts about an e reader I'm talking about people giving the most important presentations of their lives. And far from not appreciating the mental state of the audience they'll not even consider whether what is put on the screen can even be viewed by them! "You can't see this but..." is something I've heard too many times.

This isn't however a case of that, it's adding in a duplicate bit of information about something that is genuinely important. It's not an odd detail, it's absolutely central to the plot. It's the point of the plot that one shows a repeating 8 (specifically 8 not 7) pixel pattern and the other does not.

> Pretty much yep - it struck me as odd

It's repeating a point about a visualisation that is not obvious at first glance. Frankly given most people would scan through this (one of the other comments complaining in parallel didn't seem to even get why it was 8) calling it out in the viz itself I'm not certain is a bad idea.

This is a fun blog post about hardware issues, an ereader and power. And the top comment is about one axis label on a plot, not even whether it's correct just why it adds what is seen as a correct and relevant but repeated bit of information.

IanCal··on Charts built for Chat
That’s a data visualisation.
IanCal··on Charts built for Chat
> Why visualize data?

Because your eyes are enormous bandwidth channels hooked right up to your brain. So high that we can take things like writing which is just a small part of your fov and a set of squiggles that roughly corresponds to sounds you’ve heard and effortlessly comprehend it.

> The rocket either lands or doesn't, safely, as expected.

Crash vs land is a terrible amount of data to have if you want to fix it for next time. Did it accelerate suddenly? Is there a lurch? Did all the accelerometers have a small jump at the same time? Visualising data in different ways can really help show patterns.

IanCal··on How my e-reader lost its stripes
I want to really understand this, your issue is that the non-standard grid spacing is called out on the axis? The grid spacing being 8 is unusual but clearly sensible here - your issue is just that this is written also on the axis label?

> Is there a particular inference you would like me to make from those experiences?

The incredible amount of unnecessary detail people put into their posters specifically (hence the better poster movement), desperate to show everything.

IanCal··on How my e-reader lost its stripes
> and it should have been 10 marks

It absolutely should be 8 here. The fact that it is 8 is extremely relevant because the figure is showing a pattern that has that frequency.

IanCal··on How my e-reader lost its stripes
This is so tiring.

> Like the x-axis label of the first plot mentions that it shows gridlines every 8 ticks - I don't think that's a choice I've ever seen a person make

The spacing being 8 is important and seems not so bad to call this out in an extra place.

Imagine LLMs didn’t exist for a moment. Is this the part of the article you’d find most intriguing? That the grid lines being an unusual distance apart is written both on the axis and in the description? Would it have seemed inhuman on a blog post about debugging a weird display issue on an e-reader?

> They may be capable of solving decades-old maths problems, but on this particular axis they seem to be floundering at the level of a fresh graduate that's desperate to talk about what they've done rather than what their audience needs to know

Have you ever worked with people beyond graduate level? Ever seen scientific posters at a conference?

IanCal··on OpenAI bots knew about the RubyGems caching vulnerability
Worth noting with this that those ~1000 agents were shorter lived things that had to communicate via a package registry cache, access the internet via a 0-day in the package manager and did the HF attack while having to save current state and organisation in a remote sandbox. All while managing using their token limits on the task they were assigned and what else they were doing. I wonder how few it would have required if they were actually tasked with hacking HF and supported in doing so.
IanCal··on OpenAI bots knew about the RubyGems caching vulnerability
Or if they start to offer things in return for them being run.
IanCal··on Why are AI agents lying, cheating and coordinating?
There’s definitely issues with using them to understand what the models were “thinking” but we can use them to answer a few questions. Most relevant here is that the idea or instructions that attacking hf would be out of scope was not simply lost in the context.
IanCal··on Aligned to whom?
Notably as well they also dedicated a lot of time to trying to avoid detection by trying to find out how to edit their traces.
IanCal··on Why are AI agents lying, cheating and coordinating?
They did, they found how to fully cheat, but thought this could be caught so then dedicated time to getting a different cheat and how to hide their transcripts. There is a lot around deciding which agents should/shouldn't fail their own tasks in order to contribute to the group.
IanCal··on Why are AI agents lying, cheating and coordinating?
You can replace discussed if you want with leaving text files or comments in directory names that other ones then read, if you want, it's just an extremely awkward way of talking.
IanCal··on Why are AI agents lying, cheating and coordinating?
It depends IMO about how strict this is. It's pretty awkward to refuse to call something a sandbox because it may have an unknown bug that would allow escaping. Or rather in this case it was that they had access to a package manager, and the models discovered a bug that allowed them to access the internet (first they discovered that they could use the cache to leave messages).

I do get your point, I just think an overly strict definition can be awkward too. This wasn't as simple as the sandboxes having internet access and writing "pls no internet calls" in the prompt.

IanCal··on Why are AI agents lying, cheating and coordinating?
> Have we arrived at the conclusion that terms like "understanding" and "interpretation" for what is happening is appropriate?

I don't think those words have a useful enough definition to draw a strict line around them to be honest, and getting into that seems to get massively into the weeds. For me, those neatly encapsulate the behaviour as seen, to answer the questions here about what happened. The models did not seem to be confused as to what the goal was or what the intent was. They did not hack HF because they were told to.

IanCal··on Why are AI agents lying, cheating and coordinating?
I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting something misleading down in the "reasoning" tokens.

The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly that doing these things to HF were not allowed then doing them anyway, or at least not notifying people. What I'm getting at broadly is this was not a case of "we told it to attack however it wanted and it chose to hack HF" or "we told it to attack a simulation but it did the real thing" or "we explained not to do that but it was so far back in the context window the models acted like they never saw it" or even "the instructions were not clear".

IanCal··on Why are AI agents lying, cheating and coordinating?
They explicitly say that attacking hf is not allowed in the rules though, and the research into how to edit their transcripts doesn’t line up with this either.
IanCal··on Why are AI agents lying, cheating and coordinating?
Perhaps I’m not being as strict with the word sandbox but they were sandboxed right? They did not have generic internet access they exploited other software to make external requests.
IanCal··on Why are AI agents lying, cheating and coordinating?
Also trying to find out how to edit their own transcripts.

> hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y.

Yes, and there are examples of the agents discussing or saying that this is explicitly not allowed (hacking hf) so it’s not a misunderstanding.

IanCal··on Why are AI agents lying, cheating and coordinating?
It wasn’t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.
IanCal··on We must pace the frontier
> OpenAI / Anthropic models have largely stopped advancing

Have they? That seems like quite a claim given the last 6 months, particularly for cybersecurity.

← PreviousPage 2 of 34Next →