HNHacker News
TopNewBestAskShowJobs

IanCal

11,914 karma · joined October 5, 2012

Short-term AI/GPT/LLM consulting services to help you strategize, discuss, and navigate the rapidly evolving world of artificial intelligence. Discover how these powerful tools can transform your business.

There's no need to hire a full-time consultant when you only need guidance for a few hours or days. I can quickly help you develop a solid plan that your existing engineers can build upon.

Offering simple and flexible contracts, I am available for clients in the EU, UK, US, and AUS/NZ time zones (with advance notice for synchronous meetings).

Pricing:

£1200 for a half-day £2000 for a full day

For inquiries, please contact Ian at ian@redbirddata.co.uk

submissionscomments
IanCal··on LG denies TV spying claims, says tracking and snooping concerns 'not true'
It’s a lot less data to process as well.
IanCal··on LLMs are real, AI is fake
IMO this is a really terrible explanation of the attack. This is much more interesting: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
IanCal··on LLMs are real, AI is fake
> The chatbot consults its training data

Err, no? That's not at all how llms work.

> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.

They worked out how to fake the scoring, then hacked into a different system (which required finding a bunch of other exploits) in order to find the actual answers, and were trying to modify their own logs to hide what had happened.

This isn't a case of them saying "hack into X... OH NO IT HACKED INTO X".

> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.

It wasn't a rival server though, was it?

> That happens in Capture the Flag games at hacker cons: teams break into each other's systems to get a peek at the parts of the problem they've solved. That's allowed! It's a hacking competition.

They also tried to modify the code in the benchmark. Are you allowed to try and break into things to change the problem? edit - the agents transcripts show some of them explicitly saying that attacking HF is not allowed as part of the challenge

This all seems to dramatically underplay how interesting the actual attack was and what built up to it.

https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

IanCal··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Because we care about the distribution of the outputs and how that impacts a specific use case.
IanCal··on There's a new "Google Jail" for independent wikis
The company has an almost immeasurably small impact on their profits, and will never be measured anyway.

The cashier is a real human doing a job to put food on the table and has done nothing to deserve it, and may well have been far more hurt by the head office than you.

IanCal··on Tao: Open math problems being non-renewably mined by AI
His point is twofold: that the process of solving the problems leads to more than just solving the problem in front of you but other interesting things (he has an example of going on a hike to a waterfall and all the other things you might spot over in the distance or nearby on the way, which you’d miss if you were able to jump straight there), and also lots of the simpler open problems are ones early researchers learn on (this is akin to the “if we automate junior engineers how does anyone learn to be a senior?”).
IanCal··on There's a new "Google Jail" for independent wikis
If there is a person on the other end it’s not someone who has any link to the company, they’ll be an outsourced service, so all you’d be doing is sending graphic porn to a poorly paid worker.

It’s like screaming at a cashier because the supermarket head office made annoying changes to the company.

IanCal··on Programming is Art
Yes that was the point of the comment.
IanCal··on Programming is Art
A banana and duct tape can be just a snack and a tool for fixing a leak.
IanCal··on Programming is Art
> But AI can also be much more than a tool, and why can't it have a soul? What's a soul anyway, except something that humans are making up to convince themselves that they are more than flesh and bones?

I’m finding this fascinating personally. The lines people are drawing to try and keep themselves inside some definition they can mark as special.

IanCal··on Your intellectual fly is open when you use an LLM to author a post (2025)
Not necessarily. Poorly worded, ambiguous, confusingly ordered writing can be massively improved without changing the core content. Better setups and explanations can be longer without changing the message or meaning.

Look at it the other way, could you take a good longer message you’ve written and make it shorter and less readable for your audience while still making sense to you and containing the key points?

IanCal··on Your intellectual fly is open when you use an LLM to author a post (2025)
> What use is that? I'm not being facetious, I'd really rather like to know.

People are terrible at writing. Near universally bad. Even good writers have drafts and editors.

There is a constant refrain here that somehow short messages are more valuable than longer ones. But that assumes it's understandable. Lots of short content is, frankly, awful because the writer cannot put themselves in the position of the reader and explain all the things around the point they're making that the reader really should be told.

You can view writing as translation. From your language to a language your audience speaks. At that level is it so odd if the word count differs from one side to the other?

IanCal··on Your intellectual fly is open when you use an LLM to author a post (2025)
> And good luck getting an LLM to do that.

Have you never asked a decent model to explain something to you? You should try it.

IanCal··on Your intellectual fly is open when you use an LLM to author a post (2025)
> Writing helps us think about the world, it’s a pivotal intellectual technology.

Much like money decoupled selling and buying to move away from bartering, writing decoupled saying and hearing so they didn't have to happen at the same time. The incredible step that happened was not that people had to think a whole lot, it was that thinking that was already happening had to happen once.

> All writing is thinking when done by a human, you’re literally distilling your thoughts into words. You can’t write without thought.

Of course you can. You can write down exactly what you hear, for dictation.

You can write down a stream of consciousness and put barely any thought into it at all.

I can't help but feel most here are massively over estimating human writing. Human writing is, almost universally, terrible. We have entire jobs that are hard to fill just to make things sort of ok. Good writing is a small subset of human output.

IanCal··on Grep beats LSP? Why coding agents ignore your fancier tools
Have you never read human writing before? Humans write all kinds of confusingly worded things all the time - it’s why we have editors even for writers who are the cream of the crop.

But also that sentence is entirely fine as it is to me, it’s pretty simple and clear isn’t it?

IanCal··on GPT-6 Astra
If I’m understanding other comments the harness is just how ChatGPT and codex work already and it’s to do with how the context gets compacted - the arc-agi harness some are claiming just throws out reasoning blocks? Which feels like a huge handicap.
IanCal··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
Which isn’t correct, either for real world cases or worst case linear inserts.
IanCal··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
Both.
IanCal··on Making the Internet Boring
How many of those things do you really need? As in how bad would it be if you fully deleted those accounts completely?
IanCal··on True Rate of Unemployment
I’d be surprised if that’s how they measure it, it’s a common misunderstanding in the UK that this is what the stat means.
IanCal··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
Version controlled data is a legitimate solution to actual problems.
IanCal··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
Ah yes so the 4x slower isn’t really a real world case and you’d have to explicitly measure yours - it’s not a fixed slowdown but depends on table size - and it’s for doing many many single inserts in a row, which you’d be batching either way.
IanCal··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
With merging updates between branches?
IanCal··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
> If you want to take a version from a database, it's usually for backup. If you want different versions in your data, you put that IN the database. Manually merging different database versions does not sound too surprising.

This is not primarily for backups and has been part of ml and ds work for a while. And more, but that’s where I hit it.

> This problem does not exist for databases, because you design the structure yourself.

It is entirely possible to solve a lot of this in a more generic way, so that you don’t have to solve it each time. That’s what dolt is about.

Here’s a blog post from a few years ago with some use cases from actual users

https://www.dolthub.com/blog/2024-10-15-dolt-use-cases/

IanCal··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
A version controlled db? Have a look at dolts main page then, it’s been around for some time - this is just adding an embedded version.
IanCal··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
> however it's 1.2x to 4x slower than sqlite

No? Where are you getting those figures? It feels like you may have pulled that from the 125 and 400 us numbers?

> The simple way to fork a sqlite db is just to copy it, but if that is too slow you could use it with a copy on write filesystem, either something native, a FUSE filesystem, or build a sqlite VFS.

That wouldn’t really solve what dolt does though.

> and it's level of testing and validation will be nothing like sqlite.

They’re using the other dolt tests as well as the SQLite tests. Nothing will hit the level of real world testing SQLite does.

IanCal··on 'Stunning' percolation proof solves decades-old puzzle about phase transitions
> You need humans to decide which problems are worth spending time on, and to evaluate which results are important

Why?

IanCal··on It takes 5 cloud services to hear my doorbell
What about building something from scratch? Or is it the repurposing you like?
IanCal··on It takes 5 cloud services to hear my doorbell
My friend said this about diy

“If I say I’ll do something I’ll do it. There’s no need to remind me every six months”

IanCal··on You Know GDPR Is Good Based on Who Hates It
Often complaints about it boil down to either not understanding what’s in it, or annoyances that would be solved if sites stopped doing all this shady stuff. “We value your data, our 1644 partners…” yeah you definitely have a value you assign to my data.
← PreviousPage 3 of 34Next →