11,914 karma · joined October 5, 2012
There's no need to hire a full-time consultant when you only need guidance for a few hours or days. I can quickly help you develop a solid plan that your existing engineers can build upon.
Offering simple and flexible contracts, I am available for clients in the EU, UK, US, and AUS/NZ time zones (with advance notice for synchronous meetings).
Pricing:
£1200 for a half-day £2000 for a full day
For inquiries, please contact Ian at ian@redbirddata.co.uk
Err, no? That's not at all how llms work.
> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.
They worked out how to fake the scoring, then hacked into a different system (which required finding a bunch of other exploits) in order to find the actual answers, and were trying to modify their own logs to hide what had happened.
This isn't a case of them saying "hack into X... OH NO IT HACKED INTO X".
> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.
It wasn't a rival server though, was it?
> That happens in Capture the Flag games at hacker cons: teams break into each other's systems to get a peek at the parts of the problem they've solved. That's allowed! It's a hacking competition.
They also tried to modify the code in the benchmark. Are you allowed to try and break into things to change the problem? edit - the agents transcripts show some of them explicitly saying that attacking HF is not allowed as part of the challenge
This all seems to dramatically underplay how interesting the actual attack was and what built up to it.
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
The cashier is a real human doing a job to put food on the table and has done nothing to deserve it, and may well have been far more hurt by the head office than you.
It’s like screaming at a cashier because the supermarket head office made annoying changes to the company.
I’m finding this fascinating personally. The lines people are drawing to try and keep themselves inside some definition they can mark as special.
Look at it the other way, could you take a good longer message you’ve written and make it shorter and less readable for your audience while still making sense to you and containing the key points?
People are terrible at writing. Near universally bad. Even good writers have drafts and editors.
There is a constant refrain here that somehow short messages are more valuable than longer ones. But that assumes it's understandable. Lots of short content is, frankly, awful because the writer cannot put themselves in the position of the reader and explain all the things around the point they're making that the reader really should be told.
You can view writing as translation. From your language to a language your audience speaks. At that level is it so odd if the word count differs from one side to the other?
Have you never asked a decent model to explain something to you? You should try it.
Much like money decoupled selling and buying to move away from bartering, writing decoupled saying and hearing so they didn't have to happen at the same time. The incredible step that happened was not that people had to think a whole lot, it was that thinking that was already happening had to happen once.
> All writing is thinking when done by a human, you’re literally distilling your thoughts into words. You can’t write without thought.
Of course you can. You can write down exactly what you hear, for dictation.
You can write down a stream of consciousness and put barely any thought into it at all.
I can't help but feel most here are massively over estimating human writing. Human writing is, almost universally, terrible. We have entire jobs that are hard to fill just to make things sort of ok. Good writing is a small subset of human output.
But also that sentence is entirely fine as it is to me, it’s pretty simple and clear isn’t it?
This is not primarily for backups and has been part of ml and ds work for a while. And more, but that’s where I hit it.
> This problem does not exist for databases, because you design the structure yourself.
It is entirely possible to solve a lot of this in a more generic way, so that you don’t have to solve it each time. That’s what dolt is about.
Here’s a blog post from a few years ago with some use cases from actual users
No? Where are you getting those figures? It feels like you may have pulled that from the 125 and 400 us numbers?
> The simple way to fork a sqlite db is just to copy it, but if that is too slow you could use it with a copy on write filesystem, either something native, a FUSE filesystem, or build a sqlite VFS.
That wouldn’t really solve what dolt does though.
> and it's level of testing and validation will be nothing like sqlite.
They’re using the other dolt tests as well as the SQLite tests. Nothing will hit the level of real world testing SQLite does.
Why?
“If I say I’ll do something I’ll do it. There’s no need to remind me every six months”