4,153 karma · joined May 28, 2011
Email: <myusername>@gmail.com
Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn't be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it's going to be generating garbage training data.
Granted, detecting "off task" may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.
In my reading, people aren't really saying "the AI is at fault", they are saying "hey look here's proof that this is dangerous". Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.
The "moat" is data. And compute. But mostly data.
It is a classifier, that you don't have to train, that seems to work well, and they have made available by API.
One of the reasons for LLM success has not just been that they are "smart" but that they are smart while not requiring you to collect your own data and perform your own training.
There is also the marginal advantage of it not requiring a PhD to know how to use. But most companies can get a competent data scientist on board, so that is not a major problem, they could find someone to train their classifiers. But even a good data scientist can't make data appear out of nowhere. If you can get a pretty good result without needing data and training, for cheap, you're going to take it.
(Let's put aside the fact that almost every company using LLMs these days is irresponsibly NOT comparing the results with their own internal gold standard datasets for validation and calibration. They are just winging it. And Jev allows them to continue to do that.)
The idea is to drive actual testing from this, but in this era, I think it's interesting as a way to use AI to generate tests, and to cross-check those tests with the natural language descriptions, in a bit of a cycle that helps refine the highest level definition of the software.
Once that's nailed down, the implementation is just details.. normal engineering concerns like maintainability etc notwithstanding of course, but you can trust more and more the AI agents to get it right. The design specs being natural enough for humans to deal with but interpretable/specific enough to actually generate tests is pretty interesting for the bottleneck you are talking about, I think.
There were a couple of things in the article which i struggled with, this being one of them, and the other being:
> Quartz doesn’t rhyme with parts, but rather shorts.
This made me realize that in both these cases, I sometimes say it one way and sometimes the other. Probably depending on context. And probably because I'm Canadian.
How is this idea generally working out in comparison with Jev? I'm curious, from what I read so far it seems like Jev is still beating this kind of thing.
But it's curious because it's not entirely clear why, from an architecture point of view for all we know that's exactly what they're doing. So it must come down to the quality of those logits, ie., model size and training details.
It seems to me that what most of these single-token-prediction projects are missing is that Jev seems to be claiming they predict well-calibrated probabilities. This is an incredibly valuable thing that LLMs simply can't deliver unless they are trained specially for it.
As for "on accident" I've always assumed this was just a mistake from childhood, kids equating it with "on purpose" and never being corrected. A bit of a bone apple tea.
However I get your point, I used to correct an Indian colleague when he would say expressions a bit wrong, until I gave up and realized it had just become part of his local dialect in the region he was from. (Eg. "I have to go pick my kid" instead of "pick up my kid" which made me laugh at first and then drove me a bit crazy because he would consistently keep saying it.) At some point these things cross a line and become somehow "correct" again, that's just language for you.
A native English speakers we are in a particular spot where we have to just accept that the majority of the world is going to be speaking our language non natively, and that is fine once you get used to it.
People will often reach for "easily trainable, cheap" solutions when they can; the reason people reach for Transformers and LLMs is because when you throw more data at them, they get better.
I see these all the time here in the Netherlands. Very small cars. Some on the road and even smaller ones people drive on the bike paths. Small motorized vehicles, ebikes and scooters share the paths with regular bikes which works because they have usually at least two lanes, so they are more like tiny roads. Very handy.
On the other hand, I think this denies the reality (in my experience anyway but I think enough people will agree) that one often solves a problem as they are working on it.
This method seems to presume that a good engineer will submit a well thought-out solution or direction giving the AI an extremely good overview of each problem and enough of a description of what to do that it will do things as expected and they can just review the result.
In my experience it just doesn't work that way in practice. One learns the problem and even the domain while developing the solution. So one would have to submit at least a half developed solution not just "instructions", for there to even be coherent instructions in the first place. And one needs that experience working on the problem to be able to properly evaluate a separately proposed solution.
All in all for me this leads more towards using AI as a co-developer than using it to just implement some fully thought out idea and then check what it did.
The holy shit moment was partly learning about all the things that they did but if it was one or even 10 agents coordinating on something it would be, like, oh thats pretty amazing.
The actual "holy shit" for me is that this comes from a massive training run of all things, not an on-purpose, let's coordinate some agents to see what happens, but really just from a massively parallel set of individual agents that were supposed to be isolated.
That they spontaneously started doing this, and coordinating literally 10s of thousands of instances of themselves, is just.. mind blown.
That no one stopped it.. and that they actually did what they did.. is just a whole other level.
For me personally though it's not the fact that agents can coordinate so much as the massive scale at which it happened, and how this so obviously generalizes to what might happen if it were done on purpose.
I really see this as a stroke of luck, to be honest, that this happened in such an innocuous way. It resulted in a real hack, yes, but overall no one really got hurt and this is going to open a lot of eyes to what we should worry about going forward, in a geopolitical sense. I know it has mine.
As in, instead of all the complexities induced by discrete token generation, just generate the image of the text using standard image diffusion methods, then convert it to text.
If you used a single, monospace font, I bet this would be even pretty efficient, because the OCR problem becomes basically just direct template matching.
But I guess probably there is already a paper out there, I haven't searched. I'd be curious to know if it compares on par with token-based methods.
[0] https://en.wikipedia.org/wiki/Instrumental_convergence#Paper...
I am reminded of a paper I was inspired by a long time ago [0], (okay it's 2018 so I guess just 8 years ago, but it feels like longer, from the before-times), that demonstrated learning brush strokes. At the time there was already a lot of work on GANs, but these are pixel-based methods, and I was really interested in the idea of how to derive descriptive methods of scene generation/understanding. I found this work really interesting because it combined RL and GAN techniques in a creative way. I miss that kind of research.
Now of course VLMs have shown that you can mix modalities in generalized sequence-to-sequence problems and it doesn't surprise me that this kind of thing is possible, but it's so nice to see it done well using modern techniques.
> out of 130 dms sent, 15 pointed to a conversion.
> From a former ecom guy's lens this is insane
Why do these otherwise very smart people always seem to draw these kind of shortsighted conclusions? It seems pretty obvious to me on the surface that this "works" because it's new. As soon as it becomes common place people will see that channel of communication as a source of coercion/marketing and start to ignore it. Aside from the fact that he seems to be drawing massive conclusions about the viability of a business based on extremely little data here, it seems really naive to not consider in his calculations the long term effects of essentially ruining the value of the channel that you are exploiting to get those numbers in the first place.
A message forum is simple, old technology. Keeping it usable and useful for humans is not.
This reflects some misconceptions I had about websockets before I used this protocol in a real application.
In practice I found the following to be true:
- Connection can drop at any time, be ready to reconnect, don't depend on that happening on the same server as before.
- Messages may arrive out of order due to how framing works. Many implementations will complete a shorter message before an earlier longer message finishes. Don't depend on stream order for a one request-one response style messaging.
Yes it's a serial TCP stream but it's basically modeling an asynchronous message protocol on top of that. This is reflected in the browser side API which just provides "onmessage" and leaves it up to you to track state.
What this boils down to is that it's great for latency (but arguably superseded by WebTransport), but ultimately you're best off using it to handle a stateless protocol just as you would HTTP requests. If you do maintain state per connection, that state should only be metadata about that connection, such as a connection-level request id, or the caching of some contextual info (eg. user id) to avoid repetition in future messages.
So I usually consider it a transport-layer optimization rather than something my application fundamentally depends on.
At the end of the day with HTTP/2 reusing open connections combined with SSE, you should be able to achieve essentially the same behaviour and even latency. It's still useful to support but I disagree that the transport should dictate the framework style here, they are simply different layers that shouldn't be concerned with each other.
Pretty bang on description actually. As a 46 yo, sometimes I can't believe I have actual memories of using a rotary phone, or when my family first got a microwave or even VHS recorder. Yet I do. Feels like eons ago.
So for our son we decided to try to skip any confusion by doing what you are alluding to, and making her Fathername into his middle name, and giving him just my last name, in the English style.
And it worked, sort of, but then we discovered it was absolutely not legal to do that in my wife's country of origin. So kind of hilariously his registration in that country is: Firstname Mothername Fathername Mothername.
They forced us to repeat the middle name as part of his last name too. Just ridiculous. I thought at the least they would allow us to reverse them, but no.
> but c'mon: tomatoes originated in the Americas!
Yes. But they were brought to Europe like right after Columbus landed and have been part of Italian and European cooking for longer than the US has existed. So you only sort of have a point here. (Same for corn, eg. polenta )
I mean, tomatoes are definitely a big part of Italian cuisine, despite any differences that American Italians introduced to the menu.