"Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a connection there yourself".
"Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a connection there yourself".
There is no fire at all in the painting, only some smoke.
https://en.wikipedia.org/wiki/The_Trojan_Women_Set_Fire_to_t...
I don't see AI going any differently. Some companies will figure out where and how models should be utilized, they'll see some benefit. (IMO, the answer will be smaller local models tailored to specific domains)
Others will go bust. Same as it always was.
God help us all
Possibly.
I am reminded of the dotcom boom and bust back in the 1990s
By 2009 things had recovered (for some definition) and we could tell what did and did not work
This time, though, for those of us not in the USA the rebound will be lead by Chinese technology
In the USA no-one can say.
I suck at and hate writing the mildly deceptive corporate puffery that seems to be in vogue. I wonder if GPT-4.5 can write that for me or if it's still not as good at it as the expert they paid to put that little gem together.
It'd be funny if it's actually full-automated, closed-loop automation of capital allocation markets.
"Why are we doing this? How much money are we getting?" -> "I dunno. It's what the models said."
> "I dunno. It's what the models said."
The obvious human idiocy in such things often obscures the actual process:
"What it [capitalism] is in itself is only tactically connected to what it does for us — that is (in part), what it trades us for its self-escalation. Our phenomenology is its camouflage. We contemptuously mock the trash that it offers the masses, and then think we have understood something about capitalism, rather than about what capitalism has learnt to think of the apes it arose among." [0]
The link shows a significant reduction.
grep hallucination, or, https://imgur.com/a/mkDxe78.
You need to remove your shoe and drive with like two toes to get the speed just right, though.
Test drivers I have done this with takes off their shoes or use ballerina shoes.
And keeping steady state speed is not that hard.
(And even that was a downgrade compared to the more uncensored pre-release versions, which were comparable to GPT-4.5, at least judging by the unicorn test)
in our case, the bump was actually from gpt-4-vision to gpt-4o (the use case required image interpretation)
It got measurably better at both image cases and text-only cases
Releasing GPT 4.5 might simply be a reaction to Claude 3.7.
Wow, I'm old.
Now get off my lawn ))
o1 preview. o1 mini. o1. sora. o3-mini <- very good at code
OpenAI has had many releases since gpt4. Many of them have been substantial upgrades. I have considered gpt4 to be outdated for almost 5-6 months now, long before claudes patch.
People _seriously_ underestimate just how much stuff is online and how much impact it can have on training.
[0] https://the-decoder.com/openai-quietly-funded-independent-ma...
We can assume they’re lying too but at some point “everyone’s bad because they’re lying, which we know because they’re bad” gets a little tired.
2. We know that “open”ai is bad, for many reasons, but this is irrelevant. I want processes themselves to not depend on the goodwill of a corporation to give intended results. I do not trust benchmarks that first presented themselves secret and then revealed they were not, regardless if the product benchmarked was from a company I otherwise trust or not.
If a scientific paper comes out with “empirical data”, I will still look at the conflicts of interest section. If there are no conflicts of interest listed, but then it is found out that there are multiple conflicts of interest, but the authors promise that while they did not disclose them, they also did not affect the paper, I would be more skeptical. I am not “offended”. I am not “rejecting” the data, but I am taking those factors into account when determining how confident I can be in the validity of the data.
This isn't what happened? I must be missing something.
AFAIK:
The FrontierMath people self-reported they had a shared folder the OpenAI people had access to that had a subset of some questions.
No one denied anything, no one lied about anything, no one said they didn't have access. There was no data obtained under the table.
The motte is "they had data for this one benchmark"
The bailey is "they got data under the table"
Bailey: "one is free to trust their "verbal agreement" that they did not train their models on that, but access they did have."
Sigh.
> Bailey: "one is free to trust their "verbal agreement" that they did not train their models on that, but access they did have."
1. You’re confusing motte and bailey.
2. Those statements are logically identical.
Motte and Bailey refers to an argumentative tactic where someone switches between an easily defensible ("motte") position and a less defensible but more ambitious ("bailey") position. My example should have been:
- Motte (defensible): "They had access to benchmark data (which isn't disputed)."
- Bailey (less defensible): "They actually trained their model using the benchmark data."
The statements you've provided:
"They got caught getting benchmark data under the table" (suggesting improper access)
"One is free to trust their 'verbal agreement' that they did not train their models on that, but access they did have."
These two statements are similar but not logically identical. One explicitly suggests improper or secretive access ("under the table"), while the other acknowledges access openly.
So, rather than being logically identical, the difference is subtle but meaningful. One emphasizes improper access (a stronger claim), while the other points only to possession or access, a more easily defensible claim.
It was not public until later, and it was actually revealed first by others. So the statements seem identical to me.
What is "this"?
> obviously the problem with getting "data under the table" is that they may have used it to training their models
I've been avoiding mentioning the maximalist version of the argument (they got data under the table AND used it to train models), because training wasn't stated until now, and it would have been unfair to bring it up without mention. That is that's 2 baileys out from "they had access to a shared directory that had some test qs in it, and this was reported publicly, and fixed publicly"
There's been a fairly severe communication breakdown here, I don't want to distract from ex. what the nonense is, so I won't belabor that point, but I don't want you to think I don't want to engage on it - just won't in this singular posts.
> but the only reassurance being some "verbal agreement", as is reported, is not very reassuring
It's about as reassuring as it gets without them releasing the entire training data, which is, at best, with charity marginally, oh so marginally reassuring I assume? If the premise is we can't trust anything self-reported, they could lie there too?
> People are free to adjust their P(model_capabilities|frontiermath_results) based on their own priors.
Certainly, that's not in dispute (perhaps the idea that you are forbidden from adjusting your opinion is the nonsense you're referring to? I certainly can't control that :) Nor would I want to!)
And FFS I assume the dispute is about the P given by people, not about if people are allowed to have a P.
Or is the assumption that the training set is so big it doesn't matter?
Perhaps they are/were going for stealth therapy-bot with this.
Empathy done well seems like 1:1 mapping at an emotional level, but that doesn’t imply to me that it couldn’t be done at a different level of modeling. Empathy can be done poorly, and then it is projecting.
Does one of these have a higher EQ, despite both being ink and paper and definitely not sentient?
Now, imagine they were produced by two different AIs. Does one AI demonstrate higher EQ?
The trick is in seeing that “EQ of a text response” is not the same thing as “EQ of a sentient being”
This is a designed system. The designers make choices. I don’t see how failing to plan and design for a common use case would be better.
= not to say that the people that work on AI are not incredibly talented, but more that it's not human
trainimg it topretend to be a feelingless robot or sympathetic mother are both weird to me. it should state facts with us.
You are confusing a specific geographical sense of “greater” (e.g. “greater New York”) with the generic sense of “greater” which just means “more great”. In “7 is greater than 6”, “greater” isn’t geographic
The difference between “greater” and “better”, is “greater” just means “more than”, without implying any value judgement-“better” implies the “more than” is a good thing: “The Holocaust had a greater death toll than the Armenian genocide” is an obvious fact, but only a horrendously evil person would use “better” in that sentence (excluding of course someone who accidentally misspoke, or a non-native speaker mixing up words)
A lot of folks here their stock portfolio propped up by AI companies but think they've been overhyped (even if only indirectly through a total stock index). Some were saying all along that this has been a bubble but have been shouted down by true believers hoping for the singularly to usher in techno-utopia.
These signs that perhaps it's been a bit overhyped are validation. The singularly worshipers are much less prominent and so the comments rising to the top are about negatives and not positives.
Ten years from now everyone will just take these tools for granted as much as we take search for granted now.