106 karma · joined May 3, 2023
Consider yourself lucky. It’s the people who haven’t run into something like this that will end up placing too much trust in these tools.
Why does almost everyone act as if this is a valid thing to do? We all know that these models cannot verify that something is well substantiated. The mass delusion is crazy making.
The guy two comments up wrote a great book on using pandas effectively. It’s called Effective Pandas
- "A new series of reasoning models for solving hard problems. Available now." - "They can reason through complex tasks and solve harder problems than previous models in science, coding, and math." - "In a qualifying exam for the International Mathematics Olympiad (IMO), GPT-4o correctly solved only 13% of problems, while the reasoning model scored 83%." - "But for complex reasoning tasks this is a significant advancement and represents a new level of AI capability." - "As part of developing these new models, we have come up with a new safety training approach that harnesses their reasoning capabilities to make them adhere to safety and alignment guidelines. By being able to reason about our safety rules in context, it can apply them more effectively. " - "These enhanced reasoning capabilities may be particularly useful if you’re tackling complex problems in science, coding, math, and similar fields."
there are a few more in that post, but clearly OpenAI is pushing the reasoning thing A LOT
2. Plausible != Correct - ie, would someone notice if they were reading a slightly different but still logical story?
It is freaking amazing that this works almost perfectly. Seriously, it’s mind blowing. The problem is, when you keep in mind how the technology works, you realize that the “almost” can never be removed. That’s fine for some use cases but not for others. I understand that human translators make mistakes but they have a conception of truth and correctness, that matters.
This is not true. It’s not possible to guarantee this. One can be certain that two files DON’T match if they have different hashes, but one cannot be certain that two files DO match based ONLY on the fact that they have the same hash.