Then I got the music, which promises Unlimited lossless downloads for Pro tier, whereas the text says only 400 per month! Isn't this false advertising?
849 karma · joined November 4, 2020
Then I got the music, which promises Unlimited lossless downloads for Pro tier, whereas the text says only 400 per month! Isn't this false advertising?
https://developers.openai.com/api/docs/guides/agents-api/env...
That makes this much more enticing, and potentially eases transition between providers.
Clearly that would move things around.
none 35.2%, $49,791 96.7%, $23,457
35.2% on the standard harness, that's above Opus 5 on high.And then a local swarm noticed and disagreed and took it down.
Somehow codex shows 100%, so it was only chat web interface?
What exactly is the premium that you're getting for paying these prices?
I still haven't found generic solutions to selecting and copy-pasting text using the keyboard only though, when the text is not in a textbox/area.
I had used a vim-like plugin in firefox that let you do that somewhat, but nothing OS level.
That transparency alone could change the market dynamics a lot.
Just hope that as an individual you wouldn't be taxed for this though...
But alas
Surely stripe if anyone have learned to harness the cash flowing through their system. Hell, they could be emitting bonds on expected token consumption bills!
Shrinkflation is a thing and the bun has been its victim. They could slowly shrink the sausage now until it fits the bun again.
people interested in the discounted -contributor model would increase their harness beta testers. But it'd be a terrible strategy to capture the better paying customer base through openrouter and similar...
Not sure if anyone would bother.
Then the larger LLM gets all the right lights on, yields better outputs and we translate back into user domain.
I kinda thought the chain-of-thought reasoning already did this, no?
You could try doing the high level design yourself at least. Ask for its review and iterate without asking it to do it all.
Once it has generated some implementation, critique it and ensure you understand its approach And you agree with it, steer it otherwise.
If there’s anything unclear to you say so and have it rewrite it in an easier way to understand.
That's where I'd like to see this sort of checkup. Yell at me please if i just said anyone can ssh as root from any node!
> Some of the solutions here may even be simple fixes;
They are still throwing ideas. Why have they not made those simple fixes yet before disclosing? > closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access.
> This led them to believe—arguably reasonably—that the real environments they encountered were simulations.
That the AI lab most typically preaching for alignment does not consider this an obvious misalignment is a clear red flag.