It’s verbose with a layer of looking legit play first glance sloppy UI.
I don’t mind vibe code ui but at least some ui efffort would be nice. Not just 1 shot.
871 karma · joined November 2, 2021
It’s verbose with a layer of looking legit play first glance sloppy UI.
I don’t mind vibe code ui but at least some ui efffort would be nice. Not just 1 shot.
Thanks for reminding me to try this new round of models.
I wonder if there are some good tools/apis that can help here.
A lot of information is often login gated
Feel like latest models have really got hung up on “facts” and “evidence” in a really unnatural way.
And especially Claude seems to have a lot of trouble separating it’s internal thoughts, from prompts/instructions, to output.
So you get all of these weird comments and artifacts in the output.
“neighboring versions do not inherit certification. No tag, npm publication, Pages deployment, or release authorization is implied.”
It’s sort of like memoizing or distilling the knowledge. Works really well for certain type of problems.
In a world where steam allows instant returns with under 2 hours play time.
I have a few games now on ps where I barely played or found out the controls were terrible and have no way to return.
There can be whole UIs and tooling and infrastructure to manage around them and that’s what the sass offer
Funnily enough it’s extremely relevant to agents/subagents.
Are there any new dance style games?
As an aside wonder if there is something to adding game/gambling mechanics to exercise.
I wonder if it’s an artifact of OpenAI’s values or rl training approach.
Also, it prob does make it perform better just not more efficient.
Great for the OpenAI employee working on security scanning who doesn’t have to pay for their tokens.
Not so much for the dev building their web app who is trying maximize their subscription.
Like hiring an aerospace engineer to build you a shed.
I feel like codex/sol is better at well scoped hard technical problem.
Where it can sort of run this brute force analytical loop.
Like doing performance optimization or other search type problems. I think the math proofs are good examples of this.
One thing I don’t love about codex/sol is I find it tends to overengineer and be overly cautious.
I was using it to do create some scraping + data processing.
It went kind of crazy on the provenance, need at least 3 sources of consensus before promoting facts type bullshit.
defined a bunch of enums and gates.
I just wanted scrape some site data and put it into a SQLite dB. Like chill codex.
I feel like Claude is better at that.
Intermittent nudges
Injects a reminder to use ASD-STE100 Simplified Technical English which I picked up from a suggestion in another thread.
Honestly it’s working pretty well. Except for I need to check how often it’s actually firing.
https://en.wikipedia.org/wiki/Simplified_Technical_English
I feel like hooks aren’t utilized enough. Really nice for being the sort of auto steering as long as you can encode some pattern to detect the bad behaviour.
Prompts and skills just don’t cut it.
You can believe whatever you want to believe. Or maybe you live in a magical world where all your tests are pure and side effects nor IO exist.
In the other thread you think running tests in parallel doesn't count because it has documented gotchas...
You say assured? You think it’s impossible to write 2 tests that interact each other?
All you gotta do is google “test isolation”. But if you want to pretend it’s not a thing the I have no idea what to tell you?
Here’s one example since junit came up
It’s the implementation of each test that is responsible for its isolation. Every time you write or make changes to your tests. You have to think, can they mess with each other.
you can share some things and be isolated across other dimensions.
You could share a Postgres dB connection but just isolate the data logically.
You can isolate or share across individual tests, across suites, across envs, across runs, across time.
More isolation is generally better unless the complexity or cost is too high.
But I guess I was just confused as to how a test requires isolation rather than have it.
The isolation comes from the test implementation not the framework. There isn’t any framework out there that can guarantee/give you isolation.
If I create a new in mem dB in the test there’s nothing stopping me from running it in parallel?
Nothing about that “requires” isolation. It is isolation.
Wondering if you also support some non realtime models.
E.g. prompt Claude to write all sentence in reverse, or swap every 2 words etc. Then use a script to put reorder in the right ordering?
Pretty cool technique honestly. You could do it the other way as well right?
If you had a list of categories you have the model to generate a sample query and then do embedding on that?
I sometimes just leave some goals or something running before I go to bed or out and I don’t want to pay the cache text when I come back.
like if you have a long cli command or something will it still try to watermark it ?
Is there some way you can know which tokens are required to be syntactically correct vs not?
My process/playbook for determining if I go to the gym is I go if I didn’t go the day before.
This is both a premade decision and a process that requires discipline to execute on.
So why isn’t a decision made in advance an example of a simple process?
Opinionated linters is an example.
Sometimes you just gotta turn your brain off and move forward.