I just added evals for existing datasets and some classic controls problems.
Will likely look for more commercial datasets and synthetic data too.
I’m funding personally in short term but if I can find some compute can see how well this scales!
485 karma · joined January 19, 2017
I just added evals for existing datasets and some classic controls problems.
Will likely look for more commercial datasets and synthetic data too.
I’m funding personally in short term but if I can find some compute can see how well this scales!
I believe the way we advance the capabilities for general computer use is to to attempt harder and harder tasks with it, and build up the necessary infrastructure to accomplish it along the way.
Even when we don’t achieve the goal, the infrastructure we build is often useful for easier and more obtainable goals as well.
If your computer use environment can pilot Minecraft, it can also probably do your taxes
There are plenty of other features to explore, but this one alone has resulted in significant improvements in the speed of execution in many tasks including minecraft playing agents, computer use agents, and for querying and aggregating information from my telemetry systems in a lightweight low risk way.
edit: after reading a bit more of description looks like yall are taking a similar approach, kudos!
I want to be able to give agents access to computation in a secure way without giving them full access to a computer
a common use case i run into is i want to be able to configure corporate vpn software on windows machines. is there a link for a getting started guide i could try this out with?
ran into this when writing agents to fix unit tests. often times they would just give up early so i started writing the verifiers directly into the agent's control flow and this produced much more reliable results. i believe claude code has hooks that do something similar as well.
From 1:14:55-1:15:20, within the span of 25 seconds, the way Demis spoke about releasing all known sequences without a shred of doubt was so amazing to see. There wasn't a single second where he worried about the business side of it (profits, earnings, shareholders, investors) —he just knew it had to be open source for the betterment of the world. Gave me goosebumps. I watched that on repeat for more than 10 times.
It matters proportionally to the amount of time I intend to maintain it for, and the amount of maintenance expected.
> How does the agent know which to keep and which to discard for a conflict? This would that it has deep contextual information about the codebase it's looking at.
There is no guidance regarding this in the agent. For some conflicts this is important but I have found that for the conflicts I've tried this on it does not need the additional context to do the right thing! There is a option to add additional prompting before it solves a conflict. I will add notes in here to guide it such as "ignore auto generated files, i will generate them again later" so that it doesn't get stuck on generated data.
- How does training with RL differ from fine tuning?
- When would it make sense to fine tune instead of using RL?