It was categorically, undeniably better--that's where the nostalgia is coming from.
There are lots of reasons why this happened (and those are interesting questions.) But simply denying that it was better doesn't seem particularly reasonable to me.
466 karma · joined October 7, 2008
joshalbrecht@gmail.com
It was categorically, undeniably better--that's where the nostalgia is coming from.
There are lots of reasons why this happened (and those are interesting questions.) But simply denying that it was better doesn't seem particularly reasonable to me.
I don't have a simple, perfect solution. We're just trying to make it possible for individuals and smaller companies to have access to the same kinds of tooling that the largest companies already have access to, and hopefully equalize the playing field at least a little bit...
If anyone has better ideas, I'd love to hear them!
But our reaction to it has been to say "ok, well the best practice in software engineering is to make small, well-isolated components anyway, so what if we did that?"
We've been trying to really break things apart into smaller pieces (and that's even evident in mngr, where much of the code is split out into separate plugins), and have been having a ton of success with it.
I realize that that might not be an option for more brownfield / existing / legacy projects, but when making something new, I've really been enjoying this way of building things.
I believe we can use these types of tools to make software more understandable, and mngr is an example of how to do that.
In our case study, we're using AI to increase our test coverage, and if you look at it, I would argue that we are making it more understandable--now instead of just having 100's of tests, we simply have a document that describes how the software is supposed to work, and the tests are linked to that document, and checked to ensure that they conform.
That means that anyone--not just the author of the software--is now able to read through the high level tutorial description of how the commands work in order to understand what the program should do!
And as for the tests themselves, we've been able to make nice testing infrastructure--like the transcripts and recordings that were highlighted in the post--to make it even easier for us to verify the behavior of the software.
We also have an incredibly detailed style guide and set of tests and guidelines to ensure that the entire code base is consistent, and high quality. You can drop into any of the code and pretty quickly understand what is happening. And if not, claude will do an excellent job of describing how any given component works, and how it relates to the others.
Finally, mngr itself is designed to be fully transparent when it is running--you can literally attach to the coding agent you are running and see exactly what is happening, and the program makes extensive log outputs for everything it does (feel free to open a PR if you'd like to see more!)
It's not perfect formal verification, but it does feel like we're making meaningful progress on making it easier to understand software--not harder.
Then your primary project (from Sculptor's perspective) is simply whatever contains the devcontainer / dockerfile that you want it to use (to pull in all of those repos)
It's still a little awkward though -- if you do this, be sure to set a custom system prompt explaining this setup to the underlying coding agent!
(I'm a founder of Imbue, the company behind Sculptor: https://imbue.com/sculptor/ )
Obviously not the most user friendly or usable, but we found that people often got pretty confused when this was a browser tab instead of a standalone app (it's easy to lose the tab, etc)
If you have any specific questions that aren't covered there, please let us know in Discord!
I'd give it a good 20% chance of working if you set the right environment variables in there :) Feel free to experiment in the "Terminal" tab as well, you can call claude directly from there to confirm if it works.
Eventually what we want is for the whole thing to be open -- Sculptor, the coding agent, the underlying language model, etc.
Our approach is a bit more custom and deeply integrated with the coding agents (ex: we understand when the turn has finished and can snapshot the docker container, allowing rollbacks, etc)
We do also have a terminal though, so if you really wanted, I suppose you could run any text-based agent in there (although I've never tried that). Maybe we'll add better support for that as a feature someday :)
> What if my application is not dockerized?
Then claude runs in a container created from our default image, and any code it executes will run in that container as well.
> Can Claude Code execute tests by itself in the context of the container when not paired?
Yup! It can do whatever you tell it. The "pairing" is purely optional -- it's just there in case you want to directly edit the agent's code from your IDE.
> Do they all share the same database?
We support custom docker containers, so you should be able to configure it however you want (eg, to have separate databases, or to share a database, depending on what you want)
> Running full containerized applications with many versions of Postgres at the same time sounds very heavy for a dev laptop
Yeah -- it's not quite as bad if you run a single containerized Postgres and they each connect to a different database within that instance, but it's still a good point.
One of the features on our roadmap (that I'm very excited about) is the ability to use fully remote containers (which definitely gets rid of this "heaviness", though it can get a bit expensive if you're not careful)
> the feature I was most looking forward to is a mobile integration to check the agent status while away from keyboard, from my phone.
That's definitely on the roadmap!
Please feel free to join discord if you run into any bugs or have any issues at all, we're happy to help: https://discord.gg/sBAVvHPUTE
Suggestions welcome too!
This is from our fundraising post 2 years ago:
> Our goal remains the same: to build practical AI agents that can accomplish larger goals and safely work for us in the real world. To do this, we train foundation models optimized for reasoning. Today, we apply our models to develop agents that we can find useful internally, starting with agents that code. Ultimately, we hope to release systems that enable anyone to build robust, custom AI agents that put the productive power of AI at everyone’s fingertips.
- https://imbue.com/company/introducing-imbue/
We have trained a bunch of our own models since then, and are excited to say more about that in the future (but it's not the focus of this release)
Someday we'll probably have paid plans and business / enterprise licenses available as well, but our focus right now is on making it really useful for people.
To me, the whole point of our company is to make these kinds of systems more open, understandable, and modifiable, so at least as long as I'm here, that's what we'll be doing :)
We accomplished this using our hyperparameter optimizer, CARBS. We’re open-sourcing CARBS today so that other small teams experimenting with novel model architectures can experiment at small scale and trust performance at large scale.
• 11 sanitized and extended NLP reasoning benchmarks including ARC, GSM8K, HellaSwag, and Social IQa • An original code-focused reasoning benchmark • A new dataset of 450,000 human judgments about ambiguity in NLP questions
We're sharing open-source scripts and an end-to-end guide for infrastructure set-up that details the process of making everything work perfectly, and ensuring that it stays that way.
This is one of a three-part toolkit on training a 70b model from scratch. The other two sections focus on evaluations and CARBS, our hyperparameter optimizer; you can find them here: https://imbue.com/research/70b-intro/
Thoughts and questions welcome! :)
We’re an AI research company directly building human-like general machine intelligence. We have significant funding that will last a decade from investors including YC, researchers from OpenAI, and a number of high profile individuals. Work with our researchers on cutting-edge deep learning research — running experiments, implementing architectures from papers, experiment tracking tools, model debugging methods, automated hyperparameter optimizers, developer tooling & visualization, etc.
Learning & growth is a core part of our culture, and we will invest in yours. For example, the whole engineering team worked through Pieter Abbeel’s CS294 last year, and some of us also completed CS287 and CS285.
No prior machine learning experience required. For more example projects and benefits, see the full job description: https://generallyintelligent.ai/#careers
To apply: Email: jobs@generallyintelligent.ai
We’re an AI research company directly building human-like general machine intelligence. We have significant funding that will last a decade from investors including YC, researchers from OpenAI, and a number of high profile individuals.
Work with our researchers on cutting-edge deep learning research — running experiments, implementing architectures from papers, experiment tracking tools, model debugging methods, automated hyperparameter optimizers, developer tooling & visualization, etc.
Learning & growth is a core part of our culture, and we will invest in yours. For example, the whole engineering team worked through Pieter Abbeel’s CS294 last year, and some of us also completed CS287 and CS285.
No prior machine learning experience required. For more example projects and benefits, see the full job description: https://generallyintelligent.ai/#careers
To apply: Email: jobs@generallyintelligent.ai
We’re an AI research company directly building human-like general machine intelligence. We have significant funding that will last a decade from investors including YC, researchers from OpenAI, and a number of high profile individuals.
Work with our researchers on cutting-edge deep learning research — running experiments, implementing architectures from papers, experiment tracking tools, model debugging methods, automated hyperparameter optimizers, developer tooling & visualization, etc.
Learning & growth is a core part of our culture, and we will invest in yours. For example, the whole engineering team worked through Pieter Abbeel’s CS294 last year, and some of us also completed CS287 and CS285.
No prior machine learning experience required. For more example projects and benefits, see the full job description: https://generallyintelligent.ai/#careers
To apply: Email: jobs@generallyintelligent.ai
We’re an AI research company directly building human-like general machine intelligence. We have significant funding that will last a decade from investors including YC, researchers from OpenAI, and a number of high profile individuals.
Work with our researchers on cutting-edge deep learning research — running experiments, implementing architectures from papers, experiment tracking tools, model debugging methods, automated hyperparameter optimizers, developer tooling & visualization, etc.
Learning & growth is a core part of our culture, and we will invest in yours. For example, the whole engineering team worked through Pieter Abbeel’s CS294 last year, and some of us also completed CS287 and CS285.
No prior machine learning experience required. For more example projects and benefits, see the full job description: https://generallyintelligent.ai/#careers
To apply:
Email: jobs@generallyintelligent.ai
We’re an AI research company directly building human-like general machine intelligence. We have significant funding that will last a decade from investors including YC, researchers from OpenAI, and a number of high profile individuals.
Work with our researchers on cutting-edge deep learning research — running experiments, implementing architectures from papers, experiment tracking tools, model debugging methods, automated hyperparameter optimizers, developer tooling & visualization, etc.
Learning & growth is a core part of our culture, and we will invest in yours. For example, the whole engineering team worked through Pieter Abbeel’s CS294 last year, and some of us also completed CS287 and CS285.
No prior machine learning experience required. For more example projects and benefits, see the full job description: https://generallyintelligent.ai/#careers
To apply:
Email: jobs@generallyintelligent.ai
We’re an AI research company directly building human-like general machine intelligence. We have significant funding that will last a decade from investors including YC, researchers from OpenAI, and a number of high profile individuals.
Work with our researchers on cutting-edge deep learning research — running experiments, implementing architectures from papers, experiment tracking tools, model debugging methods, automated hyperparameter optimizers, developer tooling & visualization, etc.
Learning & growth is a core part of our culture, and we will invest in yours. For example, the whole engineering team worked through Pieter Abbeel’s CS294 last year, and some of us also completed CS287 and CS285.
No prior machine learning experience required. For more example projects and benefits, see the full job description: https://generally-intelligent.breezy.hr/
To apply:
Email: jobs@generallyintelligent.ai