This is tough though because enterprises go absolutely ballistic over “training on our data” - which is understandable, but will also hold us back.
260 karma · joined March 11, 2017
This is tough though because enterprises go absolutely ballistic over “training on our data” - which is understandable, but will also hold us back.
Pre AI when engineers couldn’t find the answer in commit messages or documentation they would ask the author “why” and that human would “compute” the summary on demand.
I think that’s what I expect to do with these agent sessions - I don’t want more markdown, I want to ask it questions on demand. Git AI (https://github.com/git-ai-project/git-ai) uses the prompts that way. I think that model will win out. Save sessions. Read/ask questions relevant to the current agent’s work.
On asking peers. This is regrettably on the way out today - I’ll ask engineers about complex code they generated and they can’t give good answers. I think it’s because it all happened so fast — they didn’t sit with the problem for 48 hours. So even if they steered the agent thoughtfully it’s hard to remember all the decisions they made a week later.
For example…had to build my own tool to extend git blame and track the AI generated code in our repository and save prompts:
It’s interesting how a relatively small # of synapses can do all abstract reasoning when free from those concerns.
Take the pre-frontal cortex, leave the rest.
And I trust (quite a bit) that whatever he brought to light should be followed up on - if no other reason than to respect his memory. I hope it is taken seriously and those who retaliated find themselves w/o their positions of responsibility and power over other faculty.
+ the experiments may already be in the dataset so it’s really testing if it remembers pop psychology
Definitely the place to go these days if you have a public API and want it to be as developer friendly as they come.
If China / TikTok hadn’t overplayed the hand, they wouldn’t lose this weapon. Now they likely will (good thing)
Whatever sail catches the right wind, and gets you where you are going is the right said.
It unlocked the whole thing. I could take responsibility for getting the life I wanted, but that did not mean the bad things that came my way were my fault. Bad shit always comes, not always as a result of my actions, but it's always my responsibility to change the situation.
Instead of generating API designs, we’re using LLMs to check if a team’s API is “good” by their own standards.
Many large companies that use OpenAPI have custom Spectral + Optic rules to enforce API security practices, consistent styles, versioning policies, SLA adherence, and to prevent breaking change. These tests are hard to write and cannot verify a lot of the more abstract aspects of an API design. It’s really hard to write code to check if an OpenAPI operation allows batch creation, but when you give an LLM YAML and the rule “No endpoint should allow users to create multiple resources at once” — it does a surprisingly good job at testing the design and writing a nice error message when it fails.
The hardest technical challenge has been: - figuring out how to fit OpenAPI files (sometimes 10k lines or more) into the context window. Optic already breaks up API specs into their parts: ie parameters, response bodies, headers, etc. We’ve been putting just the relevant lines of the OpenAPI file into the LLMs and resolving all their dependencies beforehand so the model has all the context it needs. - Minimizing API calls. The first time you run LintGPT it is pretty slow because it has to run every rule across every part of the API specification (1000s of calls). But we shouldn’t have to repeat that work. Most of the time parameters, properties, etc don’t change and neither do the rules. We’re building caching into our web app to make this fast / save $ for end users.
Happy to answer any questions. I really think there’s a huge use case here for linting all kinds of code, config, database schemas, policies in ways that were never possible before. And personally, I like the idea of having these smart tools guiding me towards making my work better vs generating it all for me — idk something about that just feels good.
Ok then -- so these forward passes should take about the same amount of time per inference of the next token. That matches my intuition. When does the model know to stop? Is there a stop token?
Also aside -- totally crazy this thing can run a full forward-pass many times a second. I naively would have bet each pass token a lot longer even on powerful GPUs
What we really need is better tooling to help developers maintain the spec and built generators on top of it.
I’ve been building open source API version control tools built on top of OpenAPI. It’s an easy way to keep your spec up-to-date just by looking at test traffic. If it detects a diff, it will update the OpenAPI sort of like a snapshot test. https://www.useoptic.com/cli
Some of the economics of the underworld are remarkably sophisticated. Everyone has a role and there are some parts of the supply chain that require different kinds of risk. Sometimes it’s risking cash, physical safety or personal-criminal repercussions.
A friend suggested trying to read 3 days into a 1 week vacation. I couldn’t.
And from that moment I realized it was a muscle like anything else. Consuming media that requires active participation and work, that you are not rewarded for (professional learning) takes practice.
I come home tired most days but I still manage to crank out time to read.
Maybe it won’t work for you the way it worked for me, but give it a try. You might be surprised.
I used to think “letting go” meant “giving up”…
“Letting go” in the Alan Watts sense means being willing to give up on the means — not the outcome. I’ll let go of the way I’ve been doing things, let a small part of me die, so I have a better chance of meeting the outcome I am looking for. The psychological establishment might call that flexibly minded or low aversion to loss.
Giving up is convincing yourself a difficult the outcome/goal wasn’t worth it — instead of finding new means to reach it.
That won’t make sense written down unless you’ve done a whole lot of meditating.
That's exactly why it is the way it is.
You do not want a brand associated with the highest quality and seamless integration to be associated with self-repair kits. Even under the best of circumstances, with the best of expertise, one must expect self-repair kits sold to the mass market to fail 5%? 10%? of the time?