115,911 karma · joined July 27, 2010
Contact me at hn@jacek.zlydach.pl.
(I mean it. Don't hesitate, I'm always happy to continue interesting conversations after their HN thread disappears off the front page.)
Homepage: http://jacek.zlydach.pl. More: https://keybase.io/temporal_pl.
I do software nowadays, but I'm constantly looking for opportunities to do something actually useful for humanity. If you know of some, please let me know - especially if they involve biotech, cleantech or NewSpace industry.
(see also: http://jacek.zlydach.pl/blog/2018-01-06-going-far-going-small.html)
I can already see the border shift even for mundane tasks I have Claude working on. Increasingly, I'm just setting a high-level goal, and then checking progress and occasionally answering questions or doing something like configuring a system Claude can't easily reach itself (e.g. recording a bunch of traces through my normal use of a system that Claude deemed too fragile to risk operating on its own). Of course, I get detailed instructions to help me - "go there, do this and that, then press this to capture recording, run through this script here to process, attach result to next message". In those cases, Claude is effectively using me as a tool to call.
There's this thing called "scissor effect", where when you have two wavefronts colliding at an angle, you get an apparent wave that moves much faster than either of the input waves - if you do that with light, you can trivially exceed the speed of light. Now this isn't an information-carrying wave, but I believe you can still use this phenomenon as a "virtual clock", to synchronize a bunch of independent receivers to act on a frequency faster than the real, interfering clock signals.
Well, that, or the brain is just such a mess that it acts as a spread-spectrum source and none of the high frequencies are above the noise floor.
EDIT: and of course a little background research after writing this comment revealed it's already settled (through more invasive probing) that the human brain has elements operating in kilohertz range - the mystery isn't whether this happens, but how it adds up to cognition and consciousness.
It starts with what they already claim to be doing - increasingly relying on existing models in non-trivial work related to training, evaluating and optimizing the next, more capable generation of models. As long as the proportion of work keeps shifting towards agents doing more and more of it, and humans less and less, that's RSI at play.
It may be that it turns out LLMs lack some fundamental level of judgement and it plateaus, but frankly I find this notion absurd; LLMs already show better judgement than most people. The alternative is, at some point LLMs will show the ability to futz their way into improvement of the next generation of models even without humans in the loop - even if much less efficient at first, if generation N+1 is more capable than generation N, it'll either take off or burn out.
The currently well-served market for AI is not software engineering, it's approximately all of white-collar work, ranging from accounting and law, through medicine, general office work, to school administration, education, NGOs and governance.
Not everyone is going to just publicly brag about their AI use, but it's an open secret everyone is either using LLMs for half their work, or - if for some reason they're not busy enough to arrive at this idea on their own - under pressure to start using them.
And that's at day 0. Because LLMs are not going away, thank to open weight models, and the race for capabilities has left so, so many low-hanging fruits unpicked, we'd have a decade or more of useful R&D to do even if LLM capabilities suddenly plateaued tomorrow and never improved.
I mean, look at Jev making round in the industry now - this is just a single example of a low-hanging fruit that took many years for someone to bother to pick up and market a bit. There's many, many more of these just laying around.
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
You could view the bounty prices as market evidence that most of this is rightfully treated as nothingburgers. I.e. the alternative to paying $5000 to some random person for this class of vulnerability research is not risking a trillion dollar hack the next day - it's just risking shmaybe some kerfuffle down the line, followed by fixing it through normal triage process. The bounty program is as much marketing as security, and $5000 is probably about the right price for marginal effort into sustaining the "we are treating security seriously" message.
In a way, the very existence of those bug bounty programs in large companies is evidence they don't see a reason to treat vulnerabilities seriously enough to proactively find and fix them in-house.
If security vulnerabilities would be anywhere serious as most commenters on-line seem to think, companies would pay hundreds of thousands for serious vulnerabilities, just to save a day before they get hit by them - on top of spending millions in-house to try and stay ahead of the attackers.
But they don't. Because most exploits are inconsequential and/or aren't being exploited much.
For example, with Claude, I have an "operations" project that naturally grew to cover daily use of shared family calendar, sweeping my mail inbox, and my current personal todo lists, but also a lot of the latter made it deal with my Home Assistant instance. I have separate project for specific things to do with Home Assistance (e.g. one that's about "life support" - HVAC controls, dashboards, monitoring, etc.), one about phone specifically (front-loaded with dumps of specs of my phone's hardware, OS, etc.). Each of them has its distinct set of memories accumulated over months.
And so every couple sessions, I hit a situation in which the agent has to interact with tools and rulebooks that are focus of a different project, and it fumbles a lot. E.g. HA Life Support needs to add some tasks to the todo list, or the Ops project needs to look up climate stats for some reason or other, etc. In these moments, I really wish project memories could mix - but they can't, the boundary is high.
The most annoying case is when I tell Claude that it's wrong, and we literally worked out a solution (or consensus on ethics) in a recent conversation - and then it spends couple minutes looking through past history, burning a chunk of my 5-hour limit, only to come back empty. Yep, that conversation happened in another project. *sigh*
Probably this.
> I guess the intent was to pull together the Codex, claw and ChatGPT Work paradigms and simplify them?
Probably this in part, too. They're trying to figure out how to decouple agent lifetime from conversation lifetime, without accidentally making it too useful for end users.
I noticed in the past that tooling - both for AI and in general - tends to miss the features that would make it most useful. Like, look how long did it take for the AI vendors to supported "scheduled runs", and they're still offering only toy-level configuration for that[0]. Wonder how many years it'll take to allow users to configure external triggers, and whether it'll be sooner than forced re-authentication will become frequent enough to make the feature useless in the first place.
--
[0] - I understand they don't want users to run this too often, or to accidentally end up spawning a job every 10 minutes - but with limits on frequency in place, there is otherwise no good reason for this to be a limited dropdown, instead of the "repeats every" configuration that every calendar app and phone app already uses.
It most definitely is not, hasn't been for a while now.
I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API/UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.
Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
There's a lot of things to tune there, and just as many reasons to do it.
Datacenters have massive economies of scale. Everything from cheaper electricity to having specialized, more efficient hardware to simply being able to run it continuously at near-100% utilization, all adds up.
Many things in the economy - most notably, manufacturing of most consumer goods - only makes economic sense once you're producing for/serving millions of people. This is not unusual.
> In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
Many say that, but I sincerely doubt they'd actually follow through. People might get more conservative about how they spend their tokens, but AI today is just too good at eliminating drudgery and boring / bullshit parts of daily work to give up on merely 3-5x price increase.
Are they though? Or is it just what some companies would want them to be?