Best of luck!
12,566 karma · joined April 20, 2008
Best of luck!
There are two reasons it’s used
1) it’s easier to type (*)
2) Placate people who are strongly convinced it’s the same thing. To them “functional” means “nearly the same but not yet understood”. For others it’s just a way to sidestep the first group and have a conversation.
When I see a world like this consider the intended audience. When you and I talk, we drop the word because we both know we’re are talking about (*) “this system is exhibiting goal-seeking behavior similar to other systems that are understood to pursue goals”. If I don’t know the person I will use the word and focus on the subject.
Different laws require different degree of awareness and intent for actions to qualify as a crime. Computer hacking laws are, as I’m learning from tptacek, set very high bar for intent, which is a choice by the legislature. They made a different choice for a death of a human - manslaughter crime does not require intent to kill.
Personally I’m happy they set high bar for hacking. Imagine you copy-pasted sample code with default root user name and password, and it worked. You were negligent. And you are clearly performing unauthorized access. If intent was not needed that would be jail time.
More broadly, we should as a society be very biased towards requiring intent across the board. Where clearly lacking, as is probably here, there should be a different law to discourage creating volatile situation where unintentional action can wreck havoc. Such laws exist for handling hazardous materials, for example, and it should be created for handling hazardous goal-seeking algorithms.
Which option seems preferable to you?
Starting from a business POV one should inflate terminology, hack together an MVP, and see if the market demands it before doing hardcore R&D.
But starting from technical/craftsman POV all you see is a hack and a lot of big words, so it’s easy to become jaded.
Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?
Or a combination of all those things.
Verifying sources is a recursive problem - where do you stop? Humans have intuitive feel for it, but agents don’t or at least not yet (I wonder if intuition is just a secondary neural net which is currently being added to the agents as we speak).
Also as a human you are able to examine agents erroneous trajectory, real or imaginary, without contaminating your own. Agent have a problem with that - as soon as someone else’s thought is in the context it can lose track of provenance and veracity. Sometimes I think we need a bloom filter to retroactively assign “dirty” flag to invalidated or questionable token spans already in the context.
In turn it makes me wonder how do humans do those things? Perhaps it is our human job to apply discretion going forward.
If I have this idea one day I will search for it and then think “how am I different fro that which already failed?”.
Test-Time Learning: The model updates its own memory weights while running an inference task.
Disconnecting “smart tv” from internet to use Apple TV is possible and might work, but you will get nagged to reconnect all the time. And TV mfr will eventually try to find workarounds - purchase access to wireless connections from large-footprint connectivity providers (Xfinity WIFI, 4g/5g networks, etc) or create a p2p network among their own connected devices. I doubt they do it right now as it’s less profitable to focus on niche audience, but eventually it might become profitable enough.
That said, abdicating to an LLM is the worst of all worlds - you’re not thinking and the product is not tailored.
The solution is obvious - write as much detail as you need and allow readers to interrogate the virtual you with an LLM, maybe not even reading what you write.
Are they noticeable happier?
Would be nice to hear from someone who lived both sides of the pond for a long while.
I keep thinking about various ways of “pushing back” on an agent, shortening feedback loop and extending what we can grantee about results.
At the most low level we can nullify probability of the next token if that token is not desirable (eg json schema enforcement under constrained inference), this is the fastest pushback. Various compiler checks, linters, unit tests, exotic type systems, e2e tests, production traces. Wondering what else is out there.
On a tangent, iirc pascal allowed single-pass compilation, so I wonder if we can embed compiler directly into inference, sort of constrained inference on steroids.
Or… the site will serve a random seed and the device must compute 4gb of pseudo-random data, then supply a value at a random server-demanded offset.
I should do that myself, come think of it.
I wonder if we need a list of things to tell AI to stay sharp, like this one. Sometimes I tell the model existing design is stupid and then it explains reasoning to me.
This won’t make the tech secure, but it will nullify models ability to breakout by making a controlled breakout first. Kinda like controlled forest burn.