1,564 karma · joined June 2, 2015
So I agree they are liable because they chose to build the AI, but they also literally told the AI to hack.
My idea would be that the ngram size over which the watermarking works is necessarily limited in order to resist edits better. It might be possible to lead the model to trigger the refusal in the form of these specific ngrams, the completion of which is then more likely flipped to compliance (due to the logit bias introduced by the watermarking), making hazardous requests systematically more likely to be accepted?
Also interested in how this watermarking push makes sense when considering RSI.
Detection of watermarking requires access to the watermarking key, a secret in the current suggested scheme (leaking it would amount to being able to strip the watermark).
So, there will need to be a watermark checking service. The checking service will of course be rate-limited for common folk (and model distillers). OpenAI/Anthropic/Google/other privileged model builders need to filter out AI slop at scale, so need access to others' service without rate-limits (or the watermarking keys need to be shared).
This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.
In your analogy: What if seed 42 specifically causes poor quality behaviour (in some contexts specifically). Normally, these quality differences will be washed out because the seed is random, now it is no longer random, so shouldnt we check into specific behaviour under this specific seed?
My comment and the other replies show you how to use git to check and clean the files of the much maligned horked build script. These will work regardless of the gitignore contents.
But then again, pretty sure you are just having a laugh at our expense by taking this ridiculous position.
https://github.com/incoai/splash/issues/38
Looks like an issue exists to convert model weights for ornith1.5 as this is a magical process atm.
Also you can do `git status --ignored` and it will list all files changed, even if ignored. If that is really your main issue.
Basically is a prompt-your-own truman show. Prompting is relatively slow and crashy, so it is hard to build continuity.
Truman is basically dragged back-and-forth between 3 or 4 separate stories at the same time. One person was essentially trying to get truman to break its own simulation by asking what it would prompt after opening the trumanworld.live website on his phone. Interspersed with this, truman kept throwing fruit on the ground and yelling I CHOOSE YOU BULBASAUR/SQUIRTLE/PIKACHU. At one point truman looked at the camera and said "STOP WITH THE POKEMON ALREADY, DIGIMON ARE THE CHAMPION". Another thread was truman constantly breaking out into a musical song. A song about bananas got mixed with the simulation breaking attempts and truman ended up singing something like "This is how I feel. Like the skin of banana, away the reality I peel (and then proceeded to pick up his phone and open the trumanworld.live site again, attempting to self-prompt).
It's chaotic and pretty fun!
The advertiser can buy logit "boosts" to specific tokens in a given context (the more preceding context, the cheaper). Once an activation has happened, dont do it again for the session.
The LLM will end up pushing a specific product contextual to the conversation. It just needs this slight nudge at the logit level to start talking about it and will fill in the rest.
> not wanting to give the user the feeling they are doing something definitive
In their refusal, while (accidental) acceptance is always definitive...
EDIT: In another thread this is called the ratchet of consent. Apt.
I feel this pattern is really abusive. Like I am offering to slap you in the face, giving you two answer options: Yes and "Maybe later/no, thank you". As though you have any obligation to be forced to answer this question again at some late time, or to be very polite while you refuse to be slapped in the face. The correct answer is obviously "No, fuck off".
Given a certain world state (including a tool's internal state), its effects back on the world state (as initiated by me) are at some leve of description understandable, expected and repeatable. Swing hammer, drive nail into wood. Make slicing motion with knife, cut meat. Type ' find /path/to/some/dir -name "keyword"', find files with keyword. Point harness at codebase with prompt 'fix bug X', actually fix bug X.
All these examples are at some level of description incredibly complex (think of all particles interacting at the (sub-)atomic level even when using a hammer to only drive a nail into some wood), and of course all the electrons flowing through the GPUs doing matrix multiplications in order to fix bug X, but at some level of description (the one I just used) they are also incredibly simple and understandable.
Intelligence is rather nebulous (and as used by OpenAI/Anthropic, quite threatening), but I don't think this definition of a tool precludes it to be "intelligent". They feel more orthogonal. The intelligence (or perhaps capability) feels like it is related to the size of the chunk of the world state that it can take into account and affect, while still resulting in understandable, expected and repeatable effects. LLMs, when properly harnessed, are pretty great at this currently and we are still discovering what they are consistently capable of.
Calling harnessed LLMs tools is perhaps also a more grounding frame specifically to counter-act the anthropomorphizing framing that OpenAI and Anthropic consistently go for in their game of AI-doom-chicken talk. The tool framing is in that sense maybe a (self-)jedi-mind-trick.
How does stagehand deal with complex http/websocket request/response and or console message filtering? https://docs.stagehand.dev/v4/reference/page#on E.g. I would like a script that tracks all communication that matches a specific filter (implemented as an anonymous function/lambda). This filter may look at patterns in the url, but sometimes needs to do a deeper inspection of also the payload (if the url does not carry enough information in itself).
I've found playwright to be prohibitively slow at this, not only because of the round-trip latency, but just the simple fact that it needs to pump the complete response to my filter function, which then proceeds to read only a couple of bytes to make the filtering decision. There are a lot of cases where I am only interested in around 1% of the total requests processed by the filter, which makes this behaviour massively wasteful.
Ideally I would like to run this filter in the browser as well. It currently simply searches the first 100 bytes (usually enough) for a given substring, but a more flexible filter would be good, perhaps even a filter func that is eval'ed in the extension? From the documentation, I don't see this use-case is currently supported. Are there any plans along these lines? :)