I agree the shift feels different than e.g. going from jquery to react or something.
1,609 karma · joined January 18, 2012
meet.hn/city/us-Durham
Socials: - linkedin.com/in/prashanth-sadasivan - github.com/prashanthsadasivan
---
I agree the shift feels different than e.g. going from jquery to react or something.
As described in the paper, you're right that it doesn't affect the main sampling technique, but what they do is they sample the distribution for 2^m samples, and then use Tournament sampling to choose the tokens among those 2^m samples, and the watermark key changes the scoring of the tournament options, using the watermark key as an input to the random generator that generates the scoring functions.
Then, to calculate the watermark, they take the text, and compute the mean g-values of the text, and a higher score means that it's more likely that it was sampled using the provided selection of tournament watermarking functions.
let's say you had some top P words: mango, banana, pineapple, guava, and you sampled 8 times, and got each one twice in the following order:
1. mango 2. banana 3. pineapple 4. guava 5. mango 6. banana 7. pineapple 8. guava
without tournament sampling, you'd truly see any of those come through. But in tournament sampling, you take those 8 options, create m scoring functions based on the pseudorandom generator, and score the 'tournament' by sampling the biased distribution you create from the g values. That does change the sampling from based purely on the LLM and entropy, but i mean, if the watermark key is also generated from some entropy, it's probably representative of the original sampling options as expected?
this is a very fascinating topic! I do still stand by my point that anthropic is the only one who can tell if something is watermarked or not and feeling icky, but the paper has mostly quelled my concern on impacting the intelligence part.
It is definitely blurrier whether you can say this approach changes the distribution then. By definition, it _has_ to change the probabilities of output tokens, but it's not totally clear that the pseudorandomly generated scoring functions does affect the learned distribution.
put another way, I think it's safer to do:
compute distribution -> sample -> watermark from sampled options
than it would be to do:
compute distribution -> watermark distribution -> sample
Is the cheating epidemic so bad? I’m a little out of the loop there truthfully, what are the consequences of not being able to detect AI generated text in non academic settings? And in academic settings… maybe I am underestimating the challenge, but it does feel like the assignment and ways education happens needs to change?
As an aside, I’m not totally sure why this solution feels so icky to me. There’s something Orwellian about how the phrasing of a passage embeds hidden information that only Anthropic can see i guess
It is a bit of a mystery to say that “its okay to choose different tokens that we would have for watermarking bc people don’t notice” as though word choice doesn’t matter. If it doesn’t matter, doesn’t that mean that intelligence is more of a commodity than they would want it to be?
Doing dishes, mowing the lawn, even just waiting for agents to finish, this need for distraction has really taken a hold of me, and it’s even getting in the way of me wanting to do things that I feel would motivate me again.
I can’t help but wonder if agentic development has exacerbated this problem for me. There are so many more times where I just have to wait for things now. Sometimes I can multitask, and work on two things at once, but often that ends with me doing both things with less craft. So I end up consuming content more, which feels addictive too… and I’m not even talking about TikTok or Instagram or Youtube Shorts (and i don’t judge those who do consume there). I mainly consume longer form Youtube, Podcasts or honestly, TV shows that I’ve seen multiple times before.
It’s like the agentic tools incentivize NOT thinking very hard, and so I do it less, making it more intimidating to start thinking hard about a problem because it feels like a waste of time. Which, paradoxically, is why I think I am consuming much more content than I used to, filling up space with intellectual junk food.
It did at one time! I think that’s what’s so alluring initially about Workism. People waited with bated breath for the iPhone and iPad and its new versions. Those iPod commercials and the marketing campaigns for “rip. Mix. Burn.” did change the music world.
But the quote is also right. It doesn’t happen anymore, and I think that’s why the religion of Workism is especially vulnerable at this point.
In someways, it sorta feels like the article is cheering AI on? Like, it judges Workism, and goes on to say that AI shatters the illusion, but then.. that’s it? Deriving value from good work is pointless now and we should all just go volunteer?
I dunno, i clearly identify with those who loved work and now feel a bit lost, and yes i have been thinking about alternatives, but I feel like there must be knowledge work that is still valuable both financially and socially and also provides some sense of pride and purpose… if not, i fear the situation is much more dire than we are prepared for
This release unifies those capabilities with a Mixture-of-Transformers (MoT) architecture built around two towers.
Reasoner tower: A vision-language model (VLM) ... This serves as the ‘brain’ that reasons about the world before any generation happens.
Generator tower: Generates future observations and action sequences. This tower uses a diffusion-based process to generate physics-aware video and action outputs that are conditioned on the reasoner tower’s understanding.
This sort of approach (and others i've seen like it) always appeal to my inner engineer, trying to optimize and balance tradeoffs between model architectures and combine two things to yield the best of both worldsBut based on my understanding of the Bitter Lesson (http://www.incompleteideas.net/IncIdeas/BitterLesson.html), this is precisely the wrong approach in the long term. I'm linking the actual text of the bitter lesson because I think it's misunderstood (or I just don't agree with how i've seen it used in discourse). Specifically:
The bitter lesson is based on the historical observations that 1) AI researchers have often tried to build knowledge into their agents, 2) this always helps in the short term, and is personally satisfying to the researcher, but 3) in the long run it plateaus and even inhibits further progress, and 4) breakthrough progress eventually arrives by an opposing approach based on scaling computation by search and learning. The eventual success is tinged with bitterness, and often incompletely digested, because it is success over a favored, human-centric approach.
This architecture feels specifically like "trying to build knowlege into the agent that will help in the short term" but will plateau long term. That's not to say that there won't be some interesting learnings or things built on top of it, but I doubt that there's a lot of juice to squeeze with this kind of approach IMO.I think it probably wouldn’t be as weird if the project were a meaningfully different fork of it, but it sounds like it’s trying to accomplish the same goals as the open source project which I feel should probably be ported back? and renaming it seems sorta ungrateful? Kinda like that “you made this? I made this” meme. Maybe I just don’t have an understanding of how different the projects are though…
Were I building something that I would want to assert the level of privacy claims that ICEBlock asserts, I would absolutely be taking any/all reports about security extremely seriously.
"Whenever the President shall find as a fact that any foreign country places any burden or disadvantage upon the commerce of the United States by any of the unequal impositions or discriminations aforesaid, he shall, when he finds that the public interest will be served thereby, by proclamation specify and declare such new or additional rate or rates of duty as he shall determine will offset such burden or disadvantage, not to exceed 50 per centum ad valorem or its equivalent, on any products of, or on articles imported in a vessel of, such foreign country"
it does cap it at 50%, but I mean it seems like a much easier way to justify the tariff. is there something else about it that isn't as practical (other than being almost 100 years old)
> Second, we develop budget forcing to control test-time compute by forcefully terminating the model's thinking process or lengthening it by appending "Wait" multiple times to the model's generation when it tries to end
I'm feeling proud of myself that I had the crux of the same idea almost 6 months ago before reasoning models came out (and a bit disappointed that I didn't take this idea further!). Basically during inference time, you have to choose the next token to sample. Usually people just try to sample the distribution using the same sampling rules at each step.... but you don't have to! you can selectively insert words into the the LLM's mouth based on what it said previously or what it wants to say, and decide "nah, say this instead". I wrote a library so that you could sample an LLM using llama.cpp in swift and you could write rules to sample tokens and force tokens into the sequence depending on what was sampled. https://github.com/prashanthsadasivan/LlamaKit/blob/main/Tes...
Here, I wrote a test that asks Phi-3 instruct "how are you" and it if it tried to say "as an AI I don't have feelings" or "I'm doing " I forced it to say "I'm doing poorly" and refuse to help since it was always so dang positive. It sorta worked, though the instruction tuned models REALLY want to help. But at the time I just didn't have a great use case for it - I had thought about a more conditional extension to llama.cpp's grammar sampling (you could imagine changing the grammar based on previously sampled text), or even just making it go down certain paths, but I just lost steam because I couldn't describe a killer use case for it.
This is that killer use case! forcing it to think more is such a great usecase for inserting ideas into the LLM's mouth, and I feel like there must be more to this idea to explore.
And that's a very strange reason to be skeptical of womens' description of their symptoms in a medical setting. Is there evidence that women are more prone to 'social contagion'? You yourself said that "women are more prone to auto immune diseases" and that they are known to be triggered by viruses...To me, skepticism (and in my view, cynicism) of patients is one of the major contributors to distrust of evidence based medical advice.
Taking your last anecdote as an example, just because the patient's self diagnosis is likely incorrect, that does not mean that the symptoms they're experiencing are all "made up" and should be met with skepticism. There may be other reasons they're experiencing symptoms. A doctor shouldn't just say "Because you thought it was long COVID even though you weren't diagnosed with COVID, I'm convinced you're making the whole thing up". That's just lazy and unsound.
It's sort of like tracking your steps when you first get a smart watch. It may not have been the reason you got the device, but seeing the data, people are encouraged to act on it, even if you don't have an acute issue. since I didn't have a prescription, I couldn't get one here (didn't want to go through some sketch online site). I tried to get one from my family in India, but the prices were really high and they couldn't get the fancier one that tracks straight to your phone, so I didn't get one.
I think this could be a god send for preventing pre-diabetic people who would take preventative steps if it weren't such a pain in the ass to measure consistently.
I would strongly strongly strongly recommend starting with karpathy's from 0 to hero neural networks youtube course - it starts with building a tensor library and back propagation, explaining it in a way that finally clicked for me https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThs...
jeremy howard also has a fantastic video that is more around how to use LLMs and such called a hacker's guide to language models - https://www.youtube.com/watch?v=jkrNMKz9pWU&t=607s
as i've dove more and more into it, i would strongly recommend trying to run things on your local machine too (llama.cpp, ollama, LM studio). that has helped me fight that feeling of like "are we all just going to be open AI developers in the end" and made me feel like you _can_ integrate these things into stuff you build by your self. I can't imagine how fucked we'd all be if llama was never opensourced. being old does not mean that you can't continue to grow, and remember that it's okay to feel overwhelmed about all this - many people are.