HNHacker News
TopNewBestAskShowJobs

mangoman

1,609 karma · joined January 18, 2012

https://prashanth.world

meet.hn/city/us-Durham

Socials: - linkedin.com/in/prashanth-sadasivan - github.com/prashanthsadasivan

---

submissionscomments
mangoman··on The job ain't quite the same
I don't think it's the same as management. As a manager, you're not responsible for the code your team creates. You're responsible for their output, but that's fundamentally different. I also think you're managing people's emotions, expectations, holding people accountable, etc. Maybe those who enjoy working with agents would actually rather be PMs?

I agree the shift feels different than e.g. going from jquery to react or something.

mangoman··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
Based on the SynthID-Text paper https://www.nature.com/articles/s41586-024-08025-4 I agree that the LLM's learned distribution isn't modified, but I don't think it's correct to say that the sampling process is not modified. Also I just read the paper today so I could be misinterpreting things.

As described in the paper, you're right that it doesn't affect the main sampling technique, but what they do is they sample the distribution for 2^m samples, and then use Tournament sampling to choose the tokens among those 2^m samples, and the watermark key changes the scoring of the tournament options, using the watermark key as an input to the random generator that generates the scoring functions.

Then, to calculate the watermark, they take the text, and compute the mean g-values of the text, and a higher score means that it's more likely that it was sampled using the provided selection of tournament watermarking functions.

let's say you had some top P words: mango, banana, pineapple, guava, and you sampled 8 times, and got each one twice in the following order:

1. mango 2. banana 3. pineapple 4. guava 5. mango 6. banana 7. pineapple 8. guava

without tournament sampling, you'd truly see any of those come through. But in tournament sampling, you take those 8 options, create m scoring functions based on the pseudorandom generator, and score the 'tournament' by sampling the biased distribution you create from the g values. That does change the sampling from based purely on the LLM and entropy, but i mean, if the watermark key is also generated from some entropy, it's probably representative of the original sampling options as expected?

this is a very fascinating topic! I do still stand by my point that anthropic is the only one who can tell if something is watermarked or not and feeling icky, but the paper has mostly quelled my concern on impacting the intelligence part.

mangoman··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
huh, thanks for commenting this! I trusted the linked explanation post https://declaude.org/watermarking/ but actually reading the the synthid paper showed me that my understanding was wrong: https://www.nature.com/articles/s41586-024-08025-4

It is definitely blurrier whether you can say this approach changes the distribution then. By definition, it _has_ to change the probabilities of output tokens, but it's not totally clear that the pseudorandomly generated scoring functions does affect the learned distribution.

put another way, I think it's safer to do:

compute distribution -> sample -> watermark from sampled options

than it would be to do:

compute distribution -> watermark distribution -> sample

mangoman··on Anthropic's 'watermark' text adulteration in Claude is a perversion of writing
I agree, but thats also learned, and it is done to encourage specific responses for a task, different than applying a mask to the distribution based on a key.
mangoman··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
10% is a lot! And why would this slow down the cheating epidemic? There’s tons of ways like declaude etc to get around the check. Also Claude’s style is very distinctive (e.g. “load bearing”) AND has changed since 4.5 quite dramatically. I might not be able to tell on a specific response, but I can tell you that I went from canceling my ChatGPT plan in November, to now reaching for it first and considering cancelling Claude because its style is getting really groan-inducing. It’s weird because I thought ChatGPT was really annoying about a year ago, and now codex is my first choice.

Is the cheating epidemic so bad? I’m a little out of the loop there truthfully, what are the consequences of not being able to detect AI generated text in non academic settings? And in academic settings… maybe I am underestimating the challenge, but it does feel like the assignment and ways education happens needs to change?

As an aside, I’m not totally sure why this solution feels so icky to me. There’s something Orwellian about how the phrasing of a passage embeds hidden information that only Anthropic can see i guess

mangoman··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
I’m surprised by the comments here being so favorable to anthropic. The comments are right about there being no “best” token, and yeah Gruber may have an agenda here. But I think the fundamental principle is that this approach messes with the distribution in ways that deviate from the trained model. Take the “gray” and “overcast” choices. And lets say before applying synthid the percentages were 48% and 52%. Those percentages were learned from the training data and RL. To change those percentages to 45% and 55% in a non learned way makes it seem like training wasn’t important? Or more likely they dont have the data that shows the failure modes? also, don’t these choices compound the changes to the distribution in later sampling choices?

It is a bit of a mystery to say that “its okay to choose different tokens that we would have for watermarking bc people don’t notice” as though word choice doesn’t matter. If it doesn’t matter, doesn’t that mean that intelligence is more of a commodity than they would want it to be?

mangoman··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
What are you talking about? I understand OP has the perspective of writers, but let’s say you’re asking an LLM to explain a concept and it uses green words that are actually more difficult to understand. Or you ask for an analogy to explain a topic and the analogy doesn’t quite land because the description used “gray” instead of “overcast”.
mangoman··on Hello, me. It's been a while
I’ve been trying to start this “I don’t need any distraction for this task” challenge, but honestly it’s been REALLY tough for me. I’ve even noticed apprehension to start a particular chore like folding laundry if I don’t have something to distract me. Its like… because folding laundry takes so long, without a distraction it feels like I’m wasting my time, or just running out of time in the day?

Doing dishes, mowing the lawn, even just waiting for agents to finish, this need for distraction has really taken a hold of me, and it’s even getting in the way of me wanting to do things that I feel would motivate me again.

I can’t help but wonder if agentic development has exacerbated this problem for me. There are so many more times where I just have to wait for things now. Sometimes I can multitask, and work on two things at once, but often that ends with me doing both things with less craft. So I end up consuming content more, which feels addictive too… and I’m not even talking about TikTok or Instagram or Youtube Shorts (and i don’t judge those who do consume there). I mainly consume longer form Youtube, Podcasts or honestly, TV shows that I’ve seen multiple times before.

It’s like the agentic tools incentivize NOT thinking very hard, and so I do it less, making it more intimidating to start thinking hard about a problem because it feels like a waste of time. Which, paradoxically, is why I think I am consuming much more content than I used to, filling up space with intellectual junk food.

mangoman··on What happens if an entire class of workers loses faith in their careers
> The world does not wait with bated breath for product launches. A marketing campaign will not change the world. Almost no one will remember all the work we do.

It did at one time! I think that’s what’s so alluring initially about Workism. People waited with bated breath for the iPhone and iPad and its new versions. Those iPod commercials and the marketing campaigns for “rip. Mix. Burn.” did change the music world.

But the quote is also right. It doesn’t happen anymore, and I think that’s why the religion of Workism is especially vulnerable at this point.

In someways, it sorta feels like the article is cheering AI on? Like, it judges Workism, and goes on to say that AI shatters the illusion, but then.. that’s it? Deriving value from good work is pointless now and we should all just go volunteer?

I dunno, i clearly identify with those who loved work and now feel a bit lost, and yes i have been thinking about alternatives, but I feel like there must be knowledge work that is still valuable both financially and socially and also provides some sense of pride and purpose… if not, i fear the situation is much more dire than we are prepared for

mangoman··on Nvidia Cosmos 3

  This release unifies those capabilities with a Mixture-of-Transformers (MoT) architecture built around two towers. 
  Reasoner tower: A vision-language model (VLM) ... This serves as the ‘brain’ that reasons about the world before any generation happens.
  Generator tower: Generates future observations and action sequences. This tower uses a diffusion-based process to generate physics-aware video and action outputs that are conditioned on the reasoner tower’s understanding.
This sort of approach (and others i've seen like it) always appeal to my inner engineer, trying to optimize and balance tradeoffs between model architectures and combine two things to yield the best of both worlds

But based on my understanding of the Bitter Lesson (http://www.incompleteideas.net/IncIdeas/BitterLesson.html), this is precisely the wrong approach in the long term. I'm linking the actual text of the bitter lesson because I think it's misunderstood (or I just don't agree with how i've seen it used in discourse). Specifically:

  The bitter lesson is based on the historical observations that 1) AI researchers have often tried to build knowledge into their agents, 2) this always helps in the short term, and is personally satisfying to the researcher, but 3) in the long run it plateaus and even inhibits further progress, and 4) breakthrough progress eventually arrives by an opposing approach based on scaling computation by search and learning. The eventual success is tinged with bitterness, and often incompletely digested, because it is success over a favored, human-centric approach. 
This architecture feels specifically like "trying to build knowlege into the agent that will help in the short term" but will plateau long term. That's not to say that there won't be some interesting learnings or things built on top of it, but I doubt that there's a lot of juice to squeeze with this kind of approach IMO.
mangoman··on iPhone 17 Pro Demonstrated Running a 400B LLM
I dunno, I thought that too for a while too, but there are a lot of new ideas in terms of architecture that may warrant massive training runs. Mamba and state space models are pretty interesting, but haven’t had their transformer moment yet because I haven’t really seen anyone go for broke on training it with a huge data set and model size. Even some of the more fundamental changes too like Kolmogorov–Arnold Networks or some of the ideas behind continuous back propagation haven’t really had the opportunity to be pushed to the limit. I think it’s still early days on what these models can do. And I say this as someone who bought a Mac m3 max 128gb ram, based on the hope that the on device training and inference work would eventually move locally. It’s encouraging to see the progress though and I hope it does move locally though.
mangoman··on Minions: Stripe’s one-shot, end-to-end coding agents
There’s something off-putting about making a blog post about some splashy tech that’s is a fork of an open source project, and that tech not also being open source? It reads to me like “Hey, we thought the open source goose project was just okay, so we forked it to do it better. But we’re not going to contribute it back to and instead rename it.”

I think it probably wouldn’t be as weird if the project were a meaningfully different fork of it, but it sounds like it’s trying to accomplish the same goals as the open source project which I feel should probably be ported back? and renaming it seems sorta ungrateful? Kinda like that “you made this? I made this” meme. Maybe I just don’t have an understanding of how different the projects are though…

mangoman··on OpenClaw is what Apple intelligence should have been
I guess what’s wrong with it? Let’s say it has read only access, new messages and calendar invites need approval. I’m not sure I understand the harm? I suppose data exfiltration, but like you could start with an allowlist approach. So the first few uses and reads take a while with allowing the ai to read stuff , but it doesn’t seem that crazy given it’s what we basically do with ai coding tools?
mangoman··on ICEBlock handled my vulnerability report in the worst possible way
I’ve never built something like ICEBlock that puts me personally in the crosshairs of not just normal hacking attempts, but also the political will of the federal government. I can’t imagine the cess pool that is Joshua’s DMs. I think OP makes all the right assessments when examining how seriously ICEBlock is taking the risks here. The Android push notifications assertion is proof enough to make me raise a pretty big question, let alone the other issues raised.

Were I building something that I would want to assert the level of privacy claims that ICEBlock asserts, I would absolutely be taking any/all reports about security extremely seriously.

mangoman··on Ollama and gguf
no that's incorrect - llama.cpp has support for providing a context free grammar while sampling and only samples tokens that would conform to the grammar, rather than sampling tokens that would violate the grammar
mangoman··on US Trade Court finds Trump tariffs illegal
My bringing up it's age was mainly about it not being used in a LONG time, so using it now would seem like a hail mary. I made no mention about current day laws, so I'm not sure where your impression came from.
mangoman··on US Trade Court Finds Trump Tariffs Illegal
Right. this whole episode is about foreign policy, why not use the act that is meant for it?
mangoman··on US Trade Court Finds Trump Tariffs Illegal
...sigh... i mean i guess it really is just about saying they _can_ do it instead of actually _trying to do it_....
mangoman··on US Trade Court finds Trump tariffs illegal
I don't think it requires the courts to agree - just that there's a burden or disadvantage and that it's in the "public interest" which seems like a pretty low bar to make up a story that sounds plausible. i think the idea that a trade deficit is a disadvantage is kinda brain dead, but it's plausible sounding enough to argue in court. throw in unequal tariff rates and it seems like an easier win than the IEEPA's emergency justification.
mangoman··on US Trade Court finds Trump tariffs illegal
I'm not a lawyer or even close to it, but why wouldn't the trump admin use the tariff act of 1930? quote:

"Whenever the President shall find as a fact that any foreign country places any burden or disadvantage upon the commerce of the United States by any of the unequal impositions or discriminations aforesaid, he shall, when he finds that the public interest will be served thereby, by proclamation specify and declare such new or additional rate or rates of duty as he shall determine will offset such burden or disadvantage, not to exceed 50 per centum ad valorem or its equivalent, on any products of, or on articles imported in a vessel of, such foreign country"

it does cap it at 50%, but I mean it seems like a much easier way to justify the tariff. is there something else about it that isn't as practical (other than being almost 100 years old)

mangoman··on Chain of Recursive Thoughts: Make AI think harder by making it argue with itself
a paper with a similar idea on scaling test time reasoning, this is sorta how all the thinking models work under the hood. https://arxiv.org/abs/2501.19393
mangoman··on S1: A $6 R1 competitor?
From the S1 paper:

> Second, we develop budget forcing to control test-time compute by forcefully terminating the model's thinking process or lengthening it by appending "Wait" multiple times to the model's generation when it tries to end

I'm feeling proud of myself that I had the crux of the same idea almost 6 months ago before reasoning models came out (and a bit disappointed that I didn't take this idea further!). Basically during inference time, you have to choose the next token to sample. Usually people just try to sample the distribution using the same sampling rules at each step.... but you don't have to! you can selectively insert words into the the LLM's mouth based on what it said previously or what it wants to say, and decide "nah, say this instead". I wrote a library so that you could sample an LLM using llama.cpp in swift and you could write rules to sample tokens and force tokens into the sequence depending on what was sampled. https://github.com/prashanthsadasivan/LlamaKit/blob/main/Tes...

Here, I wrote a test that asks Phi-3 instruct "how are you" and it if it tried to say "as an AI I don't have feelings" or "I'm doing " I forced it to say "I'm doing poorly" and refuse to help since it was always so dang positive. It sorta worked, though the instruction tuned models REALLY want to help. But at the time I just didn't have a great use case for it - I had thought about a more conditional extension to llama.cpp's grammar sampling (you could imagine changing the grammar based on previously sampled text), or even just making it go down certain paths, but I just lost steam because I couldn't describe a killer use case for it.

This is that killer use case! forcing it to think more is such a great usecase for inserting ideas into the LLM's mouth, and I feel like there must be more to this idea to explore.

mangoman··on Trump wins presidency for second time
That would be missing the forest for the trees in my view. I could see it having an impact, but when 60% of people say that the country is headed in the wrong direction, putting up a candidate who was in power the last four years just isn’t going to work. Biden would not have won a primary, and neither would she have
mangoman··on 1374 Days – My Journey with Long Covid (2023)
on 2) are you referring to https://en.wikipedia.org/wiki/Dancing_plague_of_1518 ? I don't know very much of the history, but the veracity of the claims is specifically called out in the wikipedia page. Given the year, I'd wager that the truth is that they died of something else...

And that's a very strange reason to be skeptical of womens' description of their symptoms in a medical setting. Is there evidence that women are more prone to 'social contagion'? You yourself said that "women are more prone to auto immune diseases" and that they are known to be triggered by viruses...To me, skepticism (and in my view, cynicism) of patients is one of the major contributors to distrust of evidence based medical advice.

Taking your last anecdote as an example, just because the patient's self diagnosis is likely incorrect, that does not mean that the symptoms they're experiencing are all "made up" and should be met with skepticism. There may be other reasons they're experiencing symptoms. A doctor shouldn't just say "Because you thought it was long COVID even though you weren't diagnosed with COVID, I'm convinced you're making the whole thing up". That's just lazy and unsound.

mangoman··on Show HN: I've spent nearly 5y on a web app that creates 3D apartments
This is pretty neat! I’m on mobile right now, but you mentioned that it’s VR ready - does the landing page work with WebXR? I’d love to try it out on my meta quest 3
mangoman··on Elon Musk suing the mods of Cyberstuck for libel
Update, this is extremely well written satire. I’m laughing at myself for believing it but it also sounds so real.
mangoman··on Elon Musk suing the mods of Cyberstuck for libel
Not sure if this is satire or not… couldn’t find other sources yet but it’s relatively new
mangoman··on Running LLaVA on iOS with Llama.cpp and TinyLlama
Author here, happy to answer any questions as best as I can! A fun project that I hope will help others get models running on mobile devices.
mangoman··on FDA clears first over-the-counter continuous glucose monitor
I recently had an unusual health event that resulted in me passing out. My wife, who is a physician, thought it might be hypoglycemia, since i'm at high risk for diabetes. She found a super friendly endocrinologist who put me on a CGM for two weeks. I never hit the hypoglycemia range during those two weeks, so it didn't really explain what my issue... but honestly the data was SUPER interesting. Just observing the various spikes made me make healthier choices, or noticing when I was feeling extra tired and seeing if that correlated to not having eaten for little while, or eating something sugary before.

It's sort of like tracking your steps when you first get a smart watch. It may not have been the reason you got the device, but seeing the data, people are encouraged to act on it, even if you don't have an acute issue. since I didn't have a prescription, I couldn't get one here (didn't want to go through some sketch online site). I tried to get one from my family in India, but the prices were really high and they couldn't get the fancier one that tracks straight to your phone, so I didn't get one.

I think this could be a god send for preventing pre-diabetic people who would take preventative steps if it weren't such a pain in the ass to measure consistently.

mangoman··on I'm an Old Fart and AI Makes Me Sad
I understand the sadness around not understanding it, it's fucking hard. however, there are more and more resources online getting published for how to get started understanding it that help with understanding the math at an abstraction that helps with learning how to build with it.

I would strongly strongly strongly recommend starting with karpathy's from 0 to hero neural networks youtube course - it starts with building a tensor library and back propagation, explaining it in a way that finally clicked for me https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThs...

jeremy howard also has a fantastic video that is more around how to use LLMs and such called a hacker's guide to language models - https://www.youtube.com/watch?v=jkrNMKz9pWU&t=607s

as i've dove more and more into it, i would strongly recommend trying to run things on your local machine too (llama.cpp, ollama, LM studio). that has helped me fight that feeling of like "are we all just going to be open AI developers in the end" and made me feel like you _can_ integrate these things into stuff you build by your self. I can't imagine how fucked we'd all be if llama was never opensourced. being old does not mean that you can't continue to grow, and remember that it's okay to feel overwhelmed about all this - many people are.

Page 1 of 4Next →