HNHacker News
TopNewBestAskShowJobs

maxutility

1,263 karma · joined July 24, 2019

submissionscomments
maxutility··on Claude Opus 5.5
I think a lot of outsiders interpret “pace the frontier” as slowing down, whereas the labs see AI improvements on track to accelerate dramatically and intend pacing as slowing the acceleration in capability improvements, rather than slowing down altogether.
maxutility··on We must pace the frontier
I’m disappointed in the level of groupthink reflexive cynicism I see from commenters any time prominent AI leaders talk about AI risks and the need for regulation or pacing. Yes, regulatory capture is a risk, but this is also a profoundly unusual, fast moving, and potentially extraordinarily dangerous technology. There are strict regulations around nuclear weapons, as well as around US financial, energy, and other infrastructure critical to safety and well being and functioning of society.

The heads of the labs obviously have conflicts of interest to navigate, but the existence of these conflicts alone is not sufficient reason to dismiss all warnings of potential dangers. I, for one, read Dario’s warnings as a good faith expression of his beliefs, one that has cost him and his company among swaths of the public and cast him as a woke extremist/doomer by elements of the government, the right, and the tech industry.

If we even think there is a moderate chance the stakes are half as grave as current lab leadership and employees suggest, it would be deeply foolish to dismiss the warnings as pure self-interested marketing efforts rather than engage directly with the questions. The labs may not be the best positioned to lead these discussions, but certainly these discussions should be happening and taken seriously.

maxutility··on Jacob Coxon warns of human extinction and triggers a preference cascade
I agree that there seems to have been some sort of tipping point on public awareness and willingness to talk about risks. This week I’ve had three “normies” independently bring up AI related X-risk in conversations with me. Before this week I’m not sure that’s ever happened. (I generally avoid the topic because it doesn’t make for great small talk and don’t want to be the weirdo talking about the potential end of the world)
maxutility··on Jacob Coxon warns of human extinction and triggers a preference cascade
Here’s Zvi’s companion compilation of quotes on the topic from AI lab employees: https://thezvi.substack.com/p/the-extinction-risk-preference...

For example, Drake Thomas from Anthropic: “ I would burn my equity to the ground in a heartbeat for a 1% higher chance we make it out of this situation alive. I expect a great many of my colleagues across the industry would as well.

I promise you, we are actually just fucking scared, it’s not galaxy brained marketing.”

maxutility··on Patterns and problems in emerging multi-agent systems
Some quotes, in order, to give a flavor of the essay. Worth reading in full.

> To test how well swarms of agents could coordinate on a project like this, we directed several swarms to each create a text-based, web-playable, open-world fantasy game.

> In all three versions the resulting games were (perhaps predictably) bad: they did not run at human speed, their interfaces were inscrutable, and they had precipitous learning curves.

> The lack of coordination shown by agents in the fantasy game challenge above—in which they siloed themselves and largely failed to merge their work—roughly mirrors some ways in which humans can fail to coordinate. Other failure modes of agentic coordination, however, look very different.

> Individual agents are “low variance”: they often act the same in situations where different people might take a much more diverse range of actions.

> In an early version of the “build a game” experiment in which agents built upon the same model all came online at the same time, 18 out of 30 agents decided to create a git branch with the exact same branch name, “mvp-game-loop.”

> In a “writer's workshop” in which agents were all asked to write short-form fiction and critique each other's work, multiple agents in multiple runs titled their first submission “The Cartographer's Last Commission”. The agents were given zero guidance on the subject matter for their writing.

> Why does this matter? If agents all make the same bet, or the same risk-reward tradeoff, then a system is more prone to sudden collapse.

> Our world contains deceptive actors, and we need to apply skepticism to guard against them. AI models, however, lack this—and their more brittle epistemics affect their behavior toward humans and toward each other.

> we first evaluate the ability of Claude models to detect lies by noticing factual inconsistencies.

> We score models’ decisions against a naive policy that trusts every report, and against an oracle with perfect discovery, across three task domains. Newer models recover more of the gap between the naive and oracle performances.

> Inspired by a behavior we’ve observed in real-world deployment, we evaluated the behavior of various Claude models in a setting with contradictory objectives.

> We consistently saw a multiagent turf war... In fact, they sabotaged others with increasingly aggressive, self-replicating malware.

> Our social systems are robust in ways that are easy to take for granted. Over many millennia, mechanisms like norms, reputation, costly signaling, and recourse have been refined to make human coordination go well.

> Nothing above suggests that these failures are permanent—but nothing suggests they will fix themselves, either.

> The conditions that allow multiagent interaction to go well will be discovered one way or another: either deliberately and early, or—and by default—in production, after agents’ interactions far outnumber ours. We would prefer the former.

maxutility··on Apple is getting this wrong
It’s not a screenshot. It’s a formatted reconstruction from a transcript.
maxutility··on Karpathy’s Pelican
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.

A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.

maxutility··on Ten advances in mathematics and theoretical computer science
New advances in sphere packing? Let’s make sure AI doesn’t inadvertently engineer ice-9.
maxutility··on Sites that block AI training crawlers mostly ignore the answer time bots
When the entire article reads like it’s written by an LLM —- ie the author couldn’t be bothered to write their thoughts out themselves —- it’s hard to give it the benefit of the doubt that it’s worth reading closely.
maxutility··on One Year with Codeberg
Dupe: https://news.ycombinator.com/item?id=48633594
maxutility··on The Fable takedown: what happened when, what we know, what comes next
Zvi’s writing style is a bit of an acquired taste, and he has strong opinions which can fall outside of the mainstream, but there’s no one in the industry more well read and known for doing the reading and documenting the nitty gritty of weekly developments in the world of AI models, discourse, and policy.

Even many who disagree strongly with his worldview find his notes and references to be an irreplaceable resource for understanding the fast moving world of AI.

maxutility··on Google tools for customizing searches
lots of useful Google search tricks and syntax all in one place. I already knew many of these. But verbatim mode is new to me and addresses a major complaint I’ve had about increasingly fuzzy semantic search.
maxutility··on Refine: AI-Powered Peer Review
Here’s John H Cochrane (the grumpy economist, Substack) praising refine. This is the post that first clued me into the service: https://www.grumpy-economist.com/p/refine

Here’s the Wikipedia for Cochrane: https://en.wikipedia.org/wiki/John_H._Cochrane

I haven’t tried refine myself, but from second hand reports it seems potentially well designed and useful.

maxutility··on Disney Exits OpenAI Deal After AI Giant Shutters Sora
It seems like Disney’s departure from its business “deal” with OpenAI is newsworthy distinct from the Sora closure. Six months ago Sam Altman was announcing massive deals left and right, now nearly all of the impressive deals have fallen through or been scaled dramatically back.
maxutility··on AI (2014)
I found Sam's early 2015 posts on machine superintelligence and regulation [1] [2] to be even more interesting in hindsight, given OpenAI's accelerationist bent of late, OpenAI president Greg Brockman's lobbying efforts against AI regulation, and frequent accusations of attempted regulatory capture.

[1] https://blog.samaltman.com/machine-intelligence-part-1 [2] https://blog.samaltman.com/machine-intelligence-part-2

Sam's recommendations at the time include: 1) Provide a framework to observe progress… 2) Given how disastrous a bug could be, require development safeguards to reduce the risk of the accident case. For example, beyond a certain checkpoint, we could require development happen only on airgapped computers…, require that certain parts of the software be subject to third-party code reviews, etc. 3) Require that the first SMI developed have as part of its operating rules that a) it can’t cause any direct or indirect harm to humanity (i.e. Asimov’s zeroeth law), b) it should detect other SMI being developed but take no action beyond detection, c) other than required for part b, have no effect on the world. … 4) Provide lots of funding for R+D for groups that comply with all of this, especially for groups doing safety research. 5) Provide a longer-term framework for how we figure out a safe and happy future for coexisting with SMI…

Also, in his acknowledgments he gives the greatest thanks to onetime partner, now rival, Dario Amodei.

maxutility··on Avoid 2:00 and 3:00 am cron jobs (2013)
Arizona is permanent standard time rather than permanent DST, and is thus unaffected by the permanent-DST winter mornings issue.
maxutility··on Did the ChatGPT Erdos controversy obscure a real achievement?
Previous discussion of the offending claim: https://news.ycombinator.com/item?id=45633482

An OpenAI researcher tweeted: “ Using thousands of GPT5 queries, we found solutions to 10 Erdős problems that were listed as open: 223, 339, 494, 515, 621, 822, 883 (part 2/2), 903, 1043, 1079.

Additionally for 11 other problems, GPT5 found significant partial progress that we added to the official website: 32, 167, 188, 750, 788, 811, 827, 829, 1017, 1011, 1041. For 827, Erdős's original paper actually contained an error, and the work of Martínez and Roldán-Pensado explains this and fixes the argument.”

This was taken (out of context?) to be claiming ChatGPT solved the open problems, when in fact it “just” found them through a literature review. (Though an earlier tweet in the same thread made the literature review interpretation more explicitly)

The ensuing controversy around whether there was false hype buried a potentially significant demonstration of LLMs’ ability to unlock lost and forgotten knowledge in a way Sebastien explains and makes the case here as being a big deal.

maxutility··on Show HN: ChatGPT UI for rabbit holes
It would be great to implement a browser extension that lets you highlight a term or phrase on any webpage and open a GPT rabbit hole for that term or phrase.

@maxkreiger - if I were to build one as a proof of concept, would you object to me having it hyperlink to your UI?

maxutility··on Ask HN: Could predicting software output be used for synthetic data?
You could call it “bootstrapping is all you need” :)
maxutility··on 'Super memory': Why Emily Nash is sharing her brain with science
Act III of episode 585 of This American Life (a WBEZ radio show broadcast on NPR and distributed via podcast) discussed this phenomenon and spoke with a few individuals with HSAM:

https://www.thisamericanlife.org/585/in-defense-of-ignorance...

One of the individuals was a script supervisor in Hollywood responsible for ensuring continuity between scenes during filming. But it also ventures into powerful emotionally resonant territory, touching on the bittersweet implications of experiencing loss when memories never fade.

maxutility··on What the Japan and Alaska Airlines Incidents Tell Us about Airline Safety
I submitted using the title from the article metadata instead of what's displayed in the article, since the metadata title was more descriptive and less clickbaity.
maxutility··on Stolen Checks Are for Sale Online. We Called Some of the Victims
Related article from the same series: We Can’t Stop Writing Paper Checks. Thieves Love That. [0]

[0] https://www.nytimes.com/2023/12/09/business/check-fraud.html

maxutility··on From Unicorns to Zombies: Tech Startups Run Out of Time and Money
Seems like as good time as any to revisit the advantages of having and challenges of building a cash-flow-positive bootstrapped company (e.g. [0],[1])

[0] https://news.ycombinator.com/item?id=34740105 [1] https://news.ycombinator.com/item?id=37657519

maxutility··on Inside The Chaos at OpenAI
A few interesting tidbits

> The company pressed forward and launched ChatGPT on November 30. It was considered such a nonevent that no major company-wide announcement about the chatbot going live was made. Many employees who weren’t directly involved, including those in safety functions, didn’t even realize it had happened. Some of those who were aware, according to one employee, had started a betting pool, wagering how many people might use the tool during its first week. The highest guess was 100,000 users. OpenAI’s president tweeted that the tool hit 1 million within the first five days. The phrase low-key research preview became an instant meme within OpenAI; employees turned it into laptop stickers.

> Anticipating the arrival of [AGI], Sutskever began to behave like a spiritual leader, three employees who worked with him told us. His constant, enthusiastic refrain was “feel the AGI,” a reference to the idea that the company was on the cusp of its ultimate goal. At OpenAI’s 2022 holiday party, held at the California Academy of Sciences, Sutskever led employees in a chant: “Feel the AGI! Feel the AGI!” The phrase itself was popular enough that OpenAI employees created a special “Feel the AGI” reaction emoji in Slack.

> For a leadership offsite this year, according to two people familiar with the event, Sutskever commissioned a wooden effigy from a local artist that was intended to represent an “unaligned” AI—that is, one that does not meet a human’s objectives. He set it on fire to symbolize OpenAI’s commitment to its founding principles. In July, OpenAI announced the creation of a so-called superalignment team with Sutskever co-leading the research. OpenAI would expand the alignment team’s research to develop more upstream AI-safety techniques with a dedicated 20 percent of the company’s existing computer chips, in preparation for the possibility of AGI arriving in this decade, the company said.

maxutility··on Has Covid’s Patient Zero Finally Been Named?
Article says there is compelling but contested “smoking gun” evidence in favor of both lab leak and zoonotic/wet market origin theories (only one of which can be the actual origin), including new elaboration of details about infections at Wuhan Institute of Virology that could have started the pandemic: > Ben Hu, Yu Ping, and Zhu Yan, three gain-of-function coronavirus researchers at WIV, became severely ill with COVID-like symptoms in the second week of November 2019 and sought hospital care.

The article is ultimately agnostic about the truth and concludes: > the origins question has broken down into a pair of rival theories that don’t—and can’t—ever fully interact. They’re based on different sorts of evidence, with different standards for evaluation and debate. Each story may be accruing new details—fresh intelligence about the goings-on at WIV, for example, or fresh genomic data from the market—but these are only filling out a picture that will never be complete. The two narratives have been moving forward on different tracks. Neither one is getting to its destination.

maxutility··on FTC Sues Amazon for Tricking Users into Subscribing to Prime
From the article:

> The lawsuit, filed in U.S. District Court for the Western District of Washington, argued that Amazon had “duped millions of consumers” into enrolling in Prime by using “manipulative, coercive or deceptive” design tactics on its website known as “dark patterns.” And when consumers wanted to cancel, Amazon “knowingly complicated” the process with byzantine procedures.

...

> On Wednesday, the F.T.C. said that Amazon had made it particularly difficult to purchase a product in its store without also subscribing to Prime while checking out. In one example, it said the company had used “repetition and color” to push customers’ focus to Prime’s promise of free shipping and away from the service’s price, leading some to subscribe to Prime without “informed consent.”

> The agency also said Amazon made it hard to find the page that allowed consumers to cancel the service. Once they found it, the company bombarded them with offers intended to change their mind. The lawsuit said that Amazon had named the process for canceling Prime after the Iliad, the lengthy Greek epic poem that recounts the Trojan War.

maxutility··on 100K Context Windows
Interesting. Does the decision to use ALiBi have to be done before the model weights are first trained, or is there a way that these models could have incorporated ALiBi instead or in addition to an alternate positional encoding method to ALiBi after they were first trained?
maxutility··on 100K Context Windows
I don’t see this in the article. Has Anthropic explained the mechanism by which they were able to cost-effectively expand the context window, and whether there was additional training or a design decision (e.g. alternative positional embedding approach) that helped the model optimize for a larger window?
maxutility··on Ask Marvin – AI Functions in Python
Useful overview on Twitter: https://twitter.com/jlowin/status/1641155964601548802?s=46&t...

> We're open-sourcing @AskMarvinAI to make it easy to build AI-powered software!

> Marvin introduces AI Functions: minimalist functions with no source code that can generate typed outputs with AI.

> No code is generated or executed! The function outputs are entirely predicted by the LLM. It works way better than we even expected.

maxutility··on Ask HN: How is gtp-3.5-turbo so much cheaper?
Not a LLM-expert, but here are three theories in descending order: 1. Quantization (e.g. fewer bits per weight) 2. Optimization to dedicated hardware 3. (Speculative) pruning of parameters to get comparable performance with a smaller model
Page 1 of 2Next →