HNHacker News
TopNewBestAskShowJobs

alextheparrot

2,198 karma · joined June 22, 2016

submissionscomments
alextheparrot··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
https://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_green...
alextheparrot··on Confidential submission of draft S-1 to the SEC
The for-profit is a PBC with the sane mission at the nonprofit [0]

[0] https://openai.com/index/built-to-benefit-everyone/

alextheparrot··on ChatGPT Images 2.0
The paper they published last year goes over some of these transformations: https://arxiv.org/pdf/2510.09263
alextheparrot··on ChatGPT Images 2.0
> Integrating an imperceptible, robust, and content-specific watermark

From the system card someone linked elsewhere in the discussion

alextheparrot··on Ask ChatGPT to pick a number from 1-10000, it generally selects from 7200-7500
No LLMs are calibrated?
alextheparrot··on Court orders restart of all US offshore wind power construction
I mean, you could also frame this as an issue the electorate could actually prioritize instead of just hoping the courts work it out
alextheparrot··on Recursive Language Models
The derivative being a grad(ient) student sampling scaffolds against evals + qualitative observations: most prompt-based llm papers
alextheparrot··on New research reveals longevity gains slowing, life expectancy of 100 unlikely
Cancer is a parasitism that kills the host (or the host dies from other causes and it is not self-sufficient). Just because something is defined by uncontrolled self-replication doesn’t mean it is stable to live forever (Which is as much a comment on homeostasis as self-renewal)
alextheparrot··on AI agent benchmarks are broken
It isn’t actually very wrong. Your example is tangential as graders in school have multiple roles — teaching the content and grading. That’s an implementation detail, not a counter to the premise.

I don’t think we should assume answering a test would be easy for a Scantron machine just because it is very good at grading them, either.

alextheparrot··on AI agent benchmarks are broken
> Good evaluations write test sets for the discriminators to show when this is or isn’t true.

If they can’t write an evaluation for the discriminator I agree. All the input data issues you highlight also apply to generators.

alextheparrot··on AI agent benchmarks are broken
I wish the other replies and this would engage with the sentence right after it indicating that you should test this premise empirically.
alextheparrot··on AI agent benchmarks are broken
If this sort of error isn’t acceptable, it should be part of an evaluation set for your discriminator

Fundamentally I’m not disagreeing with the article, but also think most people who care take the above approach because if you do care you read samples, find the issues, and patch them to hill climb better

alextheparrot··on AI agent benchmarks are broken
LLMs evaluating LLM outputs really isn’t that dire…

Discriminating good answers is easier than generating them. Good evaluations write test sets for the discriminators to show when this is or isn’t true. Evaluating the outputs as the user might see them are more representative than having your generator do multiple tasks (e.g. solve a math query and format the output as a multiple choice answer).

Also, human labels are good but have problems of their own, it isn’t like by using a “different intelligence architecture” we elide all the possible errors. Good instructions to the evaluation model often translate directly to better human results, showing a correlation between these two sources of sampling intelligence.

alextheparrot··on How we’re responding to The NYT’s data demands in order to protect user privacy
in the app: Settings ~> Data Controls ~> Improve the model for everyone
alextheparrot··on Retailers will soon have only about 7 weeks of full inventories left
That’s a premise that would make me consider the wiseness of my actions.
alextheparrot··on Lawmakers are skeptical of Zuckerberg's commitment to free speech
Quippy, but off the cuff: - I don’t go to my present town square(s) socially because it is full of a-social behavior. Same reason to avoid certain bars or clubs, prefer certain parks, or why some are wary of public transit.

- I don’t feel a right to decide the vibe of how a business curates its space. My bakery, coffee shop, local library, etc. all curate a space with an opinion. I don’t feel I have standing to assert that my preferences should dominate their choices.

As an aside, businesses are also an extension of the people, the best ones tend to just not be mode collapsed

alextheparrot··on The hacking of culture and the creation of socio-technical debt
Really enjoyed the piece.

A passing thought: the ethe of individuals in the 70s and 80s is important because of the people it informed in subsequent years. While many people still like to hack, code, etc., the relative proportion of people doing this and working in tech continues to diminish as the popularity and importance of the sector grows. I wonder if debt without values / a more cohered zeitgeist is better or worse?

alextheparrot··on Safe Superintelligence Inc.
Glibly, I’d also love your definition of the education system writ large.
alextheparrot··on OpenAI and Apple Announce Partnership
Bit of a detail, but where are you deriving “with hundreds of terabytes of unified GPU memory” from?
alextheparrot··on σ-GPTs: A new approach to autoregressive models
No, but it makes more conceptual sense given the model can consider what was said before it
alextheparrot··on I should have loved biology (2020)
I love the romance of this piece, but in my experience he’s just describing the difference in expectations of learning biology at a high school versus advanced undergraduate to graduate level.

Romance is for those who care, and most don’t. But it is so, so beautiful once you do.

alextheparrot··on Gemma: New Open Models
> “Our best shot at making the quarter is if we get an injection of at least [redacted]% , queries ASAP from Chrome.” (Google Exec)

Isn’t there a whole anti-trust case going on around this?

[0] https://www.nytimes.com/interactive/2023/10/24/business/goog...

alextheparrot··on A decoder-only foundation model for time-series forecasting
You could probably consider learning a sign wave to be a “grammar” related to periodic variations. Grammar in this context feels like “What are the core conceptual heuristics that help guide towards faster understanding”
alextheparrot··on Multi-database support in DuckDB
I'd assume they mean users interacting with the chart vs first load. So the user sees the base chart (Let's say 1MB of data on the server, less depending what gets pushed to the user) and then additional filters, aggregations, etc. are pretty cheap because the server has a local copy to query against
alextheparrot··on Multi-database support in DuckDB
Exactly what I want from DuckDB. Was playing with using it as a quicker cache for a data app backed in Snowflake, wonder how hard it’d be to write the attach for that vs doing it at the client level
alextheparrot··on TextDiffuser-2: Unleashing the power of language models for text rendering
The comment was directed at “doesn't this method add another cost and overhead for calling Text-to-Image models”
alextheparrot··on TextDiffuser-2: Unleashing the power of language models for text rendering
LLM + Text-to-Image model is exactly how DALL·E 3 is deployed, fwiw
alextheparrot··on We have reached an agreement in principle for Sam to return to OpenAI as CEO
Users here often get the narrative and motivations deeply wrong, I wouldn’t take it too personally (Speaking as a peer)
alextheparrot··on Is an All-Meat Diet What Nature Intended?
Not to address the applicability of this, but you can have advantageous fallback behavior and still say “the system isn’t intended to be in the fallback behavior”.
alextheparrot··on Run LLMs at home, BitTorrent‑style
Exactly, litigation has never been applied to content delivered over BitTorrent-style networks
Page 1 of 20Next →