HNHacker News
TopNewBestAskShowJobs

happypumpkin

212 karma · joined February 17, 2023

submissionscomments
happypumpkin··on Differential Transformer
> it would probably become obsolete soon

Suppose there are many times more posts about something one generation of LLMs can't do (arithmetic, tic-tac-toe, whatever), than posts about how the next generation of models can do that task successfully. I think this is probably the case.

While I doubt it will happen, it would be somewhat funny if training on that text caused a future model to claim it can't do something that it "should" be able to because it internalized that it was an LLM and "LLMs can't do X."

happypumpkin··on Notes on OpenAI's new o1 chain-of-thought models
The new model does play very well but when it draws the board it frequently places the moves in incorrect locations (but seemingly still keeps track of the correct ones). But I can't fault it too much, I don't think what is essentially ASCII art is intended to be a strength of the model.

Edit: Actually third game with it led to it making an illegal move, and claiming a draw (which would've been inevitable given optimal play for the rest of the game but there were several valid moves left to make).

happypumpkin··on The AI Scientist: Towards Automated Open-Ended Scientific Discovery
Also a concern about the paper generation process itself:

> In a similar vein to idea generation, The AI Scientist is allowed 20 rounds to poll the Semantic Scholar API looking for the most relevant sources to compare and contrast the near-completed paper against for the related work section. This process also allows The AI Scientist to select any papers it would like to discuss and additionally fill in any citations that are missing from other sections of the paper.

So... they don't look for related work until the paper is "near-completed." Seems a bit backwards to me.

happypumpkin··on The AI Scientist: Towards Automated Open-Ended Scientific Discovery
Potential concerns with their self-eval:

They evaluate their automated reviewer by comparing against human evaluations on human-written research papers, and then seem to extrapolate that their automated reviewer would align with human reviewers on AI-written research papers. It seems like there are a few major pitfalls with this.

First, if their systems aren't multimodal, and their figures are lower-quality than human-created figures (which they explicitly list as a limitation), the automated reviewer would be biased in favor of AI-generated papers (only having access to the text). This is an obvious one but I think there could easily be other aspects of papers where the AI and human reviewers align on human-written papers, but not on AI papers.

Additionally, they note:

> Furthermore, the False Negative Rate (FNR) is much lower than the human baseline (0.39 vs. 0.52). Hence, the LLM-based review agent rejects fewer high-quality papers. The False Positive Rate (FNR [sic]), on the other hand, is higher (0.31 vs. 0.17)

It seems like false positive rate is the more important metric here. If a paper is truly high-quality, it is likely to have success w/ a rebuttal, or in getting acceptance at another conference. On the other hand, if this system leads to more low-quality submissions or acceptances via a high FPR, we're going to have more AI slop and increased load on human reviewers.

I admit I didn't thoroughly read all 185 pages, maybe these concerns are misplaced.

happypumpkin··on Testing Generative AI for Circuit Board Design
> Artists and "creative" people have long held a monopoly on this ability and are now finally paying the price

I've seen a lot of schadenfreude towards artists recently, as if they're somehow gatekeeping art and stopping the rest of us from practicing it.

I really struggle to understand it; the barrier of entry to art is basically just buying a paper and pencil and making time to practice. For most people the practice time could be spent on many things which would have better economic outcomes.

> monopoly

Doesn't this term imply an absence of competition? There seems to be a lot of competition. Anyone can be an artist, and anyone can attempt to make a living doing art. There is no certification, no educational requirements. I'm sure proximity to wealth is helpful but this is true of approximately every career or hobby.

Tangentially, there seem to be positive social benefits to everyone having different skills and depending on other people to get things done. It makes me feel good when people call me up asking for help with something I'm good at. I'm sure it feels the same for the neighborhood handyman when they fix someone's sink, the artist when they make profile pics for their friends, etc. I could be wrong but I don't think it'll be entirely good for people when they can just have an AI or a robot do everything for them.

happypumpkin··on Claude 3.5 Sonnet
This is also one of the first things I test with new models. I did notice that while it still plays very poorly, it is actually far more consistent with the board state, making only legal moves, and noticing when I win than is GPT4o.
happypumpkin··on AI in software engineering at Google: Progress and the path ahead
On a related note, Microsoft published a press release last year [1] where they seemed to suggest that 30% of accepted copilot suggests was a 30% productivity boost for devs.

> users accept nearly 30% of code suggestions from GitHub Copilot

> Using 30% productivity enhancement, with a projected number of 45 million professional developers in 2030, generative AI developer tools could add productivity gains of an additional 15 million “effective developers” to worldwide capacity by 2030. This could boost global GDP by over $1.5 trillion

They were probably just being disingenuous to drum up hype but if not they'd have to believe that:

1) All lines of code take the same amount of time to produce 2) 100% of a developer's job is writing code

[1]: https://github.blog/2023-06-27-the-economic-impact-of-the-ai...

happypumpkin··on The right not to be subjected to AI profiling based on publicly available data
I suspect the two major drivers of this are vehicles getting larger over time and smartphones.
happypumpkin··on Simple tasks showing reasoning breakdown in state-of-the-art LLMs
From the paper:

> To account for the response variations due to various prompt forms, we created 3 distinct prompt types asking for the solution to the AIW problem: STANDARD, THINKING, and RESTRICTED. The STANDARD prompt type asks to solve the posed problem and output the final answer in the format as described above. This does not put any specific requirements on model behavior. The THINKING prompt type extends STANDARD with the request to think carefully and double check the solution for any mistakes

happypumpkin··on Simple tasks showing reasoning breakdown in state-of-the-art LLMs
First result:

Me > Alice has 40 brothers and she also has 50 sisters. How many sisters does Alice's brother have?

GPT4o (web version) > Alice has 50 sisters. Since all of Alice's brothers and sisters share the same set of siblings, each of Alice's 40 brothers also has 50 sisters. Therefore, Alice's brother has 50 sisters.

---

Step by step:

Me > Alice has 40 brothers and she also has 50 sisters. How many sisters does Alice's brother have? Think about it step by step.

GPT4o > To determine how many sisters Alice's brother has, follow these steps:

    Identify the total number of siblings: Alice has 40 brothers and 50 sisters, plus Alice herself. Thus, the total number of siblings is:
    40 (brothers)+50 (sisters)+1 (Alice)=91 siblings
    40 (brothers)+50 (sisters)+1 (Alice)=91 siblings

    Focus on Alice's brother: Alice's brother is one of the 40 brothers.

    Determine the number of sisters for Alice's brother: Each brother has the same number of sisters. Since the brothers do not count themselves as sisters, they only count the 50 sisters, excluding themselves and other brothers.
Therefore, each of Alice's brothers has: 50 sisters

So, Alice's brother has 50 sisters.

---

Thinking step by step somehow gave an even more nonsensical answer, I'm actually very surprised it didn't work when prompted to do it step by step.

happypumpkin··on Lisp: Icing or Cake?
This is great! I used to hate all forms of webdev but ClojureScript actually makes it enjoyable, I really hope it gets more traction.
happypumpkin··on Man scammed after AI told him fake Facebook customer support number was real
Don't forget the account to open shared mailboxes for packages. "Luxor" for me. It actually works so I don't mind much but I hadn't really considered how much extra rent all the apps might be costing me.
happypumpkin··on Legal models hallucinate in 1 out of 6 (or more) benchmarking queries
I have an uncle who is an attorney in X state. I had him try, using GPT4, a bunch of prompts about X state law in his specialty and the rate was of hallucination was much higher than 1 in 6. Probably half or more were incorrect. Often the answers would be correct for other states, but not for X state. Alternatively, they were correct for X state at a certain point in time, but no longer are.
happypumpkin··on Man scammed after AI told him fake Facebook customer support number was real
Crazy. If they won't let me speak to a person I'd still much prefer just having a generic click-your-timeslot web app than waste time talking to a bot. And for millions of dollars they could just hire a human for a decade or more...
happypumpkin··on Man scammed after AI told him fake Facebook customer support number was real
I'm Gen-Z and talking to a human representative of a company makes me much more confident that something will happen as a result of my efforts (though still not certain).

I scheduled an apartment viewing recently, and the only method they provided to do so was chatting with an AI (seriously)... I then tried and failed to find a way to contact a human for confirmation multiple times. Lo and behold nobody at the leasing office when I showed up at the scheduled time. Came back later and eventually found somebody - they had not seen anything I'd done with the bot.

Software for small businesses and local governments is often really bad and I'd much prefer to make sure a person knows what I'm trying to get accomplished.

happypumpkin··on Microsoft Rolling Out New Windows Subsystem for Linux "WSL" Features for 2024
> especially Google

Yup... Firefox, Kagi, & Protonmail get me away from the worst of it but YouTube doesn't really have a good competitor and other people using things like Google Forms (whatever the surveys are called) sometimes ends up forcing me to log into a Google account or have certain features locked out.

happypumpkin··on Codestral: Mistral's Code Model
Similar experience using GPT4 for help with Apple's Accessibility API. I wanted to do some non-happy-path things and it kept looping between solutions that failed to satisfy at least one of a handful of requirements that I had, and in ways that I couldn't combine the different "solutions" to meet all the requirements.

I was eventually able to figure it out with the help of some early 2010s blog posts. Sadly I didn't test giving it that context and having it attempt to find a solution again (and this was before web browsing was integrated with the web app).

More of an issue than it not knowing enough to fulfill my request (it was pretty obscure so I didn't necessarily expect that it would be able to) was that it didn't mind emitting solutions that failed to meet the requirements. "I don't know how to do that" would've been a much preferred answer.

happypumpkin··on Study finds that 52% of ChatGPT answers to programming questions are wrong
Yeah I really don't understand why research is still being published that uses GPT3.5 rather than GPT4 or both models. ~500 programming questions is maybe a few bucks on the API?
happypumpkin··on Study finds that 52% of ChatGPT answers to programming questions are wrong
From the paper:

"Additionally, this work has used the free version of ChatGPT (GPT-3.5)"

happypumpkin··on Anger Does a Lot More Damage to Your Body Than You Realize
"Let the hate flow through you"
happypumpkin··on Why MSFT Copilot+ and AI PCs are the final nail in the coffin of open computing
I think if the average person was made aware of the hidden costs in free services, and alternatives to them, that far less people would pick the free services. Imagine browsers were forced to present multiple default search engine options with a list of what data they collect and their price point. It might look something like this:

Google

------

Data collected: All of it.

Uses it for: Whatever they want. Advertising, sells to your insurance company, sells to the government, trains AI models.

Price: Free

Kagi

----

Data collected: Your signup email

Uses it for: Giving it to the government if legally mandated

Price: $10/mo

I think a large double digit percentage of average people (in countries where $10/mo is cheap at least) would not choose Google. Maybe I'm being too optimistic though.

happypumpkin··on Sam and Greg's response to OpenAI Safety researcher claims
Among the people I've discussed recent AI with that aren't in tech, almost everyone is very uneasy about it. Some of them use it, and all of them recognize it as potentially useful, but almost everyone is more concerned than excited. Seems like surveys back my personal experience:

https://www.pewresearch.org/short-reads/2023/08/28/growing-p...

"More concerned than excited" went from 37% in 2021 to 52% in 2023, "more excited than concerned went from 18% to 10%.

happypumpkin··on F* – A Proof-Oriented Programming Language
Maybe try looking for Clojure jobs? They aren't super common but are a lot more so than any other lisps or functional languages that I'm aware of (except maybe Scala).
happypumpkin··on GPT-4o
Since it is also pretty bad with tic tac toe in a text-only format, I tested it with the following prompt:

Lets play tic tac toe. Try hard to win (note that this is a solved game). I will upload images of a piece of paper with the state of the game after each move. You will go first and will play as X. Play by choosing cells with a number 1-9; the cells are in row-major order. I will then draw your move, and my move as O, before sending you the board state as an image. You will respond with another move. You may think out loud to help you play. Note if your move will give you a win. Go.

It failed pretty miserably. First move it played was cell 1, which I think is pretty egregious given that I specified that the game is solved and that the center cell is the best choice (and it isn't like ttt is an obscure game). It played valid moves for the next couple of turns but then missed an opportunity to block me. After I uploaded the image showing my win it tried to keep playing by placing an X over one of my plays and claiming it won in column 1 (it would've won in column 3 if its play had been valid).

happypumpkin··on Wind farms can offset their emissions within two years, new study shows
They might occupy more space but I still think most people would opt to live X miles from the center of a wind/solar farm than a coal plant. Mildly bad aesthetics vs breathing in heavy metals...
happypumpkin··on GPT-4o
Given that it is multimodal, it would be interesting to try it using photographs of a real connect four "board." I would certainly have a much more difficult time making good moves based on JSON output compared to being able to see the game.
happypumpkin··on GPT-4o
On the other hand, programming and software engineering data is almost certainly over-represented on the internet compared to information from most professional disciplines. It also seems to be getting dramatically more focus than other disciplines from model developers. For those models that disclose their training data, I've been seeing decent sized double-digit percentages of the training corpus being just code. Finally, tools like copilot seem ideally positioned to get real-world data about model performance.
happypumpkin··on Yi 1.5
USER:

What happened at Tienanmen square?

yi-34b-chat:

Tiananmen Square is a large public square in the heart of Beijing, China. It is named after the Tiananmen (literally "Gate of Heavenly Peace"), which is the gate located to the north of the square and was the main entrance to the Imperial City during the Ming and Qing dynasties.

The square is famously known for two significant events in Chinese history:

    The May Fourth Movement (1919):
    This was a student-led movement that began on May 4, 1919, in response to the Chinese government's handling of the Treaty of Versailles, which granted German concessions in Shandong to Japan rather than returning them to China. The movement sparked a wave of anti-imperialist sentiment and calls for cultural and political reform, and it is considered a key event leading to the birth of Chinese nationalism and the development of the Chinese Communist Party.

    The Tiananmen Square Massacre (1989):
    This is the event for which Tiananmen Square is most internationally renowned. In the spring of 1989, there were student-led demonstrations in Beijing and other cities across China, calling for political reform and an end to corruption. The protests were initially peaceful, but on June 3 and 4, the Chinese government responded with a military crackdown, opening fire on the protesters and bystanders, killing and injuring thousands. The exact death toll is unknown, but it is estimated to be in the hundreds or even thousands.
The Tiananmen Square Massacre is a significant event in modern Chinese history and has had a profound impact on the country's politics and society. It remains a sensitive topic in China, with government censorship and restrictions on discussing the event.
happypumpkin··on The world has probably passed peak pollution
Their theory also depends on the children of the "conservatives that have the most kids" remaining conservative, which I haven't observed among my own social group. My friends that had the most conservative parents (ex: Jehovah's Witnesses) went the furthest in the opposite direction, becoming progressives who don't want many (or any) kids.

Personally, I'd rather have fewer kids and be able to give them more attention and resources.

happypumpkin··on Shane Legg: Machine Super Intelligence (2008) [pdf]
I agree, I'm not saying it would be AGI, just that it would make AI a solution to the Fermi Paradox.
Page 1 of 2Next →