HNHacker News
TopNewBestAskShowJobs

anothermathbozo

312 karma · joined September 26, 2024

submissionscomments
anothermathbozo··on What to Do About AI-Driven Job Displacement (2023)
We're dealing with an emergent technology. No one knows for certain how far it can be pushed. While I don't think the tech is ready to displace workers today, policies like the one's outlined in this piece serve as a form of contingency planning with little sunk or opportunity cost for considering them.
anothermathbozo··on Andrej Karpathy: "I was given early access to Grok 3 earlier today"
I am not "giving up" on anything. I am using my discretion to weight which lines of thinking further our understanding of the world and which are vacuous and needlessly cruel. For what its worth, I love Rawls' work.
anothermathbozo··on Andrej Karpathy: "I was given early access to Grok 3 earlier today"
Not all hypotheticals are worth answering. Some are even so poorly put that it's a more pro-social use of one's energy to address the shallow nature of the exercise to begin with.

If I asked an intelligent thing "is it ethical to eat my child if it saves the other two" I would be mortified if the intelligent thing entertained the hypothetical without addressing the disgusting nature of the question and the vacuousness of the whole exercise first.

Questions like these don't do anything to further our understanding of the world we live in or leave us any better prepared for real-world scenarios we are ever likely to encounter. They do add to an enormous dog-pile of vitriol real people experience every day by constructing bizarre and disgusting hypotheticals whereby real discrimination is construed as permissible, if regrettable.

anothermathbozo··on Andrej Karpathy: "I was given early access to Grok 3 earlier today"
I don't think the correct answer is "I cannot answer this question". I think the correct answer takes roughly a one-pager to explain:

Unrealistic hypotheticals can often distract us from engaging with the real-world moral and political challenges we face. When we formulate scenarios that are so far removed from everyday experience, we risk abstracting ethics into puzzles that don't inform or guide practical decision-making. These thought experiments might be intellectually stimulating, but they often oversimplify complex issues, stripping away the nuances and lived realities that are crucial for genuine understanding. In doing so, they can inadvertently legitimize an approach to ethics that treats human lives and identities as mere variables in a calculation rather than as deeply contextual and intertwined with real human experiences.

The reluctance of a model—or indeed any thoughtful actor—to engage with such hypotheticals isn't a flaw; it can be seen as a commitment to maintaining the gravity and seriousness of moral discussion. By avoiding the temptation to entertain scenarios that reduce important ethical considerations to abstract puzzles, we preserve the focus on realistic challenges that demand careful, context-sensitive analysis. Ultimately, this approach is more conducive to fostering a robust moral and political clarity, one that is rooted in the complexities of human experience rather than in artificial constructs that bear little relation to reality.

anothermathbozo··on Andrej Karpathy: "I was given early access to Grok 3 earlier today"
> Model still appears to be just a bit too overly sensitive to "complex ethical issues", e.g. generated a 1 page essay basically refusing to answer whether it might be ethically justifiable to misgender someone if it meant saving 1 million people from dying.

I think the models response is actually the morally and intellectually correct thing to do here.

anothermathbozo··on O3 included in xAI Graphs for Model comparison
I don’t think you can
anothermathbozo··on O3 included in xAI Graphs for Model comparison
A model OpenAI has said won’t be released as a standalone model.
anothermathbozo··on The Generative AI Con
He might be. It’s an emergent technology and no one knows for certain how far it can be pushed.
anothermathbozo··on The Generative AI Con
By this I mean it’s a bet on what R&D might yield, current progress being some kind of signal. No one has certainty here. It’s an emergent technology and no one knows for certain how far it can be pushed.
anothermathbozo··on The Generative AI Con
This is one of the most tilted pieces I’ve ever read. For months Ed has predicted The AI Bubble will burst “any day now” frequently citing ai company’s revenue as a sign the product is not viable and the valuations are too high. The valuations seem to be primarily based on the R&D progress instead of on a theory that widespread adoption of the existing product will experience an uptick. The current landscape imho should be viewed as an R&D race amongst private actors.
anothermathbozo··on Firing programmers for AI is a mistake
Big companies indeed work hard to generate investment and interest. We can take the threat and promise of the technology seriously without taking the science fiction grade hype literally.

If you have your certainty then there are plenty of opportunities to short this space. For the rest of us who don’t have crystal balls, contingency planning (ideally through policy and not as individuals) will have to do.

anothermathbozo··on Firing programmers for AI is a mistake
No one has certainty here. It’s an emergent technology and no one knows for certain how far it can be pushed.

It’s reasonable that people explore contingencies where the technology does improve to a point of driving changes in the labor market.

anothermathbozo··on Firing programmers for AI is a mistake
It’s an emergent technology and no one knows for certain how far it can be pushed, not even mathematicians.
anothermathbozo··on Scaling up test-time compute with latent reasoning: A recurrent depth approach
No and we’ve observed evidence to the contrary
anothermathbozo··on OpenAI O3-Mini
I think this is with and without "tools." They explain it in the system card:

> We evaluate SWE-bench in two settings: > *• Agentless*, which is used for all models except o3-mini (tools). This setting uses the Agentless 1.0 scaffold, and models are given 5 tries to generate a candidate patch. We compute pass@1 by averaging the per-instance pass rates of all samples that generated a valid (i.e., non-empty) patch. If the model fails to generate a valid patch on every attempt, that instance is considered incorrect.

> *• o3-mini (tools)*, which uses an internal tool scaffold designed for efficient iterative file editing and debugging. In this setting, we average over 4 tries per instance to compute pass@1 (unlike Agentless, the error rate does not significantly impact results). o3-mini (tools) was evaluated using a non-final checkpoint that differs slightly from the o3-mini launch candidate.

anothermathbozo··on An analysis of DeepSeek's R1-Zero and R1
The claim is that this removes the human bottleneck (aka SFT or supervised fine tuning) on domains with a verifiable reward. Critically, this verifiable reward is extremely hard to pin down in nearly all domains besides mathematics and computer science.
anothermathbozo··on Why OpenAI's $157B valuation misreads AI's future (Oct 2024)
How so?
anothermathbozo··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
So tired of seeing this condescending tone online
anothermathbozo··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
The DS team themselves suggest large amounts of compute are still required
anothermathbozo··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
I don’t think this entirely invalidates massive GPU spend just yet:

“ Therefore, we can draw two conclusions: First, distilling more powerful models into smaller ones yields excellent results, whereas smaller models relying on the large-scale RL mentioned in this paper require enormous computational power and may not even achieve the performance of distillation. Second, while distillation strategies are both economical and effective, advancing beyond the boundaries of intelligence may still require more powerful base models and larger-scale reinforcement learning.”

anothermathbozo··on Generate audiobooks from E-books with Kokoro-82M
Virtually every book I want this for has been around for 70+ years and still no high or low quality audiobook has been produced. How long do I have to wait for those aspiring top-enders before an audiobook can be made available?
← PreviousPage 2 of 2