HNHacker News
TopNewBestAskShowJobs

angusturner

406 karma · joined September 9, 2020

ML Engineer

Interested in generative models, programming, music and philosophy of mind.

https://github.com/angusturner/

submissionscomments
angusturner··on We must pace the frontier
This seems highly optimistic to me.

What's to say current AI won't be highly power-concentrating by default?

The best models are owned by a few companies, and displacement of knowledge-workers mainly seems to benefit the capital class.

On-device / edge computing makes sense in a few very limited scenarios. And economically, price or watt/token (or watt/task completed) might always be better in large data centers.

No reason to think we are on a trajectory towards broad empowerment right now

angusturner··on GPT-6 Astra
Define LLMs. Current models look pretty different to the vanilla transformer from 2017. And its not even known what architecture and training innovations the frontier is using, right?
angusturner··on GPT-6 Astra
Its turtles all the way down
angusturner··on GPT-6 Astra
People want human art. Probably they will want it more than ever in post AGI world.
angusturner··on GPT-6 Astra
What does this even mean. There is strong evidence of LLMs doing in context learning.

Some of the linear RNN layers in recent models are provably doing SGD in hidden space during inference

angusturner··on Our decision on Cursor following its acquisition by SpaceX
What unique differentiator could musk possibly offer though?

Is there really a high demand for racist chatbots with no safety controls?

And even if there is, will governments even allow it to keep operating?

angusturner··on Explorative modeling: Train on the best of K guesses
I find the presentation of this as "factoring training" very confusing.

Factoring is something we do to distributions, which we then parameterize with our neural net.

An autoregressive model pertains to a certain factorisation of a joint distribution. And if we choose this factorisation it may have consequences for our training and sampling etc.

But to say "most models on factor generation".. I dont quite get this.

This would be better work if it didnt claim it was a new pretraining axis also. Its another way of scaling compute right?

If it were a new axis, IWAE would have an equal claim to it (as others have pointed out).

angusturner··on Explorative modeling: Train on the best of K guesses
I immediately thought of IWAE as well.
angusturner··on Ask HN: What was your "oh shit" moment with GenAI?
In 2017 I worked tirelessly with my colleagues to implement and replicate the first transformer paper.

Yesterday I left Opus 4.8 to go do some architecture research, with GPU access.

It replicated and trained a credible baseline. It implemented some ideas I'd been thinking about, and wrote custom CUDA kernels for them. It read and summarised dozens of related papers.

It has since run dozens of experiments, with minimal supervision. When a model is unstable it kills it, documents why, fires off a new configuration.

The realisation that frontier labs are doing this at scale with unlimited GPU and token budgets.

It actually scares me a bit. The realisation that the next big breakthroughs will only have light human involvement.

The prospect of recursive self improvement feels more to real to me all of sudden

angusturner··on Zerostack – A Unix-inspired coding agent written in pure Rust
For small to medium projects, an LLM can write functional (if not well crafted) Rust.

Considering how easy this is now, why choose a heavier, slower and less typesafe language?

angusturner··on Want to Write a Compiler? Just Read These Two Papers (2008)
why read that, vs an actually well-written compiler though?
angusturner··on Grok 4 Fast now has 2M context window
I thought exceptions tended to be made when its highly relevant to the technical topic at hand and also non controversial.

Outside a few weird online bubbles and pockets of the US, hardly anyone disputes the claim you are objecting to.

angusturner··on The quality of AI-assisted software depends on unit of work management
I feel this. I've had a few tasks now where in honest retrospect I find myself asking "did that really speed me up". Its a bit demoralising cause not only do you waste time, you have a worse mental model of the resulting code and feel less sense of ownership over the result.

Brainstorming, ideation and small, well defined tasks where I can quickly vet the solution : these feel like the sweet spot for current frontier model capabilities.

(Unless you are pumping out some sloppy React SPA that you don't care about anything except get it working as fast as possible - fine, get Claude code to one shot it)

angusturner··on The quality of AI-assisted software depends on unit of work management
I think most SWEs do have a good idea where I work.

They know that its a significant, but not revolutionary improvement.

If you supervise and manage your agents closely on well scoped (small) tasks they are pretty handy.

If you need a prototype and don't care about code quality or maintenance, they are great.

Anyone claiming 2x, 5x, 10x etc is absolutely kidding themselves for any non-trivial software.

angusturner··on Introduction to GrapheneOS
Interesting... Maybe I need to investigate PayPal as an option here. Best case would be my bank eventually adds tap to pay natively
angusturner··on Introduction to GrapheneOS
I recently made the shift to graphene from iOS and am mostly enjoying it.

The user profiles was slow to set up and not having shared filesystem between the user profiles creates friction. But I love that I can effectively sandbox my work apps, sandbox the Zuck apps etc, with different VPN profiles for each user.

Getting a burner google account (for gplay services) is a PITA if you are determined to get a clean slate from Googles tracking. Gplay is the only safe way to get certain apps at the moment, and make certain apps pass the device integrity checks.

I suspect one of the biggest barriers to mass adoption will be the fact that tap to pay doesn't work. IIUC apple/google pay are generally considered a privacy and security improvement over physical cards, since you don't give every merchant your actual card number.

Overall love the project and really nice to see such high quality open source software.

angusturner··on Stripe Launches L1 Blockchain: Tempo
Yeah, worlds slowest and most in-efficient write-only database. And as soon as you need to interact with goods or services in the real world, then you still need trust anyway.

All these people harping on about: "Bro I just need to move my money without trusting anyone!, I just need a trust-less way to send currency bro!"

Trust is a good thing! Banks and financial middlemen aren't the devil. Look at how many TPS the visa network can do thanks to trust.

If it weren't for some minimum of social/institutional trust the whole of society would collapse anyway and your digital coins would finally converge to their true value (zero - or actually negative once you add in the externalities).

angusturner··on Google will allow only apps from verified developers to be installed on Android
Fuck google for this. Awful decision. Guaranteed to be abused when Google or government despots decide that certain apps (or developers) aren't aligned with their interests.

Feeling very frustrated with the way the internet is going lately. This plus OSA + chat control. And compounded by the imperative for AI companies to keep hoovering up any and all data they can get their hands on, wiring it into "agentic" workflows and such.

angusturner··on Google will allow only apps from verified developers to be installed on Android
You mean the guy that bans people from twitter for disagreeing with him? And has made a chatbot that spouts right-wing conspiracies in the name of being "anti-woke"?
angusturner··on UK government states that 'safety' act is about influence over public discourse
As an Aussie, I was feeling somewhat consoled about the state of the US by the fact that the EU and UK still seem to have their heads screwed on.

OSA and chat control have made me seriously rethink that…

Has everyone lost their mind?

angusturner··on How large are large language models?
I wish people would stop parroting the view that LLMs are lossy compression.

There is kind of a vague sense in which this metaphor holds, but there is a much more interesting and rigorous fact about LLMs which is that they are also _lossless_ compression algorithms.

There are at least two senses in which this is true:

1. You can use an LLM to losslessly compress any piece of text at a cost that approaches the log-likelihood of that text under the model, using arithmetic coding. A sender and receiver both need a copy of the LLM weights.

2. You can use an LLM plus SGD (I.e the training code) as an lossless compression algorithm, where the communication cost is area under the training curve (and the model weights don’t count towards description length!) see: Jack Rae “compression for AGI”

angusturner··on How large are large language models?
There is an excellent talk by Jack Rae called “compression for AGI”, where he shows (what I believe to be) a little known connection between transformers and compression;

In one view, you can view LLMs as SOTA lossless compression algorithms, where the number of weights don’t count towards the description length. Sounds crazy but it’s true.

angusturner··on My experiment living in a tent in Hong Kong's jungle
By the definition you have provided though, someone that has access to stable, safe or functional housing but then chooses to not to use it (eg opting to camp instead), is not homeless.

Edit: the word “lack” really is the key word. This implies no choice, right?

angusturner··on Claude 4 System Card
Agree the media is having a field day with this and a lot of people will draw bad conclusions about it being sentient etc.

But I think the thing that needs to be communicated effectively is that these these “agentic” systems could cause serious havoc if people give them too much control.

If an LLM decides to blackmail an engineer in service of some goal or preference that has arisen from its training data or instructions, and actually has the ability to follow through (bc people are stupid enough to cede control to these systems), that’s really bad news.

Saying “it’s just doing autocomplete!” totally misses the point.

angusturner··on Gemini Diffusion
One under appreciated / misunderstood aspect of these models is they use more compute than an equivalent sized autoregressive model.

It’s just that for N tokens, autoregressive model has to make N sequential steps.

Where diffusion does K x N, with the N being done in parallel. And for K << N.

This makes me wonder how well they will scale to many users, since batching requests would presumably saturate the accelerators much faster?

Although I guess it depends on the exact usage patterns.

Anyway, very cool demo nonetheless.

angusturner··on Gemini Diffusion
You assume that for small steps (I.e taking some noisy code and slightly denoising) you can make an independence assumption. (All tokens conditionally independent, given the current state).

Once you chain many steps you get a very flexible distribution that can model all the interdependencies.

A stats person could probably provide more nuance, although two interesting connection I’ve seen: There is some sense in which diffusion generalises autoregression, because you don’t have to pick an ordering when you factor the dependency graph.

(Or put otherwise, for some definitions of diffusion you can show autoregression to be a special case).

angusturner··on Continuous Thought Machines
Hm suppose for argument sake that feeding a batch of data through some moderately large FF architectures takes on the order of 100ms (I realise this depends on a lot parameters - but this seems reasonable for many tasks / networks).

Now suppose instead you have an CTM that allocates 10ms on the standard FF axes, and then multiplies it out by 10 internal “ticks” / recurrent steps?

The exact numbers are contrived, but my point is : couldn’t we conceivably search over that second arch just as easily?

It just boils down to whether the inductive bias of building in some explicit time axis is actually worthwhile, right ?

angusturner··on A critical look at MCP
I’m really glad to see people converging on this view because I feel a bit insane for not understanding all the hype.

Like, yeah, we need a standard way to connect LLMs with tools etc, but MCP in its current state is not a solution.

angusturner··on A critical look at MCP
I think this is article is too generous about the use of stdio - I have found this extremely buggy so far, especially in the python sdk.

Also if you want to wrap any existing code that logs or prints to stdout then it causes heaps of ugly messages and warnings as it interferes with the comms between client and server.

I just want a way to integrate tools with Claude Desktop that doesn’t make a tonne of convoluted and weird design choices.

angusturner··on Secret Deals, Foreign Investments: The Rise of Trump’s Crypto Firm
Yeah this is what I don’t really get about the stable coin hype.

In Australia for example, we are almost cashless now and most bank transfers are instant and cheap - no blockchain required.

Digital currency is a good idea - unclear what value the blockchain part actually provides?

Page 1 of 4Next →