HNHacker News
TopNewBestAskShowJobs

tzone

251 karma · joined December 14, 2019

submissionscomments
tzone··on Why I'm still bearish on LLMs after Navier-Stokes
It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine. It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).

On that note, I actually had an overall harness (for experimenting) that was essentially like this: "for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer".

It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff.

Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend an infinite amount of money.

tzone··on Why I'm still bearish on LLMs after Navier-Stokes
It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine.

It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).

On that note, I actually had an overall harness (for experimenting) that was essentially like this: "for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer".

It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff.

Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend infinite amount of money.

tzone··on The Google Play app review process now regularly takes longer than a week
This is one part where it should really just be done by AI for 99% of the cases. Just have a security focused AI model that is biased towards flagging things a bit more conservatively.

Only apps that get flagged by the AI reviewer then have to go through a separate, slower human review process.

tzone··on Shopify moves back to Native from React Native
With latest AI models, we will most likely see shift back to just fully Native apps even for smaller teams. This is type of stuff that AI models do really well, if you get a particular feature or even a bigger app change done in one platform, you can just tell Opus or Fable to replicate to the other platforms and in almost all cases it will just do it as well as most human teams would have done it.
tzone··on On the Navier–Stokes Millennium Prize Problem
Physics alone is more than enough to “further human understanding”. All current mathematicians can move to other sciences, closest being fields in physics, and it will all continue to progress just fine.
tzone··on On the Navier–Stokes Millennium Prize Problem
Well, if in future we do end up with a magical tool that can solve any formal mathematical problem on a whim, we really won’t need field of mathematics anymore as it is today.

There would be no need to deliver new mathematical insights by solving problems. You would just have a magical math problem solving machine and that’s it.

tzone··on On the Navier–Stokes Millennium Prize Problem
Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.

Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .

Wild times

tzone··on Tao: Open math problems being non-renewably mined by AI
While AI companies have almost infinite money, they still don’t want to blow million dollar budgets on problems if there isn’t high likelihood that it will be successful.

But within next 10 years as costs drop significantly and even more improvements are made, yes it is very likely that almost every single existing math problem will get a serious AI cracking done on it

tzone··on Navier-Stokes – Tristan Buckmaster [pdf]
You are glossing over the fact that there is reasonable suspicion that OpenAI used privileged information to do “the scoop”. I.e. they used the fact that researcher used OpenAI tools to get advantage .

Imagine if OpenAI opened up a high frequency trading arm and suddenly stole all the prompts and research that other HfT firms are doing through OpenAI tools and start making bank based on that . Wouldn’t that be straight up insane?

tzone··on On the Navier–Stokes Millennium Prize Problem
This Tristan guy's statement reads like something a normal, reasonable human being would write.

Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win". https://x.com/sama/status/2097385167002415140

OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.

tzone··on On the Navier–Stokes Millennium Prize Problem
It is so disappointing that we can't have such a monumental moment in history without the controversy. OpenAI leadership clearly doesn't seem to care too much about ethics. Is it a requirement to completely lack integrity to have a ground breaking company?

The reality is clear though. The chances of AI models overtaking majority of mathematics within next 10 years is becoming very high. Especially if it becomes cheaper to run these models.

As math formalizations improve, AI can have faster progress in math, compared to even computer science or software engineering.

It is simultaneously the best and the worst time to be a mathematician right now.

tzone··on Reflections on Americans' Net Worth
The issue for Americans is that living a pretty "normal" lifestyle is way more expensive compared to living very similar "normal" lifestyle in most other developed cities.

Easiest example would be groceries. If you live in a crazy expensive place like Manhattan, you could easily spend 4-5x more on groceries compared to a city like Warsaw. But it's not like you are getting better produce or products in Manhattan, you are getting same stuff, just paying way more.

And this is true for other day to day stuff too like Housing, Healthcare, Schools, etc.

tzone··on Reflections on Americans' Net Worth
No sane person would "feel wealthy" if they are living in a shitty 2 bedroom apartment, even if the place costs 1m$. In places where shitty apartments cost that much, a place that would be considered "luxurious" can easily cost 5m$ or more.

Of course rational thing to do would be to take the money and move to a less expensive city, but with the same logic, anyone who has 1m net-worth could liquidate their assets, move to a different country, retire and live a proper wealthy lifestyle.

But we don't see that happening all that much at all. Mostly what happens is that people are choosing to live middle class lifestyles in most expensive areas that they can afford.

tzone··on Reflections on Americans' Net Worth
Of course people in US are very wealthy if you are looking at dollar amounts. It would be impossible for cost of living to be this high, without a lot of people having decent amount of wealth.

But people's perception of their own wealth depends way more on relative value of their wealth compared to cost of living. If you live in a city where some shitty 2 bedroom apartments cost 1m$+, you aren't going to feel all that rich even if you have 10 million dollar net worth.

tzone··on Cursor removed cost information from the usage page and CSV export
I have tried multiple times to switch from Cursor to VSCode + command line tools or VSCode plugin for Claude, but it just doesn't work as well as Cursor itself.

Especially if you are doing "remote" development through SSH. If you are doing stuff where you still have to write some parts of the code manually or you have to fix few things here and there that the AI outputs, you still need a real editor.

tzone··on Golang proposal: container/: generic collection types
Garbage collector alone can't do what happens in Go language+runtime. Simplest example is that there is no way to allocate fixed arrays with non-pointer structs in Java. That is just not possible in the language.

Design of Go allows programmer to rewrite their program to make it as GC efficient as needed. You can even have essentially zero GC overhead and do stuff manually for really high performance needs. And you can also write regular code when GC isn't a big deal.

Those options simply don't exist in Java language. Your only bet in Java would be to embed some C code which is a nightmare of its own.

There is a reason why almost all backened systems that run on JVM are a huge pain in the ass even at moderate scale. They all end up rewriting parts in other languages and at the end just rewriting the whole thing.

tzone··on Golang proposal: container/: generic collection types
People use Go vs Java not because of the language differences but because of massive improvement that GO runtime is compared to JVM.

It is crazy that people still don't understand why JVM sucks, and why GO's approach to minimize GC pause latency is far superior design decision compared to whatever JVM has been trying to do with its GC iterations for god knows how many years with gazzillion different variants that all suck in different ways.

tzone··on The mean means nothing: data visualization to debug a latency problem
Imo, average has to be one of the biggest net negative stats you can collect. It is much better to collect just total throughput stat since that is the only thing that average is useful for, but unlike average stat, it doesn't cause confusion and tons of mistakes.

There are so many issues with average stats, not just the fact that people misunderstand it often, but also in collection and graphing, people often end up having final graph that ends up being some type of average of averages which becomes even more useless and completely meaningless.

If your metrics/monitoring system doesn't allow you to collect and graph proper distributions/percentiles, you really should change that.

tzone··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Latest models like Opus 5.0 work well in both scenarios. You can just give it a goal and let it make a plan, review the plan and then let it go ham.

Or you can also use it piece-meal and have it do parts of work that you want. It works in both ways and works really well.

I don't think there is any model right now that you can use with 0 oversight, but you can do pretty complex stuff without writing single line of code at this point.

tzone··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Which model are you using? There is monumental difference between models. Even between "frontier models". When people tell these stories, it would be great to also add which model you were using.

As an example Opus 5.0 is in completely different class compared to Cursor Grok 4.5 even if the benchmarks don't show such massive difference. Not even talking about regular stuff like Sonnet or Composer or stuff like that.

tzone··on Write code like a human will maintain it
I think people have outdated ideas on AI code quality. Latest ones (i.e. ones after Claude Opus 4.8) are a game changer. They now write better code than a lot of engineers. They are also able to understand much larger overall context and make less mistakes then most engineers would make.

Difference between something like Opus 4.8 and even 4.5 is massive, not even talking about difference between Opus 4.8 and Sonnet or Composer.

What i see online is that a lot of people are using cheaper models, stuff like sonnet or cursor composer or even latest cursor grok. But they aren't the same thing as using the best models. The difference in quality and correctness is huge.

tzone··on Write code like a human will maintain it
have been doing "racing to debug" things with AI for quite some time too. Have to say, before Opus 4.8, there wasn't even a comparison for anything slightly complex.

Opus 4.8 has been the turning point though. I am not winning many races compared to Opus 4.8, especially if it is something more complex.

I think if everyone was using Opus 4.8 (or better) for every single task/question/etc, you would have much different overall sentiment. I don't think you would have many AI disbelievers left.

tzone··on Claude Desktop spawns 1.8 GB Hyper-V VM on every launch, even for chat-only use
Because 90% of those UIs is written by these new AI models. That is how they are able to churn so much new user facing stuff all the time. The fact that it works at all is proof that these new AI models are actually pretty decent.
tzone··on $30B for laptops yielded a generation less cognitively capable than parents
Trend is pretty clear pretty much across all western countries. Even among ones that have supposedly highest quality public education like Norway, Sweden, Germany, etc...

People are getting too stuck on US specific issues and missing that this is a pretty global problem.

tzone··on The unbearable joy of sitting alone in a café
I think it depends on what counts as doing nothing. Every time I cut my hair, I sit in a chair for ~30minutes silently without doing anything. My barber knows I don't like small talk so he just cuts my hair and that's it, there is no conversation.

I would say it is very enjoyable 30 minutes every time I do it. I don't think anyone would describe that kind of experience as hard to do?

tzone··on How SQLite is tested
I am surprised to see that there isn't a lot of information about performance regression testing.

Correctness testing is important but the way SQLLite is used, potential performance drops in specific code paths or specific type of queries could be really bad for apps that use it in critical paths.

tzone··on A crypto founder faked his death. We found him alive at his dad's house
Most important property of blockchains is verifiable transparency. That is a huge deal and yes, most financial and government systems would be much better for people if they were built with that type of verifiable transparency.

Decentralization is important but isn't as important. There are many successful chains that aren't decentralized at all and are quite useful.

Most top/real projects in blockchains aren't immutable. Pretty much everything is upgradable/changeable, including blockchains themselves.

Trustless - even this part isn' true for most blockchain systems. They all contain various levels of trusting different entities.

It is the transparency and verifiability that is the key idea and most important improvement that blockchains bring.

tzone··on Jepsen: MySQL 8.0.34
That is sort of my point. I think "repeatable read" is a fools gold. You think you wont need to do locking, but it is too easy to make incorrect assumptions about what guarantees "repeatable read" provides and you can make very subtle mistakes which leads to rare, extremely hard to diagnose correctness issues.

Repeatable read type of setups also make it much easier to accidentally create much longer running transactions, and long running transactions/too many concurrent open transactions/etc can create really unexpected, very hard to resolve performance issues in the long run for any database.

tzone··on Jepsen: MySQL 8.0.34
Here is my issue with this. We are assuming you need to read a "consistent snapshot" in some type of real time application. Because if it isn't real time, you can always have those snapshot type of querying on "replicas", since that is a lot easier to implement correctly without sacrificing performance.

So assuming you are looking at reading "consistent snapshot" in the context of a real time transaction. If the data that you want to read as a "consistent snapshot" is small, locking + reading is good enough in most cases.

If the data to read is too large (i.e. query takes long time to execute, and pulls a lot of data), you are going to have ton of scaling issues if you are depending on something like "repeatable read". Long running transactions, long running queries, etc are bane of all the database scaling and performance.

So you really want to avoid that anyways, you would almost always be much better of changing your application logic to make sure you can have much shorter, time bounded transactions and queries and setup better application level consistency scheme. Otherwise you will at some point hit scaling/performance problems and they will be an absolute nightmare to fix.

tzone··on Jepsen: MySQL 8.0.34
I have been advocating for the longest time that "repeatable read" is just a bad idea. Even if implementations were perfect. Even when it works correctly in the Database, it is still very tricky to reason about when dealing with complex queries.

I think two isolation levels that make sense are either:

* read committed

* serializable

You either go all the way to have a serializable setup, where there are no surprises. OR, you go in read committed direction where it is obvious that if you want have a consistent view of the data within a transaction, you have to lock the rows before you start reading them.

Read committed is very similar to just regular multi-threaded code and its memory management, so most engineers can get a decent intuitive sense for it.

Serializable is so strict that it is pretty hard to make very unexpected mistakes.

Anything in-between is a no man's land. And anything less consistent than Read Committed is no longer really a database.

Page 1 of 3Next →