HNHacker News
TopNewBestAskShowJobs

igorkraw

1,774 karma · joined January 26, 2018

hackernews ( a ) krawczuk (point ) eu
submissionscomments
igorkraw··on Show HN: Whiteboard (YC W26) – An open-source IDE for thoughtful software design
Thanks! Sorry if I missed it. I had clicked ok the download on the page and it only pointed me at Mac.

If you can create an aur it'd be awesome for the arch crowd :-)

igorkraw··on Show HN: Whiteboard (YC W26) – An open-source IDE for thoughtful software design
Looks cool :-)Linux wen plz so I can play with it?
igorkraw··on Space Force Uniform Is Inspired by Quasi-Satirical Film 'Starship Troopers'
Well that tracks, he was in the Nazi party and produced (but not designed) the uniforms https://en.wikipedia.org/wiki/Hugo_Boss_(businessman)
igorkraw··on How accurate have Ed Zitron's AI skeptic predictions been?
To shamelessly shill a passion project: A friend of mine and I try to be skeptical-but-reasonable on our podcast https://kairos.fm/muckraikers/ we aggregate papers and reporting and try to contextualize it with our own (hopefully useful) perspectives and takes
igorkraw··on Deutsche Bank becomes first foreign yuan clearing bank in Europe
Are humans without agency then?

There's more ways of neutralizing power than meeting it with equivalent threat. You can also make it not worth exerting the power, or introduce enough friction to let it evaporate itself.

And so far, the highest historical power concentrations have always been un-stable, while the polities that put checks on that power and embed themselves in cooperative trade networks lasted for much longer

igorkraw··on Deutsche Bank becomes first foreign yuan clearing bank in Europe
Sure, and now compare polities which manage to keep defection to a minimum or structure the process of "defecting" to serve the polities purpose instead of being truly adversarial. Bruce Bueno de Mesquitas "the logic of political survival" and various works by daron acemogolu have the data on how much better these polities do.

If there's no pure play dominant strategy, you will always get a fraction of defectors willing to try their luck. But coopération still emerges, and the polities that manage to cooperate more still tend to survive and thrive more.

igorkraw··on Deutsche Bank becomes first foreign yuan clearing bank in Europe
Yes, I messed that up based on phonetic recall I think. Thanks for correcting
igorkraw··on U of Michigan drops first-semester grades to ‘curb mental health crisis’
Sure , it's difficult, and not always possible,but maybe we should treat it as what it is (an unfortunate reality) rather than creating false valor.

I don't think people are unaware of the reality of real life, and the ones that are won't wake up if you terminate them in college. If we want to make sure, make them do mandatory work Placements instead of the grades, to experience the pressure of real world without risk of career failure

igorkraw··on Deutsche Bank becomes first foreign yuan clearing bank in Europe
You should learn about chimpanzee wars and how dolphins treat tortoises, and about the enslaving ants
igorkraw··on Deutsche Bank becomes first foreign yuan clearing bank in Europe
In game theory, cooperation is a dominant strategy and evolves consistently. Slave owners and other injustice committing humans often invoke "human nature" out of trying to cope with their knowledge that it is wrong but whatever trauma or other thing makes them do it anyway, to absolve themselves of agency.

(To be clear, not targeting you personally here, but targeting the rhetoric and fallacy)

igorkraw··on Deutsche Bank becomes first foreign yuan clearing bank in Europe
Not really, kindness also exists. It is the escape hatch from the evolutionary zero sum game and let's you pick a better equilibrium (bargaining vs pure Nash).

We can just not fight. There is enough room and resources,we can get more of them and build more efficient tech, and we are all more alike than different.

igorkraw··on Deutsche Bank becomes first foreign yuan clearing bank in Europe
Which of I understand the Wikipedia article correctly would mean it's a euphemism for "there are no just wars"?
igorkraw··on Deutsche Bank becomes first foreign yuan clearing bank in Europe
The how of the reference seems important, it's not like he brought it up first apparently

>Xi has referenced Allison’s term before. In a speech in Seattle in 2015, Xi said, “There is no such thing as the so-called Thucydides Trap.” And in an October 2023 meeting with Senate Majority Leader Chuck Schumer, Xi said, “The ‘Thucydides’ Trap’ is not inevitable, and Planet Earth is vast enough to accommodate the respective development and common prosperity of China and the United States.

igorkraw··on U of Michigan drops first-semester grades to ‘curb mental health crisis’
Genuine take: maybe it would be more efficient for the economy and society, as well as better for everyone involved individuall, to make room for the ease in period as well, with a clearly communicated end towards the same, so people can acclimatize?

Speaking as someone who grew up in a slightly more cruel world, but still received way more kindness than previous generations (did not die of previously deadly respiratory complications in childhood, got accepted into gifted children program for enthousiasm, not IQ, got recommendation letters for the same reasons, short sighted, ADHD) in my self evaluation, each time I got a break it spurned me to work harder and try more to "pay it back", and every time I powered through strict evaluation and scrutiny it left scars that overall impede my productivity/would do so if I had not gotten therapy to overcome them .

Nature is cruel, but also inefficient. We have long since started building a "rela world" full of safety bubbles that let us thrive. Why not continue with this one?

igorkraw··on TheoremDB – A public workspace for machine mathematics
Fun, I've been building a similar idea for the last year in my evenings, but focusing on human curation and human consumption and ensuring the humans understand the concepts correctly. Gonna be nice to have all these orthogonal projects complementing one another
igorkraw··on I am retiring from fulltime writing (& pseudonymity) to launch Guardian Angel
That's the best answer I've seen to that question so far, my honest respect. I am much more skeptical than you, for me I have seen a sea change in _tool capabilities_ (mainly pre opus 4.5 to post 4.5) and harness engineering but no strong change in the type of errors made and the pattern of harness engineering (the pattern of "set things up for the LLM to see when it fucks up and let it flail till the verifier tells it to stop").

I would actually expect the sea changes as you describe it in your first criteria to continue with 1) vision, audio and video natively integrated 2) continued scaling of e2e rlvf for workflows with large scale labeling efforts 3) ASICs and widescale deployment of diffusion models leading to speed ups

But as of right now, I still expect these models to need humans to prune the output to the gold and set up the harness right for both the novel bits, and for the boilerplate to be cohesive with the global intent.

Which is of course an amazing potential boost in productivity, but still a sigmoid flattening.

As for your second criteria that includes cost, I think we might every well see this coming soon, but it's difficult to estimate with the efficiency gains still possible.

Thanks for engaging:-)

igorkraw··on I am retiring from fulltime writing (& pseudonymity) to launch Guardian Angel
Honest question: What _would_ convince you that we are not heading towards a singularity? Like, is this belief falsifiable?

(Secondary question: what do you mean with singularity?)

As an offering of me engaging in good faith, a controversial opinion of mine: I legit think arpanet going online and starting the networking all of humanity into a massive coupled complex system fits the definition of singularity of "the moment after which predicting what will happen becomes hard to impossible", although that of course heavily depends on your definitions of prediction and hard/impossible.

igorkraw··on Bunny Database
I'd like to be a added, who should I email?
igorkraw··on I fed 24 years of my blog posts to a Markov model
Another term used is identifiable (although learnable and identifiable are not synonyms, I think identifiability is one precondition for learnability).

Identifiability means that out of all possible models, you can learn the correct one given enough samples.causal identifiability has some other connotations

See here https://causalai.net/r80.pdf as a good start (a nose in a causal graph is Markov given its parents, and a k-step Markov chain is a k-layer causal dag)

igorkraw··on Just 0.001% hold 3 times the wealth of poorest half of humanity, report finds
I see this argument often but for me it misses something.

The difference is about power. The wealth being this concentrated means the power is concentrated.

If people are okay with the idea of an ETF, or a wealth manager (or any type do fund manager/investment bank) then they should be okay with sovereign wealth funds/national ETFs that provide dividends with a guaranteed single share single vote setup.

If you want competition, then the US government used to be good at creating and sustaining artificial compétition in military procurement - similar to how Amazon let's teams compete on the same projects internally.

Because the competition would be artificially and enforced by laws, there's just as much as potential for massive efficiency gains as there is potential for corruption (the Norwegian national wealth fund has gone swimmingly for them)

igorkraw··on What I don’t like about chains of thoughts (2023)
You need to think about 1) the latent state 2) the fact that part of the model is post trained to bias the MC towards abiding by the query in the sense of the reward.

A way to look at it is that you effectively have 2 model "heads" inside the LLM, one which generates, one which biases/steers.

The MCMC is initialised based on your prompt, the generator part samples from the language distribution it has learned, while the sharpening/filtering part biases towards stuff that would be likely to have this MCMC give high rewards in the end. So the model regurgitates all the context that is deemed possibly relevant based on traces from the training data (including "tool use", which then injects additional context) and all those tokens shift the latent state into something that is more and more typical of your query.

Importantly, attention acts as a Selector and has multiple heads, and these specialize, so (simplified) one head can maintain focus on your query and "judge" the latent state, while the rest can follow that Markov chain until some subset of the generated+tool injected tokens give enough signal to the "answer now" gate that the middle flips into "summarizing" mode, which then uses the latent state of all of those tokens to actually generate the answer.

So you very much can think of it as sampling repeatedly from an MCMC using a bias, A learned stoping rule and then having a model creating the best possible combination of the traces, except that all this machinery is encoded in the same model weights that get to reuse features between another, for all the benefits and drawbacks that yields.

There was a paper close when OF became a thing that showed that instead of doing CoT, you could just spend that token budget on K parallel shorter queries (by injecting sth. Like "ok, to summarize" and "actually" to force completion ) and pick the best one/majority vote. Since then RLHF has made longer traces more in distribution (although there's another paper that showed as of early 2025 you were trading reduced variance and peak performance as well as loss of edge cases for higher performance on common cases , although this might be ameliorated by now) but that's about the way it broke down 2024-2025

igorkraw··on Reasoning LLMs are wandering solution explorers
I'd encourage everyone to learn about Metropolis Hastings Markov chain monte carlo and then squint at lmms, think about what token by token generation of the long rollouts maps to in that framework and consider that you can think of the stop token as a learned stopping criterion accepting (a substring of) the output
igorkraw··on GPT-5: Overdue, overhyped and underwhelming. And that's not the worst of it
I have a tiny tiny podcast with a friend where we try to break down what parts of the hype are bullshit (muck) and which kernels of truth are there, if any, startedpartially as a place to scream into the void, partially to help the people who are anxious about AGI or otherwise bring harmed by the hype. I think we have a long way to go in terms of presentation (breaking down very technical terms to an audience that is used to vague-hype around "AI" is hard), but we cite our sources, maybe it'll be interesting gpr you to check out out shownotes

https://kairos.fm/muckraikers/

I personally struggle with Gary Marcus critiques because whenever they are about "making ai work" it goes into neurosymbploc "AI" which o have technical disagreements with, and I have _other_ arguments for the points he sometimes raises which I think are more rigorous, so it's difficult to be roughly in the same camp - but overall I'm happy someone with reach is calling BS ad well.

igorkraw··on Measuring the impact of AI on experienced open-source developer productivity
Cool, thanks a lot. Btw, I have a very tiny tiny (50 to 100 audience ) podcast where we try to give context to what we call the "muck" of AI discourse (trying to ground claims into both what we would call objectively observable facts/évidence, and then _separately_ giving out own biased takes), if you would be interested to come on it and chat => contact email in my profile.
igorkraw··on Measuring the Impact of AI on Experienced Open-Source Developer Productivity
Could you either release the dataset (raw but anonymized) for independent statistical évaluation or at least add the absolute times of each dev per task to the paper? I'm curious what the absolute times of each dev with/without AI was and whether the one guy with lots of Cursor experience was actually faster than the rest of just a slow typer getting a big boost out of llms

Also, cool work, very happy to see actually good evaluations instead of just vibes or observational stuies that don't account for the Hawthorne effect

igorkraw··on A Love Letter to People Who Believe in People
I really believe in the importance of praising people and acknowledging their efforts, when they are kind and good human beings and (to much lesser degree) their successes.

But, and I mean their without snark: What value is your praise for what is good if I cannot trust that you will be critical of what is bad? Note that critique can be unpleasant but kind, and I don't care for "brutal honesty" (which is much more about the brutality than the honesty in most cases).

But whether it's the joint Slavic-german culture or something else, I much prefer for things to be _appropriate_, _kind_ and _earnest_ instead of just supportive or positive. Real love is despite a flaw, in full cognizance if it, not ignoring them.

igorkraw··on Don’t let an LLM make decisions or execute business logic
What is your definition of "understand them well"?
igorkraw··on “Vibe Coding” vs. Reality
Check the actual paper on the type of sorts it actually got speedup on :-) (hint: a few percentage points on larger n,similar to what pgo might find, the big speedup is for n around 8 or so, where it basically enumerated and found a sorting network)
igorkraw··on Ladder: Self-improving LLMs through recursive problem decomposition
Nah, it's much simpler, the models aren't reliably able to recall the correct rule from memory - it's im the training set for sure.

This is another specialized synthetic data generation pipeline for a curriculum for one particular algorithm cluster to be encoded into the weights, not more not less. They even mention quality control still beim important

igorkraw··on Mercury Coder: frontier diffusion LLM generating 1000+ tok/sec on commodity GPUs
Please make sure aider and llm-cli can use this soon,kthx :-)
Page 1 of 19Next →