HNHacker News
TopNewBestAskShowJobs

gravypod

6,918 karma · joined July 29, 2014

I'm Joshua Katz. All opinions are my own. I do not represent my employer.

In the past I have:

  - Built multiple complex, high availability, systems to support low powered IoT deployments in healthcare and retail
  - Managed compute infrastructure on AWS, Azure, and bare metal
  - Improved prototype game engines and created course material for college classes teaching low level rendering APIs (Vulkan)
  - Launched and thoroughly tested high throughput expert systems
  - Become the "authoritative" "expert" in Markdown at a major tech company.
I know: C, Java, Python, PHP, and Golang

I love using: Docker, Kubernetes, Bazel, MongoDB

${username}@${username}.com & admin@${first_name}${last_name}.me

meet.hn/city/40.7709499,-74.2264601/Orange

submissionscomments
gravypod··on Lawmakers Introduce Multiple Laws to Curb Flock After 404 Media Coverage
Weren't many of these reports after 404's reporting?
gravypod··on GLM-5.3 and the spread of advanced cyber capabilities
I'm not an insuder but here's my thinking:

1. We can never know if we have enough safe guards

2. Banking systems are already "insecure" by some definition of this.

3. Because of this insecurity a lot of government regulations were passed on logging transactions, being able to roll them back, etc.

We've seen these trigger I'm financial markets: https://www.investor.gov/introduction-investing/investing-ba...

If you've made large purchases before or strange purchases before you have likely hit similar consumer side things.

The network banks talk on is called swift. The way it works is, as I understand it, if you say move money from account 1 to account 2 it just happens so if I knew your account number I could drain all funds. Because of this, there are systems that block certain transactions. Could ai hack around them? Maybe, but they are undocumented, battled hardened over decades, and even if they did the bank could roll back the transactions.

gravypod··on Gemini 4 Argon
If they, through standard accounting practices, cannot show they are profitable how do we know they are profitable? I assume they can't just say "we are profitable assuming you ignore the costs we paid to build up XYZ". I'm guessing they could pre - rent compute for 10 years or something and maybe that's all on the balance sheet but you'd just amortize that and I think that falls under GAAP?
gravypod··on The last time my family was replaced by technology
Could the original chart be including non-cash consumption? (Private jet flights paid for by company, stocks, bonds, etc)
gravypod··on Gemini 4 Argon
I could start a business with a similar growth trajectory. We can mail people $100 bills for a low low payment of $10. I just need $500 billion dollars of startup capital and I can show you a 10x yoy growth for a few years.
gravypod··on GLM-5.3 and the spread of advanced cyber capabilities
Yes, but no one had time to cycle through shodan and find the right things to attack. Now you can just attack them all.
gravypod··on GLM-5.3 and the spread of advanced cyber capabilities
I think our biggest advantage in hardening is inaccessibility of swift and how many circuit breakers are in place to prevent large moves of money.
gravypod··on GLM-5.3 and the spread of advanced cyber capabilities
> When someone wielding a non-safeguarded model deletes the money in everyone’s bank account, I look forward to the HN comments claiming it’s an attempt by Anthropic to pull off regulatory capture.

This statement portrays a fundamental misunderstanding of how the infrastructure which powers these systems work. Note: I am not saying there are no risks, I am just saying the risk you are focusing on is the least likely one of all I have seen people be upset by.

Far higher risks one could outline are:

1. Network-connected PLCs for big infrastructure (drinking water, sewage, power, etc) being tampered with.

2. Extremely persistent malware tailored for every permutation of hardware + software.

3. Cyber criminals improve in technical capabilities (phishing sites, scam calling, propaganda campaigns, etc).

But, the cork is out of the bottle on this one. With even basic models you can begin a loop of training specialized models on low cost hardware which can be used to do specific hacking tasks.

I don't know what the best antidote to this is but I doubt that it will be in limiting access to OSS models to people in the USA as all of the threats listed above come from *outside actors*.

gravypod··on GLM-5.3 and the spread of advanced cyber capabilities
This part of the article really stands out:

> On Sept. 17, NIST’s Center for AI Standards and Innovation (CAISI) published its own assessment of GLM-5.3’s cyber capabilities. CAISI found that GLM-5.3 is “the most cyber-capable open-weight model released to date” and that it lags the US frontier by about four months on an aggregate of CAISI’s cyber benchmarks.

To translate: "This free model, you can host yourself, is at max 4 months behind Anthropic - as confirmed by Anthropic and the US Government - and it won't reject your requests"

Interesting play before an IPO...

gravypod··on Owed a billion dollars in Nvidia stock
In an era of digital records keeping, this does not sound impossible.
gravypod··on Don't couple your Go code to GitHub
It's harder to replace in history.
gravypod··on Plan mode is dead
If your company is really interested in doing AI stuff you may be able to sell people on the idea of a "fast agentic CI" where you:

1. Track code coverage from all current test cases in your big 8hr runs.

2. Index the coverage so it goes testcase name -> methods touched.

3. On every commit run an agent which looks at the diffs, tries to predict which tests the code being edited is touching, and run just those tests (or a random subset if there are too many).

Once you have this you can also automatically kick off builds every 8hr that contain all the previously submitted patches. Once that is done, if a test begins to fail, it can automatically bisect the history to find change, and notify the author.

You could also get these benefits with a Bazel like build system which can cache test executions so others don't need to rerun them if the binary is not changed.

gravypod··on We write code by hand
LLMs will make tradeoffs without consulting you. This is sometimes a feature and often a bug. Many times a decision made up front can have non obvious implications down the line. This is one of the biggest values I bring to be table as a more experienced swe.

I think if I fully understand everything the agent is going to do and trivially verify the output quality, LLMs are a win. For one off prototypes, the same is true. For software which is unique, complex, and performing a task which has not been fully specified I find myself needing to drop into my editor more and more.

Recently I've been moved to a research focused team and I really need to have proof that something is happening. I've had Claude Opus 5 gaslight me by telling me it did something and when I read the code it obviously did not. This happens more and more with my tasks that are kind of complex.

I've found a lot of the claims by AI people have been 6mo - 2 years "ahead" of my experience. I think we are now in an era where harness engineering is highly valuable (making a test, looping an agent, manually annealing with new ideas) but the claims that no one writes code manually seems like it may not be fully there for all code.

gravypod··on We write code by hand
I feel like I have gone from 100% hand written code -> 0% hand written -> 50% and climbing.

I think there are things LLMs are really good at coding up and experimenting with but I think there's still a need to understand the btoader context of your software so you need to dive in at some point.

gravypod··on Bend 2 and the Vibe-Coding Trap
I cannot understand the hostility being directed towards you for this project. It seems very interesting. Your reply was very well thought out. I am very confused.
gravypod··on GPT-6 Astra
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

I think I have the following questions about what AGI would look like:

1. Do you expect an AGI to be able to competently do any knowledge work that an able human does today?

I think this is implied by the "General" component. I would assume that anything we call AGI would be able to do any of these tasks if given the time and reference material needed.

2. How would you expect AGI to handle edge cases? (Missing context, no known solution, under specified instructions, over specified instructions)

I would expect an agent to be able to look at the context that work exists in and correctly attenuate it's intentions for these goals. Simpler solutions, more thorough reporting, etc based on the need.

3. In my work I attend meetings, write reports, write code, research things, etc. Would AGI be able to reliably do that?

I would say that AGI would need to do this. I would classify this as the "Intelligence" component. Obtaining context, building a model of a problem, solving it, and convincing others.

4. Would it be able to inspire trust in itself? Trust can be established through verification of it's outputs, the construction of introspective tools, no hallucinations, etc.

I would say yes to this as well. It would be a component of the "Intelligence" to know that buy in is more important than the completion of a task.

To these points, will Astra be able to do these things? If not, I would hesitate to call it AGI.

gravypod··on GPT-6 Astra
My success rate for solving software engineering challenges encountered in my day jobs has been near 100% for my entire career. I can only think of a few tasks I kicked back and said they were impossible. For example, after trying to get a signal processing system working reliably I decided to sit down and calculate the actual limits of the channel we were sending the data over and found that from a basic estimation it would not be possible to do. In start ups you don't really get to get stuck in a spiral and not fix things.

I find agents often get into these cases during research tasks.

gravypod··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent.

For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces.

> 2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this.

You could also have a classification of what counts as "cheating" (coordination, accessing the internet, etc) and score the results. If you are seeing a spike in this (even in a small group of the evals) you could manually look at those. Or you could stop inference on cheating sessions.

> 3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" from "concerning AI weirdness" is not a trivial task.

If you start seeing "now I have access to the internet" or something similar, maybe that's a good signal something is going wrong?

Also, obviously, it would not be outside of OpenAI's ability to just completely air gap this training system. For example:

1. No network connection.

2. GPS based NTP for time sync for your servers.

3. Mirror of all apt, pypi, go, c++, Rust, Java, etc packages. (<5TB of data)

4. Take your training data and use that for a mirror of the web. (http://example.com -> mirror server -> local training copy).

They had systems connected to the internet connected to this system which was not air gapped. Designing an air gap system would be super easy, well within the means of openai, and betrays the assumption that they think they are actually building something dangerous.

gravypod··on Bye, Bye GitHub
It's a shame because I like GitLabs UX and runners much more than other forges. I've been considering forking gitlab and removing complexity. GitLab uses multiple GB of memory idle with no users. It's sad.
gravypod··on GLM-5.3-Flash
Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc that isn't just brute force ablations?
gravypod··on Zohran and the Short Link
People are discovering go links from first principles.
gravypod··on Walgit – a Git server that is one binary in front of an object store
Is there a market for an actual GitHub competitor now that everyone is looking around and is very angry about stability?
gravypod··on Run 290B+ frontier MoE models locally on your gaming PC
The linked paper is a much more interesting read: https://arxiv.org/pdf/2608.16157

I think there are many other significant inference improvements which an be built out that are being ignored because the API between client and inference stack would be tricky to nail down.

This is not a criticism of the research but instead the presentation but the repo looks very fishy, it isn't clear that this is from a bunch of researchers from Berkley. It also does not make very clear (on the GitHub) what optimizations or performance they are targeting.

gravypod··on There's no reason for software to be slow anymore
I recently built a piece of code which downloads a bulk set of data, indexes it for search, and then serves a pretty web UI on top of this with the help of some AIs. Normally I would have reached for sstables, sqlite, etc. This time, because the lookup patterns actually would not have been too efficient on sstables and SQLite would have been overkill, I had an agent take the data structures, pack the text effectively, and build a prefix tree for fast auto completion from the search bar. It was great. I could have done this all before but I wouldn't have. I would have felt sqlite was fast enough. The resulting web server is significantly faster feeling (because the optimized lookup speeds) than an sqlite implementation would feel like.

I think engineers building very complex systems now have a lot of performance knobs to twiddle that would have just been too costly for human effort. Since we constrain the responsibilities of the agent slop is less of a problem. We relegate it to defined tasks with clear API boundaries and test harnesses.

gravypod··on AI companies destroy physical books – let's scan rare books before it's too late
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

That's what I keep saying about the van Goghs I burn to heat my home but everyone is still mad at me!

gravypod··on Two wheels, a few tradeoffs, and gas prices
We will not reverse climate change by consuming less. The solutions are equally political and technological. Driving a smaller vehicle does help but real solutions would be: building homes next to jobs, public transportation, electrification, and carbon capture.

> On a motorcycle, you’re in the environment instead of sealed off from it. Lavender above Grasse in June, a 5-degree temperature drop when the road cuts through a forest, leaning into corners with your whole body

Very romantic, but not very realistic. Most commuters need to show up for work in heat, rain, snow, etc.

gravypod··on New cars are getting 1.2cm longer and 0.5cm taller each year
I have taken a few cars on that drive. Adding in traffic, I usually have to fill up at the gas station just outside of Boston that has the burger king in it. In my Mirage can get me to Boston, drive around a bit, and usually just need to fill up on my way back.
gravypod··on New cars are getting 1.2cm longer and 0.5cm taller each year
Mitsubishi just discontinued the car line that I drive [0]. The Mirage was great. I can drive from NYC to Boston on ~10 gallon. The only downsides were: No arm wrest, no sunglasses holder, and the media console won't turn on if you leave the car in the hot sun for a while.

I'm hoping to make my next car an electric car but all of the US ones are massive.

[0] - https://www.mitsubishicars.com/mirage-history

gravypod··on DeepSeek-V4-Flash Update
If they did have a super powerful secret AI why would avoid:

1. Selling access to the US? 2. Ship more and better software? 3. Talk about it publicly?

If you had a secret AI better than anything currently available the mere mention of this would crater the US tech investment sentiment.

They could even host these Chinese models through AWS and sell at-cost inference on Bedrock and obliterate the US model companies.

gravypod··on DeepSeek-V4-Flash Update
What would the benefit of this be? If China stops open weights, US labs still have the intelligence frontier. Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't.

Undermining out entire economy by subsidizing the release of DIY versions of our main economic drive sounds like a huge win for China.

Page 1 of 34Next →