HNHacker News
TopNewBestAskShowJobs

bfeynman

780 karma · joined November 10, 2022

submissionscomments
bfeynman··on Cursor Introduces Composer 2.5
That only seems plausible if whatever corpse of xAI is around is giving them engineering time. I don't know if they hired a bunch of ex frontier lab staff but its unlikely they have the technical capability to train their own frontier models especially the pretraining. Because the thing is if its not competitive with claude/codex it will be panned.
bfeynman··on New York to tax luxury second homes in NYC
is that really what people want? The fact that people say why not have 50 story concrete blocks everywhere to get more people feels like exact thing that would destroy what makes living in the city nice... Tenement housing sucked, why add thousands of people to crammed parts of city. We should be incentivizing sprawl and better transportation.
bfeynman··on Cutting inference cold starts by 40x with LP, FUSE, C/R, and CUDA-checkpoint
probably just AI slop and using wrong semantics, they mean speedup ratio.
bfeynman··on Bun Rust rewrite: "codebase fails basic miri checks, allows for UB in safe rust"
it's more straightforward to write safe rust when rust owns everything, In real world you often are interfacing with underlying libs or systems etc, which you need to treat as invariants but also handle yousrelf manually to make guarantees to compiler. unsafe exists in tons of codebases it's just you have to make sure you encapsulate it properly, which is what this bug is.
bfeynman··on Quack: The DuckDB Client-Server Protocol
Very hyped for this and updates. Have been using my own workarounds for a while with own WAL things and then sort of generating snapshots which with duckdb is so cheap was simpler than really implementing concurrent writes and mutations but this will make it so much easier.
bfeynman··on Launch HN: Voker (YC S24) – Analytics for AI Agents
do you have experience as PMs? Looking at website, it looks like you just use llms to guess what categories are? Seems like trap for garbage in garbage out. Otherwise you would need someone technical to figure out how to setup the proper KPI monitoring things...
bfeynman··on The 'Hidden' Costs of Great Abstractions
A lot of software practices are not based on engineering but rather psychology, and based on the fact that software costs $$$ to develop and maintain. There has been a complete paradigm shift with AI and it has underlined all of this. It's plausible there is no need (in high level programming at least) for any sorts of best practices, design patterns, code hygiene because not only will persons not even be looking at it, but AI can just rewrite an entire service if it needs to add a small feature and refactor everything, nothing needs to be written to be scalable or extensible if this cost becomes free. This is getting closer and closer to reality every month.
bfeynman··on We gave an AI a 3 year retail lease and asked it to make a profit
How is it not analogous to data leakage? The claim is that the system works autonomously, or at minimum could, but there is effectively signal via human in the loop feedback. That's leakage into test time evaluation. Also the coding analogy is malappropriated, in that the llm is using its own signals autonomously in the environment. Using a kalman filter on a ICBM with its own sensors is analogous to the coding agent and is autonomous. A system where a human is course correcting based on signals/sensor data is what's presented here, that is not autonomous.
bfeynman··on We gave an AI a 3 year retail lease and asked it to make a profit
that is ... not correct? This is classic example of data leakage, the yes/no things are signals feeding back to the model influencing (and here, basically guiding) future decisions.
bfeynman··on We gave an AI a 3 year retail lease and asked it to make a profit
I feel bad that people have to read this. It's complete puffery, made up for clicks, and the biggest thing is the pure bravado with which a company says, "Hey, let's just waste a ton of money, all for a potential blog and marketing piece." This is not really automated in any fashion. I was dubious at first, but then I saw the screencaps showing the devs interacting with Luna via a Slack workflow with a human in the loop — meaning they're literally just proxying their own behavior through an LLM. This is no different than anyone who consults AI for any decision with context. To get even more technical on the fallacy: this is not automation, as there is data leakage at every step where there is a human in the loop. A broken clock is right twice a day; an LLM could cycle through 100 guesses to pick a number, but don't market that as an oracle. Aside from that, you could just look at the pictures and context (retail in SF) and assume making a profit here would be near impossible. An actual AI ceo would probably have immediately cancel the lease.
bfeynman··on Next Grok model training with 10T parameter model
Isn't what the leading labs are currently chasing after is not pretraining and massive parameters but enriched and deep fine tuning and post training for agentic tasks/coding? MoE with just new post training paradigms lets smaller models perform quite well, and much more pragmatic to scale inference with. Given that, this choice seems super odd, as the frontier labs seem to stay neck and neck, and I don't even see Grok being used in any benchmarks because of how poorly it performs
bfeynman··on Moving fast in hardware: lessons from lab to $100M ARR
Nice read but falls into a vast reductionist trap, a lot of survivorship bias dressed up as design philosophy or strategic bets. The context of decisions made decades ago != now, people were working under different constraints etc. Trying to frame the avionics example as the "subtractive" innovation is the most egregious, transistors were over 1000x times smaller, weight wasn't even a consideration.
bfeynman··on OpenAI Acquires TBPN
Robinhood did exact same thing, it's more for marketing reach and distribution stuff. Wouldn't be surprised in few years they let it go or spin it down, just paying for a funnel/some narrative control
bfeynman··on Tell HN: Litellm 1.82.7 and 1.82.8 on PyPI are compromised
pretty horrifying. I only use it as lightweight wrapper and will most likely move away from it entirely. Not worth the risk
bfeynman··on Launch HN: Captain (YC W26) – Automated RAG for Files
not to mention they are using 3p apis for everything.. gemini, reranking etc...
bfeynman··on Launch HN: Captain (YC W26) – Automated RAG for Files
I think I've lost count of how many of these start ups I've seen. But what I really cant fathom is that pricing which is completely out of band. You can already talk to files directly with gemini, just wrapping other apis etc makes no sense. This is even stuff now you can easily codegen entire solutions for esp object storage based ones. Don't see actual any value add or differentiators here. It's obviously not that secure, and ingestion pipeline/connectors are also commodity.
bfeynman··on How Codex Is Built
Does anyone else find the way they are writing this full marketing hubris that definitely misconstrues how most people would interpret this. Codex isn't a "model" that is self improving, it's using GPT to write code that is in a wrapper program that also uses GPT. Sure it's kind of neat loop for development, but why are they anthropomorphizing it so much? People designing chips don't say that computers are self evolving, even anthropic just says that claude (the model) writes most of claude code. Heck, you could use claude or any llm to write code for codex..
bfeynman··on Launch HN: Vela (YC W26) – AI for complex scheduling
Lot of puffery in this describing constraint and actual messy problems that you are all most likely just being thrown into the context for an llm agent... None of the case studies demonstrate complex scheduling at all and are just all individual serial threads. buffers, preferences and options are all simple. The hard part of scheduling is when you have multiple pending invites or invitations that have to resolve and track it down, if someone asks for a meeting on a day that you currently already have a pending invite for, and how far away that day is, and how important the relationship is etc...
bfeynman··on OpenClaw surpasses React to become the most-starred software project on GitHub
openclaw while cool just allowed a larger tranche of technophiles who didn't necessarily have all the skills/understanding or time to do a bunch of things that have been readily available for like over 1.5 years. There is value in that, but there is huge surge in the number of people who are even able to take advantage of the novelty. Reminds me of when hugging face came out with transformers and all of a sudden you no longer needed to wrestle with anaconda and order of installation for all the deps.
bfeynman··on Infrastructure decisions I endorse or regret after 4 years at a startup (2024)
All in on AWS and using GitOps with TF instead of much more feature rich CDK...
bfeynman··on DOGE Track
Basic knowledge of civic history and political science makes this point very salient. Anyone with a clue would know this from the beginning - that's why it was so terrifying to see what actually was motivating people, feels like the ultimate recipe for unchecked power and disaster with bad actors employing fools to do their bidding.
bfeynman··on I’m joining OpenAI
the ability to almost "discover" or create hype is highly valued despite most of the time it being luck and one hit wonders... See many of the apps that had virality and got quickly acquired and then just hemorrhaged. Openclaw is cool, but not for the tech, just some of the magic of the oddities and getting caught on somehow, and acquiring is betting that they can somehow keep doing that again.
bfeynman··on A sane but bull case on Clawdbot / OpenClaw
What part is new? the thing that took off is that it allows technophiles who couldn't probably flash a raspberry pi to feel like they are hackers. All of this stuff exists in tons of random AI apps that exist already it just wasn't really that much of a value add, there is just a virality of it reaching audiences that previously only knew how to use a chat app.
bfeynman··on Qwen3-TTS family is now open sourced: Voice design, clone, and generation
Not a great metric, research in academia doesn't necessarily translate to value. In the US they've poached so many academics because of how much value they directly translate to.
bfeynman··on Using an expensive model made our agent 75% cheaper
This stuff is funny because it shows how worthless the actual service is on top of it and its mostly a wrapper in race to bottom. Same thing when cursor talks about how growth was flatlining until gpt-4 came out... So many of these companies need to disappear.
bfeynman··on Show HN: I built a tool to create AI agents that live in iMessage
I feel like Apple makes this very hard and against their ToS right or in a gray area, aka shut it down at any moment. There are businesses that solely try to manage provisioning of verified apple accounts and having worked with a bunch of them reliability is pretty hard.
bfeynman··on Show HN: I built a tool to create AI agents that live in iMessage
Having worked with mobile messaging and AI and trying to wrestle with iMessage and WhatsApp because I'm pretty sure this goes against ToS as they monitor how many new contacts and messages you're sending and get flagged to be shut down... If this is not the case, it would be very interesting....
bfeynman··on Claude Code CLI was broken
this is funny in context of their main dev advocate constantly bragging about how claude writes all of his code for claude code cli....
bfeynman··on 27M Fewer Car Trips: Life After a Year of Congestion Pricing
In what world do you get that conclusion? Dense cities in other parts of the world rely on mass transit to move people. FSD so you can have self driving cars in the street -> ? -> increasing congestion. The point is you should have more effective volume transit not optimizing random ones. 1000 cars on FSD are an optimization better than 1000 taxi drivers, compared to a train or a few buses.
bfeynman··on Show HN: Use Claude Code to Query 600 GB Indexes over Hacker News, ArXiv, etc.
The manifold structure of embedding spaces isn't semantically uniform, you've found a nice little novelty thing but it's not rigorous, and using AI slop to name this vector algebra instead of finding or running a benchmark to show that its actually works better.
← PreviousPage 2 of 11Next →