HNHacker News
TopNewBestAskShowJobs

DetroitThrow

1,118 karma · joined March 19, 2020

submissionscomments
DetroitThrow··on Gemini 4 Argon
It was never built with an open source community in mind. It was always DOA in a world where Rust existed.
DetroitThrow··on Laya the open source version of Jev
It would be amazing to have big BERTha with per-token pricing on GCP or AWS. There are many times I am reaching for a cheap classifier with the general behavior of an LLM.
DetroitThrow··on GitHub is having trouble counting things
I've switched off to runners and event scheduling inside AWS, and am working on moving my company off of GitHub. I know engineers inside GitHub and it doesn't sound like the talk they've been putting out about reliability is actually being addressed with many resources internally; the best people in the company are chasing AI product features (which are apparently quite profitable).
DetroitThrow··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Yes exactly. Automatic model downgrade seems horrible for a lot of production workloads, even if you are deterministically constraining the behavior of your agents.
DetroitThrow··on On the Navier–Stokes Millennium Prize Problem
I don't think the person describing the paper by Buckmaster as the same as the 2023 paper by Córdoba and Martínez-Zoroa is really discussing this in good faith fwiw. There are some massive advancements within it and if the person was participating wasn't just regurgitating something to "win the argument" in their eyes, they wouldn't describe it that way.
DetroitThrow··on On the Navier–Stokes Millennium Prize Problem
>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.

If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.

Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.

Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.

DetroitThrow··on On the Navier–Stokes Millennium Prize Problem
It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the question is dodged and words gain new meaning. The question can be answered with "Yes, we trained on their conversations" or "No, we did not train on their conversations".

I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.

DetroitThrow··on OpenAI fought dirty on career-making math problem
Given that he has other former collaborators corroborating this horrific behavior, it seems like this a career spanning pattern, and it's interesting to see just how much @sama is willing to lend his support to someone like Bubeck.

Stains an important moment in the history of AI progress for me. The future seems bleak with people like this at the reins.

https://x.com/dheeraj_nagaraj/status/2097266146445774924

DetroitThrow··on GPT 6 Astra
404?
DetroitThrow··on Isitdoneyet.gg is a website I made to figure out if games are complete
Can't check whether star citizen is done yet :/
DetroitThrow··on The OG Creator of Task Manager on Windows Built a New Task Manager
I don't mind color, the glow effects just hurt the ability to read what's there (this reminds me of the gaming setups teenagers are drawn to)...but there's no reason a utility like a task manager should be closed source or have a pro version.
DetroitThrow··on The OG Creator of Task Manager on Windows Built a New Task Manager
Wow, thought the guy would be a MSFT OG but this looks like it sucks? Vibed, closed source, paid extensions.. Why on earth would anyone use this?
DetroitThrow··on Previewing the Model Hardware Standard
It's a bit of both for Anthropic I think, sometimes cutting edge and quite interesting or just good improvements, sometimes ignoring best practices either recently established or known for decades. Obvious to see where the smart people are in high places at the company.
DetroitThrow··on GLM-5.3-Flash
Lately I've been throwing tasks at Qwen and a frontier or recently-frontier model (as well as Kimi, GLM, etc) and the smaller parameter models are not really comparable to Opus when it comes to making intelligent decisions about greyer areas of good software architecture.

Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).

DetroitThrow··on Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
It's still not as good as GPT5.6 or Opus5 but it's better than KimiK3. Good job xAI team.
DetroitThrow··on Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.

That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.

DetroitThrow··on We replaced Redis with MySQL for inventory reservations and it scaled
>Aren't the LLMs trained on a massive corpus of human written texts? If that stands, then they are doing what they were asked, kind of?

I think you might be interested in reading in training data generation, training, and post training papers/articles. I think you might be surprised at how much intervention there is on some of these levels.

DetroitThrow··on Not hiring junior engineers won't solve the problem you think you have
This blogpost didn't really address the elephant in the room that increasingly the bottleneck is no longer the part that jr devs could help out with.

Every startup I know that hired some level of jr's and encouraged them to use AI found themselves in a code review bottleneck for basic code quality and architecture decisions. Many startups I know are largely forgoing jr devs.

Code quality regarding evolvability, reliability, maintainability is still part of the development process. All of the people who claimed src code is just assembly on Twitter some months ago have started chiming in that they were wrong.

I do wonder what this does to the talent pipeline like mentioned in the post.

DetroitThrow··on Bending Spoons makes first post-IPO acquisition with $1.3B Airtable deal
They also hike the price so significantly most people stop using it. See meetup.com for example.
DetroitThrow··on Bending Spoons makes first post-IPO acquisition with $1.3B Airtable deal
yes, _every_ event I go uses Luma or something else nowadays.
DetroitThrow··on SpaceX Starship Flight 13 livestream [video]
How many times have they managed to catch it in the last 8 or so missions? There have been a few misses like this already with the boosters
DetroitThrow··on India's first privately-developed rocket reaches orbit on debut launch
>FFCS is called holy grail of liquid engines

Particularly for reusable liquid engines :)

DetroitThrow··on GPT-5.6
Looks like they reset everyone's Fable usage.
DetroitThrow··on GPT-5.6
DeepSWE seems to strongly, strongly prefer ChatGPT models. There were also major flaws in its methodology pointed out recently, that overlap strongly with the flaws OpenAI pointed out in its SWE Verified report.

I use both ChatGPT and Claude for engineering work on a daily basis, touching performance critical code to application backends to frontend work, and I've found that DeepSWE scores don't reflect my reality when I assess high quality output from the models/harnesses.

Not that Opus always beats GPT 5.5., but that 5.5 is ahead of Opus on a general benchmark smells off to me.

DetroitThrow··on Amazon without the knockoffs
>I'll be more peeved if they monetize it

FSL (vs a copyleft license or just plain old OSS) implies they want to turn this into a revenue source for themselves ultimately, unfortunately.

>Maybe I should put one of those buy me a coffee links on the repo

Absolutely :) Cool project.

DetroitThrow··on Amazon without the knockoffs
On top of that, it has a more restrictive license than AmazonBrandFilter. Given this appears to be a very simple AI project, why not just reimplement any missing functionality from AmazonBrandFilter into something under a free license? The most difficult to duplicate component is MIT.
DetroitThrow··on US Supreme Court rules geofence warrants require constitutional protections
She's not as big on some of the broader interpretations of the 4th amendment that more civil liberty minded justices would lend credence to.
DetroitThrow··on The CEO of Mullvad is the main financer of the Swedish Örebro party
He's entitled to his political views and just as we're entitled to potentially use or not use his service because of them :)

Not sure why it's such an issue to discuss the political views of the beneficiaries of services we use. I understand it's mostly uninteresting as far as comment sections go, but it's always bizarre to see a defense of political association when often the impetus for sharing this type of information is for people/consumers to exercise their right to associate with business based on their political outlook.

DetroitThrow··on Michigan bill would bar employers from requiring after-hours coms with workers
Everyone gets to share but it's also completely within the forum rules to call out irrelevant anecdotes as uninteresting to the discussion.

I have no idea why you're making a comparison to a TV show; nothing that was described was anything akin to that. I just made examples out of insufferable and clueless forum comments, that very clearly detract from discussion more than they contribute to it.

I don't think you should assume that describing meaningless and unrelated anecdotes as "uninteresting" is equivalent to users calling for a forum ban, which is seemingly what you're doing when you point to forum rules when encountering a critique.

DetroitThrow··on Fintech Engineering Handbook
When performance isn't a concern, I largely agree! Not every financial system can use big decimal as their base, though, too. And HFT isn't the only place in the financial sector where this performance concern might pop up.
Page 1 of 18Next →