HNHacker News
TopNewBestAskShowJobs

practice9

539 karma · joined March 28, 2018

submissionscomments
practice9··on AI handles incidents, engineers lose touch with their systems
You rather generously assume engineers are in touch with their systems.

Even before layoffs many teams just maintained things org has long lost coherent knowledge of

After layoffs and typical org knowledge churn - you can either rewrite it (but how? Product team responsible for original implement requirements is long gone too) or recoup (reverse document) some of that lost knowledge with AI and actually learn

practice9··on I believe there are entire companies right now under AI psychosis
It's because HN is in AI meta-psychosis :)

Our experience is very similar except we didn't really have a review process before, and now LLMs find bugs before PRs get merged in main.

We had 5x-100x speedups in some legacy but important pipelines, with no regressions (validated after extensively by humans). It's not that the code was actively bad. It's just only 1-5% people in the local SWE market would be able to write code that runs so fast and efficient and benchmark it correctly.

We found a subtle correctness bug that was in production for half of the decade (both GPT-5 and Claude Opus were able to find it), confirmed by human after.

And we keep finding subtle bugs that have been introduced by humans before (despite the human reviews, the particular domain is just difficult no matter how many docs and comments and tests one writes)

practice9··on Claude for Excel
LLMs are getting quite good at reviewing the results and implementations, though
practice9··on Building the mouse Logitech won't make
I find it hilarious/sad that the 0.5x cheaper Ergo M575 has much better design in that regard (just plastic that doesn’t degrade)
practice9··on I gave the AI arms and legs then it rejected me
They should have used Claude Code for reviews
practice9··on AI coding agents are removing programming language barriers
Humans cannot reason about code at scale. Unless you add scaffolding like diagrams and maps and …

Things that most teams don’t do or half-ass

practice9··on Supreme Court's ruling practically wipes out free speech for sex writing online
A variation of “no taxation without representation”?
practice9··on Why Claude's Comment Paper Is a Poor Rebuttal
The human is a bad co-author here really.

I deployed lots of high performance, clean, well documented etc code generated by Claude or o3. I reviewed it wrt requirements, added tests and so on. Even with that in mind it allowed me to work 3x faster.

But it required conscious effort on my part to point out issues and inefficiencies on LLMs part.

It is a collaborative type of work where LLMs shine (even in so called agentic flows)

practice9··on Sycophancy in GPT-4o
Well the system prompt is still the same for both models, right?

Kinda points to people at OpenAI using o1/o3/o4 almost exclusively.

That's why nobody noticed how cringe 4o has become

practice9··on Whistleblower details how DOGE may have taken sensitive NLRB data
Kinda similar in a way to China or Russia “disappearances”
practice9··on Gemma 3 Technical Report [pdf]
But who is the target group?

Last time only some groups of enthusiasts were willing to work through bugs to even run the buggy release of Gemma

Surely nobody runs this in production

practice9··on Why LLMs still suck at OCR
I tried the square example from the paper mentioned with o1-pro and it had no problem counting 4 nested squares…

And the 5 square variation as well.

So perhaps it is just a question of how much compute you are willing to throw at it

practice9··on Order Declassifying JFK and MLK Assassination Records [pdf]
With LLMs that even might be automated
practice9··on GPT-4o with scheduled tasks (jawbone) is available in beta
Well none of the labs have good frontend or mobile engineers or even infra engineers

Anthropic is ahead in this because they keep their UIs simplistic so the failure modes are also simple (bad connection)

OpenAI is just pushing half baked stuff to prod and moving on (GPTs, Canvas).

Find it hilarious and sad that o1-pro just times out thinking on very long or image-intense chats. Need to reload page multiple times after it fails to reply and maybe answer will appear (or not? Or in 5 minutes?). Kinda shows they’re not testing enough and “not eating their own food” and feels like chatgpt 3.5 ui before the redesign

practice9··on Learning not to trust the All-In podcast
One of those guys needs to be fined for the pump & dump scheme (with SPCE: Virgin Galactic), and the other one should be investigated if he was receiving money from the Russian government or influence agents.
practice9··on UK will give sovereignty of Chagos Islands to Mauritius
More like reparations
practice9··on Floating megabomb heaves to near the English coast
The interesting thing is that ship was damaged almost immediately after leaving the port, had a chance to stop in Russian ports along the way but instead is doing a tour near EU countries.

The crew is either amazingly incompetent or malicious/complicit.

You don’t put a ship that can blow up near: a. Gas&oil terminals, b. military air base

practice9··on GPT-4o
It is cringe overenthusiastic, but a proper instructions/system prompt will fix that mostly
practice9··on GPT-4o
I always double-check even the most obscure facts returned by GPT-4 and have yet to see a hallucination (as opposed to Claude Opus that sometimes made up historical facts). I doubt stuff interesting to kids would be so out of the data distribution to return a fake answer.

Compared to YouTube and Google SEO trash, or Google Home / Alexa (which do search + wiki retrieval), at the moment GPT-4 and Claude are unironically safer for kids: no algorithmic manipulation, no ads, no affiliated trash blogs, and so on. Bonus is that it can explain on the level of complexity the child will understand for their age

practice9··on OpenAI Bought Chatgpt.com
Wasn't it pointing to https://x.ai just a few months ago? Interesting
practice9··on World_sim: LLM prompted to act as a sentient CLI universe simulator
teknium / Nous released Mistral finetunes (Hermes) that are quite great, and even published the datasets used for training.

But for the worldsim I think they are really using Claude (probably Haiku or Sonnet) via openrouter (https://openrouter.ai/).

practice9··on Meta outage
Interesting. I had a problem a few months ago with DNS not resolving Meta servers on my Starlink internet connection, but I was able to use the UI and the apps nonetheless, just couldn't open the store or update firmware.

Seems like they really did change something in the latest firmwares.

practice9··on Meta outage
Unless it's a recent change, it works perfectly fine offline (wifi turned off).

As for alternatives, there is Pico, but Quest 3 may be superior in games selection. Or go wired which is of course less portable

practice9··on Russian cosmonaut sets record for most time in space – more than 878 days
Here is an interesting and relevant context: there is a huge amount of evidence that Roscosmos is taking an active part in the war effort. There is a great source about it from Eric Berger of Ars Technica: https://arstechnica.com/space/2023/06/it-appears-that-roscos...

Take that into account when reading news made from state press-releases like the one in the post.

P.S.: For the full context, the Roscosmos ex-boss also has had his own private military company for a while.

practice9··on How to get coworkers to stop giving me ChatGPT-generated suggestions?
But you can distill a wall of PR fluff to a concise message with ChatGPT... The problem is the average user thinks "long text = good" and their bosses often agree.
practice9··on Contra Wirecutter on the IKEA air purifier (2022)
Aren't most websites like Wirecutter in the business of paid reviews / affiliate marketing?
practice9··on BYD's YangWang U8 launched, can float on water for 30 minutes and sail 3km/h
That car's exterior is uglier than Cybertruck
practice9··on OnnxStream: Stable Diffusion XL 1.0 Base on a Raspberry Pi Zero 2
Wait, but the new title doesn't seem to be correct
practice9··on Can generalist foundation models beat special-purpose tuning?
Yeah the Medprompt name is misleading
practice9··on Does GPT-4 Pass the Turing Test?
It was trained to respond like that to not alienate groups of people. But fine-tuned and with another pre-prompt, it would give absolutely different answer.
Page 1 of 9Next →