HNHacker News
TopNewBestAskShowJobs

dcl

429 karma · joined April 15, 2014

submissionscomments
dcl··on Gemini 4 Argon
What harness you use? Have you tested more than 1?
dcl··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Is it though? I have found it much easier to fit/diagnose/improve simple classifiers on embeddings/hash vectorized features/etc than iterate on prompts, especially true when using the more modern LLM's (like gpt 5.6 Luna) where you can't even set the temperature to get any sense of determinism.

I do note though, that is infinitely easier to 'deploy' a Jev/LLM based solution than a data/model pipeline.

dcl··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
That is indeed reasonable. But you still need a bunch of trusted data to validate the accuracy of Jev.
dcl··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
I've used LLM's to classify things for numerous projects and it's always really hard to beat linear classifier/decision tree over simple embeddings or even the sklearn HashingVectorizer. However, you DO need to trust your validated data, ground truths for this to work - but you should have these anyway to validate a Jev or similar solution.
dcl··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
What is...?
dcl··on What heraldry and Japanese mon can teach about visual-identity generators
no scroll bar :(
dcl··on GPT-6 Sol and Luna
Interesting. I have been asking it to review code from Opus 5 and it found tonnes of issues, Opus agreed with the findings too.
dcl··on Pentium II at 600Mhz with Voodoo 3 Emulated on 86Box with M6 Mac Mini
A bunch of PS2 cards had this effect on me. Fifa 2001, Gran Turismo 3, etc.
dcl··on GPT-6 Sol and Luna
How have you found Muse Spark 1.3? It doesn't get much mention, despite pretty good benchmarks. I've been using a bit at home and find it quite good, often finding mistakes made by Opus 5.
dcl··on Why do we need human mathematicians anymore?
the smallest problem that cannot fit in a brain would be pretty interesting
dcl··on Why do we need human mathematicians anymore?
What do you think all the PhD's and post-doc students are doing...? It's not called graduate-descent for nothing :p
dcl··on Brood War Bench
In the early days of SC2, I remember people using genetric programming to optimize build orders. I remember a slightly unorthodox Zerg Roach Rush which was _really_ fast.
dcl··on Ask HN: What are you working on? (September 2026)
Using AI to catch up on reinforcement learning advances, trying to 'solve' Gin Rummy. Will move on harder games after that.
dcl··on GPT-6 Astra
That's kind of wild. Our first PC was a an IBM PS/2 486SX 33Mhz, 4MB RAM, that was purchased in 1993.
dcl··on GPT-6 Astra
They overclocked well though, I think you could run the 300Mhz chips at >400Mhz.

I also believe you could get motherboards that supported 2 Celeron chips. I have no idea how effective/useful it was, but it was certainly a cheap/interesting way to get multiple CPU's.

dcl··on Muse Spark 1.3
thank you
dcl··on Muse Spark 1.3
Thats not exactly 'validated'. Feels very noisy, it is not a good bar for either - does this code do what the user actually asked - is this code actually 'good'

There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.

I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.

That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.

dcl··on Muse Spark 1.3
Any more info on this?
dcl··on Muse Spark 1.3
Well thats very interesting. Thank you. Will be interesting to see how hard/easy it is to translate my Claude skills, loop design, etc to the new harness.

This kind of raises another question to me regarding the coding benchmarks, how much of it is model versus harness?

dcl··on Muse Spark 1.3
Very keen to try this after using Claude Code over the last few months. Should I just point Claude Code to Muse Spark endpoint (because I'm familiar with Code)? What do people think of Muse Code or other coding agent harnesses?
dcl··on Muse Spark 1.3
Hopefully, 'validated' AI code
dcl··on 2004 RuneScape fit a multiplayer RPG into 56k dial-up
I played this as a young teen in 1998 and it's still the most powerful gaming experience and memories I have.
dcl··on Transfer files over an Ethernet patch cable
I believe a lot of NIC's could autodetect and adapt to the cable? I'm pretty sure I had 2 computers connected to each other using a regular network cable.
dcl··on Apple introduces M6 and M5 Ultra
iPods were a very popular Apple product once upon a time...
dcl··on Memelang: Token-Terse Query Language
Question: For things like this to be truly worthwhile, do you have to ensure the training, post-training of LLM's is filled with brilliant memelang code? The models are probably already good at 'thinking' in SQL at the moment, but if they aren't trained on memelang, surely they will have to use more effort to translate in context to produce equivalent quality queries?

Wouldn't surprise me if in the next few generations we start seeing more LLM generated languages that LLM's prefer to use for expressability, conciseness, etc.

dcl··on Agent Skill to Force Docs in ASD-STE100 Simplified Technical English
One question I have about stuff like this: How does it affect the models intelligence or thinking ability. If this modifies the output or chain of thought in any way, it may impact what the model is capable of right? Especially if it's not trained to use this kind of language during training.
dcl··on Benchmarking Opus 5 on SlopCodeBench
finally the benchmark for me
dcl··on Simulate cassette tape audio profiles using FFmpeg
vaporwave fans rejoice
dcl··on The unreasonable difficulty of time series forecasting
I have to explain this to managers, execs and stakeholders all the time. ML models work great for systems where the rules/dynamics do not change over time. With Forecasting, in a lot of domains where you want a forecast, everything is subject to change - laws, policies, regs, customers appetites, competitors behaviours, etc.
dcl··on GLM 5.2 and the coming AI margin collapse
That is going to be absolutely wild for whoever can access/afford it.
Page 1 of 7Next →