HNHacker News
TopNewBestAskShowJobs

daemonologist

2,585 karma · joined August 9, 2023

knçhwl7tg@mozmail.com (remove the cedilla from the 'c')
submissionscomments
daemonologist··on Waymo says can't avoid bike lanes because riders want to be dropped off in them
Going around a parked car is not merely an inconvenience, it introduces an extra risk of being hit from behind (obviously you should check over your shoulder before moving into the lane, but this is the imperfect real world, and even the act of checking over your shoulder is a small risk) or by a vehicle pulling out of a cross street which didn't see you through the stopped car.

However I agree that there isn't an obvious solution without making major improvements to infrastructure - right now where the bike lane is just paint everyone parks in it (Uber, taxis, delivery drivers, etc.).

daemonologist··on DeepSeek-V4 Technical Report [pdf]
$1.47/M input, $3.48/M output, open weights (MIT license), and competitive with the frontier on their selected benchmarks. Big if it holds up on real-world tasks.
daemonologist··on 3.4M Solar Panels
I'm in the US and it's showing a 100W panel for USD 37.21 (free shipping, including tariffs but not state/local taxes).

Also the panels Carter installed were solar water heaters - in 1979 solar photovoltaics were just starting to expand beyond satellites and cost like $40/watt.

daemonologist··on ChatGPT Images 2.0
SynthID survives basic transforms including screenshots/photos, although it can of course be defeated. Even still it helps with the laziest fakes, which there seem to be a lot of - I've seen several quite widespread misinformative images over the past couple months that failed a synthID check.

Anyways I think approaching the problem from both directions is probably good.

daemonologist··on ChatGPT Images 2.0
There are a couple of AI-esque misspellings - in the More Myth than Menace wolves image, on the right in the "at a glance" section, it reads "wolves aarely approach people," and in the Typography image the text in the top right is "Type connncts us all."

But yeah the quality is remarkable, and rather scary.

daemonologist··on Claude Code to be removed from Anthropic's Pro plan?
I have to assume they're compute constrained and thus need to either raise prices or cut their lowest-margin products (which amounts to more or less the same thing, but with different optics), or turn away new users.

My assumption is that people are able to very easily saturate Pro with Claude Code and therefore even though the quotas are lower (more than proportionally) the utilization of those quotas is higher enough that Pro is less profitable.

daemonologist··on Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
Not that I'm aware of. It's kind of like building a PC or a bicycle - you're putting mostly-standardized parts together rather than starting from first principles, but there are so many permutations that you can either use a single known-good configuration or immerse yourself in forums and tinker until you can fit things together yourself. Plus both the inference engines and models are of course moving really fast.

I use Opus 4.7 in Claude Code lol, plus Zed (as a text editor, not a harness). Open-weights models that I can run are for me not useful for multi-turn ("agentic") tasks. I do use Qwen 3.6 for one-off tasks like "write a function to pretty-print this weird data structure" or "explain this config file," and Gemma 4 26B for non-coding tasks like "create a timestamped table of contents from this podcast transcript."

daemonologist··on Show HN: Ctx – a /resume that works across Claude Code and Codex
I'd use it if I hit the 5 hour quota mid-change and then came back later in the day in a new terminal (depending on the input/output ratio of my now un-cached context, of course).
daemonologist··on Framework Laptop 13 Pro
The touchscreen is backward-compatible with the old/regular FW13, so I imagine the regular FW13 screen is forward-compatible with the Pro. (Of course, I don't know if they'll sell that configuration or if you'd have to cobble it together from the marketplace.)
daemonologist··on Framework Laptop 13 Pro
All four ports support Thunderbolt 4 - if you scroll down to "Interfaces" on the product specs page there's a graphic showing everything that's supported.
daemonologist··on Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
First of all nothing you can run locally, on that machine anyways, is going to compare with Opus. (Or even recent Sonnet tbh - some small models benchmark better but fall off a bit in the real world.) This will get you close to like ~Sonnet 4 though:

Grab a recent win-vulkan-x64 build of llama.cpp here: https://github.com/ggml-org/llama.cpp/releases - llama.cpp is the engine used by Ollama and common wisdom is to just use it directly. You can try CUDA as well for a speedup but in my experience Vulkan is most likely to "just work" and is not too far behind in speed.

For best quality, download the biggest version of Qwen 3.5 27B you can fit on your 4090 while still leaving room for context and overhead: https://huggingface.co/unsloth/Qwen3.5-27B-GGUF - I would try the UD-Q5_K_XL but you might have to drop down to Q5_K_S. For best speed, you could use Qwen 3.6 35B-A3B (bigger model but fewer parameters are active per token): https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF - probably the UD-Q4_K_S for this one.

Now you need to make sure the whole model is fitting in VRAM on the 4090 - if anything gets offloaded to system memory it's going to slow way down. You'll want to read the docs here: https://github.com/ggml-org/llama.cpp/tree/master/tools/serv... (and probably random github issues and posts on r/localllama as well), but to get started:

  llama-server -m /path/to/above/model/here.gguf --no-mmap --fit on --fit-ctx 20000 --parallel 1
This will spit out a whole bunch of info; for now we want to look just above the dotted line for "load_tensors: offloading n/n layers to GPU" - if fewer than 100% of the layers are on GPU, inference is going to be slower and you probably want to drop down to a smaller version of the model. The "dense" 27B will be slowed more by this than the "mixture-of-experts" 35B-A3B, which has to move fewer weights per token from memory to the GPU.

Go to the printed link (localhost:8080 by default) and check that the model seems to be working normally in the default chat interface. Then, you're going to want more context space than 20k tokens, so look at your available VRAM (I think the regular Windows task manager resource monitor will show this) and incrementally increase the fit-ctx target until it's almost full. 100k context is enough for basic coding, but more like 200k would be better. Qwen's max native context length is 262,144. If you want to push this to the limit you can use `--fit-target <amount of memory in MB>` to reduce the free VRAM target to less than the default 1024 - this may slow down the rest of your system though.

Finally, start hooking up coding harnesses (llama-server is providing an OpenAI-compatible API at localhost:8080/v1/ with no password/token). Opencode seems to work pretty reliably, although there's been some controversy about telemetry and such. Zed has a nice GUI but Qwen sometimes has trouble with its tools. Frankly I haven't found an open harness I'm really happy with.

daemonologist··on Soul Player C64 – A real transformer running on a 1 MHz Commodore 64
You can chat with the model on the project page: https://indiepixel.de/meful/index.html

It (v3) mostly only says hello and bye, but I guess for 25k parameters you can't complain. (I think the rather exuberant copy is probably the product of Claude et al.)

daemonologist··on All phones sold in the EU to have replaceable batteries from 2027
I believe part of the legislation is that manufacturers must make spare parts available for five years.
daemonologist··on GitHub's Fake Star Economy
I remember talking to some of the folks running UIUC's hackathon (probably ten years ago) and they'd built a sort of page-rank for Github - hand-identifying the most prominent and reputable projects/individuals and then using follows and stars to transfer that reputation. I don't know how well it worked in practice or if it was every published, but it might be more effective than pure star count.

(This was for admissions iirc - they had limited slots and a portion of them were allocated to people with a strong github rank.)

daemonologist··on NSA is using Anthropic's Mythos despite blacklist
> This puts the US government into a loose / loose position.

You might even call it... a tight spot

daemonologist··on The RAM shortage could last years
AMD has built some consumer GPUs in the recent past with HBM - RX Vega and Radeon VII (although I assume not all "HBM" is created equal).
daemonologist··on Changes in the system prompt between Claude Opus 4.6 and 4.7
Note that these are the "chat" system prompts - although it's not mentioned I would assume that Claude Code gets something significantly different, which might have more language about malware refusal (other coding tools would use the API and provide their own prompts).

Of course it's also been noted that this seems to be a new base model, so the change could certainly be in the model itself.

daemonologist··on Claude Opus 4.7 Model Card
The benchmark GP mentioned is measuring at 128k-256k context (there's another at 524k-1024k, where 4.6 scored 78.3% and 4.7 scored 32.2%).

The longer the context the worse the performance; there isn't really a qualitative step change in capability (if there is imo it happens at like 8k-16k tokens, much sooner than is relevant for multi-turn coding tasks - see e.g. this old benchmark https://github.com/adobe-research/NoLiMa ).

daemonologist··on Qwen3.6-35B-A3B: Agentic coding power, now open to all
I didn't see a direct comparison, but there's some overlap in the published benchmarks:

                           │ Qwen 3.6 35B-A3B │ Haiku 4.5               
   ────────────────────────┼──────────────────┼──────────────────────── 
    SWE-Bench Verified     │ 73.4             │ 66.6                    
   ────────────────────────┼──────────────────┼──────────────────────── 
    SWE-Bench Multilingual │ 67.2             │ 64.7                    
   ────────────────────────┼──────────────────┼──────────────────────── 
    SWE-Bench Pro          │ 49.5             │ 39.45                   
   ────────────────────────┼──────────────────┼──────────────────────── 
    Terminal Bench 2.0     │ 51.5             │ 61.2 (Warp), 27.5 (CC)  
   ────────────────────────┼──────────────────┼──────────────────────── 
    LiveCodeBench          │ 80.4             │ 41.92                   

These are of course all public benchmarks though - I'd expect there to be some memorization/overfitting happening. The proprietary models usually have a bit of an advantage in real-world tasks in my experience.
daemonologist··on Qwen3.6-35B-A3B: Agentic coding power, now open to all
The 3B active is small enough that it's decently fast even with experts offloaded to system memory. Any PC with a modern (>=8 GB) GPU and sufficient system memory (at least ~24 GB) will be able to run it okay; I'm pretty happy with just a 7800 XT and DDR4. If you want faster inference you could probably squeeze it into a 24 GB GPU (3090/4090 or 7900 XTX) but 32 GB would be a lot more comfortable (5090 or Radeon Pro).

122B is a more difficult proposition. (Also, keep in mind the 3.6 122B hasn't been released yet and might never be.) With 10B active parameters offloading will be slower - you'd probably want at least 4 channels of DDR5, or 3x 32GB GPUs, or a very expensive Nvidia Pro 6000 Blackwell.

daemonologist··on Qwen3.6-35B-A3B: Agentic coding power, now open to all
No - this model has the weights memory footprint of a 35B model (you do save a little bit on the KV cache, which will be smaller than the total size suggests). The lower number of active parameters gives you faster inference, including lower memory bandwidth utilization, which makes it viable to offload the weights for the experts onto slower memory. On a Mac, with unified memory, this doesn't really help you. (Unless you want to offload to nonvolatile storage, but it would still be painfully slow.)

All that said you could probably squeeze it onto a 36GB Mac. A lot of people run this size model on 24GB GPUs, at 4-5 bits per weight quantization and maybe with reduced context size.

daemonologist··on Taking on CUDA with ROCm: 'One Step After Another'
ROCm usually only supports two generations of consumer GPUs, and sometimes the latest generation is slow to gain support. Currently only RDNA 3 and RDNA 4 (RX 7000 and 9000) are supported: https://rocm.docs.amd.com/projects/install-on-linux/en/lates...

It's not ideal. CUDA for comparison still supports Turing (two years older than RDNA 2) and if you drop down one version to CUDA 12 it has some support for Maxwell (~2014).

daemonologist··on Taking on CUDA with ROCm: 'One Step After Another'
Among consumer cards, latest ROCm supports only RDNA 3 and RDNA 4 (RX 7000 and RX 9000 series). Most stuff will run on a slightly older version for now, so you can get away with RDNA 2 (6000 series).
daemonologist··on 20 years on AWS and never not my job
Good domain name.
daemonologist··on Helium is hard to replace
It's often found alongside natural gas because the rock structures that can trap methane can also trap other gasses, but the original source is different - thermal decomposition of organic matter for natural gas and radioactive decay, mostly of uranium and thorium, for helium.

I agree that the "accumulation over millions of years" is similar (and similarly a potential problem if we burn through all that accumulation).

daemonologist··on I still prefer MCP over skills
Obvious example is a corporate chatbot (if it's using tools, probably for internal use). Non-technical users might be accessing it from a phone or locked-down corporate device, and you probably don't want to run a CLI in a sandbox somewhere for every session, so you'd like the LLM to interface with some kind of API instead.

Although, I think MCP is not really appropriate for this either. (And frankly I don't think chatbots make for good UX, but management sure likes them.)

daemonologist··on Study found that young adults have grown less hopeful and more angry about AI
The idea is the opposite - "nobody" can make money selling software anymore, because software can be cheaply created by an LLM, so you want to start a business that previously would have had to buy software/software engineers in order to support some other product.

However, even if that holds true (which is a big if - right now I wouldn't want to run a business backed by vibe software), and even if there are enough such business ideas to go around, there's going to be quite a lot of turmoil in the meantime.

daemonologist··on Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
Whisper is still old reliable - I find that it's less prone to hallucinations than newer models, easier to run (on AMD GPU, via whisper.cpp), and only ~2x slower than parakeet. I even bothered to "port" Parakeet to Nemo-less pytorch to run it on my GPU, and still went back to Whisper after a couple of days.
daemonologist··on AI singer now occupies eleven spots on iTunes singles chart
It's interesting to me that all AI music sounds slightly sibilant - like someone taped a sheet of paper to the speaker or covered my head in dry leaves. I know no model is perfect but I'd have thought they'd have ironed out this problem by now, given how pervasive it is and how significantly it degrades the end product.
daemonologist··on F-15E jet shot down over Iran
If you scroll to the bottom of that page, they discuss possible evidence of damage to the radar from satellite imagery.
← PreviousPage 4 of 24Next →