HNHacker News
TopNewBestAskShowJobs

ac29

7,346 karma · joined January 20, 2012

email: hn@imap.cc
submissionscomments
ac29··on Codex on AWS bedrock bug causing 10x charges
There haven't been any free resets in the past week, there were 4 in the first half of the month
ac29··on GLM-5.3 Artificial Analysis Benchmarks
Cost per task:

  Model                        Score    Cost / Task    Output Tokens / Task
  -------------------------------------------------------------------------
  Muse Spark 1.2 (xhigh)        56.8          $0.40                  30,430
  Gemini 3.7 Flash (high)       56.0          $0.40                  36,847
  GPT-5.6 Terra (max)           56.6          $0.51                  20,838
  GPT-5.6 Sol (high)            57.3          $0.52                   7,545
  GLM-5.2 (max)                 53.0          $0.56                  32,200
  GLM-5.3 (max)                 59.5          $0.68                  41,107
  GPT-5.5 (xhigh)               56.3          $0.69                  16,893
  Grok 4.6 (high)               60.9          $0.84                  21,735
  Kimi K3 (max)                 59.7          $0.84                  25,474
  GPT-5.6 Sol (xhigh)           59.0          $0.87                  11,098
  Qwen3.8 2.4T A95B             57.7          $0.95                  32,472
  Claude Opus 5 (medium)        58.6          $0.98                  12,459
  Qwen3.8 Max                   58.1          $1.13                  38,287
  GPT-5.6 Sol (max)             60.9          $1.23                  16,879
  Claude Opus 5 (high)          61.5          $1.52                  21,353
  Claude Opus 4.8 (max)         57.3          $1.65                  33,557
Benchmark score:

  Model                        Score    Cost / Task    Output Tokens / Task
  -------------------------------------------------------------------------
  Claude Opus 5 (high)          61.5          $1.52                  21,353
  GPT-5.6 Sol (max)             60.9          $1.23                  16,879
  Grok 4.6 (high)               60.9          $0.84                  21,735
  Kimi K3 (max)                 59.7          $0.84                  25,474
  GLM-5.3 (max)                 59.5          $0.68                  41,107
  GPT-5.6 Sol (xhigh)           59.0          $0.87                  11,098
  Claude Opus 5 (medium)        58.6          $0.98                  12,459
  Qwen3.8 Max                   58.1          $1.13                  38,287
  Qwen3.8 2.4T A95B             57.7          $0.95                  32,472
  Claude Opus 4.8 (max)         57.3          $1.65                  33,557
  GPT-5.6 Sol (high)            57.3          $0.52                   7,545
  Muse Spark 1.2 (xhigh)        56.8          $0.40                  30,430
  GPT-5.6 Terra (max)           56.6          $0.51                  20,838
  GPT-5.5 (xhigh)               56.3          $0.69                  16,893
  Gemini 3.7 Flash (high)       56.0          $0.40                  36,847
  GLM-5.2 (max)                 53.0          $0.56                  32,200
ac29··on Gemini 3.7 Flash
> Using Google products in general is an effing nightmare as soon as you have to give them money

Spending money via Google Pay on Android is extremely easy, Google does know how to accept customer's money (in the consumer space)

ac29··on Gemini 3.7 Flash
Last time I used the AI studio free tier it was limited to one or two requests, effectively useless. I think people are complaining its too hard to set up a paid API key (no reason you should make paying customers spend more than a few clicks and a minute of their time to pay you)
ac29··on Grok 4.6
Muse Spark 1.2 benchmarks just shy of Opus/Sol and is significantly less expensive than Sonnet (which admittedly is overpriced). Haven't personally used it though
ac29··on Grok 4.6
China has as least much engineering talent as the US, the claims that Chinese models must just be distilled US models feels like xenophobia.

There are more plausible explanations for why the models are similar - all the labs are buying the same datasets from third parties

ac29··on Grok 4.6
I was under the impression labs released post-training checkpoints fairly often?

So Model N/N+1 might literally have had the exact the same pretraining run and only differ on how much/what kind of postraining they got

ac29··on ChatGPT Desktop (Codex Desktop) for Linux
I haven't had bluetooth issues in years, but I did have an agent reverse engineer a smartphone app that was required for programming some BT headphones. Now I can push my desired runtime settings automatically to the headphones when they connect to my computer

So, yes, I would say agents are pretty good at working with Bluetooth on Linux

ac29··on Codex in ChatGPT desktop app for Linux is now in preview
Your earlier post is misleading, since "it" sounds like the model, not the installer. (of course if the latter was created by the former, its basically the same complaint, but I suspect what you are talking about was a deliberate design decision by a human)
ac29··on Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks
Any idea what pricing will be like? The fact that you call out Drone and Defense industries on your page suggests to me this is extremely expensive.

I recently priced out getting a PCB done in China and its insanely cheap. My design isnt complicated, a few ICs, a few connectors, an a dozen or so passives. PCB manufacturing, parts, and soldering/assembly is on the order of $10-20 total per board and that is with zero volume discounts. The parts alone would cost that much in the US, and I suspect the sort of companies that you contract out to also aren't really interested in tiny orders so they would probably quote huge setup fees (looks like your average order so far is ~$7000).

ac29··on Born Against, or why hobby programming communities are against LLM usage
One can piss away a lot of money gardening with little to show for it, especially if you buy bagged soil. Buy a raised planter, some soil, some fertilizer, some seedlings, a couple garden tools and you can easily be in for $150+ which is definitely more than you'll get out of a season worth of growth even if nothing fails.

On the other end of the spectrum, if you grow in fertile native soil from seed your costs can be close to zero

ac29··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
For reference, IBM was founded in 1911. At the time "computer" was a profession, not a machine, and the majority of US households had no electricity or telephone service
ac29··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
Not so much hardware these days, about half of IBM's revenue comes from software (Red Hat related stuff alone probably brings in a ton), and about 20% comes from consulting
ac29··on Muse Code and Muse Spark 1.2
Doesnt seem suspect to me, training runs have checkpoints and there is no reason you cant release a checkpoint even if you are still training the model
ac29··on Muse Code and Muse Spark 1.2
> They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.

If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2

ac29··on The Warp Agent CLI
Yeah Anthropic is the only company I am aware of that only allows subscription usage in first party products
ac29··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
> Apple has worked very hard to make unified memory a feasible approach, and the benefits of that are pretty clear in apple silicon- that efficiency not only results in power and therefore thermal gains, but also in a significantly faster full loop per process: or a faster time to token. This is why even their single core mobile chips in the budget line Neo out perform PC processors with several times more threads and RAM[1]

Unified memory has existed for decades in the PC space, Apple didnt invent it.

And the test you linked to has nothing to do with unified memory, its a web browser benchmark (almost entirely constrained by single threaded CPU performance that Apple better than competitors at).

ac29··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
llama.cpp has used mmap by default for years
ac29··on Alphabet's cash burn raises alarm for Big Tech as AI spending climbs
The article notes Google Cloud revenue grew 82% YoY
ac29··on Arduino Launches Plug-and-Play Modules for Long-Range Sensor Projects
I took an embedded C/C++ class online 5-10 years ago and hated it. I dont want to be doing bit shifting math to put some device's registers in the correct state for me to write text to a screen
ac29··on Arduino Launches Plug-and-Play Modules for Long-Range Sensor Projects
Yeah I did a project with ESP32 and micropython recently and getting the proof of concept up was very very quick. The agent I was using also had no problem writing drivers for peripherals that didnt already have micropython drivers (I seem to recall it ported C or Arduino reference code)
ac29··on M 3.9 Experimental Explosion – 147 Km ENE of Ponce Inlet, Florida
> I’d really rather we didn’t kill animals tbh, but when it comes to growing our food we have little choice right now

Several hundred million Indians would disagree

ac29··on Kimi K3: Open Frontier Intelligence
Its just a single benchmark, but Luna 5.6 xhigh scores within the margin of error the same as Opus 4.8 max on DeepSWE for 8x cheaper. Luna max is quite a bit higher than Opus and still 4x cheaper
ac29··on Kimi K3: Open Frontier Intelligence
OpenCode Go is a great deal but I recently dumped my subscription because I found myself rarely reaching for it over my Anthropic sub (I can get 40 hours of work a week out of the $20 sub and almost never hit weekly limits). Subscribed to OpenAI as my secondary and I've been really impressed with that too so far.

I expect if they add Kimi 3 to Go the limits are going to be really low since 2.7 is already one of the most limited models and 3 is much larger.

ac29··on Kimi K3: Open Frontier Intelligence
You do understand that the "frontier" people are usually talking about is the cost-intelligence frontier right?

By definition there is no model that is both cheaper and as intelligent or better than another on the frontier.

ac29··on The real prices of frontier models. Tokens * Price, right?
> You will see people claim Claude uses 2x to 4x the tokens of GPT.

My guess is something to this effect was in the prompt and the LLM made a point of correcting it

ac29··on Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
Try an 8 bit quant of Qwen 35B, but temper your expectations. Those Qwen 3.6 models are impressive for the size, but you need an order of magnitude more parameters to actually be useful for more than trivial work in my opinion.
ac29··on QuadRF can spot drones and see WiFi through my wall
Phased array radars are export controlled in the US. It doesnt mean its illegal to build or own, but it might be illegal to sell in some cases
ac29··on Show HN: Getting GLM 5.2 running on my slow computer
llama.cpp supports a wide variety of 4-bit and smaller quants and mmap's models by default, so you dont need to be able to hold the weights in memory (the OS will handle bringing them in from storage as needed)

Its cool to see this implemented in a tiny amount of code without dependencies, but does it actually bring more performance?

ac29··on Anthropic's Method to Losing Goodwill in a Few Easy Steps
> So... maybe we can still use third party harnesses with Claude Code subscriptions... for now?

The way I read this is: yes, if the third party harness uses Anthropic's Agent SDK. Most of them do not, AFAIK, and are still against ToS (though maybe its not enforced for now)

← PreviousPage 2 of 34Next →