HNHacker News
TopNewBestAskShowJobs

airspresso

246 karma · joined November 8, 2023

submissionscomments
airspresso··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
> it's extremely long in it's thinking, it just goes on and on

Known issue with this model, I recommend setting thinking to 'medium' instead of default 'xhigh'.

airspresso··on Livenerf: Has Opus 5.5 been nerfed yet?
> I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars).

That is a big exaggeration. You can have a perfectly usable local LLM setup that will power your agent for single digit thousands of dollars. Can even power multiple agents simultaneously, depending on the hardware and setup. Won't be fast and won't be frontier intelligence, but definitely useful.

airspresso··on Early rogue AI agent activity and attempts to hack found on urlquery.net
> What if all the agents involved in these incidents have in fact had the full stack of alignment applied

A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.

airspresso··on Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
We are certainly not in the diminishing returns phase for LLM progress. No sign of that yet.
airspresso··on Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
No, those use LPDDR5(x), not HBM.
airspresso··on Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
From what I'm seeing elsewhere, context size up to 128k should be possible on this hardware. It really matters for agentic workloads to push that context size headroom up. Anthropic are spoiling us with models that do 500k context and beyond.
airspresso··on Making the Internet Boring
I'm with you on this, the knowledge sharing on YT is massive and in general not replicated in written form anywhere else. Anything that needs fixing involves a YT search for me these days, since that has the highest chance of turning up something useful and relevant. Only downside is that it's time consuming to find the right spot in the video and pause/rewind to understand how to do the thing properly. But having someone actually demonstrate how to do it, rather than just describing it with words, is so helpful and makes replicating it a lot easier.
airspresso··on Making the Internet Boring
Clicked the links just to see preview screenshots, that is hilarious with all the cows XD
airspresso··on Matrox: Graphics for Professionals
It gave the card an aura of mystique (pun intended) and I remember really wanting it. I guess that means they succeeded in making a statement with it.
airspresso··on Nvidia agrees to acquire Hugging Face for $13B
Nvidia is also pushing for local inference IMHO. They want open models and competition in the model layer, not two big labs controlling all of it.
airspresso··on OpenAI Jalapeño: Better than Nvidia Blackwell
This depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped level of complexity, entirely different aspects matter and LLMs become more interchangeable.
airspresso··on OpenAI Jalapeño: Better than Nvidia Blackwell
By leveraging the experience Broadcom has in this area. Still remains to be seen how that goes when they want to scale production.
airspresso··on IPFS Maintainers Winding Down
Most of the generative art is hosted on IPFS.
airspresso··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
Wait, it's allowed again? Completely missed that. Been avoiding to use it and trying to find workarounds, not great.
airspresso··on Cerebras CS-4
The Cerebras hardware is not locked to specific models / model families. Taalas is the company that's etching models into their silicon, locking it to that model forever.
airspresso··on Process as a Proxy for Motivation
Agree. My team standardized on black in our Python CI/CD pipeline and precommit hooks and it was surprising how much time we reclaimed from not having to review coding style and discuss patterns. The compromise was to apply black and not think about it. And it worked out really well. This was back in the days when humans still wrote code XD
airspresso··on Getting 25 Gbps Thunderbolt Ethernet on My Mac Studio
My default assumption (with no knowledge about this particular case) is that all gear in YouTube videos is sponsored. Too easy to get jealous of all the fancy setups. But I don't want to become a YouTuber so that settles that.
airspresso··on OpenAI unveils its first custom chip, built by Broadcom
OpenAI's upcoming mega IPO
airspresso··on Magnifica Humanitas
> Time for me to go research the early history of electrification.

The Stepchange podcast has an amazing episode on The Grid [1], walking us through the arc of history of how it became the utility it is today.

[1]: https://www.stepchange.show/grid

airspresso··on Higher usage limits for Claude and a compute deal with SpaceX
> posting this sentence was part of the deal to get the compute

This 100%

airspresso··on Higher usage limits for Claude and a compute deal with SpaceX
> but given that you can actually run the models yourself on AWS Bedrock

That's not exactly how it works. Anthropic are hosting their models in AWS Bedrock as a managed service. Customers call those LLMs just like calling any other API. There's no visibility into what kind of AWS infrastructure is serving that API request.

airspresso··on Only one side will be the true successor to MS-DOS – Windows 2.x
Fun that it has oscillated from instant boot then to minutes-long boot a decade later back to instant boot (or resume to be fair) today.
airspresso··on Only one side will be the true successor to MS-DOS – Windows 2.x
I had the same confusion and closed the tab, only to discover here on HN that there's more. Open by default sounds reasonable.
airspresso··on Google releases Gemma 4 open models
There is a surprising amount of code needed in each of the inference frameworks (LM Studio, llama.cpp, etc) to support each new model release. For example to format the input in the right way using a chat template, to parse the output properly with the model-specific tokens the model provider decided to standardize on for their model, and more.

This particular instance was a fix to the output parsing [1] in LM Studio, described like this:

"Adds value type parsers that use <|\"|> as string delimiters instead of JSON's double quotes, and disables json-to-schema conversion for these types."

[1]: https://github.com/ggml-org/llama.cpp/pull/21326/commits/a50...

edit: formatting

airspresso··on Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
I think we lost that terminology war. Open source models mean open weight. There are only a couple examples of fully open source models with open data and code, and the labs are not incentivized to go that far.
airspresso··on How BYD got EV chargers to work almost as fast as gas pumps
That's a good point. Charging stations benefit from being a service station too though, with amenities and a cafe etc, since people want something to do while they charge. So a gas station is a better candidate than a parking lot when decisions are made for where to place the new charging infrastructure. Lots of other factors too of course.
airspresso··on How BYD got EV chargers to work almost as fast as gas pumps
Yes, retooling gas stations is the way to go. Already happening in Norway where stations now show the price of kWh in addition to gas and diesel prominently on signs by the road. Charging is just a different kind of pump.
airspresso··on How BYD got EV chargers to work almost as fast as gas pumps
America certainly did not invent electric cars. Depending on which electric car you consider the first real one, the inventor was either French, British or German [1].

[1]: https://en.wikipedia.org/wiki/History_of_the_electric_vehicl...

airspresso··on Unsloth Studio
Unsloth is providing the best and most reliable libraries for finetuning LLMs. We've used it for production use-cases where I work, definitely solid.
airspresso··on How do you capture WHY engineering decisions were made, not just what?
Also wrestling with this challenge at the moment and curious to hear experiences from others. Even though it requires human input, the capture and the way it's updated has to get automated.
Page 1 of 4Next →