If the drivers are upstreamed and good, it should work well.
7,256 karma · joined May 7, 2013
meet.hn/city/gb-Cardiff
If the drivers are upstreamed and good, it should work well.
Prefill (input tokens) is heavily compute bound. And the ratio of input to output continues to rise, as typically in agentic sessions you have a few tokens output for a tool call and (many) thousands of input from the tool result.
Then you have cached input tokens, which is a totally different issue, system RAM or NVMe bound.
Obviously output tokens is VRAM memory bandwidth bound, but this is less and less of the bottleneck these days for overall agentic speed.
1) most law requires intent, especially criminal. OpenAI certainly didn't "intend" to hack these companies given they did sandbox them etc.
2) Given the agent hacked them, not a human, a lot of law requires a person/employee to have done it to hold the company liable if it was part of their work duties.
I think the only real potential ground is negligence (in not sandboxing them correctly and being reckless with running these tests at all), but this requires not taking reasonable precautions. They could argue that they _did_ but it was so novel the precautions failed. But it's important to say if this happens again in the future it's arguably much harder to try and make this case.
Interestingly this was solved with new laws for self driving cars, most of which assign the company that is operating the car as the "person" involved explicitly.
I'm a native English speaker as well, I shudder to think what a second language speaker would do (even if they were very confident with the language!).
You'll notice buildings in central London that are listed _but_ have housing association ownership nearly all have fibre to each apartment, as they did portfolio wide deals with hyperoptic etc.
And pipes don't help. Openreach (who owns the copper network) are not allowed in 99% of cases to "fix" copper with fibre under the agreements they have with building owners, they can only make like for like repairs.
The worst affected apartment buildings are 90s and pre 2015ish. Everything after that got fibre installed at build time.
Btw it is worth checking if you have an altnet like hyperoptic, community fibre or g network available. The majority do and if you are just checking for openreach or VM broadband it won't show up.
But yes, I think we are in agreement :).
Can you tell I have mental scars from nextjs bundle optimisation?
Also, removing dependencies is not easy. I've seen many corporate that have a huge UI lib for example that everyone should use for brand consistency. But it's many MBs of JS, because it has to cover every possible use case.
This doesn't even get into 3rd party vendors who _also_ ship react et al and have other bundles.
I'm not saying the article is wrong, but if you want to free main thread time especially at the most critical point (when the user has initially loaded the page) you _probably_ will find that most of the opp is in bundle size and hydration improvements.
The issue is though that it's too focused on interactivity. In reality, 90%++ of slow sites are not slow because of interactivity really, they are slow because they ship enormous react/nextjs bundles and have extremely heavy hydration work to do.
_so many_ sites have bundles >10MB that need to be downloaded, parsed and hydrated.
I've even seen (many) sites which have multiple SPAs stacked inside of them.
If you're on a slow internet connection and/or CPU the page is basically unusable for many tens of seconds and no amount of yielding post bundle hydrate will really solve that.
https://martinalderson.com/posts/watch-out-for-cache-read-co...
Btw I still haven't came across any decent model that is <$0.01/MTok cache costs apart from deepseek thru their official API (even with the price increases).
Seems like a bit of an opportunity for someone to take - drop cache read costs significantly.
Also, IME they are dosed completely wrong. So many people seem to be on very low doses, which has no improvement on placebo in the studies I've read.
Whereas, higher doses are _hugely_ better than placebo, especially for anxiety.
What's worse is a lot/most studies on SSRIs in general often don't adjust for dose. Which seems like an enormous oversight to me.
Omarchy is integrating agents much more closely with the UI.
For example, there is a bug in the logitech mouse driver (combined with a specific AMD USB chipset) that occasionally does not restart properly after suspend. Coding agent fixed it (two lines of code) with a kernel module that automatically rebuilds after kernel upgrades.
Another example was I was getting hard crashes, because by default swap on fedora was in memory (IIRC, or it could have been me doing it wrong), so when there was a lot of memory pressure on my system from running larger LLMs it'd OOM and go down. This would have been painful to debug normally, but in <2mins a coding agent can read the kernel logs and fix it.
I also had it set GNOME up the way I like it, which would have taken hours manually, because I don't know what I'm searching for extension wise.
There's many other examples of just slightly broken stuff that coding agents can fix that was a total pain before and I just didn't have the patience to fix them manually.
You should get far better performance on 6GHz, though there are harsh power limits in most countries with various ways to enable it which may complicate things. But if you can get full power on 6GHz it should reach 10m no issues at all, and with 320MHz should be at least twice as fast.
If you can also get MLO working (which is not simple) you'd be able to bond another band with it.
GitHub Actions is 'simpler' to scale, it's just a bunch of runners, and Azure has a lot of capacity for VMs (mostly). Git itself is a compute intensive process and is far more interlinked.
Even if it did, it still doesn't make much economic sense running a model locally vs on a datacentre.
For example, I managed to just about squeeze a Q2 quant of Qwen 3.7 27b on my 9070XT. I get around 60tps decode (slightly faster prefill). _but_ it uses 300W of power to do so. At UK electricity rates of 30c/kWh this works out at something like 42c/MTok. I can get far far better models on openrouter cheaper than that, plus I'm not horrendously constrained on context length.
I'm not sure how true this is, but when using "forced" json output it def had a big drop off in quality - https://arxiv.org/html/2408.02442v3.
I think you're better not fighting it with hacks like this and find a different model.