4,143 karma · joined July 10, 2012
See this as a list of merely the most popular varieties you can find at stores: https://waapple.org/varieties/all/
The "module name is network path" is a convenient convention but not at all some "limitation" of the tooling.
1. Agents usually depend on output of the commands they're running in order to make decisions about what to do next, so how do they behave while they're "waiting around" for the output they need?
2. Agents can _already_ run software async, via multiple mechanisms: raw CLI tools like "nohup", literally running tools in parallel (I see Sol do this often in the Opencode TUI harness), and using parallel sub-agents to e.g. research in parallel.
Thus I wonder, how much does this really improve speed vs only improving the "appearance" of getting more done faster?
I am already seeing my "normie" friends getting locked out of accounts due to not understanding passkeys. If they don't have their phone, or it's dead, or it breaks, or is stolen, they just can't access their account anymore. They have no idea how they work or what they're trading off, nor do they understand that they should have prepared for this scenario ahead of time somehow. Upon telling them "yeah you have to use your phone now that you have a passkey" they all universally say "wtf, that's stupid, I never want to have that happen again, I will never use a passkey again."
Passkeys should never have been built for general audiences, they are a huge mistake, I hope they cease to be relevant and die due to everyday folks realizing they're inconvenient and the "more secure" gains ain't worth it for the usability nightmares.
It's in the parent article under a section named "Doom" in case that asset URL ever changes.
That's not quite true. The only thing they don't show per-provider is benchmark data, cause I don't think they are doing continuous benchmarking of each model from each provider, as I assume they feel that's too expensive. You can see hugely detailed breakdowns for near-time metrics per provider for any model by visiting the page for that model on Openrouter. For example see the page for Qwen 3.8 27B: https://openrouter.ai/qwen/qwen3.8-27b
Some of the killer stats they show per provider:
- Pricing: Effective price accounting for cache hit rate, by provider
- Performance: Throughput in tok/s, latency, E2E latency, tool call error rate, structured output error rate, and more; all per provider.
- Uptime: You have to click on the provider to see their specific uptime, but doing so does show the last-7-days uptime, and you can click to see more.
We can alreay write notes down for every conversation and then send them to the person involved saying "we talked about X, Y, and Z." That last step is the key because it lets them object in writing if you mischaracterize things. From a "catching someone in a lie" the most important step is that one, because you form a paper trail where the other party can correct or contest what was written and bring that up now, and the fact that they didn't is itself evidence in case of a dispute later. The apple watch feature doesn't do that, it just dragnets everything. Even if it recorded the audio, we are in a faked-audio world so unless you have some signal that they agreed that they said a thing ahead of time, they can always deny it later.
I also have been using the same LSP config for approximately 7 years.
I don't get the pain.
Being unable to turn on the windshield wipers (not engage for a swipe, turn on) without fully taking eyes off the road to interact with the giant touch screen, that's going to get people killed. The #1 thing that convinced me Teslas are absolutely unfit to operate safely for many common actions was driving one at all.
> Advertising as inference not influence.
People buying ads want to influence customers, but the platforms can instead focus on finding the folks who are already going to buy the product and show them the ad in order to gain the sweet sweet attribution saying "I am responsible for the sale".
You can test this by taking your ad-spend and flatly multiply it: if you spend 5 times as much money to advertise to 5 times as many people, but your sales don't also go up by 5x, then the money you were spending at the 1x rate probably wasn't influencing the buyers either, instead the ad companies were merely targeting all the customers you were already going to get sales from.
This is on things like "we have to migrate how we do TLS termination via coordinated infra changes and DNS changes for a couple thousand customers, many of which are fortune 500 companies". It's maddening. If there's stakes, and it's not a small enough problem that there's not already a StackOverflow post/library for how to do XYZ, then you have to be so vigilant.
> Thus: over the next two years, major pieces of software are likely to run out of remotely-exploitable bugs.
> While I think this is great, for law enforcement and offensive intelligence agencies, it’s going to be a nightmare.
> So what do we do about it? I honestly have no idea. [...] it’s just occurring to me that we’re on a long greasy slide to a place that will look different than where we are today. [We’re] just going to have to hope that this time we make the right choices.
llama-server.exe ^
-m "Qwen3.8-27B-UD-Q2_K_XL.gguf" ^
--presence-penalty 0.0 ^
--repeat-penalty 1.0 ^
--fit-ctx 128000 ^
-ctk q4 0 ^
-ctv q4 0 ^
--reasoning-budget -1 ^
--chat-template-kwargs "{\"preserve thinking\": true}" ^
--host 0.0.0.0 ^
--port 8033
[0] https://huggingface.co/unsloth/Qwen3.8-27B-GGUF (UD-Q2_K_XL)[1] https://github.com/ggml-org/llama.cpp/releases
Instructions if you want to do the same:
1. download two files llama-b10434-bin-win-cuda-13.3-x64.zip and cudart-llama-bin-win-cuda-13.3-x64.zip from that llama.cpp Github releases page, and extract both into the same folder.
2. Download the Qwen3.8-27B-UD-Q2_K_XL.gguf file from huggingface and put it into the same folder beside the `llama-server.exe`.
3. Create a file named "RUN_QWEN_3.8.bat" next to `llama-server.exe` and put the text above into that bat file. Double-click the bat file, then open http://localhost:8033 in your browser to see a chat window.
You can use it with any agents by pointing them at http://localhost:8033/v1 which is a working OpenAI compatible endpoint (it doesn't use a token, if you give one it's ignored).
Congratulations, you're now running Qwen 3.8 27B.
Note: I built the computer in question for playing games, yes it needed to be Windows 11 for anticheat reasons to play games with family, I didn't want to dual boot so here I am. I figure I should share instructions for folks who may also have a Windows PC around for such purposes. Specs for this are AMD 9800X3D, 32GB of system RAM, RTX 5070Ti 16GB
The thing about most sites is they're public and you don't need to sign a real contract to use them. Can't have it both ways.
https://indigene.app/plants/symphyotrichum-subspicatum
https://indigene.app/plants/juncus-patens
Very helpful for choosing plants.
It also, unfortunately, means it's not possible (via most passkey implementations) to back those passkeys up to paper. Which is quite unfortunate: backing up to paper is one of the most stable and human accessible ways of ensuring redundancy and continuity, an inevitable but also oft-ignored part of credential management.
Security folks would like to pretend "solving continuity" isn't a problem, or is a problem that doesn't need to be accessible.
It's pretty amazing to write your own agent BTW. I've got a zero-dependency all-in-one-file agent harness I wrote myself. I use it all the time now because I can get it from anywhere and I can know EXACTLY what it'll do (as much as you can with any model), what it's been told vs not. Using it as a harness for models I'm hosting myself makes me feel like some kind of LLM homesteader: it's a set of tools I'll always have that will only change as much as I want it to change.
My inspiration was the mayor idea from Gastown, plus wanting to formalize the informal workflow I used with agents and Jira at $dayjob.
As for the Orchestrator, it's pretty simple. In essence, it's like "Jira/Trello/Kanban on autopilot". Work items have states, a state machine defines how those work items transition between states, states are todo, in progress, retrying, reviewing, code reviewing, done. work items also have connections, allowing the LLMs to specify a dependency graph, and the dependency graph informs the dispatch order/parallelism, as well as when branches have to be merged. I talk to the steward, the steward has tool calls for interacting with all the data, and the orchestrator auto-dispatches all the work that comes in. I can generate work as fast as I can describe it to the steward, and that's usually the bottleneck.
So far I haven't had to deal with "how do you get the LLM to re-organize the work mid flight due to a worker finding something not accounted for by the planning", but I assume it'll come soon. The most complicated digraph I've tossed at it was 9 items and 4 layers deep. The kind of work I've given it hasn't been scoped large enough yet, so we'll see how it tackles that.
I'm still using Claude at work (they're the only approved provider), but wow are the smaller models starting to SMOKE the big ones. At this point, all I'd consider paying out of my own pocket for is the lowest-limit Anthropic/GPT plan to get a big model as the Steward, but I wouldn't pay for ANY of the Anthropic models as the workers who do all the work. And as time passes, I don't know if I'd even do that; the open models are serving SO well.