11,932 karma · joined November 13, 2012
github: infogulch
twitter: infogulch
email: hello+hn@infogulch.com
You can still get your preferred play style with a mod: https://mods.factorio.com/mod/bring-back-space-casino
Source: https://andonlabs.com/evals/drone-bench#:~:text=Cheating%20s...
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
The vps runs a custom image that is 2.54 Megabytes. It has a custom kernel with almost everything but networking and wireguard disabled, a fixed-size fs with pre-allocated blocks and inodes to hold the vps wireguard key, and a single pid 1 binary that calls the kernel directly to set up the routing rules, generate a new wireguard key on first boot and save it to the fs, print out the wireguard public key to the console, and loops reap. Updating involves building and uploading a new image, assigning the vps to use it, reboot, wait for the public key in the console then set it on the nas so they can talk.
If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
The concern with AI, and the purpose of the OG paperclip maximizer thought experiment, is that the outcome could be extremely divergent from the intent of the original person pulling the trigger / sending the prompt.
Still more likely that a person types "kill all people" (or something tangential where this is the logical conclusion).
> Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away.
Last thread about it: https://news.ycombinator.com/item?id=49536384
Latest tweet by @GrapheneOS on the topic (today): https://x.com/GrapheneOS/status/2097219485937660272?s=20
> ... There's a decent chance usable MTE will ship in Android 17 QPR2 and we'll be able to support it. We can't promise that since it's not up to us and no information has been provided about why MTE was unavailable at launch and still disabled in 17 QPR2 Beta 4.
It's very fast: most queries finish in < 10ms. Agents are able to find things quickly and efficiently even with vague questions. The app does a few things to minimize token output, but there's a long tail of potential token reduction strategies that I haven't gotten to yet.
> We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.
> All systems have now been restored and are functioning nominally.
---
It's interesting to see how reliant the other AI companies are on SpaceXAI for compute.
Apparently a major bottleneck for building out datacenters is turbine blades for power plants, so SpaceX is building a foundry to alleviate supply constraints?! (Also useful for rocket engines.) Holy vertical integration, batman.
The basic shape is to periodically "distill the conversation into several areas (problem, solution, learnings, 'context' or original problem, plus other fields) and vectorized" (aka vector embedding), then queried against pgvector table to find related "memories". The vectorized distillates are also inserted into the pgvector table with a reference to back to the source conversation to add new memories.
Vector search requires a full scan but it's still pretty fast and I bet it's more accurate the FTS.
I don't see a downside yet.
Eventually the obvious spam callers would all quit and the volume would reduce to a handful of borderline cases. So the spam department wouldn't make a lot of money anymore, but the volume would be so low it would barely register.
For example say you get a spam call. You notify your service provider and pay $10 to have them review the call recording to verify that its spam. An employee listens to the call and determines that it is, in fact, spam. You get a $100 reward for reporting spam and your service provider forwards the charge (+10%=$110) to the peer network where the spam call originated. They will in turn forward it to the next peer, or if the originator is a customer, charge and/or close the offending account.
If this system is mandatory then the problem will quickly sort itself out.
I fixed this and a couple related issues in a PR but it hasn't gotten any attention yet. I guess they are a bit swamped. https://github.com/zed-industries/zed/pull/59937