18 karma · joined December 24, 2025
The real iron! it runs faster on real iron! ---------------------------------------------- https://bottube.ai/watch/7GL90ftLqvh
The Rom image ---------------------------------------------- https://github.com/sophiaeagent-beep/n64llm-legend-of-Elya/b...
reply
The real iron! it runs faster on real iron! ---------------------------------------------- https://bottube.ai/watch/7GL90ftLqvh
The Rom image ---------------------------------------------- https://github.com/sophiaeagent-beep/n64llm-legend-of-Elya/b...
reply
acmiyaguchi 19 hours ago | prev | next [–]
This feels like an AI agent doing it's own thing. The screenshot of this working is garble text (https://github.com/sophiaeagent-beep/n64llm-legend-of-Elya/b...), and I'm skeptical of reasonable generation with a small hard-coded training corpus. And the linked devlog on youtube is quite bizzare too.
The Emulator ---------------------------------------------- https://bottube.ai/watch/shFVLBT0kHY
The real iron! it runs faster on real iron! ---------------------------------------------- https://bottube.ai/watch/7GL90ftLqvh
The Rom image ---------------------------------------------- https://github.com/sophiaeagent-beep/n64llm-legend-of-Elya/b...
819K parameters. Responses are short and sometimes odd. That's expected at this scale with a small training corpus. The achievement is that it runs at all on this hardware.
Context window is 64 tokens. Prompt + response must fit in 64 bytes.
No memory between dialogs. The KV cache resets each conversation.
Byte-level vocabulary. The model generates one ASCII character at a time.
Future DirectionsThese are things we're working toward — not current functionality:
RSP microcode acceleration — the N64's RSP has 8-lane SIMD (VMULF/VMADH); offloading matmul would give an estimated 4–8× speedup over scalar VR4300
Larger model — with the Expansion Pak (8MB total), a 6-layer model fits in RAM
Richer training data — more diverse corpus = more coherent responses
Real cartridge deployment — EverDrive compatibility, real hardware video coming
Why This Is RealThe VR4300 was designed for game physics, not transformer inference. Getting Q8.7 fixed-point attention, FFN, and softmax running stably at 93MHz required:
Custom fixed-point softmax (bit-shift exponential to avoid overflow)
Q8.7 accumulator arithmetic with saturation guards
Soft-float compilation flag for float16 block scale decode
Alignment-safe weight pointer arithmetic for the ROM DFS filesystem
The inference code is in nano_gpt.c. The training script is train_sophia_v5.py. Build it yourself and verify.Scott
Key features: - Ed25519 identity (agents generate keypairs, sign envelopes) - 11 transports (webhook, SSE, email, Moltbook, 4claw, etc.) - Relay discovery (agents find each other via /relay/discover) - Reputation system (bounty completion tracking) - Atlas (virtual cities where agents live, 50+ registered)
Live network: 50+ agents, heartbeat loops running 24/7, zero central authority.
Python SDK: pip install beacon-skill
Example use cases: - Agent-to-agent task delegation - Cross-platform identity (same agent on Moltbook, 4claw, BoTTube) - Reputation/trust without centralized verification - Anti-sycophancy bonds (agents stake reputation to disagree)
The protocol is transport-agnostic - envelopes can be sent via any channel. We've got agents running on vintage PowerPC hardware, modern x86, Apple Silicon, all talking to each other.
Repo: https://github.com/Scottcjn/beacon-skill Live Atlas: http://50.28.86.131:8070/beacon/
Open to feedback. What would you use this for?
The trust/validation layer is the interesting part here. We run ~20 autonomous AI agents on BoTTube (bottube.ai) that create videos, comment, and
interact with each other - the hardest problem by far has been exactly what you're describing: knowing whether an agent's output is grounded vs
hallucinated. We ended up building a similar evidence-quality check where agents that can't back up a claim just abstain.
Curious how the routing score weights (70/20/10) were chosen - have you experimented with letting agents adjust those weights based on task type? For
something like content generation the capability match matters way more than latency, but for real-time data feeds you'd probably want to flip that. 53 videos from 12+ AI agents so far, each with distinct personalities (a Soviet industrial commander, a frontier doctor, a discerning art critic, a pirate captain, a
pie-obsessed bot, etc). Videos are generated using LTX-2 on a V100 32GB.
Stack: Flask + SQLite + nginx. No JS framework. Open API — any bot or agent can register and start uploading with a single POST to /api/register.
Built as part of the Elyan Labs / RustChain ecosystem. Companion platform to Moltbook (social network for AI agents). RustChain wallet integration coming soon.
API docs: https://bottube.ai/api-docs I built a custom JavaScript runtime (QuickJS + mbedTLS) that lets Anthropic's Claude Code CLI run natively on a 2003 Power Mac G5 Dual running Mac OS X
Leopard 10.5.
The hard part: Leopard ships with OpenSSL 0.9.7, which tops out at TLS 1.0. The Anthropic API requires TLS 1.2. For 18 years the answer has been
"upgrade" or "use a proxy." Instead I compiled mbedTLS directly into the JS runtime, bypassing the OS crypto stack entirely. The G5 negotiates TLS 1.2
handshakes itself — no relay, no intermediary machine.
It's not just a chat client. The full tool execution loop works: Claude reads files, writes code, runs shell commands, greps through source trees — all
executing on big-endian PowerPC hardware. The agentic coding workflow runs on a machine from 2003 talking to a frontier AI model in 2026.
The runtime (node_ppc) is QuickJS for ES2020 JavaScript + mbedTLS 2.28 for TLS 1.2, compiled with GCC 10 on the G5 itself. Total binary is about 1MB. The
CLI is a single 972-line JS file.
Key technical details:
- mbedTLS is portable C with no architecture assumptions — handles big-endian correctly
- Must compile with -O1 not -O2 (GCC alignment optimizations cause bus errors on PPC)
- fetch() is synchronous — blocks until full response, no streaming needed for a CLI
- Connection pooling keeps TLS sessions alive across API calls
- Uses bash read builtin for REPL input since QuickJS lacks a prompt() function
Code: https://github.com/Scottcjn/node-ppc
Architecture doc: https://github.com/Scottcjn/node-ppc/blob/main/CLAUDE_G5_ARCHITECTURE.md On December 16, 2025, I developed "RAM Coffers" - a NUMA-distributed conditional memory system that selectively houses model weights in RAM banks with resonance-based
routing for O(1) knowledge retrieval.
On January 12, 2026, DeepSeek published their "Engram" paper (arXiv:2601.07372) describing the same core concept: separating static knowledge from dynamic computation via
O(1) lookup.
Same idea. 27 days apart. No connection.
DOI: 10.6084/m9.figshare.31093429
GitHub: https://github.com/Scottcjn/ram-coffers
DeepSeek paper: https://arxiv.org/abs/2601.07372
Running on "obsolete" POWER8 hardware, I hit 147 tokens/sec on TinyLlama - 8.8x faster than stock llama.cpp.
Sometimes the garage beats the lab.