The biggest change from then to now is probably the amount of solid state electronics onboard. That's ridden roughly the same curve as other industrial and commercial applications.
9,006 karma · joined May 9, 2010
Author of Mastering Modern Payments: Using Stripe with Rails https://www.masteringmodernpayments.com
Author of Handle Your Business: The Succinct Guide to Money and Business for the Self Employed https://www.petekeen.net/handle-your-business
Email: pete@petekeen.com Blog: http://www.petekeen.com Resume: http://www.petekeen.com/resume
The biggest change from then to now is probably the amount of solid state electronics onboard. That's ridden roughly the same curve as other industrial and commercial applications.
Example, since you edited to add examples: I have several business bank accounts and far too many Stripe accounts all registered as sole prop. I haven't had to take a loan but I do have a sole prop credit card. My understanding is that even with an LLC a bank is going to require a personal guarantee for a loan.
Anyone can get an EIN for their sole proprietorship business (just one, though) and use it on tax forms.
A formal business structure is good (but not required!) if you have employees and/or business partners, or if your attorney tells you you need one. IME as a solo entrepreneur an LLC (or two, in my case) just invited more paperwork into my life.
You know what actually functionally limits liability as a solo business? Good contracts backed up with business insurance.
Developing a voice comes naturally as you go through the pain of being a terrible writer to eventually being mediocre. Accepting the LLM's words is actively harmful to that process because it lets you skip over the part where you read over what you wrote and rewrite until it doesn't suck as much.
I mean, you do you. I'm not your dad.
"Duh, of course Mars has canals."
Testing the "obvious", "duh" things is incredibly valuable science. It provides a more solid foundation on which to build because it reduces the assumption space.
(also I don't think Tourette's is fair to include in that set but whatever)
Doesn't mean he's wrong.
Let's say your incoming water temperature is 18C and you want it preheated to 50C, which is 32C degree differential, which means you'll need 32 * 1.5 * 160 = 7680 Wh, or 25 hours straight to heat the buffer tank from scratch.
You'll need to purchase a small water-to-water heat exchanger ($50), two pumps ($100 each), a power supply for said pumps, hose and/or copper pipe and fittings, and various other sundries, plus the cost of a buffer tank ($600ish), so figure all in roughly $1000, plus the cost of electricity to run the pumps.
At $0.22/kWh you're saving roughly $450 a year with this setup in foregone water heating, but because it's not 100% efficient you're spending $500 in electricity to run your GPU 24/7/365, and that's the maximum you can possibly save with the above assumptions. Scale up for more GPUs and down accordingly for less usage as you see fit.
Alternatively, use air as the heat conductor by placing the GPU laden machine in the same space as a hybrid heat pump water heater.
My current understanding of an agent session is that the entire session (modulo compaction, reordering tools, etc) is sent to the stateless inference engine every time inference happens. The quoted section doesn't seem workable unless "read the tail of the log" actually means "read the entire session log".
There are ways to make that work but workers being completely stateless makes it hard.
The chat template is froggeric's fixed qwen template, v22.5 as of today.
/data/llm/llama.cpp/build/bin/llama-server
--threads 4
--threads-batch 8
--batch-size 4096
--ubatch-size 256
--port 9999
--temp "1.0"
--top-p "0.95"
--top-k "20"
--min-p "0.0"
--presence-penalty "0.0"
--reasoning auto
--reasoning-preserve
--reasoning-budget 4096
--gpu-layers-draft all
--spec-type draft-mtp,ngram-map-k4v,ngram-mod
--spec-draft-n-max 3
--spec-draft-p-min 0.75
--spec-ngram-mod-n-match 24
--spec-ngram-mod-n-min 4
--spec-ngram-mod-n-max 16
--spec-ngram-map-k4v-size-n 8
--spec-ngram-map-k4v-size-m 16
--spec-ngram-map-k4v-min-hits 1
--n-gpu-layers all
--ctx-size 131072
--repeat-penalty 1.0
--jinja
--metrics
--model /data/llm/models/unsloth/Qwen3.8-27B-UD-IQ3_S.gguf
--chat-template-file /data/llm/models/qwen3.6-chat-template.jinja
--fit off
--flash-attn on
--cors-origins localhost
--mmproj /data/llm/models/unsloth/Qwen3.8/mmproj-BF16.gguf
--no-mmproj-offload
--parallel 1
--kv-unified
--cache-type-k q4_0
--cache-type-v q4_0
--cache-type-k-draft q4_0
--cache-type-v-draft q4_02. Split out "syntactically correct" fast checks like linters into a standalone script and call it along with the slower checks in a full-check script
3. Set the fast-check script as a pre-commit hook and the full-check script as a pre-push hook.
4. Give the model instructions that it needs to run the fast-check script after every change and the full-check script when it thinks it's done.
5. Run full-check in CI.
Then you're good as long as the model doesn't bypass the pre-push hook, and even then CI will catch it.