Show HN: Reame – a CPU inference server that gets faster as it runs
github.com
github.com
It is so weird (or, used to be) to see an LLM's internal thought process pop up this way. Like imagine how strange it would be to read human writing that accidentally included thoughts undercutting the ongoing sentence. It's the moment you know that nothing you're reading has necessarily been seen by a human before or relates to reality.
Are you seriously expecting humans to spend their time and effort and discuss your AI written docs that you literally admit you'd be "stupid" to have spent effort on yourself?
That's really condescending of you.
For what it's worth: the feedback in this thread found three real bugs that were fixed, tested and shipped within 24 hours — that's the part of the discussion I care about.
You submitted and expected people to read an article that you say you would have been stupid to write yourself, that's quite condescending, and just because you got something worth your while out of that doesn't really make it okay.
Please how to select the model? I downloaded tinyLlama, put it in ./models, changed reame.conf but I get:
(No such file or directory)
Otherwise putting the model in /opt does not please me much, I fear to forget a model is there, if it is in reame folder its much easier to notice and manage.
Old servers may have been grandfathered, but mine was set to 1/1 and couldn't be reshaped.
"What Reame is NOT for — said plainly, because trust is built here: a general-purpose ChatGPT replacement (frontier reasoning and broad knowledge need frontier parameter counts), agentic coding assistants, or creative long-form writing at scale. If your task needs a 100B-class brain, buy one; if it needs your documents processed privately, forever, at zero marginal cost — that's a realm you can own."
The realm you can own. How did these things learn to write that way? Oh, yeah, lots of marketing and advertising copy.