It’s been an absolute gamechanger.
the docs are good. when creating the initial CA make absolutely sure you set the CA expiration to 10-30 years, the default is 1 which means your whole setup explodes in a year without warning.
It’s end to end encrypted, and with tail lock enabled, nodes can not be added without user’s permission.
I didn’t start with tailscale because the only way you could log into it was with Google or GitHub or something. I don’t trust Microsoft or Google with auth for my internal network. I thought about running Headscale but Nebula was faster/easier for me.
There's also potential for malicious updates to compromise a network (as there is with most software unless you're auditing the source for each update).
E2EE is only as meaningful as where the keys reside, and how easily those keys are abused.
The metadata is generally public information, I don’t care about that.
The malicious updates and key abuse are more concerning. It’s true for all software, and probably better done with OS, like on iOS.
The VPN could steal the keys, but that’s a lawsuit!
But with a malicious update, they could ship them to their infra, targeting some users. The product then becomes malware!
totally different use case.
brew install ollama; ollama serve; ollama pull llama3: 8b-v2.9-q5_K_M; ollama run llama3: 8b-v2.9-q5_K_M
https://ollama.com/library/dolphin-llama3:8b-v2.9-q5_K_M
(It may need to be Q4 or Q3 instead of Q5 depending on how the RAM shakes out. But the Q5_K_M quantization (k-quantization is the term) is generally the best balance of size vs performance vs intelligence if you can run it, followed by Q4_K_M. Running Q6, Q8, or fp16 is of course even better but you’re nowhere near fitting that on 8gb.)
https://old.reddit.com/r/LocalLLaMA/comments/1ba55rj/overvie...
Dolphin-llama3 is generally more compliant and I’d recommend that over just the base model. It's been fine-tuned to filter out the dumb "sorry I can't do that" battle, and it turns out this also increases the quality of the results (by limiting the space you're generating, you also limit the quality of the results).
https://erichartford.com/uncensored-models
https://arxiv.org/abs/2308.13449
Most of the time you will want to look for an "instruct" model, if it doesn't have the instruct suffix it'll normally be a "fill in the blank" model that finishes what it thinks is the pattern in the input, rather than generate a textual answer to a question. But ollama typically pulls the instruct models into their repos.
(sometimes you will see this even with instruct models, especially if they're misconfigured. When llama3 non-dolphin first came out I played with it and I'd get answers that looked like stackoverflow format or quora format responses with ""scores"" etc, either as the full output or mixed in. Presumably a misconfigured model, or they pulled in a non-instruct model, or something.)
Dolphin-mixtral:8x7b-v2.7 is where things get really interesting imo. I have 64gb and 32gb machines and so far the Q6 and q4-k_m are the best options for those machines. dolphin-llama3 is reasonable but dolphin-mixtral is a richer better response.
I’m told there’s better stuff available now, but not sure what a good choice would be for for 64gb and 32gb if not mixtral.
Also, just keep an eye on r/LocalLLaMA in general, that's where all the enthusiasts hang out.
Just download a single file and run it.
I think port forwarding configuration is a pain that does not offer value over just poking a hole in your firewall to do an authenticated connection over ssh.