1,430 karma · joined December 16, 2009
Contact: gergely@imreh.net
Based at http://gergely.imreh.net
Taipei Hackerspace co-founder https://taipeihack.org
[ my public key: https://keybase.io/imrehg; my proof: https://keybase.io/imrehg/sigs/Vqc30FUDkAzeR6b4jTTQown3Qjned_Xju7DqvvT6DfU ]
Codeberg Terms of Use[^0] makes it clear that it's only for "projects covered by a licence for free and open source software, free and open source hardware, or free cultural works", and "You must not share projects that mostly consist of code written by "generative AI"-tools". They also only provide private repos where the FOSS system need it for their infra or what not, and not just for any purpose.
These are totally fine for Codeberg to do, and it's wonderful for FOSS!!
However it is not a viable place for many projects that are on GitHub/GitLab now and would want to migrate, say: any commercial project, that would need "permission-free" private repos, or when the teams want to decided themselves how much generative AI their developers should use, not their Git-hosting provider...
So yeah, different ballpark.
[0]: https://codeberg.org/Codeberg/org/src/branch/main/TermsOfUse...
And it's just funny for what details the brain is unable to suspend disbelief...
- earthquakes: these can arrive indeed before you even feel them (a few seconds before); they filter it to areas where it's expected to be felt, there are discussions around how sensitive to make it (how wide area it should cover)
- extremely heavy rain: this was new recently, and very useful, when locally it all went cats & dogs, more than maybe I've ever seen. See an example here: https://fosstodon.org/@imrehg/117058657161401667
- military drills: the yearly "if there's an invasion drill", they had sent reminder a week before, and then at the start/end; example here: https://fosstodon.org/@imrehg/117030180601973598
- Chinese rocket launches that will fly over the island (not sure what could one do about it, and never seen anything, so maybe this is more borderline)
Interestingly (for me), the headers seems to change over time. Started with "Presidential Alert", there were "Government Alert" then, now "Emergency Alert" and just "Alert". I'm not sure if there area list of these or it evolved over time...
I do wonder what would have happened if some months later they do actually find a contract, maybe slid between a filing cabinet and a wall... Absence of evidence and evidence of absence, aye?
Also enjoy Stallman's title, President and Chief GNUisance, which he definitely tried very hard to live up to, with that reply. (And I won't fault him for that. :)
The second I didn't say it's any of llama.cpp's "fault", but it is _related_ to llama.cpp since it's being shipped in another system, aye?
Can't stick to the old hash either, because older version have different bugs. E.g. on older versions the same Qwen3.6 model reliably fails to call specific tools due to template issues, while just having the newer llama.cpp version has that fixed. So different versions - different bugs, rather than no bugs.
Why the beating you are trying to gimme, mate? :)
As much as I can tell, the ROCm version of llama.cpp would be a bit faster on prompt processing, but about the same on the token generation as Vulkan. Real life benchmarks don't seem to give any "ROCm or nothing" sort of vibes. And the difference between the performance of different models are way bigger than the difference between the llama.cpp versions (and versus different runtimes like the llama.cpp/GGUF and the MLX runtimes on Mac for the same models)...
I've tinkered enough with the serving, that I'd rather do something with them with, say 10% slower speed, than spending hours on seting things up again... YMMV
Two examples:
- https://github.com/ggml-org/llama.cpp/pull/25863 Someone's few lines change broke the native (ROCm) support for the AMD GPU inside Framework (and other integrated systems), and any rollback or proper fix is pending for almost a month. Fortunately there's workaround (switching to Vulkan rather than ROCm devices), but both the way the bug was introduced and the way it is not fixed just doesn't give much confidencen
- LM Studio is using llama.cpp internally for GGUF, they ship their own build with their closed source system as "runtimes". Their ROCm runtime does not enable the the AMD GPU inside the Framework, even thought the llama.cpp version would support it. So their runtime keeps telling me that there's no supported AMD GPU -- again, the solution is to use the GPU with the Vulkan devices. Not fixed since Jan at least https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1...
I guess overall it's the worst runtime I've seen so far, except for all the other runtimes out there... I'm a fan, though in some cases I don't have enough knowledge, or I don't have access to fix things, and that feels like a bummer...
The current LLM-driven stuff seems to break down the expectations, and now there are a lot of places which just ignore anything I send in (the same tinkers as before), generally, for a while. Though there are some tools that picked up the pace and actually react faster (so YMMV here too).
But there are few things more frustrating as being half-way. Case to point is LM Studio. It's closed source, has bugs (duh!), and there's at least a GitHub issue tracker to report the bugs -- but then by and large nothing happens to those reported things. It's almost worse than not having an issue tracker (then I could justify never to really touch LM Studio again, this way I keep hoping against hope that reports will turn into fixes and thus I keep using and keep reporting...)
Also, haven't realised that it was off for almost a month (the email is from 13th July, so I guess few others have realised it either way, if it just makes its way to HN).
If it was a truly world representation, this might be different. But if things like health are sacrificed (ie. no WHO access either), I don't think they really deserve the benefit of the doubt.
[0]: https://en.wikipedia.org/wiki/Taiwan_and_the_United_Nations
As others mention it - it seems to shows the Watts used as well :) (and network, and GPU, and disks,....)
This is a blast from the past. Centaur was doing VIA Technology's CPUs, and I was at VIA (in Taiwan) while this author was in Centaur (in the US). I was on the embedded side, but I remember some distinct collaborations with the US team, so there's a non-zero chance to have crossed path.
I've done a couple of checks, and it seems very marginally better on some local benchmarks I'm running, but it's not super scientific evaluation.
I still use the MTP version as it _feels_ slightly better quality, and because the unsloth quantizations I can get have more variety to fit into the various systems at hand... but that's not for the MTP aspect, unfortunately.
In the article they did have ~2x performance on the 27B (which might be something to retry, though on my Framework that would bring it from 5 -> 10 token/s so still "excrutiating" speed, probably).
YMMV for sure.
On a 2021 M1 Pro (32GB RAM) I can get either of them as `IQ4_NL` quantized models (the first with reduced context, around 160k; the second can do the whole 264k with RAM left over), running something like 30tokens/s.
On a Framework 13 AMD AI HX370 it can use the same, but both on Q8_0 quantization, full context window, parallelism. Speed is just ~15tokens/s so slower, but definitely smarter than the lower quantized siblings.
Both of them are good developer partners for an engineer who wants more of a second pair of eyes and a rubber duck, rather than a model to just do everything for them. Pretty good for my brain dumping, some commit reviews, sanity checks, just always assume that every claim has to be checked and re-checked.
The only problem is really the context loading, that's pretty slow (starts off around 300token/s on empty context, by the time we get to something like 70-80k which is just a bit of repo discovery, it can run around 80 prompt token/s or less, so there's always a lot more waiting around. Local tools need to bump all of their timeouts, and have to be mindful that there's unlikely to be really meaningful parallelism on these machines with local models.
I'm still figuring out how to approach these things, though. Definitely better than glorified autocomplete or search tool (and too slow for the former, pretty decent for the latter). Their limited skill and performance make it more in line with other tools like my IDE or editors, that they are still in the "tools" compartment of my thinking, rather than "independent, cognitively active entities". Which feels like a good thing.
What I see instead, really, is that most projects no longer, or very rarely look at any contribution, and e.g. any issue + PR/MR combo I make, has a much higher chance of never being looked at and some bot just closes it. Even though the rest of the project might be actually quite active.
It takes some getting used to, apparently being filed together with the "noise", when I try to go out of my way to be as much "signal" as possible. But well, if I really wanted to, I can just run my own changes locally, that's the beauty of OSS, but I hope we can get to some more balanced place over time (being forever the optimist).
In my org, as our Team account (quota-based) was growing, to onboard more people we _had to_ switch to Enterprise (and thus API-based costs). The switch and now cost visibility really drives home the "did this code review worth $15?" kinds of questions...
We are just building the monitoring and I think there will be a crack-down sooner rather than later.
I'm more "code as craft" person, so I'm mainly unaffected (plus found my favourite local LLMs to be the missing rubber ducky with no associated cost), butI cannot say the same for everyone else around.
I think, this needs the original game files to run, if I read things correctly. So probably just gonna read the dev journals, rather than fly this particular bird again...
Having said that, on this query I've seen very little difference in the quality, there's nothing to be "2x as good on" for the "2x quota usage", so shrugs?
It's a Keychron Q10 Max (Alice Layout) - looks like a split keyboard but it's one piece. It's excellent typing, has both wired (USB-C) and wireless (Bluetooth, probably also radio too? I don't use that) connectivity too. I don't normally use the LED lights, but occasionally they are fun...
It's heavy, so not portable like the one the author uses. Had something like that before, the portability was nice, but then didn't use it much. This is not (practically) portable, but still has all the flexibility, and it's a joy typing with it. I wouldn't go as far as I love a piece of euqipment, but I do look forward using it every day.
I do think more about keyboards as I use Mac, Linux, built in laptop keyboards, this stand alone one, etc... And because of the variety it's really hard to build up some muscle-memory. Ctrl, Option, Alt, Fn, ... basically all the extra keys beside the alphabet are slightly different in all systems. So it's more conscious typing than I'd hope for, but not toooooo bad (and it's not the keyboard's problem, I might have to look into remapping stuff, but it's not that level of pain yet).
Happy typing, everyone!
- get the gravitational constant with these two known masses
- then can deduct the mass of the unknown Earth by its interaction with other masses (say the "g" gravitational acceleration value)
- then from the mass and the otherwise measured size of Earth the density pops out
More details in good ol' Wikipedia: https://en.wikipedia.org/wiki/Cavendish_experiment#Derivatio...
Imho the best options there are
- GitLab (if you want someone else to host your things, and be full featured),
- cloud provider-based repos (GCP, AWS has git hosting, if e.g. you are already using them, but it's a subset of features, needing to tie in other services to be a full replacement), or
- go down the self-hosting route if you have the capacity. GitLab is pretty easy to self-host, and there's forgejo, both mentioned in the doc.
Either way you are looking at paying for hosting, one way or another, which for commercial projects should be a baseline. The question is what to pay for, versus what the team should do itself if they can.
Loads of fun in this and in that. And does indeed make me sometimes question if I know anything about programming (in a good way:)