2,577 karma · joined August 1, 2014
Yeah, I guess international law was rather incompatible with their ambitions at the time. Fair play to them though, they certainly made the most of it.
I'm 99% sure the verbosity required in the system prompt to teach non-M$ models this new ever-so-slightly-different-but-not-obviously-necessary chart def abstraction format, and the iterations required to get it right, will outweigh any supposed efficiency gains resulting from using it.
Just stating the obvious.
Interesting because again, the license Kimi shipped under [0] defines "Model as a Service" as
> giving a third party access to language model inference or fine-tuning (e.g., via API) in a manner that allows such third party to exercise meaningful control over the inputs, parameters, or training data
and then they have that clause around if you operate such a business above $20M aggregate revenue you need a separate agreement with Moonshot before commercial use, which presumably captures the majority of the larger neoclouds best placed to optimise this.
But where's the line? Say you offer infra optimised for GPU inference, warm pools, isolation per customer, exposed control plane, billed per GPU-hour rather than token, the invoice says compute rather than calls. The customer arguably 'self-hosts', you're probably fine? And if you as a provider run the serving stack and hand your customer an inference API, you're inside the definition regardless of whether you charge by the second or by the token. But what about if you give them direct hardware access, but have the weights cached on NVMe / ramdisk hyperlocal to the infra they're renting so that their hf cli pull only takes a few seconds? Sure, a managed warm pool of GPUs with K3 pre-loaded probably isn't ok, but a local hugging face lru cache holding 'whatever your customers pull down most often', superoptimised for fast weight swaps that the customer controls... is? Is it?
Again, where's the line? Is it materially different from a local docker registry mirror? What about safetensors checkpoints pre-sharded for the specific hardware topology you're offering? Does it matter whether you perform the checkpoint optimisation yourself and make it available, or merely cache one published on HF that happens to target exactly the hardware you rent out? What if you published that checkpoint yourself?
I'm definitely overthinking this, and I'm sure there's been conversation here about this already, but the other kimi threads[1] are enormous. And I am curious.
I'm also curious to know whether Moonshot would actually be against a setup like this. Guessing they would if it was AWS (not quite elastic but not that dissimilar), but what about others? Realistically I guess it'd be easier to just talk to them, especially if you were doing it in a way that targets a slice of the pie they never would have gotten anyway due to data residency requirements etc..
[0]: https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE [1]:https://news.ycombinator.com/item?id=49065752
The original Wasmtime 1.0 blog post explains it all really well: https://bytecodealliance.org/articles/wasmtime-1-0-fast-safe...
After a bunch of middle clicking I landed here [0]. So if I understand correctly, the current stack switching proposal depended on exception handling to be implemented first for resume.throw, so that bit was blocked until now.
But it's also further along than you might assume [1], i.e. you can already invoke Wasmtime with `-W stack-switching` and hit the boundaries where the experimental implementation breaks.
Sadly though, like many ambitious Wasm things that don't have revenue directly attached to them, it now seems partly a case of finding someone willing to sponsor or finish the remaining work [2].
I also found the list of open stack-switching issues[3] useful.
(also, minor nit - Wasm != Wasmtime)
[0]: https://github.com/bytecodealliance/wasmtime/issues/10248
[1]: https://github.com/bytecodealliance/wasmtime/issues/12941
[2]: https://github.com/bytecodealliance/wasmtime/issues/12941#is...
[3]: https://github.com/bytecodealliance/wasmtime/issues?q=state%...
WASM is already running in production at a whole bunch of financial services orgs and government infra.
The thing is, it's not running anywhere near HTML, CSS or JavaScript. It's running serverside, mostly on Wasmtime - which, as it happens, is what this post is all about.
The Guardian is only 'left leaning' in that it provides mildly controversial but reassuringly inconsequential opinions and sound bytes for people to regurgitate so they can convince themselves they've done something about that uncomfortable feeling they're carrying around. It's like a sports team. Sponsors and all.
Hah I guess in that way it is just like Pravda.
It supports ydotool[1], wtype[2] and "clipboard fallback with clipboard restore". The first two you can probably think of as AHK equivalents - they wire in at the input layer and inject keystrokes when injecting text. wtype is wayland-only and a bit less invasive, ydotool supports non-wayland also apparently, but I haven't tried it. Neither approach provides 'instant text' - you have to watch the text get typed out, and you don't touch your keyboard while it's happening; the clipboard implementation is fallback for a reason as it's the least reliable. The first two work 'well enough' though, and are fairly tunable.
The other thing hyprvoice does in probably the most linux-friendly and universal way is the 'hotkey handling'. The server creates a socket in /tmp that the cli can then ping when the user triggers the start/stop/cancel, and they do this by binding whatever their DE's keyboard shortcut mapping mechanism is to trigger `hyprvoice toggle` as a background shell command. This works extremely well and is much cheaper than you'd intuitively think coming from Windows. This way you don't have to interface with DE-specific global keyboard listeners etc, but leave that to the WM (that's not to say that your installer couldn't prompt the user to configure the keyboard shortcut for them with their detected WM, you just wouldn't do it in the software itself).
I haven't actually looked at your project in too much depth yet as I have a solution for this already, so apologies if none of the above is news to you. Hope it helps though - happy to poke around and contribute something if the gap's still there.
[0]: https://github.com/leonardotrapani/hyprvoice [1]: https://github.com/ReimuNotMoe/ydotool [2]: https://github.com/atx/wtype
Not sure about other solutions, but one suggested workaround here would be to silently uninstall Windows without consent.
I've not used gnome for years, but I have a vague memory of gnome/mutter running on a single main thread which used to lock up quite a lot (javascript etc). And because in X it was X that used to manage things like rendering the mouse pointer every frame, whereas in Wayland it flipped to mutter having to do it directly, the stalls were way more obvious in wayland than X, which is where I think a lot of this perception came from.
Again, not sure how much of this is accurate, but that's the point I was trying to make.
Wayland has been great for me for a few years now. I don't use Gnome or nvidia though.
How does your thread-per-connection model compare to Heikki's proposal[0][1] from back in 2023?
[0]: https://www.postgresql.org/message-id/31cc6df9-53fe-3cd9-af5... [1]: https://www.youtube.com/watch?v=xLLakMmVtbY
They might as well change their name to Anthropomorphic at this point.
Anyway, don't most places legally require nearly silent electric vehicles to emit some kind of artificial noise?
In town, filtering, weaving through traffic, getting to the front at lights etc., being able to make a sound which is so ubiquitously embedded in culture that it's instantly recognisable, and so easily localised, really makes a difference. It might be audible, but it's still quieter than many bigger bikes that people ride around town on, and less obnoxious. I guess I'm not the only one who feels that way, as I get a ton of smiles and so many people make an effort to move out of my way - much more so than other bikes I see on the road.
I've been super excited for electric motorbikes for years. I nearly bought a Zero FXS/FXE during covid, and then for the last year or two i've been looking hard at a BMW CE04. But they’d change how I ride, and I’d be more hesitant using them around town simply because being almost inaudible makes me nervous in UK traffic. In saying that, I'd be a lot more comfortable riding around places with a decent cycling culture like Cambridge, where people are used to looking around for smaller quieter vehicles, so I guess this too will change over time. E-bikes are great, but there the problem isn't the ride, it's the theft/security/insurance aspects.
So yeah, I guess until a few of these things change, my buzzy Vespa, with its awesome clutch and gears and crappy little drum brake on the front, will continue to be my go-to.
I know that reads like I'm being snarky, but I'm not trying to be. Within the last decade in the UK we have had (among others):
- the 2016 Snoopers' Charter
- the 2022 Police, Crime, Sentencing and Courts Act
- the 2023 Public Order Act
- the 2023 Online Safety Act
- the 2024 Addendum to the 2016 Snoopers' Charter
Couple that with the whole push to repeal the 1998 Human Rights Act and withdraw from the European Convention on Human Rights, and the fact we've started imprisoning pensioners for holding up signs, and it's really difficult to _not_ see exactly where this leaves us.
I wish I could look away and get on with my life, but I can't - and I'm starting to realise that that is also part of the design. The increasing reluctance to express controversial opinion online isn't an accidental side effect of all of this legislation. It is intended behaviour, the desired outcome.
also thanks for my l10spuh :)