Windows AI Studio Preview
github.com
github.com
Funny that some cuda stuff works better through Windows virtualizing Linux than Windows natively, but if we're being honest even as a native Linux user, WSL probably provides a better user experience (vs having to use Nvidia drivers on Linux anyways)
This clearly shows that GNU/Linux must be superior
https://forums.developer.nvidia.com/t/545-drivers-have-bad-f...
https://forums.developer.nvidia.com/t/wayland-native-wayland...
https://gitlab.freedesktop.org/xorg/xserver/-/issues/1317
and literally me being a Linux Nvidia 1080ti user for years and having plenty of issues
Automatic kernel update? https://forums.developer.nvidia.com/t/nvidia-smi-not-working...
What about upgrading to a new version of CUDA? https://stackoverflow.com/questions/43022843/nvidia-nvml-dri...
What about trying something like enabling forward compatibility for CUDA using an older driver? https://discuss.pytorch.org/t/torch-is-unable-to-detect-cuda... (This issue was actually just posted within the last day, so clearly people still have problems.)
If you haven't run into any issues, then I'd say you're very lucky. Just don't pretend lots of others haven't run into issues.
Besides that, I don't think I ever heard that they supports container at all.
For now, but seems to be planned in the future.
Besides, it says the following:
> Windows AI Studio will run only on NVIDIA GPUs for the preview, so please make sure to check your device spec. WSL Ubuntu distro 18.4 or greater should be installed and is set to default prior to using Windows AI Studio.
So seems it'll only run in Linux, as it's a requirement to have Ubuntu WSL running beforehand.
Yes, it'll only run in Linux, but only when that Linux is also running inside Windows.
https://marketplace.visualstudio.com/items?itemName=ms-windo...
The model is technically in everyone's Windows installs, but we don't have the C++ projected WinRT headers to use the Microsoft.Windows.Vision library.
Image-to-text models, filtered by "Microsoft": https://huggingface.co/models?pipeline_tag=image-to-text&sor...
It looks like it doesn't recognize more than one line at once, but combined with another model or algorithm to detect text bounding boxes it'd be handy.
The one somewhat unique offering in Azure is the Document Layout model which gives you back the OCR with titles, headers, paragraphs, and tables all labeled.
This is really, really good for RAG since it's often useful to stuff the nearest header into the chunk of text when generating an embedding (much, much better results this way).
I do hope they are considering going in that direction though.
1. Why would they miss the opportunity to go all in and make use of the Neural engines for this? Or do they already, and I just don’t know how to interpret it?
2. At what point is Apple going to think - hmm, we have a kick ass processor on our hands. What if we run our server fleet - the ones that serve iCloud, Apple Store, all Apple services, databases etc on M* processors? It sure would help even better the economies of scale for Apple to go to TSMC and say - here is our new CPU design for servers, iPhone, iPad, watch and whatever VR thing. Why no love for the server side that must be orders of magnitude power hungrier today?
With upgraded ram (for just 230€ for each 8GB) and storage (just over 1000€ for 2tb; a samsung 990 pro is like 170€) that might be another story, but "check your specs before downloading this app" seems very un-appley. Also, no cuda support etc. Maybe if they make it exclusive to the mac studio?
As for the price - when you get to 32/64/96GB levels of RAM it’s the cheapest setup on the market that can get you this much VRAM. At least that’s what it was half a year ago when I checked the last time
That is the point. Cuda is widely used in particular for training as opposed to tuning and is just flat out not available. So you buy your nice $6000 machine and it just does not work for its intended purpose.
> it’s the cheapest setup on the market that can get you this much VRAM.
shared VRAM is not everything.
Nvidia intentionally nerfs the amount of VRAM in their consumer cards so you need to buy their ridiculously overpriced enterprise cards. I think it's fair to say Apple's lineup is the cheapest way to get >64GB of VRAM-ish memory and that's definitely not because Apple prices their products so fairly.
Those are not equivalent in speed. macs RAM are much slower than these GPUs.
You can swap memory back and forth between RAM and VRAM (with a huge performance penalty) of course, but that's not exactly usable or comparable to what Apple's VRAM sharing setup allows.
Nvidia doesn't sell a nice-but-not-amazing GPU equivalent to Apple's processing power and memory bandwidth that's also capable of operating on >80GB of VRAM at once. Apple's SoC is kind of an oddball in that regard.
I suppose you could take a regular old iGPU (for AMD, Intel, probably also Qualcom/Mediatek) and use its shared memory capabilities as a comparison. However, iGPUs are terrible at machine learning tasks, they don't come close to what Apple can do with their dedicated accelerators.
The best middle ground may be the laptop GPUs with both dedicated RAM and shared RAM, but those will start swapping memory back and forth like crazy running large ML workloads so they're not really comparable.
If you want to run a model that operates on a huge amount of memory at once, I don't think there is a desktop option that can do what Apple does without going for the massive overkill GPUs that will crush the Macbook in terms of performance (at great cost).
Perhaps you know a GPU or iGPU that's capable of running 80GB VRAM workloads at comparable speeds? Because I don't.
- you need more than 16 or 32 GB of ram.
- but less than 100.
- speed is important.
- but losing a factor 4 compared to a GPU is fine.
- money is very limited.
- but paying 6000 for a mac studio is cheap.
- this is very important professional work.
- but I don't need servers, ECC ram, raid, another OS...
Don't get me wrong, I also think Nvidia is heavily price gouging, but there are definitely trade offs and it is not clear at all why 80GB is your magical number and not, say, 108 or 28.
Year 0 of cool thing - Nothing and silence
Year 1-3 of cool thing - Maybe, if you're lucky, a mention on some hardware thing related to it
Year 3-5 of cool thing - Hardware or software launches that uses thing, no mentions of this besides the earlier one if any.
Year 5-6 of cool thing - Next part of hardware or software launches that uses previous launch, no mentions of this besides the earlier one if any.
Obviously, the time-frames differ, but that's generally how they do things.
I wonder why Microsoft helps nVidia, instead of using their own technology?
Here’s an example: https://github.com/Const-me/Cgml
So probably, the group wasn't even aware of that technology because it's far away by either professional connection, or by personal/relationship connections, or they knew about it but had another goal than "maximize use of own stuff" and made the call that the tradeoffs wasn't worth it.
I think the only reason for that monopoly is the mental inertia of everyone involved. The complexity of these AI models is contained within the data in the models (gigabytes of numbers in these tensors), the GPU-running code is rather simple, most of that code is basic BLAS stuff. Unlike traditional GPGPU applications (FEM, numerical simulations, fluid dynamics), the compute kernels used in AI are easily portable across GPU APIs.
Maybe when I have some free time, I should port my library to Linux + Vulkan, just to prove the point.
In my experience, using any non-native GPU API is asking for troubles. On Windows, the native ones are D3D 11 and 12, on Linux and Android it’s often Vulkan, and on MacOS it’s either Metal or that newer thing they have built specifically for AI. It seems ILGPU only has backends for CUDA, OpenCL, and CPU SIMD.
I have general suspicion towards custom compilers. Writing a good compiler is hard, for GPUs even harder due to the weird execution model and insufficient documentation from GPU vendors. HLSL compiler is supported by Microsoft, and compute shaders are used by many millions of gamers every day. Similarly, CUDA is supported by nVidia, usually pretty stable, my only issue with CUDA is vendor lock-in. I have an impression people are often unhappy with the quality and hardware compatibility of less popular GPU APIs like ROCm and OpenCL.
BTW, I asked a friend to test my program on their low-and laptop. The laptop has some Intel core i3 with integrated GPU. The performance wasn’t great at about 1 token/second (single-channel memory), but at least my code worked. I’m not sure it would have worked on that computer if the backend was based on OpenCL.
https://blogs.windows.com/windowsdeveloper/2023/11/15/elevat...
https://blogs.windows.com/windowsdeveloper/2023/12/14/direct...
Coming soon!
=(
But it's easy to publish something on github, and you get an bug tracker. So .. I guess they do it because it's easy to do so?
So I uninstalled it and now my wsl prompt starts with (base) and I don't know how to disable it and all my python scripts are broken because they can't find all the libraries I've installed from pip throughout the years.
0/10 would not recommend.
Try: conda config --set auto_activate_base false