I'm really salty because I "upgraded" to a 5700XT from a Nvidia GTX 1070 and can't do AI on the GPU anymore, purely because the software is unsupported.
But, as a dev, I suppose I should feel some empathy that there's probably some really difficult problem causing 5700XT to be unsupported by ROCm.
My money, looking at nothing, would be on one of the two Vulkan backends added in Jan/Feb.
I continue to be flummoxed by a mostly-programmer-forum treating ollama like a magical new commercial entity breaking new ground.
It's a CLI wrapper around llama.cpp so you don't have to figure out how to compile it
And that more or less answered it.
So compiling the correct version of llama.cpp for their hardware is confusing.
Compound that with everyone’s relative inexperience with configuring any given model and you have prime grounds for a simple tool to exist.
That’s what ollama and their Modelfiles accomplish.
Weirdly, the Python bindings built without issue with pip.
Hadn't thought about it recently. After seeing it again here, and being gobsmacked by the # of genuine, earnest, comments assuming there's extensive independent development of large pieces going on in it, I'm going with:
- "The puzzled feeling you have is simply because llama.cpp is a challenge on the best of days, you need to know a lot to get to fully accelerated on ye average MacBook. and technical users don't want a GUI for an LLM, they want a way to call an API, so that's why there isn't content extalling the virtues of GPT4All*. So TL;DR you're old and have been on computer too much :P"
but I legit don't know and still can't figure it out.
* picked them because they're the most recent example of a genuinely democratizing tool that goes far beyond llama.cpp and also makes large contributions back to llama.cpp, ex. GPT4All landed 1 of the 2 vulkan backends
#ifndef __HIP__
#include <cuda_fp16.h>
#include <cuda_runtime.h>
#else
#include <hip/hip_fp16.h>
#include <hip/hip_runtime.h>
#define cudaSuccess hipSuccess
#define cudaStream_t hipStream_t
#define cudaGetLastError hipGetLastError
#endif
Then your CUDA code works on AMD.