I've used Ollama primarily Laptops with Windows, Mac or Linux. Performance is the biggest issue here. Using what I have on hand, a laptop with Nvidia T2000 GPU. The ~3b parameter models fit in 4gb of GPU memory, so they perform decent. This machine has 64GB of system memory and Intel i7 10875H.
You'll want at least 5 tokens per second for it to be viable as chat. The T2000 with an i7 does ~20 tokens per second on Llama3.2:3b.
Llama3.3 70b is 42gb as a file, so you'll want 64gb ram minimum before you can even load it. This model does 2 tokens per second. Some hosted solutions are literally 1000x faster with 2000+ tokens per second, on the same model.
My opinion is that local inference is useful, but the larger models don't operate fast enough for you to iterate on your problems and tasks. Start small and use the larger models over lunch breaks when the smaller models stop solving your problems.
Depending on how sensitive your data is to your org, read the Terms of Use for some of the AI providers, some of them don't mine every chunk of data. DeekSeek vs Cerebras would be a good comparison.
Alternatively, you might be able to solve your problems with data that is representative of the original data. Instead of looking at payroll data directly for example, generate similar data and use that to develop your solution. Not always possible to do, but it might let you use a larger cloud hosted model without leaking that super private data.