The most profitable use case for local LLMs will be one where the end user doesn't even know a local LLM is running, in the same way that a user doesn't know what libraries Photoshop is running, to them it's just Photoshop.
For example, lets say some image editing software decided to use Stable Diffusion to fill in image data in one of their Content Aware tools or something, they would not tell the user to install and run Ollama or sdapi from their CLI. They would install the LLMs when you install the app, and talk to it when you use the app. The end user would never know an LLM is being ran locally, any more than they know DirectX is running. (some might)
I like this use case because image/music/video editing software already requires good CPU/GPU, and in the case of Photoshop, I'm used to my fans blaring when I run Filter Gallery (lol) I as the end user would not need to know that LLMs are being invoked as I use software.
I think this use case is a lot stronger than any cloud-based one as long as it's this expensive to run GPU in the cloud - and the fact that present cloud behavior is to use one of the Big 3, anyone looking for cloud AI will use an OpenAI or another major provider - in the end something from Microsoft, Google, etc.