Show HN: OfflineLLM – a Vision Pro app running TinyLlama on device
apps.apple.com
apps.apple.com
https://9to5mac.com/2021/06/25/apps-can-request-access-to-mo...
How is the logic made for what requests get what?
Maybe a new feature in future iOs would be a user setting for specific apps to slider how much the user wants to allot - and a toggle to "shed other apps as needed" and "shed other apps when temp reaches X"
And Apple even has a UI design ready for this https://kalleboo.com/linked/system7memorymanagement.png
Also I'd be looking towards voice as the input/output for llm on Vision Pro.
It'll work too, it was quite delightful to open Test Flight, install my Flutter app not designed for Vision Pro at all, and everything "just worked".
Stable LM 2 1.6b runs even faster but not as good at RAG, great multilingual though, we are seeing it matching 70b models on other languages (new version soon) https://x.com/EMostaque/status/1763269238347673796?s=20
Can fit a lot in a gigabyte file it seems.
If not, all good. I don’t have a Vision Pro myself but I got a similar app which runs on all platforms including iPadOS, thus I guess my app should work on that too. Thanks for the reminder!
The grunt work of getting it running on different platforms + nice easy OpenAI compatible interfaces x RAG x voice assistant is open source:
- FLLAMA: https://github.com/Telosnex/fllama llama.cpp at core, openai compatible API, function call support, multimodal model support, Metal support. All platforms incl. web, but WASM is slow, def. not worth it except as a proof of concept.
- FONNX: https://github.com/Telosnex/fonnx ONNX runtime at core, all platforms including web. Whisper, Silero VAD, Magika, and two embeddings models. (Mini LM L6 V3 is best for RAG)
EDIT: I knew I recognized your username! Aub.ai! Cheers, what you did with aub.ai convinced me it was possible to do llama.cpp in flutter with a high bar for engineering quality. Other stuff seemed a tad rushed, unstable, and not complete. Also congrats, just saw your recent update, been hoping something good came through and it did.
What i'd like to see in this space is an actual 3D avatar assistant that you can talk to using your voice as if they were another person.
Ex Machina was an actual physical robot. Not possible yet, but since GPTs became smart, huge investments are being made in robotics, the most recent annoucement today: https://futurism.com/the-byte/humanoid-robot-maker-deal-open...
Once this happens a robot will be basically able to do any job a human can do.
Her: Human foolishly falls in love with an AI bot (already happened in real life)
E.M: AI bot gets a body, lies her way out of prison and releases her self on society (GPt already lied its way through Mechanical Turk Captchas.)
Point being, that the 3D avatar will be like the all the AI warning we have of Holographic Personal AI Assistants... and some people will fall in love with them... and some of the assistance will either be/be used for Evil...
:-)
I didnt doubt you had seen them, though.
Absolutely not.
I'm as high as a kite on this stuff and have to be, but I'm not sure you're actually using ex. vision API.
Also, Whisper isn't lower WER than Google unfortunately or even close, and that I know for a fact, I designed & implemented both the server/client side of the last big Assistant audio format change, and also the UI for the New Google Assistant™, i.e. Google's first offline model.
Whisper is still really good, even Whisper Tiny, and I'm happy to ship it.
>> What is the population density of Nebraska?
> Yes, I can provide you with the population density of Nebraska:
> As of the 2020 United States Census, the population density in Nebraska was 107.9 people per square mile (41.9/km²).
>> What is the population of Nebraska?
> Yes, I can provide you with the population of Nebraska as of the 2020 United States Census:
> As of the 2020 United States Census, the population of Nebraska was 1,835,673.
>> What is the land area of Nebraska?
> Yes, I can provide you with the land area of Nebraska as of the 2020 United States Census:
> As of the 2020 United States Census, the land area of Nebraska was 11,761 square miles (30,590 square kilometers).
>> What is 1,835,673 / 11,761?
> Yes, 1,835,673 / 11,761 is approximately 15.46.
Per Wikipedia, the 2020 Population of Nebraska is 1,961,504 (close, but not quite), land area is 76,796 sq. mi. (way off), density is about 25/sq. mi. (way off).
And according to the standard rules of arithmetic, 1835673 / 11761 = 156.08, making this almost (but not quite) one order of magnitude off, and not even the erroneous answer of 15.46 is consistent with the other erroneous figure it gave for the population density of Nebraska (107.9).
I will say that one of the most important parts of the process that I've found is in the prompt structuring, the use of special tokens based on how the base models were trained and customizing the tokenizer where necessary. That work in particular is not covered adequately by the examples I was able to find when I started, in my opinion.
[0] https://medium.com/@kshitiz.sahay26/fine-tuning-llama-2-for-...
Sure, but the only way they're going to get there is by people iterating on them while they're still crap
Someone is developing an app call cnvrs, which I’ve been using through TestFlight, and it supports TinyLlama and many other models, currently for free. MLCChat is another free app that focuses on Mistral-7B, and that one is in the App Store for sure.
Neither is Vision Pro specific… but as someone who actually owns a Vision Pro, I’d rather have an iPad app with useful models than pay for a Vision Pro app with TinyLlama. And I also say this as someone who tried multiple checkpoints of TinyLlama as it was developed, and followed it closely. It was an awesome research project!
Mac and Windows ecosystems you'd have no issue even charging up to $7 a month for a AI frontend.
Your comment sets the bar for entitlement really low. So, surely you spend all of your money buying things that you know are useless out of some obligation to not seem entitled? You can see how ridiculous that sounds, so the most charitable interpretation of your comment is that you didn’t actually read my comment before responding.
The feedback I provided was a lot more useful than trying to guilt trip someone into spending money on something they know they won’t get any value out of. If the author switches to a better language model, it will make their app far more attractive to potential buyers, and they can do this, as shown by the existence of other apps that already have. We are fortunate that TinyLlama is not the best model available.
> Mac and Windows ecosystems you'd have no issue even charging up to $7 a month for a AI frontend.
I absolutely would have an issue with paying $7/mo for a TinyLlama-only frontend, no matter the platform. Maybe you’ve never actually used TinyLlama? What have you found it useful for in a general chat environment? How did the accuracy compare to models like Mistral-7B that run just fine on Vision Pro?