Agreed and yes would be awesome... But currently this does not infer large meta 7b model efficiently - like 1 token or lower per seconds. But the small toy story (not so useful) models are fast.
If the above mentioned API / python binding is ready, I'll make a streamlit interface demo. The streamlit demo should be simple. But I have to figure out python binding.