Similar situation happened to me too "and 24GB RAM, this took 20 seconds to start up. After posing the question "What is the most common way of transportation in Amsterdam?", the Vicuna model began to generate its response word by word, taking *15 minutes* to complete the task"
Time is money and the openai is pennies