Context is 4096? My app db DDL is 19877 tokens (using Llama2 tokenizer) long, so that means we need to do a RAG for handling the DDL prompt injection.
A model like this with a 32k long seq_len, like Mixtral, would be a killer for me.
A model like this with a 32k long seq_len, like Mixtral, would be a killer for me.