Really? For me that gives "Bonjour, comment êtes-vous?" with the default settings.
> text generation output
Yeah, text generation is really something that requires a big model. The Llama 7B param model quantized to 4bit is 13G and that is the smallest model I'd actually attempt to use for unconstrained text generation.