arguably it takes more effort/expertise to curate those examples than just pull in a bunch of data that roughly loosely covers the intended surface area. bitter lesson redux?
Most QLoRAs only target attention layers. This gives pretty good performance for minimal processing time. But this one adds adapters to all linear layers. Also I'm not sure what a "step" is but they ran the whole dataset through at least twice.