You actually don't need more RAM to batch multiple inference tasks of the same model.
(Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).
(Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).