Unless you are already on a batch size 1 and TensorFlow greets you with a nice OOM message... Try some Wide-ResNet for state-of-art classification and play with its widening parameter.
Use gradient checkpointing.
...if the batch is the thing that doesn't fit in RAM, as opposed to the model.