That’s a good observation. For this project, I found that while the base model could “read” the image, it didn’t really understand how to use it. GRPO allowed it to effectively search the solution space.
5 karma · joined April 14, 2019
GitHub Link: https://github.com/sunildkumar/model_indistinguishability
Paper Link: https://arxiv.org/abs/2402.00793
Check it out: https://code.groundlight.ai/python-sdk/blog/grime-guardian
Source code: https://github.com/sunildkumar/GrimeGuardian