Similar Items: Exploration Hacking: Can LLMs Learn to Resist RL Training?