critique is easier than generaterank: Reward modeling → generate that maximize reward like PPO: AI critic