Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 2

RoboReward: Optimizing Vision-Language Reward Models for Scalable Robotic Learning

Authors

Ananya Manoj, Shamiul Mirza, Ujjawal Gupta, Rajveer S Yaduvanshi, Manisha

Abstract

Designing robust and scalable reward functions remains a significant challenge in reinforcement learning for robotic systems. Traditional reward engineering approaches are task-specific, labor-intensive, and fail to generalize across diverse real-world environments. In this work, we introduce RoboReward, a vision-language reward model that leverages pretrained CLIP embeddings to generate scalar reward signals from visual observations and natural language instructions. To improve computational efficiency, we propose an 8-frame temporal aggregation strategy that reduces redundancy while preserving essential task-relevant information. Compared to full-frame processing, this reduces training time from 6.48 days to approximately 2 hours on a single T4 GPU. Experimental results demonstrate strong performance, achiev-ing a Pearson correlation of 0.896 with ground-truth rewards and a Mean Squared Error (MSE) of 0.2048. Furthermore, RoboReward generalizes across multiple tasks without requiring manual reward engineering. These results highlight the effectiveness of vision-language models for scalable and efficient reward modeling in robotic learning systems.