Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning | Digital Library | PAMCET | PAMCET