Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment | Digital Library | PAMCET | PAMCET