$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation | Digital Library | PAMCET | PAMCET