Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space | Digital Library | PAMCET | PAMCET