Sign in

Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder.

Yusheng DaiHang ChenJun DuXiaofei DingNing DingFeijun JiangChin-Hui Lee
Published in: ICME (2023)
Keyphrases