Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder.
Yusheng DaiHang ChenJun DuXiaofei DingNing DingFeijun JiangChin-Hui LeePublished in: CoRR (2023)
Keyphrases
- cross modal
- audio visual speech recognition
- multi modal
- audio visual
- multi stream
- visual data
- multimedia retrieval
- perceptual information
- speech recognition
- bit rate
- visual recognition
- visual speech
- visual similarity
- multimedia databases
- image retrieval
- visual information
- training set
- visual content
- computer vision
- semantic similarity
- noisy environments
- high dimensional data
- visual features
- nearest neighbor
- hidden markov models
- similarity measure