Whisper speech recognition model open-sourced
OpenAI open-sourced Whisper, a speech recognition model. An encoder-decoder Transformer trained on 680,000 hours of multilingual and multitask supervised data collected from the web, it approaches human-level robustness and accuracy on English speech recognition, and supports transcription in multiple languages as well as translation into English. The release drew attention also because the company, then focused on commercial APIs, published the models and inference code.