Iroro Orife

Featured Co-authors

Research Publications

A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation

Cinematic audio source separation is a relatively new subtask of audio source separation.

ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus

We introduce the ÌròyìnSpeech corpus – a new dataset influenced by a des...

BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus

BibleTTS is a large, high-quality, open speech dataset for ten languages.

Learning Nigerian accent embeddings from speech: preliminary results based on SautiDB-Naija corpus

This paper describes foundational efforts with SautiDB-Naija, a novel corpus.

AVASpeech-SMAD: A Strongly Labelled Speech and Music Activity Detection Dataset with Label Co-Occurrence

We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection.

Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets

With the success of large-scale pre-training and multilingual modeling...

MasakhaNER: Named Entity Recognition for African Languages

We take a step towards addressing the under-representation of the African languages in NLP.

Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages

Research in NLP lacks geographic diversity, and the question of how NLP can empower under-resourced languages continues to be crucial.

Towards Neural Machine Translation for Edoid Languages

Many Nigerian languages have relinquished their previous prestige and public acceptance due to various social factors.

Improving Yorùbá Diacritic Restoration

Yorùbá is a widely spoken West African language with a writing system rich in diacritics.

Masakhane – Machine Translation For Africa

Africa has over 2000 languages. Despite this, African languages account for a significant diversity in culture and communication.

Audio Spectrogram Factorization for Classification of Telephony Signals below the Auditory Threshold

Traffic Pumping attacks are a form of high-volume SPAM that target telephony services.

The Marchex 2018 English Conversational Telephone Speech Recognition System

In this paper, we describe recent improvements to the production Marchex Telephone Recognition System.

Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text

Yorùbá is a widely spoken West African language with a complex writing system.

Semi-Supervised Model Training for Unbounded Conversational Speech Recognition

For conversational large-vocabulary continuous speech recognition, innovative models are crucial.