# Iroro Orife

## Featured Co-authors

- [Graham Neubig](/content/profile/graham-neubig/index.html) - 243 publications
- [Sebastian Ruder](/content/profile/sebastian-ruder/index.html) - 68 publications
- [Orhan Firat](/content/profile/orhan-firat/index.html) - 68 publications
- [Herman Kamper](/content/profile/herman-kamper/index.html) - 67 publications
- [Ankur Bapna](/content/profile/ankur-bapna/index.html) - 47 publications
- [Benoît Sagot](/content/profile/benoit-sagot/index.html) - 35 publications
- [Yacine Jernite](/content/profile/yacine-jernite/index.html) - 32 publications
- [Alexander Lerch](/content/profile/alexander-lerch/index.html) - 31 publications
- [Julia Kreutzer](/content/profile/julia-kreutzer/index.html) - 30 publications
- [Vukosi Marivate](/content/profile/vukosi-marivate/index.html) - 27 publications
- [Stella Biderman](/content/profile/stella-biderman/index.html) - 26 publications

## Research Publications

### [A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation](/content/publication/a-generalized-bandsplit-neural-network-for-cinematic-audio-source-separation/index.html)
Cinematic audio source separation is a relatively new subtask of audio source separation.

### [ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus](/content/publication/iroyinspeech-a-multi-purpose-yoruba-speech-corpus/index.html)
We introduce the ÌròyìnSpeech corpus – a new dataset influenced by a des...

### [BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus](/content/publication/bibletts-a-large-high-fidelity-multilingual-and-uniquely-african-speech-corpus/index.html)
BibleTTS is a large, high-quality, open speech dataset for ten languages.

### [Learning Nigerian accent embeddings from speech: preliminary results based on SautiDB-Naija corpus](/content/publication/learning-nigerian-accent-embeddings-from-speech-preliminary-results-based-on-sautidb-naija-corpus/index.html)
This paper describes foundational efforts with SautiDB-Naija, a novel corpus.

### [AVASpeech-SMAD: A Strongly Labelled Speech and Music Activity Detection Dataset with Label Co-Occurrence](/content/publication/avaspeech-smad-a-strongly-labelled-speech-and-music-activity-detection-dataset-with-label-co-occurrence/index.html)
We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection.

### [Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets](/content/publication/quality-at-a-glance-an-audit-of-web-crawled-multilingual-datasets/index.html)
With the success of large-scale pre-training and multilingual modeling...

### [MasakhaNER: Named Entity Recognition for African Languages](/content/publication/masakhaner-named-entity-recognition-for-african-languages/index.html)
We take a step towards addressing the under-representation of the African languages in NLP.

### [Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages](/content/publication/participatory-research-for-low-resourced-machine-translation-a-case-study-in-african-languages/index.html)
Research in NLP lacks geographic diversity, and the question of how NLP can empower under-resourced languages continues to be crucial.

### [Towards Neural Machine Translation for Edoid Languages](/content/publication/towards-neural-machine-translation-for-edoid-languages/index.html)
Many Nigerian languages have relinquished their previous prestige and public acceptance due to various social factors.

### [Improving Yorùbá Diacritic Restoration](/content/publication/improving-yoruba-diacritic-restoration/index.html)
Yorùbá is a widely spoken West African language with a writing system rich in diacritics.

### [Masakhane – Machine Translation For Africa](/content/publication/masakhane-machine-translation-for-africa/index.html)
Africa has over 2000 languages. Despite this, African languages account for a significant diversity in culture and communication.

### [Audio Spectrogram Factorization for Classification of Telephony Signals below the Auditory Threshold](/content/publication/audio-spectrogram-factorization-for-classification-of-telephony-signals-below-the-auditory-threshold/index.html)
Traffic Pumping attacks are a form of high-volume SPAM that target telephony services.

### [The Marchex 2018 English Conversational Telephone Speech Recognition System](/content/publication/the-marchex-2018-english-conversational-telephone-speech-recognition-system/index.html)
In this paper, we describe recent improvements to the production Marchex Telephone Recognition System.

### [Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text](/content/publication/attentive-sequence-to-sequence-learning-for-diacritic-restoration-of-yoruba-language-text/index.html)
Yorùbá is a widely spoken West African language with a complex writing system.

### [Semi-Supervised Model Training for Unbounded Conversational Speech Recognition](/content/publication/semi-supervised-model-training-for-unbounded-conversational-speech-recognition/index.html)
For conversational large-vocabulary continuous speech recognition, innovative models are crucial.
