David Ifeoluwa Adelani
Featured Co-authors
- Graham Neubig - 243 publications
- Junichi Yamagishi - 127 publications
- Sebastian Riedel - 105 publications
- Dragomir Radev - 71 publications
- Dietrich Klakow - 69 publications
- Sebastian Ruder - 68 publications
- Genta Indra Winata - 58 publications
- Pontus Stenetorp - 45 publications
- Isao Echizen - 45 publications
- Mikel Artetxe - 41 publications
- Saif M. Mohammad - 39 publications
Research Publications
SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects
Despite the progress we have recorded in the last few years in multilingual...
YORC: Yoruba Reading Comprehension dataset
In this paper, we create YORC: a new multi-choice Yoruba Reading Comprehension dataset...
ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus
We introduce the ÌròyìnSpeech corpus – a new dataset influenced by a des...
Improving Language Plasticity via Pretraining with Active Forgetting
Pretrained language models (PLMs) are today the primary model for natural...
NollySenti: Leveraging Transfer Learning and Machine Translation for Nigerian Movie Sentiment Classification
Africa has over 2000 indigenous languages but they are under-represented...
MasakhaNEWS: News Topic Classification for African languages
African languages are severely under-represented in NLP research due to ...
SemEval-2023 Task 12: Sentiment Analysis for African Languages (AfriSenti-SemEval)
We present the first Africentric SemEval Shared task, Sentiment Analysis...
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages
Africa is home to over 2000 languages from over six language families ...
BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting
The BLOOM model is a large open-source multilingual language model capable...
BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus
BibleTTS is a large, high-quality, open speech dataset for ten languages...
TOKEN is a MASK: Few-shot Named Entity Recognition with Pre-trained Language Models
Transferring knowledge from one domain to another is of practical import...
Task-Adaptive Pre-Training for Boosting Learning With Noisy Labels: A Study on Text Classification for African Languages
For high-resource languages like English, text classification is a well...
MCSE: Multimodal Contrastive Learning of Sentence Embeddings
Learning semantically meaningful sentence embeddings is an open problem ...
yosm: A new yoruba sentiment corpus for movie reviews
A movie that is thoroughly enjoyed and recommended by an individual might...
Is BERT Robust to Label Noise? A Study on Learning with Noisy Labels in Text Classification
Incorrect labels in training data occur when human annotators make mistakes...
Multilingual Language Model Adaptive Fine-Tuning: A Study on African Languages
Multilingual pre-trained language models (PLMs) have demonstrated impressive...
Pre-Trained Multilingual Sequence-to-Sequence Models: A Hope for Low-Resource Language Translation?
What can pre-trained multilingual sequence-to-sequence models like mBART...
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis
Sentiment analysis is one of the most widely studied applications in NLP...
Preventing Author Profiling through Zero-Shot Multilingual Back-Translation
Documents as short as a single sentence may inadvertently reveal sensitive...
MasakhaNER: Named Entity Recognition for African Languages
We take a step towards addressing the under-representation of the African...
Privacy Guarantees for De-identifying Text Transformations
Machine Learning approaches to Natural Language Processing tasks benefit...
Robust Differentially Private Training of Deep Neural Networks
Differentially private stochastic gradient descent (DPSGD) is a variation...
Distant Supervision and Noisy Label Learning for Low Resource Named Entity Recognition: A Study on Hausa and Yorùbá
The lack of labeled training data has limited the development of natural...
Unsupervised Pidgin Text Generation By Pivoting English Data and Self-Training
West African Pidgin English is a language that is significantly spoken...
Generating Sentiment-Preserving Fake Online Reviews Using Neural Language Models and Their Human- and Machine-based Detection
Advanced neural language models (NLMs) are widely used in sequence generation...