Elizabeth Salesky
Featured Co-authors
- Graham Neubig 243 publications
- Ryan Cotterell 156 publications
- Isabelle Augenstein 93 publications
- Alan W Black 73 publications
- Antonios Anastasopoulos 66 publications
- Yulia Tsvetkov 64 publications
- Yossi Adi 58 publications
- Colin Raffel 57 publications
- Matteo Negri 52 publications
- Jan Niehues 48 publications
- Alex Waibel 45 publications
Research
Pixel Representations for Multilingual Translation and Data-efficient Cross-lingual Transfer 05/23/2023
We introduce and demonstrate how to effectively train multilingual machi...
A Holistic Cascade System, benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech Translation 01/25/2023
Expressive speech-to-speech translation (S2ST) aims to transfer prosodic...
Language Modelling with Pixels 07/14/2022
Language models are defined over a finite set of inputs, which creates a...
BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus 07/07/2022
BibleTTS is a large, high-quality, open speech dataset for ten languages...
UniMorph 4.0: Universal Morphology 05/07/2022
The Universal Morphology (UniMorph) project is a collaborative effort pr...
Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP 12/20/2021
What are the units of text that we want to model? From bytes to multi-wo...
Assessing Evaluation Metrics for Speech-to-Speech Translation 10/26/2021
Speech-to-speech translation combines machine translation with speech sy...
A surprisal–duration trade-off across and within the world's languages 09/30/2021
While there exist scores of natural languages, each with its unique feat...
SIGTYP 2021 Shared Task: Robust Spoken Language Identification 06/07/2021
While language identification is a fundamental speech and language proce...
Robust Open-Vocabulary Translation from Visual Text Representations 04/16/2021
Machine translation models have discrete vocabularies and commonly use s...
The Multilingual TEDx Corpus for Speech Recognition and Translation 02/02/2021
We present the Multilingual TEDx corpus, built to support speech recogni...
SIGTYP 2020 Shared Task: Prediction of Typological Features 10/16/2020
Typological knowledge bases (KBs) such as WALS (Dryer and Haspelmath, 20...
SIGMORPHON 2020 Shared Task 0: Typologically Diverse Morphological Inflection 06/20/2020
A broad goal in natural language processing (NLP) is to develop a system...
A Corpus for Large-Scale Phonetic Typology 05/28/2020
A major hurdle in data-driven research on typology is having sufficient ...
Phone Features Improve Speech Translation 05/27/2020
End-to-end models for speech translation (ST) more tightly couple speech...
Relative Positional Encoding for Speech Recognition and Direct Translation 05/20/2020
Transformer models are powerful sequence-to-sequence architectures that ...
Generalized Entropy Regularization or: There's Nothing Special about Label Smoothing 05/02/2020
Prior work has explored directly regularizing the output distributions o...
CMU-01 at the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in Morphology 07/23/2019
This paper presents the submission by the CMU-01 team to the SIGMORPHON ...
Exploring Phoneme-Level Speech Representations for End-to-End Speech Translation 06/04/2019
Previous work on end-to-end translation from speech has primarily used f...
Fluent Translations from Disfluent Speech in End-to-End Speech Translation 06/03/2019
Spoken language translation applications for speech suffer due to conver...
Towards Fluent Translations from Disfluent Speech 11/07/2018
When translating from speech, special consideration for conversational s...
Optimizing Segmentation Granularity for Neural Machine Translation 10/19/2018
In neural machine translation (NMT), it is has become standard to transl...