My research encompasses three areas: speech technology ethics, voice AI for language diversity, and multimodal understanding of speech. Each is described below.
Speech Technology Ethics
Speech and language technologies raise hard questions about fairness, privacy, transparency, and whose languages get built for. My work on speech technology ethics combines technical expertise with perspectives from law, philosophy, and public policy to address them.
Research Themes
- Linguistic diversity and minoritized languages
- Fairness and bias in speech technologies
- Privacy, consent, and speech data governance
- Explainable and accountable AI
- Ethics education for speech technology
Current Activities
Current initiatives
- Co-editing Speech Tech Ethics, a forthcoming book with Maggie Glass exploring ethical frameworks for speech and language technologies.
- Contributing to the Coalition for Advancing Research Assessment (CoARA) working group on Ethics and Research Integrity Policy in Responsible Research Assessment for Data and Artificial Intelligence.
- Developing ethical guidelines and educational resources for speech technology research and teaching.
- Integrating AI ethics into speech technology curricula and professional training.
Leadership & Professional Service
- Ethics Chair, INTERSPEECH 2025
- Working Group Member, Coalition for Advancing Research Assessment (CoARA)
- Chair, Ethics Committee, Faculty Campus FryslΓ’n, University of Groningen (2018β2020; 2025β2026)
Selected Publications & Resources
Books
- Speech Tech Ethics (with Maggie Glass), in prep.
Selected work
- Ethical guidelines for speech technology research and deployment.
- Educational materials on responsible AI and speech technology.
Education & Outreach
- Organizer, Speech Tech Summer Schools (2023, 2024)
- Developed the STQ module "GPT in the Classroom: Ethical Considerations" (2024)
- Completed professional training in AI ethics through the University of California and Lund University.
Events & Recent Initiatives
- 2026 β Who Governs the Spoken Word? Speech Tech, Minoritized Languages, and the Limits of AI Law
- 2026 β Ethical Implications of AI in Research Assessment
- 2024 β Co-organizer of Exploring the Dark Side of Future Language Technologies: Linguistic (In)security, Ethics, and Privacy in the HumanβMachine Era
- 2023β2025 β Dutch Speech Tech Day series
Voice AI for Language Diversity
Voice AI for Language Diversity is about building speech technology for languages with limited digital resources. Combining multilingual learning, transfer learning, and data-efficient methods, I work to make high-quality speech technology accessible beyond the world's major languages.
Research Themes
- Speech technology for low-resource and endangered languages
- Multilingual and cross-lingual speech models
- Transfer learning and data-efficient AI
- Speech synthesis and automatic speech recognition
- Language technology for linguistic diversity and language vitality
Current Research
Current research lines
- Text-to-speech synthesis for languages with limited speech data.
- Automatic speech recognition using multilingual and cross-lingual transfer learning.
- Phonological feature mapping between related languages.
- Data augmentation techniques for low-resource speech technologies.
- Evaluation methodologies tailored to low-resource speech technology.
Research Challenges
This research addresses several fundamental challenges in multilingual speech technology:
- Limited training data β developing robust speech systems from small amounts of labelled data.
- Cross-lingual transfer β transferring knowledge from resource-rich languages to under-resourced languages.
- Evaluation β designing meaningful benchmarks and evaluation methodologies for low-resource settings.
- Language vitality β supporting linguistic diversity through practical speech technology.
Applications & Impact
The goal of this research is to ensure that advances in Voice AI benefit speakers of all languages, not only those with abundant data.
Applications include:
- Digital support for regional and minoritized languages.
- Improved accessibility through multilingual speech interfaces.
- Educational technologies for language learning and revitalisation.
- Speech technologies that contribute to language documentation and preservation.
Selected Projects
Examples of ongoing and recent work include:
- Multilingual text-to-speech systems.
- Cross-lingual speech recognition.
- Speech technology for Frisian, Bildts, and other under-resourced languages.
- Development of evaluation frameworks for multilingual speech technologies.
Selected Publications
Amooie, R., Hao, Y., De Vries, W., Dijkstra, J., Coler, M., & Wieling, M. (2026). Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation.
Amooie, R., De Vries, W., Hao, Y., Dijkstra, J., Coler, M., & Wieling, M. (2025). Enhancing standard and dialectal Frisian ASR: Multilingual fine-tuning and language identification for improved low-resource performance. ICASSP 2025.
Do, P., Coler, M., Dijkstra, J., Marchenko, I., Verkhodanova, V., & Le Maguer, S. (2025). The Blizzard Challenge 2025.
Lux, F., Meyer, S., Behringer, L., Zalkow, F., Do, P., Coler, M., et al. (2024). Meta learning text-to-speech synthesis in over 7000 languages. Interspeech 2024.
Multimodal Voice AI: Understanding Meaning Beyond Words
Understanding speech means more than transcribing words β it means catching emotion, intent, and social context too. My work on Multimodal Voice AI pushes machines toward that fuller kind of understanding.
Sarcasm is the central test case: speakers routinely mean something other than what they say. Combining acoustic, linguistic, and multimodal contextual cues, we build AI systems that can better understand and generate expressive, socially aware speech.
Research Themes
- Multimodal understanding of speech and language
- Sarcasm detection and interpretation
- Expressive and emotionally rich speech synthesis
- Prosody and paralinguistic information
- Human-like conversational AI
Current Research
Our current research investigates how multimodal AI can recognize and generate subtle aspects of human communication.
Key research directions include:
- Multimodal sarcasm detection β developing systems that combine speech, text, and contextual information to identify sarcastic intent.
- Sarcastic speech synthesis β creating expressive AI voices capable of producing nuanced and contextually appropriate sarcastic speech.
- Cross-cultural communication β studying how sarcasm markers and pragmatic cues differ across languages and cultures.
- Conversational AI applications β exploring how socially aware speech technologies can improve human-machine interaction.
Selected Publications
Li, Z., Chen, Y., Lai, H., Gao, X., Nayak, S., & Coler, M. (2026). SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning.
Li, Z., Zhang, Y., Gao, X., Nayak, S., & Coler, M. (2025). Making Machines Sound Sarcastic: LLM-Enhanced and Retrieval-Guided Sarcastic Speech Synthesis.
Gao, X., Bansal, S., Gowda, K., Li, Z., Nayak, S., Kumar, N., & Coler, M. (2025). AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation. IEEE Transactions on Affective Computing.
Li, Z., Gao, X., Zhang, Y., Nayak, S., & Coler, M. (2024). A functional trade-off between prosodic and semantic cues in conveying sarcasm.
Raghuvanshi, D., Gao, X., Li, Z., Bansal, S., Coler, M., Kumar, N., & Nayak, S. (2025). Intra-modal relation and emotional incongruity learning using graph attention networks for multimodal sarcasm detection. ICASSP.
Li, Z., Gao, X., Nayak, S., & Coler, M. (2023). Sarcasticspeech: Speech synthesis for sarcasm in low-resource scenarios. SSW.
Media Coverage & Outreach
Research on sarcasm detection and socially aware Voice AI has received international media attention, including:
- The Guardian β Why 'emotional AI' is fraught with problems (2024)
- The Guardian β Researchers build AI-driven sarcasm detector (2024)
- Euronews β AI can detect sarcasm now. Greatβ¦ (2024)
- Popular Science β Neural network trained on Friends can recognize sarcasm (2024)
- BBC Radio 4 β The World this Weekend (2024)
- CBC Radio β As It Happens (2024)
- NOS β Computer herkent sarcasme steeds beter (2024)
Future Directions
Current and future research directions include:
- Improving multimodal fusion methods for social and emotional understanding.
- Developing more natural expressive speech synthesis.
- Creating multilingual resources for sarcasm and pragmatic understanding.
- Investigating ethical questions surrounding emotion recognition in AI.
- Understanding cultural variation in human communication.