Solutions
Speech Technology
Products
Resources
Company
Contact
← All articles
← All articles

Industry Insight

4 min read

Published

VoiceInteraction researchers share latest findings at CALLing-FLUL

Research presented by VoiceInteraction Computational Linguists Liege Mendes and Anna Havras at CALLing-FLUL shows how regional variation, sensitive language and Ukrainian–Russian code-switching challenge speech recognition systems.

By

VoiceInteraction Research Team

VoiceInteraction Computational Linguists Liege Mendes and Anna Havras shared their latest research at the first CALLing-FLUL, held in Lisbon on 8 and 9 September 2026.

Their presentations explored two different but related challenges for speech technology: how responsible speech AI can recognise potentially offensive language in context, and how Automatic Speech Recognition (ASR) can better represent Ukrainian–Russian code-switching.

Both studies highlight a broader challenge for speech recognition: language is not a fixed list of words. It varies across regions and communities, and speakers may move between languages within the same conversation.

CALLing-FLUL is organised by Master's and doctoral Linguistics students at the School of Arts and Humanities of the University of Lisbon, in collaboration with the Centre of Linguistics of the University of Lisbon, providing a peer-reviewed forum for ongoing and completed linguistics research.

Why fixed lists fall short for offensive-term recognition

Liege Mendes presented preliminary findings from her upcoming Master's thesis, Inteligência Artificial Responsável Aplicada ao Processamento Automático de Fala no Reconhecimento de Termos Ofensivos.

Her research examines why identifying potentially offensive or sensitive language requires more than maintaining a fixed list of censored terms. Meaning and offensiveness can vary between communities, regions and varieties of the same language. A term may be widespread in one context, geographically restricted in another, or carry a different interpretation across European and Brazilian Portuguese.

The study combines linguistic categorisation with applied speech processing. Preliminary work across 14 linguistic atlases of Brazilian Portuguese maps differences in the frequency and geographic distribution of potentially offensive terms. The longer-term objective is to inform linguistic filters that can identify, flag and process sensitive expressions while accounting for their origin, occurrence and context of use.

This context-aware approach matters when ASR output is used in media and other information environments. Responsible processing depends not only on recognising a word correctly, but also on understanding the limits of a category and avoiding the assumption that language is homogeneous.

Building ASR data for Ukrainian–Russian code-switching

Anna Havras presented Developing Robust Ukrainian-Russian Code-Switching ASR: Introducing the UCS-BN Corpus as part of the conference poster session.

Ukrainian remains comparatively underrepresented in publicly available ASR training data, while everyday speech in parts of Ukraine frequently includes Ukrainian–Russian code-switching or code-mixing. Existing datasets often emphasise scripted, formal or predominantly monolingual speech and therefore do not fully represent rapid, spontaneous transitions within sentences.

Anna's ongoing research introduces the Ukrainian Code-Switching Broadcast News Corpus (UCS-BN) to address this gap. The corpus contains 46 hours of broadcast audio, divided equally between Ukrainian first-language and Russian first-language subsets. Both subsets include code-switching and code-mixing. The source material comes from the Podrobytsi news channel and is licensed under Creative Commons BY 4.0.

By representing spontaneous language alternation and relevant phonetic variation, UCS-BN is intended to support more robust Ukrainian ASR research and improve language identification in complex bilingual environments. The work also establishes reproducible criteria for creating and analysing this kind of speech resource.

Representative data is part of responsible speech AI

The two studies address different languages and research problems, but they share one operational lesson: speech systems need data and linguistic analysis that reflect how people actually communicate.

A fixed vocabulary cannot fully capture the social and regional context of sensitive language. In the same way, a predominantly monolingual or scripted dataset cannot fully represent spontaneous bilingual speech. Both presentations show the importance of treating linguistic variation as a core research requirement rather than an edge case.

These are ongoing studies, not announcements of released product functionality. Their value lies in making difficult language-processing questions explicit, testing methods against real speech data and strengthening the connection between computational linguistics and responsible engineering.

From research questions to operational speech technology

CALLing-FLUL created a useful space for emerging researchers to compare methods, test assumptions and discuss work while it is still developing. Liege and Anna's participation also reflects the role of computational linguistics within VoiceInteraction: examining language closely so that speech technology can respond more reliably to real users and real communication contexts.

As both research projects continue, their findings can contribute to broader work on representative speech data, multilingual processing and responsible AI. VoiceInteraction will continue connecting research with the practical requirements of speech technology while preserving the distinction between ongoing investigation and validated product capability.

Explore VoiceInteraction's research projects.

← Back to all articles

CONTINUE READING

Related articles

Explore more articles connected to this topic, from practical use cases to product updates and speech technology insights.

Explore more articles connected to this topic, from practical use cases to product updates and speech technology insights.

Operational speech workflows require different approaches

Discuss transcription, monitoring, accessibility, or conversational analysis requirements with the VoiceInteraction team.