Filter by research area

10 publications

Preprint

How Robust Are LLMs to Vietnamese Dialects?

Minh Tran, Trinh Chau, Thanh-Nhan Le, Nam Hai Tran, Cuong Dang, Luan Thanh Nguyen, Duc Hoang

TL;DR Evaluates how reliably large language models handle regional Vietnamese dialect variation, revealing robustness gaps that matter for inclusive Vietnamese NLP.

ICASSP

A Multi-Task Approach Towards Robust Vietnamese Audio-Based Toxic Span Detection

Vy Le-Phuong Huynh†, Huy Ba Do†, Luan Thanh Nguyen

TL;DR A multi-task approach jointly detects toxic Vietnamese speech and localizes its harmful spans, improving robustness through shared supervision.

AACL-IJCNLP

Don’t Take it Literally! Idiom-aware Vietnamese Translation via In-context Learning

Luan Thanh Nguyen, Parisa Kordjamshidi

TL;DR An in-context learning framework helps LLMs translate Vietnamese idioms by accounting for figurative and culturally grounded meaning.

INTERSPEECH

ViToSA: Audio-Based Toxic Spans Detection on Vietnamese Speech Utterances

Huy Ba Do†, Vy Le-Phuong Huynh†, Luan Thanh Nguyen

TL;DR ViToSA introduces the task of locating toxic spans directly in Vietnamese speech, together with data and baseline models.

IEEE Access

ZeST: A Zero-resourced Speech-to-Speech Translation Approach for Unknown, Unpaired, and Untranscribed Languages

Luan Thanh Nguyen, Sakriani Sakti

TL;DR ZeST translates speech between languages without transcripts or parallel data by using visually grounded self-supervised speech representations.

EMNLP

Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges

Nguyen Van Dinh†, Thanh Chi Dang†, Luan Thanh Nguyen, Kiet Van Nguyen

TL;DR A new dataset and evaluation setting expose how current Vietnamese NLP models struggle across regional dialects.

ACL Findings

ViHateT5: Enhancing Hate Speech Detection in Vietnamese with Unified Text-to-Text Transformer Models

Luan Thanh Nguyen

TL;DR ViHateT5 reframes Vietnamese hate-speech tasks as text-to-text generation, providing unified models and resources for language-safety research.

SIGUL @ INTERSPEECH

VGSAlign: Bilingual Speech Alignment of Unpaired and Untranscribed Languages using Self-Supervised Visually Grounded Speech Models

Luan Thanh Nguyen, Sakriani Sakti

TL;DR VGSAlign discovers bilingual speech correspondences without transcripts by aligning visually grounded representations across languages.

PACLIC

SMTCE: A Social Media Text Classification Evaluation Benchmark and BERTology Models for Vietnamese

Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

TL;DR SMTCE consolidates Vietnamese social-media classification tasks into a common benchmark and evaluates BERT-based models across them.

IEA/AIE

Constructive and Toxic Speech Detection for Open-Domain Social Media Comments in Vietnamese

Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

TL;DR The ViCTSD dataset supports joint study of constructive and toxic language in open-domain Vietnamese social-media comments.

† Undergraduate student supervised by Nguyen Thanh Luan.