How Robust Are LLMs to Vietnamese Dialects?
TL;DR Evaluates how reliably large language models handle regional Vietnamese dialect variation, revealing robustness gaps that matter for inclusive Vietnamese NLP.
TL;DR Evaluates how reliably large language models handle regional Vietnamese dialect variation, revealing robustness gaps that matter for inclusive Vietnamese NLP.
TL;DR A multi-task approach jointly detects toxic Vietnamese speech and localizes its harmful spans, improving robustness through shared supervision.
TL;DR An in-context learning framework helps LLMs translate Vietnamese idioms by accounting for figurative and culturally grounded meaning.
TL;DR ViToSA introduces the task of locating toxic spans directly in Vietnamese speech, together with data and baseline models.
TL;DR ZeST translates speech between languages without transcripts or parallel data by using visually grounded self-supervised speech representations.
TL;DR A new dataset and evaluation setting expose how current Vietnamese NLP models struggle across regional dialects.
TL;DR ViHateT5 reframes Vietnamese hate-speech tasks as text-to-text generation, providing unified models and resources for language-safety research.
TL;DR VGSAlign discovers bilingual speech correspondences without transcripts by aligning visually grounded representations across languages.
TL;DR SMTCE consolidates Vietnamese social-media classification tasks into a common benchmark and evaluates BERT-based models across them.
TL;DR The ViCTSD dataset supports joint study of constructive and toxic language in open-domain Vietnamese social-media comments.
† Undergraduate student supervised by Nguyen Thanh Luan.