active project
LLMs for Lesser Resourced Languages
Research on tokenization and representation choices for language communities poorly served by mainstream pretrained vocabularies.
This research thread focuses on how tokenization choices affect large language model behavior for languages that are poorly served by common pretrained vocabularies. The goal is to make representation choices more inspectable and to connect tokenizer design with practical language technology needs while avoiding deficit framing.
- Role
- Research direction and implementation
- Period
- 2025-present
- Audience
- NLP researchers, language technologists, students