active project

LLMs for Lesser Resourced Languages

Research on tokenization and representation choices for language communities poorly served by mainstream pretrained vocabularies.

This research thread focuses on how tokenization choices affect large language model behavior for languages that are poorly served by common pretrained vocabularies. The goal is to make representation choices more inspectable and to connect tokenizer design with practical language technology needs while avoiding deficit framing.

Role
Research direction and implementation
Period
2025-present
Audience
NLP researchers, language technologists, students

Themes

Skills and Tools

PythonTokenizersTransformersEvaluationNLP

Back to projects