tutorials
Introductions to tokens
Overview
In this lesson, we'll look at a definition of token and see some examples in English.
Outcomes
After completing this lesson, you should be able to …
- define token
- describe the distinction between token and type
- describe patterns for delimiting tokens in one or more natural languages
Before starting
- Read the following section of Manning et al.'s Introduction to Information Retrieval:
- https://nlp.stanford.edu/IR-book/html/htmledition/tokenization-1.html
Background
What is a token?
In 2.2.1 of Introduction to Information Retrieval, Christopher D. Manning, Prabhakar Raghavan and Hinrich Schütze offer the following definition:
A token is an instance of a sequence of characters in some particular document that are grouped together as a useful semantic unit for processing.
Notice that definition doesn't use word. It aligns with the notion of lexical item. By this definition, a token could be a single word, a sequence of words, or just a part (ex. morpheme) of a word.
Let's look at an example:
I like Valeria's shoes. ["I", "like", "Valeria", "'s", "shoes", "."]
Given our definition of token, we can tokenize or split the sentence I like Valeria's shoes. in a number of different ways. Here are two possibilities:
["I", "like", "Valeria", "'s", "shoes", "."]
["I", "like", "Valeria's", "shoes", "."]
You might be wondering which representation is "correct" or preferable. As we'll see in later lessons, different problems require different solutions.
There are, however, some general considerations:
compatibility with existing tools: statistical models for part of speech tagging, parsing, etc. are trained to expect a certain tokenization strategy. If you're using an off-the-shelf tool (ex. POS tagger) to process data you've tokenized with a strategy unfamiliar to the model1, performance is likely to be poor.
extending or supplementing existing datasets: if you want to supplement some existing dataset, you probably want your tokenization strategy to match.
size: different tokenization strategies can increase or decrease the number of unique tokens encountered. When we learn about classification tasks, this will become important.
In a very large corpus, which of the two tokenization strategies presented above will results in fewer unique tokens (i.e., a more compact set of tokens)?
Types vs tokens
If tokens are roughly lexical items, then we can think of a token as an individual occurrence of a lexical item. For example, the word "the" may occur many times in a document.2
Each of those occurrences is a token.
We use type to reference the category represented by all of the occurrences of a single lexical item (ex. occurrences of "the" treated as a category).
- Identical tokens constitute a single type.
- The set of types or vocabulary of some text is its unique set of tokens.
- The type vs token distinction is that of unique vs non-unique strings.
Let's look at an example:
"his friends are also his enemies"
tokens = ["his", "friends", "are", "also", "his", "enemies"]
types = {"his", "friends", "are", "also", "enemies"}
Tokens in English
If you examine tokenized English datasets, such as those made available by the Universal Dependencies, you'll notice that …
- tokens don't contain whitespace
- possessive markers are split into separate tokens
- ex.
"Iggy's"["Iggy", "'s"] - clitic negations are split into separate tokens
- ex.
"shouldn't"["should", "n't"] - acronyms are not split
- ex.
"D.C."->"D.C."
For further guidance, see https://universaldependencies.org/en/index.html.
Next steps
Practice
- Give an example of a token that is not a word
- Consider a non-English language that you've studied. Alternatively, search for a sample of writing in non-English language.
- Does the language use whitespace to delimit word boundaries? Are there other features that appear useful for delimiting words?
- Give one reason why splitting
"couldn't"as["could", "n't"]is preferable to["couldn", "'t"]
Footnotes
- "unfamiliar" here means data that does not match or resemble the model's training data (ex. the two datasets were processed using different tokenization strategies). ↩
- To understand why, review the parts of speech lesson. ↩