Text — tokens, embeddings and batches¶
Text has no built-in numeric representation. Tokenization supplies discrete IDs; embeddings supply vectors; batching supplies a consistent input shape. These are separate operations.
Start here before text classifiers.
Code key: “Runnable toy” includes imports and inputs. Other snippets are excerpts: reuse torch, nn, and the model, loader, tokenizer or helper named in the section. Projects contain the complete runnable scripts.
Tokens are not meaning¶
clear image / bad blurry photolengths 2 and 3clear=2, image=3, bad=4, blurry=5, photo=6IDs are keys, not numeric meaning[2,3,0] / [4,5,6]mask [1,1,0] / [1,1,1][2,3,E]one vector per token position[2,E]divide each sum by its real token countToy vocabulary illustration. The runnable project has its own saved training-only vocabulary; pretrained tokenizer IDs depend on the checkpoint.
| Tokenization | What gets an ID | Trade-off |
|---|---|---|
| Word | a word such as inspection |
interpretable, but a large vocabulary and unknown words |
| Subword | reusable pieces of words | handles many rare words; may expand sequence length |
| Character | individual characters | small vocabulary; longer sequences |
A token ID is a lookup key: ID 200 is not “twice as meaningful” as ID 100. Do not feed IDs to a linear layer as ordinary numeric features.
For a tiny word-based baseline, an original vocabulary example:
training_texts = ["camera finds scratches", "camera checks panels"]
vocab = {"<pad>": 0, "<unk>": 1}
for text in training_texts:
for token in text.lower().split():
if token not in vocab:
vocab[token] = len(vocab)
def encode(text):
ids = [vocab.get(token, 1) for token in text.lower().split()]
return ids or [1] # explicitly handle an empty/cleaned-away sentence
Build your vocabulary from training data only. Reuse the same mapping for validation and new inputs. A minimum frequency removes rare vocabulary entries at the cost of more <unk> tokens. Save the tokenizer/cleaning rules and vocabulary with the model.
Lowercasing, punctuation removal and whitespace splitting are baseline choices. They can erase useful distinctions: not, numbers, accented letters, product codes and punctuation can carry meaning. Avoid custom cleaning that conflicts with a pretrained tokenizer.
Special tokens¶
<pad> fills unused positions; <unk> represents unrecognized input. Beginning/end tokens mark boundaries when required. BERT-style [CLS] and [SEP] are part of that model's input convention; the tokenizer normally inserts them. Subword vocabularies reduce the unknown-word problem but do not guarantee that every input is understood.
Use the matching pretrained tokenizer¶
The tokenizer and model must share their vocabulary, normalization and special-token convention. BERT and DistilBERT are related, but use the exact tokenizer belonging to the chosen checkpoint.
from transformers import AutoTokenizer
checkpoint = "distilbert/distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
batch = tokenizer(
["Inspect the panel.", "The camera missed a small mark."],
padding=True, truncation=True, max_length=128, return_tensors="pt",
)
print(batch["input_ids"].shape) # [N,L], integer IDs
print(batch["attention_mask"]) # 1 = real token; 0 = padding
print(tokenizer.convert_ids_to_tokens(batch["input_ids"][0].tolist()))
This may download tokenizer files. AutoTokenizer chooses the implementation for the checkpoint; BertTokenizerFast explicitly chooses a BERT fast tokenizer. .tokenize(text) shows pieces, .convert_ids_to_tokens(...) helps inspect encoded IDs, and .get_vocab() exposes the mapping. The precise pieces and IDs are checkpoint-dependent; do not memorize sample IDs.
Padding, truncation and attention masks¶
Default collation cannot stack vectors with different lengths. Choose a representation suitable for the model:
| Choice | What changes |
|---|---|
| Fixed padding | every example is padded to a chosen maximum |
| Dynamic padding | each batch pads only to its longest sequence |
| Length bucketing | similar-length sequences share batches, reducing waste |
| Truncation | cuts inputs exceeding max_length; useful text can be lost |
| Flattened IDs + offsets | EmbeddingBag pools variable-length inputs without padding |
For lengths 3, 8 and 16, dynamic padding uses 48 positions, of which 21 are padding. Padding to 64 uses 192 positions. That arithmetic is not a speed benchmark: transformer attention and other operations have additional costs, and an attention mask does not remove all padded computation.
Use padding=True for the longest sequence in the current tokenizer call, or padding="max_length" for fixed padding. truncation=True and an explicit max_length bound the input; choose the bound with the model's limit and the text length distribution in mind. A DistilBERT base checkpoint typically supports up to 512 positions, including special tokens.
An attention mask identifies valid positions, usually with 1, and padding with 0. It is not a class label. For manual embedding pooling, you must exclude padding yourself. For transformer batches, pass it to the model.
Pad at collation time¶
from transformers import DataCollatorWithPadding
from torch.utils.data import DataLoader
collator = DataCollatorWithPadding(tokenizer=tokenizer, return_tensors="pt")
loader = DataLoader(tokenized_dataset, batch_size=16, shuffle=True,
collate_fn=collator)
Each tokenized_dataset sample should contain unpadded tokenization fields and an integer labels value. The collator builds [N,L] tensors at batch time. Keep collation on CPU; move the finished batch to the model device in the training loop. This also supports DataLoader workers.
Embeddings are learned lookups¶
nn.Embedding(V,E) stores a trainable table with V rows and E values per row. It maps IDs to vectors; it does not create contextual understanding by itself.
import torch
from torch import nn
table = nn.Embedding(12, 5, padding_idx=0)
ids = torch.tensor([[2, 3, 0], [4, 1, 5]], dtype=torch.long)
vectors = table(ids)
print(vectors.shape) # [2,3,5]: batch, length, embedding dimension
The vectors start from initialization unless pretrained values are loaded. Gradients from the downstream objective teach useful entries. padding_idx=0 suppresses normal gradient updates for the padding row; masking still matters when pooling or after custom initialization.
Alternative representations¶
Counts, TF-IDF, learned word vectors and contextual encoders retain different information about text.
Code and details
| Representation | What it captures | What it loses / costs |
|---|---|---|
| One-hot | identity in a vocabulary-sized vector | sparse; no learned similarity |
| Bag of words | token occurrence/count | word order |
| TF-IDF | counts weighted down for common terms | context and order; fit on training data |
| Static vectors: GloVe, Word2Vec, FastText | learned distributional relationships | a fixed vector per token, with model-specific OOV behavior |
| Contextual model: BERT/DistilBERT | representations influenced by surrounding tokens | more computation and memory |
A small learned embedding can be sufficient for a short-label task. A pretrained language model can help with context, but validation must establish that benefit for your data.
Similarity, context and visualization¶
Cosine similarity compares direction, not vector length:
Near 1 means aligned vectors; near 0 means nearly orthogonal; near -1 means opposed directions. It does not universally mean synonyms, unrelated concepts and antonyms respectively. Antonyms can occur in similar contexts and have similar embeddings.
GloVe returns the same stored vector for a word regardless of sentence. A contextual model can distinguish bat as an animal from sports equipment. AutoModel exposes .last_hidden_state with shape [N,L,E]; select the actual token position or aggregate its subwords. Do not assume a word always occupies a fixed index, and do not include padding or special tokens by accident.
The studied toy embedding exercise predicts context words from a centre word. Repeated co-occurrence and cross-entropy updates shape the lookup table. A tiny toy corpus demonstrates the mechanism; it does not establish broad semantic knowledge. Vector analogies are illustrative relationships, not guaranteed facts.
PCA (PCA(n_components=2).fit_transform(vectors)) projects onto two directions of high variance; t-SNE emphasizes local neighborhoods. Both distort the original space. A nearby pair in a two-dimensional plot is not sufficient evidence of semantic accuracy. Check similarity in the full vector space and evaluate the downstream task.
Remember and apply¶
Check yourself: does zero-padding reduce an unmasked mean embedding? Yes: dividing by total padded length changes the vector's magnitude. The mask must determine both the sum and its denominator.
Continue with pooled and pretrained text classifiers, where the same IDs and masks connect to logits, loss and fine-tuning.
Sources: Hugging Face tokenizers, padding collator, PyTorch Embedding, GloVe.