In this article, I explain what I learned about word embeddings, why they are important in Natural Language Processing (NLP), and how FastText works at a high level. I also share a small Python experiment that helped me understand the concept better.
What Are Word Embeddings?
A word embedding is a numerical representation of a word.
Instead of giving a machine-learning model only the word:
laptop
I can represent it as a vector containing numbers:
laptop → [0.12, -0.43, 0.67, 0.21, …]
I found that a vector can contain many dimensions. When I look at one number in the vector, I cannot normally say that it represents something specific such as “technology” or “size.” The useful information comes from the whole vector and its relationship with other vectors.
The main idea I learned is that words that appear in similar contexts can develop similar representations.
For example, words such as:
laptop
computer
keyboard
software
can have relationships in an embedding space because they often appear in related contexts.
This is different from simply assigning a unique number to every word. An embedding attempts to capture useful patterns from language.
Why Do Word Embeddings Matter?
Before learning about embeddings, I looked at how words can be represented using one-hot encoding.
Suppose my vocabulary contains:
phone
laptop
printer
I could represent them as:
phone → [1, 0, 0]
laptop → [0, 1, 0]
printer → [0, 0, 1]
This tells a computer that the three words are different, but it does not naturally show relationships between them.
For example, I might consider phone and laptop to be more related than phone and printer, but one-hot encoding does not represent that relationship.
Word embeddings use dense numerical vectors that can capture patterns learned from language data.
I learned that embeddings are useful for many NLP applications, including:
Text classification
Sentiment analysis
Search systems
Recommendation systems
Machine translation
Question answering
Other language-processing tasks
The Embedding Method I Studied: FastText
One of the embedding methods I studied is FastText.
FastText is related to the Word2Vec family of methods, but one feature that makes it particularly interesting is its use of subword information.
Instead of treating every word as one complete unit, FastText can also learn from smaller pieces of words called character n-grams.
For example, when I look at the word:
developer
I can find character sequences such as:
dev
eve
vel
elo
lop__
ope
per
The exact n-grams depend on the model settings.
This becomes useful when I consider related words such as:
develop
developer
developing
development
These words share several character patterns.
FastText can use these patterns when learning word representations.
*Why Is Subword Information Interesting?*
The subword idea was one of the parts I found most interesting.
If a model has already seen words such as:
develop
developer
developing
and later encounters another related word, FastText can use the character pieces it has learned to help construct a representation.
This can be useful for:
Rare words
Related word forms
Words with similar prefixes or suffixes
Some unusual word forms
Languages with many variations of words
I also learned that this does not mean FastText automatically understands every new word. It uses patterns learned from character pieces to construct a numerical representation.
My Python FastText Experiment
To understand the concept practically, I used the Python gensim library and created a small dataset related to technology.
from gensim.models import FastText
I used a small dataset because my goal was to understand how the model works rather than build a production-level NLP system.
Understanding the settings
vector_size=50 means that each word is represented using a vector with 50 dimensions.
vector_size=50
window=3 controls how many nearby words are considered as context during training.
window=3
min_count=1 allows words that appear only once in my small dataset to be included.
min_count=1
Finally:
sg=1
selects the Skip-gram approach.
Checking Similar Words
After training my model, I could ask it to find words whose vectors are similar to another word.
For example:
model.wv.most_similar(“software”, topn=5)
The model returns a list of words together with similarity scores.
These scores are not percentages. They represent how similar the corresponding vectors are according to the similarity calculation.
Cosine similarity is commonly used for comparing word vectors. At a high level, it measures how similar the direction of two vectors is.
If two vectors point in similar directions, their similarity score can be higher.
However, I should not interpret the results from my small dataset as evidence that FastText has learned general knowledge about technology. My dataset is far too small for that. The experiment is mainly useful for demonstrating how embeddings work.
Looking at a Word Vector
I can also inspect the vector generated for a word:
model.wv[“developer”][:10]
This displays the first ten values from the vector.
The output could look something like:
[0.02, -0.01, 0.04, 0.03, -0.02, …]
The exact values can change depending on the training process.
One important thing I learned is that I should not assign a human-readable meaning to each individual number.
For example, I cannot simply say:
first number = programming
second number = technology
The meaning is distributed across the entire vector.
What matters is how the vectors relate to one another.
Testing an Unseen Word
Another part I found interesting was the idea of an unseen or rare word.
Suppose my training data contains:
develop
developer
developing
development
These words share character patterns.
FastText can use subword information when creating word representations, which makes it different from methods that rely only on complete words.
This helped me understand why character-level information can be useful in NLP.
It also showed me that a word does not always have to be treated as an indivisible object.
What I Found Difficult
At the beginning, I thought a word embedding was simply a long list of numbers assigned to a word.
After studying and experimenting with FastText, I realized that this explanation misses the most important part.
The process is closer to:
Words
↓
Context + subword information
↓
Learned numerical representations
↓
Relationships between vectors
↓
NLP applications
I also initially found the idea of a vector space difficult.
I now think of an embedding space as a mathematical map. Each word has a position on the map, and words with related usage can have related positions.
The model learns these positions from the training data rather than a person manually assigning them.
What I Learned From the Experiment
The biggest thing I learned is that the individual numbers inside an embedding are not the main story.
What matters is the relationship between the vectors.
I also found the use of character n-grams in FastText surprising. Before learning about FastText, I mostly thought of a word as one complete unit. Seeing that a model can use smaller pieces of words gave me a different way of thinking about language.
Another lesson I learned is that the size and quality of the training data matter. My small dataset was useful for learning how FastText works, but it would not be enough to create reliable word representations for a real-world application.
My Key Takeaways
-After working through this topic, these are the main points I took away:
-Word embeddings represent words as numerical vectors.
-The relationships between vectors are more important than individual
numbers.
-Words appearing in similar contexts can develop related representations.
FastText uses subword information through character n-grams.
-Subword information can help when working with rare or unseen word forms.
Similarity scores are measurements of relationships between vectors, not percentages.
- small dataset is useful for learning the concept, but it is not enough to evaluate a real NLP model.
For me, the most important change in understanding was moving from thinking of words as simple pieces of text to thinking of them as learned numerical representations with relationships to other words.
That is what made word embeddings much easier for me to understand.
References
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. https://arxiv.org/abs/1301.3781
Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching Word Vectors with Subword Information. https://arxiv.org/abs/1607.04606
Gensim Documentation. FastText. https://radimrehurek.com/gensim/models/fasttext.html
IBM. What Are Word Embeddings? https://www.ibm.com/think/topics/word-embeddings
