Words feel simple to us. To a machine learning model, though, words mean nothing at all. A model only understands numbers. So before any NLP project can move forward, we need a reliable way to turn sentences into numbers. That’s exactly where Bag of Words and TF-IDF come in, and this article walks through both, step by step, the same way we cover them in Episode 93 of the Intelevo YouTube series.
By the end, you’ll know how to count words the smart way, how to spot which words actually matter in a document, and how to build both techniques in a few lines of Python. Let’s get started.
Why We Need Numbers, Not Just Clean Words
In the previous episode, we tokenized, stemmed, and lemmatized messy text into a tidy list of words. That step matters, but it isn’t the finish line. A machine learning model still can’t read “cat” or “forecasting” directly. It needs a numeric representation instead.
This is the real question behind Bag of Words and TF-IDF: how do we turn a sentence into numbers, without losing what makes it meaningful? Two techniques answer that question in two different ways. Bag of Words counts words. TF-IDF goes a step further and figures out which of those words actually carry meaning. Together, they form the foundation of almost every classic text-processing pipeline.
The Grocery Bag Analogy
Picture a simple trip to the store. You buy apples, bread, milk, and apples again. Once everything lands in your shopping bag, the order you picked them up in disappears. You can still count exactly what’s inside, though, and how much of each item you have.
That’s precisely what Bag of Words does with a sentence. Every word gets tossed into the bag. We don’t track where a word sat in the sentence. Instead, we only track which words showed up, and how many times each one appeared.
Take this small example: “cat sat on the mat, cat slept.” Once it’s in the bag, we simply count. The word “cat” appears twice. The words “sat,” “on,” “the,” “mat,” and “slept” each appear once. That’s the entire idea behind Bag of Words, and it really is that straightforward.
Bag of Words, Defined Properly
Now let’s put a formal definition on it. Bag of Words represents text as a list of word counts. It ignores grammar completely. It ignores word order completely, too. Only the presence and frequency of each word matter.
Three things happen whenever we build a Bag of Words model:
- It builds a vocabulary. Every unique word across all your documents becomes a column.
- It counts, not orders. Each document becomes a row of numbers, showing how many times each word appears.
- It creates a matrix. Rows represent documents. Columns represent words. Cells hold the counts.
That matrix is the whole point. Once your text turns into a table of numbers, any machine learning algorithm can start working with it.
Seeing the Matrix in Action
Let’s build one by hand, since a concrete example clears up any lingering confusion fast. Consider three tiny documents:
- D1: “the cat sat”
- D2: “the dog sat too”
- D3: “cat and dog played”
Across all three, our shared vocabulary contains seven unique words: the, cat, sat, dog, too, and, played. Every document turns into a row of numbers against those seven columns. D1 gets a 1 under “the,” “cat,” and “sat,” and a 0 everywhere else. D2 and D3 follow the same logic.
That grid of numbers, documents as rows and words as columns, is a genuine Bag of Words matrix. It’s exactly what a model needs to start learning patterns from text.
Building Bag of Words in Python
Here’s the good news: you’ll never build that matrix by hand in a real project. Scikit-learn handles it in just a few lines.
from sklearn.feature_extraction.text import CountVectorizer
docs = ["the cat sat", "the dog sat too", "cat and dog played"]
vectorizer = CountVectorizer()
bow_matrix = vectorizer.fit_transform(docs)
print(vectorizer.get_feature_names_out())
# ['and' 'cat' 'dog' 'played' 'sat' 'the' 'too']
print(bow_matrix.toarray())
# [[0 1 0 0 1 1 0]
# [0 0 1 0 1 1 1]
# [1 1 1 1 0 0 0]]
The CountVectorizer class does two jobs in one call. First, it fits, meaning it learns the vocabulary from the documents you give it. Then, it transforms, meaning it converts each document into its own row of counts. Calling get_feature_names_out() returns the vocabulary in alphabetical order, and calling toarray() returns the finished matrix. Two lines of code accomplish exactly what we built by hand above.
The Catch: Not All Words Deserve Equal Credit
Plain word counts have a real blind spot, though. Words like “the,” “is,” and “and” show up constantly, and they rack up huge counts. Despite that, they tell you almost nothing about what a document actually discusses.
Compare two words directly. “The” appears in nearly every English sentence ever written. Its count runs high, yet its signal stays low. Meanwhile, a word like “forecasting” shows up rarely across an entire collection of documents. When it does appear, though, it often shows up several times within one specific document. That pattern is a genuine clue about the topic.
Plain Bag of Words can’t tell these two situations apart. This gap is exactly why Bag of Words and TF-IDF get taught together, since TF-IDF exists to solve this precise problem.
Thinking Like a Detective
Here’s a useful mental picture. Imagine a detective working through a stack of case files. A skilled detective ignores words that every file shares, like “case,” “date,” or “report.” Instead, they zero in on the word that appears constantly in one particular file, yet rarely anywhere else. That rare, concentrated word becomes the real clue.
TF-IDF turns that detective instinct into a formula. It boosts words that repeat often within a single document. At the same time, it quietly mutes words that show up everywhere. Two ingredients make this possible:
- Term Frequency (TF) measures how often a word appears in one specific document.
- Inverse Document Frequency (IDF) measures how rare that word is across the entire collection of documents.
Multiply these two ingredients together, and you get a score that rewards focus while punishing overuse.
The TF-IDF Formula, Kept Simple
Here’s the formula, written as plainly as possible:
TF-IDF = Term Frequency × Inverse Document Frequency
Term Frequency equals the count of a word in one document, divided by the total number of words in that same document. Inverse Document Frequency equals the log of the total number of documents, divided by the number of documents that contain the word.
You’ll rarely calculate this by hand in practice, since scikit-learn handles the arithmetic instantly. Still, it helps to remember what each half rewards and punishes. TF rewards a word for repeating inside one document. IDF punishes a word for appearing across too many documents.
A Tiny Worked Example
Numbers make this idea click fast, so let’s walk through one. Across our three sample documents, “the” appears in all three, five times in total. Meanwhile, “forecasting” appears in only one document, but three times within that single document.
Because “the” shows up in three out of three documents, its Inverse Document Frequency shrinks toward zero. That gives it a low TF-IDF signal overall, even though its raw count runs high. Because “forecasting” only appears in one out of three documents, its Inverse Document Frequency stays large. That gives it a high TF-IDF signal.
The takeaway matters: TF-IDF automatically demotes filler words. At the same time, it automatically promotes the words that genuinely describe a document. Best of all, this happens without any manual stop-word list.
Building TF-IDF in Python
Switching from raw counts to weighted scores takes one small change in your code. Instead of CountVectorizer, you simply use TfidfVectorizer.
from sklearn.feature_extraction.text import TfidfVectorizer
docs = ["the cat sat", "the dog sat too", "cat and dog played"]
tfidf = TfidfVectorizer()
tfidf_matrix = tfidf.fit_transform(docs)
print(tfidf.get_feature_names_out())
# ['and' 'cat' 'dog' 'played' 'sat' 'the' 'too']
print(tfidf_matrix.toarray().round(2))
# [[0. 0.6 0. 0. 0.6 0.46 0. ]
# [0. 0. 0.5 0. 0.5 0.38 0.65]
# [0.5 0.38 0.38 0.5 0. 0. 0. ]]
Notice how closely this mirrors the previous example. Same shape, same fit_transform() pattern, same overall workflow. Only the numbers change. Instead of plain integers like 1 and 0, you now get decimal scores. Look closely, and you’ll see “the” earns a smaller score than words like “sat” or “dog,” which don’t appear in every document. That’s TF-IDF quietly doing its job.
Bag of Words vs. TF-IDF: A Direct Comparison
At this point, a side-by-side comparison helps cement the difference between these two techniques.
| Bag of Words | TF-IDF | |
|---|---|---|
| What it stores | Raw word counts | Weighted importance scores |
| Common words (“the,” “is”) | Treated as important | Automatically down-weighted |
| Rare, topic-specific words | No special treatment | Boosted so they stand out |
| Best for | Quick baselines, small vocabularies | Search, ranking, most real-world tasks |
Bag of Words works fine for quick baselines and small vocabularies, since it’s fast and easy to interpret. TF-IDF, however, tends to perform better for search engines, ranking systems, and most production-grade text tasks. It naturally highlights what makes each document distinctive, which usually leads to stronger downstream results.
Connecting the Full Pipeline
Let’s zoom out and connect everything end to end, since seeing the whole pipeline together makes each individual step feel less abstract.
- Tokenize and clean. This step, covered in Episode 92, splits, stems, and lemmatizes raw text.
- Bag of Words. This step counts every word per document.
- TF-IDF. This step reweights those counts by importance.
- Numeric vector. This final step produces a vector that’s ready for a machine learning model.
Every stage feeds directly into the next one. Clean words become counted words. Counted words become weighted scores. Weighted scores become a vector that any model can genuinely learn from. Once you see this chain clearly, Bag of Words and TF-IDF stop feeling like separate topics and start feeling like two connected steps in one continuous process.
Three Things Worth Remembering
Before wrapping up, let’s lock in three simple takeaways:
- Bag of Words just counts. Dump the words into a bag, count what’s inside, and don’t worry about order.
- TF-IDF adds judgment. Words that show up everywhere get muted. Rare, meaningful words get boosted instead.
- Scikit-learn does the math. Both
CountVectorizerandTfidfVectorizerturn either idea into just two or three lines of working code.
That’s genuinely all there is to it. If these ideas feel simple by now, that’s exactly the goal.
What Comes Next
Bag of Words and TF-IDF count words well, but they still miss something important: meaning. Neither technique knows that “cat” and “kitten” relate to each other, even though any human reader would spot that connection instantly.
That gap sets up our next episode perfectly. Episode 94 introduces word embeddings, along with an overview of Word2Vec and GloVe. These techniques teach a model that related words share similar numeric representations, even when those words never appear in the same sentence. Counting gets replaced with genuine understanding, and that shift changes what’s possible in NLP.
Quick Glossary for Reference
A short glossary helps when you revisit these ideas later, so here are the core terms in one place.
- Vocabulary: the complete set of unique words found across all documents in a collection.
- Document: one unit of text, such as a sentence, a review, or an article, that gets converted into a row of numbers.
- Bag of Words: a text representation built purely from word counts, with no regard for grammar or order.
- Term Frequency (TF): how often a specific word appears within one document, relative to that document’s total word count.
- Inverse Document Frequency (IDF): a measure of how rare a word is across the entire document collection.
- TF-IDF: the product of Term Frequency and Inverse Document Frequency, used to score how meaningful a word is to a specific document.
- Stop words: extremely common words, like “the” or “is,” that usually carry little topic-specific meaning on their own.
- Vector: a row of numbers representing one document, ready to feed into a machine learning model.
Keep this list nearby whenever you revisit Bag of Words and TF-IDF in future projects. These eight terms cover almost everything you’ll need to read scikit-learn documentation confidently, and they’ll come up again once we move into word embeddings.
Watch the Full Video
This article covers the core ideas behind Bag of Words and TF-IDF, but the full video walks through every example with narration, on-screen code, and a complete build from scratch. If you learn better by watching and following along, head over to the Intelevo YouTube channel and check out Episode 93.
While you’re there, a like, a subscribe, or a quick comment genuinely helps the channel grow, and every comment gets read personally. See you in Episode 94, where words start to carry real meaning.
