Text preprocessing in NLP

Text Preprocessing in NLP: Tokenization, Stemming, and Lemmatization

Machines do not read the way humans do. A sentence like “I loved the movies!” means nothing to a computer until someone breaks it down. That breakdown process has a name: text preprocessing in NLP. It sits at the very start of every natural language project, and it quietly does most of the heavy lifting.

In this article, we walk through three core techniques step by step: tokenization, stemming, and lemmatization. Each one plays a specific role. Together, they turn messy human sentences into clean, structured data. By the end, you will understand exactly how text preprocessing in NLP works, and why it matters so much for real-world machine learning projects.

This article is the companion piece to Episode 92 of the Intelevo YouTube series. If you prefer to watch and listen, the full video walks through the same ideas with live code demonstrations.

Why Machines Struggle With Raw Text

Let’s start with the core problem. Machine learning models only understand numbers. They cannot directly interpret words, punctuation, or grammar. So before any model can learn from text, that text needs conversion into a numeric, structured format.

Here is the challenge, though. Raw text is messy. It contains punctuation, capitalization, plurals, and different verb tenses. Consider these three words: “run,” “running,” and “ran.” To a human, they clearly relate to the same action. To a machine, without help, they look like three completely unrelated strings of characters.

This is exactly where preprocessing comes in. It cleans, simplifies, and standardizes text so that a model can find patterns instead of noise. Without this step, even the most advanced algorithm struggles to make sense of language.

Think of it like prepping ingredients before cooking. You would never throw a whole, dirty onion straight into a pan. Instead, you wash it, peel it, and chop it first. Raw text works the same way. It needs cleaning and cutting down before it becomes useful input for any machine learning “recipe.”

Where This Step Fits in the NLP Workflow

Before we go further, it helps to zoom out and see the bigger picture. A typical natural language project moves through several stages. First, raw text arrives from somewhere: a review, a tweet, a support ticket, or a chat log. Next, that text gets cleaned and standardized. Then, the cleaned text gets converted into numbers. Finally, a model trains on those numbers and makes predictions.

The steps we cover in this article sit right in that second stage. They act as the bridge between messy, human writing and the structured input a model actually needs. Skip this bridge, and everything downstream suffers.

Other cleaning steps often join the pipeline too, such as lowercasing text or removing common filler words like “the” and “is.” These extra steps matter, but tokenization, stemming, and lemmatization remain the three techniques every learner should master first. They form the core of nearly every text-processing workflow you will encounter.

The Three-Step Recipe

Every text preprocessing in NLP pipeline builds on three foundational ideas. First, tokenization breaks text into smaller pieces. Next, stemming chops those pieces down to a rough root. Finally, lemmatization refines them into a proper, dictionary-accurate base form.

Let’s go through each step individually, starting with the very first cut.

Step 1: Tokenization

Tokenization means splitting a piece of text into smaller units, called tokens. Usually, these tokens are individual words. Sometimes, they are entire sentences instead.

Picture a sentence as a long chain. Tokenization cuts that chain into individual links. Once split, each link, or token, becomes something a machine can count, study, and rearrange on its own.

For example, take the sentence “I love NLP!” After tokenization, it becomes four separate tokens:

["I", "love", "NLP", "!"]

Notice that even the exclamation mark becomes its own token. Nothing gets lost. Everything just gets separated into its smallest meaningful pieces.

Two Common Types of Tokenization

There are two levels at which tokenization typically happens, and the right choice depends on the task.

Word tokenization splits text into individual words. It is the most common form, and you will see it used across nearly every NLP project. For instance, “Text data is everywhere” becomes:

["Text", "data", "is", "everywhere"]

Sentence tokenization, on the other hand, splits a paragraph into individual sentences. This becomes useful when sentence-level meaning matters, such as in summarization tasks. For example, “NLP is fun. It is also useful.” becomes:

["NLP is fun.", "It is also useful."]

Both types serve different purposes, so choosing the right one depends entirely on what you plan to build next.

Tokenization in Python

Let’s see tokenization in action using Python’s NLTK library. First, import the function you need:

from nltk.tokenize import word_tokenize

text = "Text preprocessing is simple and powerful!"

tokens = word_tokenize(text)
print(tokens)

# Output:
# ['Text', 'preprocessing', 'is', 'simple', 'and', 'powerful', '!']

Just three lines of code turn a raw sentence into a clean, structured list. That simplicity is exactly the point. Tokenization does not need to be complicated to be powerful.

Step 2: Stemming

Once text gets tokenized, the next step often involves stemming. Stemming reduces a word down to its rough root, or “stem,” by chopping off common endings like -ing, -ed, or -es. It relies on simple, fixed rules rather than any real understanding of language.

Think of stemming like trimming a plant with garden shears. It works quickly and gets the job done, but the result is not always neat. Sometimes, the output is not even a real word.

Why bother with something so rough, then? Because grouping similar words together matters more than perfect grammar in many tasks. Words like “run,” “running,” and “ran” all describe the same underlying action. Stemming groups these variations together, so a model treats them as one concept instead of three separate, unrelated words.

Stemming Examples

Let’s look at a few real examples to see stemming’s strengths and its limitations side by side.

Original WordStemmed Form
runningrun
studiesstudi
happilyhappili
fliesfli
organizationorgan

Notice something important here. “Running” stems cleanly to “run,” a valid word. However, “studies” becomes “studi,” and “happily” becomes “happili.” Neither of those results is a real dictionary word.

This inconsistency is the trade-off with stemming. It works fast, and it works well for many use cases. However, it does not check whether its output actually makes grammatical sense. Speed comes at the cost of precision.

Step 3: Lemmatization

This brings us to the third technique: lemmatization. Like stemming, lemmatization also reduces a word to its base form. Unlike stemming, though, it looks the word up properly, similar to checking a dictionary. As a result, the output is always a real, valid word, called a “lemma.”

Think of lemmatization like asking a librarian for the correct base word, rather than just snipping off an ending yourself. It takes a little more time, but it delivers far more accurate results.

Consider these examples:

  • “studies” becomes “study”
  • “running” becomes “run”
  • “better” becomes “good”

That last example deserves attention. Lemmatization can even handle irregular forms, something stemming simply cannot manage. “Better” and “good” share no common letters at all, yet lemmatization correctly connects them through grammatical knowledge, not simple pattern matching.

Stemming vs. Lemmatization: Which Should You Use?

Both techniques aim to reduce words to a common form, but they take very different paths to get there. Here is a direct comparison:

FactorStemmingLemmatization
SpeedFastSlower
MethodChops fixed endingsDictionary and grammar lookup
AccuracyRough, rule-basedPrecise, context-aware
OutputMay not be a real wordAlways a real word
Example“studies” → “studi”“studies” → “study”

So, which one should you choose? A simple rule of thumb helps here. Use stemming when speed matters most, such as in large-scale search indexing. Use lemmatization when accuracy matters most, such as in chatbots or sentiment analysis, where meaning needs to stay intact.

Consider a search engine that processes millions of queries every minute. Speed becomes critical there, so stemming often wins despite its rough edges. Now consider a medical chatbot that interprets patient symptoms. A single misread word could change the entire meaning of a sentence, so lemmatization becomes the safer, more responsible choice.

Neither technique is universally “better.” Instead, each one solves a different problem, and understanding both gives you the flexibility to choose correctly for your specific project.

Coding Both Techniques Together

Let’s compare stemming and lemmatization directly in Python, using the same word for both.

from nltk.stem import PorterStemmer, WordNetLemmatizer

stemmer = PorterStemmer()
lemmatizer = WordNetLemmatizer()

word = "studies"

print(stemmer.stem(word))            # studi
print(lemmatizer.lemmatize(word, pos="v"))  # study

Notice the pos="v" parameter in the lemmatizer call. This tells the function to treat the word as a verb, which helps it find the correct base form. Same input word, two completely different philosophies. Porter’s stemmer trims fast, while WordNet’s lemmatizer looks the word up properly.

Putting the Full Pipeline Together

Now, let’s connect all three steps into one continuous pipeline. Suppose we start with the raw sentence: “I loved the movies!”

First, tokenization breaks it into individual tokens:

[I, loved, the, movies, !]

Next, stemming or lemmatization refines those tokens further. “Loved” becomes “love,” and “movies” becomes “movie”:

[I, love, the, movie, !]

Finally, what remains is model-ready data. Clean, consistent tokens, ready for the next stage of any NLP project.

This is the core reason text preprocessing in NLP exists. Every advanced language model, from simple search engines to sophisticated chatbots, begins with this exact pipeline. Nothing works without it.

Why This Step Matters So Much

At this point, you might wonder why so much attention goes into what seems like a small preparatory step. Here is the answer. Text preprocessing in NLP directly determines how well every downstream model performs.

If tokenization splits text incorrectly, every following step inherits that mistake. If stemming or lemmatization gets skipped entirely, a model might treat “run,” “running,” and “ran” as three unrelated words, weakening its ability to find patterns. Consequently, poor preprocessing leads directly to poor results, no matter how sophisticated the model architecture becomes afterward.

On the other hand, careful preprocessing creates a strong foundation. It reduces noise, groups related concepts together, and produces clean, structured input. As a result, models trained on well-processed text tend to learn faster and perform more reliably.

In short, this unglamorous first step often makes the biggest difference in how well any NLP project ultimately performs.

Common Mistakes to Avoid

Even simple techniques can trip up beginners. Here are a few pitfalls worth watching for as you practice.

First, avoid mixing stemming and lemmatization within the same pipeline. Pick one approach and stay consistent, since blending the two often creates messy, mismatched tokens. Second, do not forget to handle punctuation deliberately. Sometimes punctuation carries meaning, such as in “!” for excitement, so decide early whether to keep or discard it. Third, remember that lemmatization needs a part-of-speech tag to work correctly. Without it, the lemmatizer may return the original word unchanged instead of its true base form.

Finally, always inspect your output. It takes only a moment to print a few tokens and confirm they look correct. That small habit saves hours of debugging further down the pipeline.

What Comes Next

With clean tokens now in hand, a natural question arises. How does a machine actually use these tokens to learn something meaningful? That is exactly where the next episode picks up.

In the upcoming Episode 93, we explore Bag of Words and TF-IDF. These techniques take the clean tokens produced through text preprocessing in NLP and convert them into numerical vectors. Only then can a machine learning model finally begin learning from the text directly.

Final Thoughts

Text preprocessing might not sound glamorous, but it forms the backbone of every successful NLP project. Tokenization breaks text into manageable pieces. Stemming offers a fast, rough way to group similar words. Lemmatization delivers a slower, more accurate alternative that always produces real words.

Together, these three techniques transform messy, human sentences into clean, structured data that machines can actually learn from. Once you understand this pipeline, working with text data no longer feels intimidating. It feels systematic, approachable, and genuinely simple.

For the full walkthrough with live code demonstrations, watch Episode 92 on the Intelevo YouTube channel. And stay tuned for Episode 93, where we take these clean tokens and turn them into numbers a machine can truly learn from.

Leave a Comment

Your email address will not be published. Required fields are marked *