This sales forecasting mini project walks you through a complete Prophet pipeline, from raw retail data to a business-ready forecast.

Sales Forecasting Mini Project: A Complete Prophet Pipeline on Real Retail Data

Forecasting theory is one thing. A working project is another thing entirely. That gap is exactly why this sales forecasting mini project exists. Instead of another isolated lesson on trend, seasonality, or holidays, this project connects every piece into one real pipeline. You move from a raw CSV file to a forecast a store manager could actually use tomorrow morning.

This article is the companion reference for Episode 90 of the Intelevo YouTube series. Watch the video for the full walkthrough with visuals, or use this article to review the code, revisit the logic, or take notes at your own pace. Either way, by the end, you will see forecasting not as a statistics puzzle, but as a simple, repeatable process.

Why a Mini Project, and Why Now

Over the last several episodes, you learned time series data, decomposition, ARIMA, machine learning approaches, and finally Prophet. Each episode covered one idea in isolation. That approach builds understanding, but it does not build confidence. Confidence comes from applying every idea together, on messy, real data, with no shortcuts.

This sales forecasting mini project closes that gap. It takes the Prophet skills from the previous episode and applies them to an actual retail sales dataset, complete with promotions, holidays, and the everyday noise that real business data always carries. Nothing here gets simplified beyond recognition. Nothing here skips the parts that usually trip people up.

Think Like a Store Manager First

Before writing a single line of code, pause and think like a store manager. A good manager never guesses how much stock to order. Instead, they quietly blend three signals.

First, they track the sales trend. Is weekend revenue climbing steadily this year? Second, they factor in seasonal rhythm. Diwali brings a rush. Monsoon season brings a lull. Third, they account for known upcoming events, like a scheduled promotion or a store anniversary sale.

That mental model is precisely what Prophet does with numbers. It blends a trend, a seasonal pattern, and known events into one forecast. Keep this analogy in mind as you move through the rest of this project. Every technical step maps back to this simple, intuitive idea.

The Real Cost of Getting a Forecast Wrong

Forecasting is not an academic exercise. It carries real financial consequences, and both directions of error hurt.

Order too little stock, and shelves go empty during a rush. Sales get lost. Customers get frustrated, and some quietly switch to a competitor. Order too much stock, and you end up with wasted shelf space, forced markdowns, and cash tied up in boxes that are not selling.

A good forecast sits right in the middle. It gives you enough accuracy to order stock with genuine confidence, rather than a hopeful guess. This is the business reason a sales forecasting mini project matters so much. It is not about building a model for its own sake. It is about making a better business decision.

Meet the Dataset

This project uses daily sales records from a multi-store retail chain. Picture the kind of spreadsheet a real analyst opens on a Monday morning: rows and rows of dates, store identifiers, and revenue figures.

The dataset includes four key columns. The date column marks one row per store, per day. The store column identifies which outlet the row belongs to. The sales column records revenue for that store, on that day. Finally, the promo and holiday columns flag any known special days.

Here is a detail worth remembering. Prophet only strictly needs two of these columns. Everything else exists to help you explain the bumps in the data, not to feed the core model directly.

What You Need Before You Start

This project stays deliberately lightweight on tooling. You need Python, along with four familiar libraries: pandas for handling the data, Prophet for the forecasting model itself, scikit-learn for the accuracy check, and matplotlib for the charts. Install Prophet with a single command, pip install prophet, and you are ready to go.

You also need one dataset: daily sales records with a date column, a store identifier, a sales figure, and any promo or holiday flags you can gather. Real business data rarely arrives perfectly clean, and that is fine. This project embraces that reality rather than avoiding it.

One Assembly Line, Seven Stations

Picture this entire pipeline as one assembly line with seven stations. First, you load the data. Next, you clean it. Then, you explore it visually. After that, you split it by time. Then, you fit the model. Next, you generate a forecast. Finally, you evaluate the result.

Skip a station, and the forecast wobbles. Complete all seven, in order, and the whole pipeline works smoothly. Let’s walk through each station in detail.

Station One and Two: Load and Prepare the Data

The first two stations move fast, and they set the foundation for everything that follows.

# 1. Load the data
df = pd.read_csv("retail_sales.csv")

# 2. Focus on one store first
store = df[df["Store"] == 1].copy()

# 3. Rename to Prophet's two columns
store = store.rename(
    columns={"Date": "ds", "Sales": "y"})
store = store[["ds", "y"]]

First, pandas loads the CSV file. Then, a simple filter narrows the data down to a single store. This step matters more than it seems. Blending every store together blurs each store’s unique pattern, so filtering first keeps the signal clean.

Finally, the columns get renamed to the two names Prophet expects: ds for the date, and y for the value you want to forecast. Notice what is missing here. There is no scaling step, no differencing, and no stationarity check. Prophet handles that complexity internally, so you do not have to.

Station Three: Always Look Before You Forecast

Before fitting any model, always visualize the data first.

# Quick visual gut-check
store.plot(
    x="ds", y="y", figsize=(10, 4))
plt.show()

This single line of code reveals a lot. Look for weekly dips. Are Sundays consistently slower than weekdays? Look for yearly swings. Does a festive season create a predictable surge every year? Look for sudden jumps. Do spikes line up with known promotion dates? Finally, look for gaps. Are there missing days that need a fix before modeling?

Five seconds spent eyeballing this chart can save an hour of confused debugging later. Never skip this step, even when a deadline is tight.

Station Four: Split by Time, Not by Chance

This station trips up more beginners than any other step in this kind of project. Testing on data the model already trained on is like grading an exam using the answer key the student copied from. The result looks great on paper, but it tells you nothing real.

# Hold out the last 90 days to test honestly
cutoff = store['ds'].max() \
         - pd.Timedelta(days=90)

train = store[store['ds'] <= cutoff]
test  = store[store['ds']  > cutoff]

The fix is straightforward. Calculate a cutoff date ninety days before the most recent date in the dataset. Everything before that cutoff becomes the training set. Everything after becomes the test set. This mirrors reality: you train on the past, and you evaluate against the future.

Remember this rule for any time series problem, not just this one. Always split chronologically. Never shuffle a time series randomly, even if that is your habit from other machine learning projects.

Station Five: Teach Prophet About Promotions and Holidays

Now the real modeling begins.

from prophet import Prophet

model = Prophet(yearly_seasonality=True,
                weekly_seasonality=True)
model.add_country_holidays(country_name="IN")
model.add_regressor("promo")

model.fit(train)

First, a Prophet instance gets created with yearly and weekly seasonality enabled. Next, one line adds an entire country’s public holidays automatically, with no manual date list required. Then comes the new piece for this project: a regressor for promotions.

A regressor is simply one more clue. It tells Prophet whether a promotion ran on a specific day. Prophet folds that clue directly into its trend and seasonality math, alongside everything else it already knows. Finally, the model gets trained, but only on the training split, never on the full dataset.

Station Six: Turn the Fitted Model Into a Forecast

With a trained model in hand, generating a forecast takes only a few lines.

future = model.make_future_dataframe(periods=90)
future["promo"] = 0   # assume no promos ahead

forecast = model.predict(future)
model.plot(forecast)

First, a future dataframe extends the timeline ninety days ahead. Since no promotions are planned yet, the promo column gets set to zero for those future dates. Then, a single call to predict produces the forecast, complete with an upper and lower confidence range.

The resulting chart looks familiar if you watched the previous Prophet episode, except now it runs on your own store’s real numbers. The line represents the central forecast. The shaded band represents the honest uncertainty range around that estimate.

Station Seven: Score the Forecast With One Honest Number

A forecast without an accuracy check is just a guess with extra steps. This project uses one metric to keep things simple: MAPE, or mean absolute percentage error.

The formula looks like this:

MAPE = average( |actual − forecast| / actual ) × 100

Break it into plain terms. “Actual” means the real sales figure for that day. “Forecast” means what Prophet predicted for that same day. “Average” means this calculation repeats across every day in the test period, then gets averaged. Finally, multiplying by 100 turns the result into a clean percentage.

from sklearn.metrics import (
    mean_absolute_percentage_error as mape)

score = mape(test['y'], predicted) * 100

A MAPE of around eight percent means the forecast typically lands within eight percent of real sales. That level of accuracy is often good enough to plan stock around confidently, especially compared to a manual, gut-feel estimate.

From Forecast Chart to Stock Order

A forecast only matters if it changes a decision. This is where the whole pipeline pays off in practical terms.

Order your base stock quantity to match yhat, Prophet’s central estimate. During weeks with high uncertainty, keep a buffer stocked up to yhat_upper, the top of the confidence range. Flag any days where a promotion is expected, and staff and stock accordingly. Finally, treat forecasting as an ongoing habit, not a one-time task. Re-forecast weekly as fresh sales data arrives, and let the model stay current.

Common Mistakes in a Sales Forecasting Mini Project

A few pitfalls show up again and again in projects like this one. Knowing them in advance saves real time.

First, avoid forecasting every store as one combined blob. Instead, model each store, or a small cluster of similar stores, separately. Second, avoid splitting train and test sets randomly. Always split by date instead, since shuffling destroys the time-based structure the model depends on.

Third, avoid ignoring promotions and holidays. Feed every known future event into the model as a regressor, rather than leaving the model to guess. Fourth, avoid reporting only the single yhat number to stakeholders. Always share the full uncertainty range alongside it, since a single number hides real risk.

Bringing It All Together

Step back for a moment and look at the whole picture. This sales forecasting mini project started with a real, messy dataset. That data got cleaned and split by time, honestly and without shortcuts. One Prophet model got trained, with holidays and promotions folded directly into the math. Finally, one forecast chart turned into a concrete stock decision, ready for a real business to use.

That is the entire point of building a mini project instead of studying isolated concepts. You now understand not just how each piece works, but how every piece fits together into something usable. You have built a forecasting project end-to-end, not simply learned the theory behind one.

Watch the Full Walkthrough

This article summarizes the key ideas and code from Episode 90 of the Intelevo YouTube series, but the video adds visual explanations, live chart readings, and a full narrated walkthrough of every step in this sales forecasting mini project. If you want to see the pipeline built in real time, with every chart and every line of code explained, watch the full episode on the Intelevo channel.

If this project made forecasting feel more approachable, consider subscribing to Intelevo for the rest of the series, and drop a comment sharing which dataset you would like to see forecasted next. Up next, Episode 91 shifts gears from numbers to words, with an introduction to text data and NLP basics.

Leave a Comment

Your email address will not be published. Required fields are marked *