Forecasting used to mean one thing: pick a formula and hope it fits. ARIMA and SARIMA gave us that formula for years. They work well, but they carry a strict assumption. They expect your data to follow one fixed mathematical pattern.
Real life rarely cooperates. Sales spike during festivals. Demand shifts with the weather. Traffic changes on weekends. A single formula cannot capture all of that at once. This is where time series forecasting with machine learning changes the game entirely.
This article is the companion piece to Episode 88 of the Intelevo YouTube series. If you prefer to watch and listen, the full video walks through every concept here with visuals and a live coding demo. This article gives you the same content in written form, so you can study it at your own pace and revisit the code whenever you need it.
By the end, you will understand why machine learning models forecast differently than statistical models. You will also learn how to engineer features, split your data correctly, train a model, and judge whether your predictions are actually good.
Let’s get started.
Why Move Beyond ARIMA?
ARIMA and SARIMA rely only on a series’ own past values. They assume the future looks statistically similar to the past. For simple, stable patterns, this assumption works beautifully.
However, most real-world series don’t behave that simply. A retail store’s sales depend on more than just yesterday’s sales. They depend on whether it’s a weekend, whether a festival is coming up, and whether a competitor just launched a discount. ARIMA cannot easily absorb all of these signals together.
Machine learning models solve this problem differently. Instead of assuming one fixed structure, they learn flexible, non-linear relationships directly from data. As a result, they can combine dozens of signals into a single prediction. This flexibility comes at a cost, though. ML models need well-designed features and enough historical data to learn from. Once you provide that, they can outperform traditional models by a wide margin.
The Core Idea: Meet Priya, the Shopkeeper
Here’s an analogy that makes the entire concept click immediately.
Imagine Priya, a shopkeeper who has run her store for twenty years. She never opens a textbook or calculates a formula before guessing tomorrow’s sales. Instead, she remembers what she sold yesterday. She also remembers what she sold last Monday. She notices if a festival is approaching. She glances outside to check the weather.
Then, without even realizing it, she blends all of these clues into a single instinctive number. That’s her forecast for tomorrow.
This is exactly what a machine learning model does, except it does the blending mathematically instead of instinctively. Every technique in this article, from feature engineering to model training, exists to teach a computer to think like Priya. Keep her in mind as we move forward, since she makes every abstract idea concrete.
Feature Engineering: Giving Your Model a Memory
A machine learning model has no built-in sense of time. Therefore, we must hand it that sense manually through features. This step, called feature engineering, is arguably the most important part of the entire pipeline.
Four feature types show up again and again in time series forecasting with machine learning:
Lag features capture what happened before. A lag of one day tells the model yesterday’s value. A lag of seven days tells it last week’s value on the same weekday. These features let the model reference recent history directly.
Rolling statistics smooth out noise. A seven-day rolling average, for example, shows the general trend without reacting to a single unusual day. Rolling standard deviation, meanwhile, tells the model how volatile recent values have been.
Calendar features extract structure from the date itself. Day of week, month, and flags like “is_weekend” or “is_holiday” give the model a sense of recurring patterns tied to the calendar.
External signals bring in the outside world. Weather, promotions, and festival indicators often explain sudden jumps or drops that the series’ own history cannot explain alone.
Once you combine these four categories, your model gains something close to Priya’s intuition. It sees the recent past, the broader trend, the calendar context, and any outside influences, all at once.
Splitting Time Series Data the Right Way
Before training any model, you need to split your data into training and testing sets. This step sounds simple, but it hides a trap that catches many beginners.
In most machine learning problems, you shuffle your data randomly before splitting it. For time series, you must never do this. Shuffling allows the model to train on future data and then get tested on the past. That gives it an unfair advantage it will never have in real deployment.
Instead, always preserve chronological order. Train your model on earlier dates, then test it on later dates. This approach is called walk-forward validation, or time-aware splitting. It mimics how the model will actually be used: forecasting the future from what it currently knows.
Remember this rule above everything else in this article. Get the split wrong, and every other step becomes meaningless, no matter how sophisticated your model is.
Meet the Machine Learning Models
Three models come up constantly in time series forecasting with machine learning. Let’s look at each one briefly.
Linear Regression is the simplest starting point. It fits one straight-line relationship between your features and the target. It won’t capture complex patterns, but it gives you a fast, interpretable baseline to compare against.
Random Forest takes a completely different approach. It builds many decision trees, and each tree independently makes its own guess. Then, the forest averages all of those guesses into one final prediction. This ensemble approach smooths out individual errors and handles non-linear patterns well.
Gradient Boosting, including popular implementations like XGBoost, builds trees sequentially rather than independently. Each new tree focuses specifically on correcting the mistakes of the tree before it. This sequential correction often produces highly accurate forecasts, though it needs careful tuning to avoid overfitting.
For our live demo, we picked Random Forest. It strikes a strong balance between accuracy and forgiveness. Even with default settings, it rarely performs poorly, which makes it an excellent starting point for anyone new to ML-based forecasting.
How Random Forest Actually Learns
Let’s return to Priya’s world for a moment, but scale her up.
Picture a hundred shopkeepers, each with slightly different experience and slightly different information. Ask every one of them to independently guess tomorrow’s sales. Some will guess high, some will guess low, and most will land somewhere reasonable.
Now average all one hundred guesses together. That average is usually far more reliable than any single shopkeeper’s opinion, even the most experienced one.
This is precisely how Random Forest works. Each decision tree trains on a slightly different, randomly sampled slice of the data and features. Individually, a tree might overfit or make a strange call. Collectively, though, their average smooths out these individual mistakes. No single tree needs to be brilliant, since the crowd naturally corrects itself. This simple idea, known as ensembling, powers some of the most reliable forecasting models used today.
Python Walkthrough: Building Features
Let’s turn theory into code. Here’s how you build the four feature types we discussed earlier, using pandas.
import pandas as pd
df['lag_1'] = df['sales'].shift(1)
df['lag_7'] = df['sales'].shift(7)
df['roll_7'] = df['sales'].shift(1).rolling(7).mean()
df['day_of_week'] = df['date'].dt.dayofweek
df['is_weekend'] = df['day_of_week'].isin([5, 6])
df = df.dropna() # drop rows without full history
Let’s break this down. shift(1) moves the sales column forward by one row, so today’s row now holds yesterday’s value. shift(7) does the same thing, but seven days back, capturing last week’s value on the same weekday.
Next, roll_7 calculates a seven-day rolling average. Notice that we shift by one day before rolling. This detail matters enormously, since it ensures the rolling window never includes today’s actual value. Skipping this shift would leak future information into your features, silently ruining your model’s real-world accuracy.
Then, we pull day_of_week directly from the date column and flag weekends with a simple boolean check. Finally, dropna() removes any rows that don’t have complete lag or rolling history yet, since a model cannot learn from incomplete features.
Python Walkthrough: Training and Predicting
With features ready, training the model takes surprisingly few lines of code.
from sklearn.ensemble import RandomForestRegressor
features = ['lag_1','lag_7','roll_7','day_of_week','is_weekend']
split = int(len(df) * 0.8) # time-ordered split, no shuffle
X_train, X_test = df[features][:split], df[features][split:]
y_train, y_test = df['sales'][:split], df['sales'][split:]
model = RandomForestRegressor(n_estimators=200, random_state=42)
model.fit(X_train, y_train)
predictions = model.predict(X_test)
First, we import RandomForestRegressor from scikit-learn. Then, we define our feature list and calculate an 80 percent split point. Because we never shuffle the data, this split naturally preserves chronological order.
Next, we separate our features and target into training and testing sets. X_train and X_test hold our engineered features, while y_train and y_test hold the actual sales values.
Finally, we create the model with 200 trees, fit it on the training data, and generate predictions on the unseen test data. That’s it. Five lines of feature engineering combined with three lines of model training give you a fully working ML forecaster.
Evaluating Your Forecast: MAE and RMSE
Once you have predictions, you need to know how good they actually are. Two metrics answer this question clearly, and thankfully, neither requires heavy math to understand intuitively.
Mean Absolute Error (MAE) measures the average size of your mistake, expressed in the same units as your data. If your MAE is 12 units, your predictions miss the actual value by 12 units on average. The formula looks like this:
MAE = average( | actual − predicted | )
Root Mean Squared Error (RMSE) measures something similar, but it punishes larger errors more heavily than smaller ones. This happens because RMSE squares each error before averaging. As a result, one big miss hurts your RMSE score far more than several small misses would.
RMSE = √ average( (actual − predicted)² )
For both metrics, lower always means better. Think of them simply as your model’s average miss distance. If you need a metric that treats every error equally, use MAE. If large errors concern you more than small ones, watch RMSE closely instead.
Visualizing Actual vs. Predicted Sales
Numbers alone can feel abstract, so plotting actual values against predicted values helps enormously. When you overlay these two lines on a chart, you don’t need them to match perfectly. Instead, watch whether they move together.
If your predicted line rises and falls in sync with your actual line, your model has learned the underlying pattern successfully. Small gaps between the lines are completely normal and expected. Large, consistent gaps, however, signal that your features or model need more work.
This visual check often reveals problems that raw metrics hide. For example, a model might have a decent overall MAE while still missing every weekend spike. A quick plot exposes that immediately.
Common Pitfalls to Avoid
Before wrapping up, let’s cover four mistakes that trip up almost every beginner in time series forecasting with machine learning.
Data leakage happens when future information accidentally slips into your features. A rolling average that includes today’s value is a classic example. Always double-check your shifts and windows to prevent this.
Shuffled splits destroy the entire purpose of time-aware validation. If you shuffle before splitting, your model effectively peeks into the future during training, making your evaluation results meaningless.
Ignoring seasonality wastes some of the strongest signals available. Calendar features like day of week and month often carry more predictive power than people expect. Skipping them leaves real accuracy on the table.
Overfitting occurs when your model memorizes noise instead of learning genuine patterns. This risk grows with model complexity, so always validate on a proper holdout set before trusting your results.
Avoid these four traps, and you’ll sidestep the majority of problems beginners face with ML-based forecasting.
Key Takeaways
Let’s bring everything together into one clear summary.
First, time series forecasting with machine learning means combining many signals into a single prediction, just like Priya does instinctively. Second, features give your model memory through lags, rolling statistics, and calendar signals. Third, always split your data chronologically and never shuffle it. Fourth, ensemble models like Random Forest average many simple guesses into one stronger, more reliable prediction. Finally, MAE and RMSE both answer the same core question: how far off were you, on average?
None of this required magic. It only required turning Priya’s instinct into a language a computer could understand: Python and well-designed features.
What’s Next
In Episode 89, we move to Facebook Prophet, a forecasting tool built specifically for business use cases. Prophet handles seasonality automatically and produces forecasts that non-technical teams can read at a glance, often with just one line of code. If today’s article felt hands-on, you’ll enjoy how much Prophet simplifies the workflow further.
Watch the Full Episode
This article covers the core ideas from Episode 88 of Intelevo, but the video adds visual walkthroughs, a live coding demo, and extra explanations that are easier to absorb by watching than by reading alone. Head over to the Intelevo YouTube channel to watch the full episode.
If you found this guide useful, please consider subscribing to the channel and sharing it with someone who’s learning data science. Your feedback in the comments genuinely shapes future episodes, so don’t hesitate to share your thoughts, questions, or requests there.
Thank you for reading, and see you in Episode 89.
