Forecasting feels like magic until you see how it actually works. A retailer predicts next month’s sales. A delivery app predicts tomorrow’s demand. A power grid predicts next week’s electricity load. Behind many of these predictions, you’ll find two workhorse models: ARIMA and SARIMA.
These two models have powered serious forecasting work for decades, long before machine learning became mainstream. Analysts trust them because they’re transparent, well-understood, and surprisingly effective on real-world business data. Once you learn the small set of ideas behind them, you’ll recognize why they remain a first choice for so many forecasting problems today.
This article is the companion piece to Episode 87 of the Intelevo YouTube series. It breaks down ARIMA and SARIMA models in plain language first, then backs every idea with the math and the Python code you need to use them yourself. By the end, you’ll understand exactly how these models turn raw historical data into a genuine forecast.
Let’s get started.
A Quick Recap: What Is Time Series Data?
Before ARIMA makes sense, you need one building block: time series data.
A time series is simply a sequence of values, ordered by time. Daily temperatures, monthly sales, and hourly stock prices all count. Unlike a regular dataset, order matters here. Shuffle the rows, and you destroy the pattern.
Time series data also has a second property worth knowing: stationarity. A stationary series behaves consistently over time. Its average and its spread stay roughly constant. A non-stationary series, on the other hand, drifts, trends, or swings unpredictably. As you’ll see shortly, this distinction sits at the heart of how ARIMA works.
With that refresher out of the way, let’s meet the acronym itself.
What Do ARIMA and SARIMA Actually Stand For?
ARIMA stands for AutoRegressive Integrated Moving Average. That’s a mouthful, so let’s simplify it immediately: ARIMA blends three separate techniques into a single model. Each technique solves one specific problem in forecasting.
SARIMA stands for Seasonal AutoRegressive Integrated Moving Average. It’s not a different model. Instead, it’s ARIMA with one additional ingredient: an explicit understanding of repeating seasonal cycles.
Once you understand the three ingredients inside ARIMA, SARIMA becomes an easy extension. So, let’s unpack ARIMA first, one letter at a time.
Ingredient 1: AR, or AutoRegression
Think about weather for a second. If today is hot, tomorrow is probably hot too. That’s the entire intuition behind AutoRegression.
AR predicts today’s value using recent past values. It looks at yesterday, the day before that, and so on, then combines them into a forecast. Naturally, more recent observations usually carry more weight than older ones.
The parameter p controls how far back the model looks. If p equals 2, the model uses the last two time steps to predict the next one. A small p keeps the model simple. A large p lets it capture longer memory, though it also adds complexity.
Formally, the AR component looks like this:
y(t) = c + φ1 * y(t-1) + φ2 * y(t-2) + ... + φp * y(t-p) + e(t)
Don’t worry about memorizing that formula. Focus on the idea instead: recent history informs the next prediction.
Ingredient 2: I, or Integrated
Here’s a relatable comparison. Driving on a flat highway is easy to predict. Driving on a bumpy, winding road is not. Raw time series data often behaves like that bumpy road: it trends upward, dips downward, and refuses to sit still.
The “Integrated” part of ARIMA fixes this problem through a technique called differencing. Instead of modeling the raw values, the model looks at the change between consecutive values. This simple trick often flattens a wandering trend into something steady, or stationary, as statisticians call it.
The parameter d tells the model how many times to apply this differencing step. Most real-world datasets need just one or two rounds. Once the data looks flat and stable, the model can finally see the true underlying pattern, instead of getting distracted by the trend.
This step matters more than it might seem. Without it, ARIMA would chase a moving target instead of learning a genuine relationship.
Ingredient 3: MA, or Moving Average
Every good forecaster learns from their mistakes. That’s exactly what the Moving Average component does.
MA works in three steps. First, it makes a prediction. Second, it compares that prediction against the actual value. Third, it uses the size of that error to adjust the next forecast. Over time, this error-correction habit sharpens the model’s accuracy.
The parameter q decides how many past errors the model remembers. A model with q equal to 1 only considers the most recent error. A model with q equal to 3 remembers the last three. Consequently, larger values of q allow the model to correct for longer-lasting patterns in its own mistakes.
Here’s the formula, again purely for reference:
y(t) = c + θ1 * e(t-1) + θ2 * e(t-2) + ... + θq * e(t-q) + e(t)
Notice something interesting. AR looks at past values. MA looks at past errors. Together, they cover two very different sources of information.
Putting the Three Ingredients Together: ARIMA(p, d, q)
Now that you know all three ingredients, ARIMA’s notation makes complete sense. Analysts typically write it as ARIMA(p, d, q).
Think of these three numbers as a street address for your model. They tell it exactly how to behave:
- p is the look-back window from AutoRegression.
- d is the number of differencing steps from the Integrated component.
- q is the error memory from the Moving Average component.
Change any one of these numbers, and you get a genuinely different model. That flexibility is precisely why ARIMA remains popular decades after its introduction. It adapts to an enormous range of datasets, simply by tuning three small integers.
How Do You Choose p, d, and q?
This question trips up most beginners, so let’s simplify it.
Analysts traditionally use two diagnostic plots to select p and q: the ACF (AutoCorrelation Function) and the PACF (Partial AutoCorrelation Function). Think of these plots as fingerprints. Each dataset leaves behind a unique fingerprint, and these plots help you read it.
The ACF plot reveals patterns useful for choosing q, since it captures how a value correlates with its own past, including indirect effects. The PACF plot, meanwhile, strips away those indirect effects and reveals patterns useful for choosing p.
For d, the answer is often simpler. You keep differencing the data until it looks stationary, which you can confirm visually or through a statistical test like the Augmented Dickey-Fuller test.
In practice, though, most working analysts skip manual plot-reading altogether. Instead, they rely on automated tools like Python’s auto_arima() function, which scans many combinations of p, d, and q, then picks the best-performing one. This automation saves significant time, especially when you’re forecasting many series at once.
ARIMA in Action: A Sales Forecasting Example
Let’s ground all of this theory in a concrete scenario. Imagine you run a retail business, and you have three years of monthly sales data.
ARIMA studies that historical pattern first. It notices the general upward trend. It notices the natural month-to-month fluctuation. Then, it combines both observations into a projection for the months ahead.
Once trained, the model produces a forecast that extends smoothly from your historical data. The forecasted values won’t be perfect, since no model predicts the future with total certainty. However, they’ll capture the underlying momentum in your data far better than a simple guess or a straight-line extrapolation would.
This is precisely why businesses lean on ARIMA for inventory planning, staffing decisions, and budget forecasts. It turns messy historical numbers into a defensible, data-driven projection.
Retailers aren’t the only ones who benefit. Call centers use similar models to predict ticket volume. Hospitals use them to anticipate patient admissions. Even city planners use them to forecast traffic flow. In each case, the underlying workflow stays the same: gather historical data, fit an ARIMA model, then extend the pattern forward with a defined level of confidence. That confidence matters, since ARIMA doesn’t just produce a single number. It also produces a range around that number, so decision-makers understand how much uncertainty to expect.
The Limitation: ARIMA Doesn’t Understand Seasons
ARIMA is powerful, but it has one significant blind spot: seasonality.
Consider ice cream sales. They spike every single summer, without fail. This pattern repeats every year, like clockwork. However, plain ARIMA only examines recent values. It has no built-in concept of “this happened last July too.” As a result, it can miss recurring seasonal spikes entirely, especially in datasets where the seasonal effect is strong.
This limitation isn’t a flaw exactly. It’s simply a gap that a different model needs to fill. That’s where SARIMA enters the picture.
SARIMA: ARIMA That Understands the Calendar
SARIMA solves the seasonality problem directly. As mentioned earlier, it takes the exact same ARIMA framework and adds one more set of parameters for the seasonal pattern.
Analysts write SARIMA as SARIMA(p, d, q)(P, D, Q, m). The first set, (p, d, q), works exactly like before. The second set, (P, D, Q, m), applies the same three ideas, but to the repeating season instead of the raw series.
Here, m represents the length of one seasonal cycle. For monthly data with yearly seasonality, m equals 12. For daily data with weekly seasonality, m equals 7. This single number tells SARIMA exactly how far to look back for the repeating pattern.
If ARIMA is a street address, SARIMA simply adds a zip code. That addition provides crucial context: the repeating seasonal rhythm that ARIMA alone would otherwise miss.
SARIMA in Python: A Working Example
Theory only goes so far. Let’s implement SARIMA in Python using the statsmodels library. The entire process takes just a handful of lines.
import pandas as pd
from statsmodels.tsa.statespace.sarimax import SARIMAX
# sales is a pandas Series indexed by month
model = SARIMAX(sales, order=(1, 1, 1),
seasonal_order=(1, 1, 1, 12))
fit = model.fit()
forecast = fit.forecast(steps=6)
Let’s walk through this code, line by line.
First, the script imports pandas, since our sales data lives in a pandas Series. Next, it imports SARIMAX from statsmodels, which handles both the non-seasonal and seasonal components in one unified class.
Then, it creates the model itself. The order=(1, 1, 1) argument sets the non-seasonal (p, d, q) values. The seasonal_order=(1, 1, 1, 12) argument sets the seasonal (P, D, Q, m) values, and that final 12 tells the model our season repeats every 12 months.
After that, model.fit() trains the model against our historical sales data. Finally, fit.forecast(steps=6) generates a forecast for the next six months.
Notice how little code this actually requires. The library handles every bit of the underlying math, from differencing to error correction to seasonal adjustment. Your real job, as the analyst, is choosing sensible values for p, d, q, P, D, Q, and m. Everything else follows automatically.
When Should You Choose ARIMA or SARIMA?
At this point, a natural question comes up. How do you decide between the two?
Start by plotting your data. If the series shows a clear, repeating seasonal pattern, such as higher sales every December or higher traffic every Monday, then SARIMA is the better fit. Its seasonal parameters exist specifically to capture that repetition.
On the other hand, if your data shows no obvious seasonal rhythm, plain ARIMA usually works just as well, and it trains faster too. Fewer parameters mean less risk of overfitting, especially on shorter datasets.
A helpful rule of thumb: always start simple. Fit a plain ARIMA model first, then examine the residuals, which are the leftover errors after fitting. If those residuals still show a repeating pattern, that’s a strong signal you need SARIMA instead. This step-by-step approach keeps your modeling process disciplined, rather than guessing which model to use from the start.
Either way, both models share the same underlying logic. Once you understand ARIMA deeply, SARIMA takes only a small extra step to learn.
Key Takeaways
Let’s condense everything into a few clear statements:
- AR means learning from the past. Recent values inform the next prediction.
- I means flattening the bumps. Differencing turns trending data into something stable.
- MA means learning from mistakes. Past forecast errors sharpen future predictions.
- ARIMA(p, d, q) combines all three ingredients into one flexible model.
- SARIMA extends ARIMA with a fourth parameter, m, to handle repeating seasonal cycles.
- In practice, tools like
auto_arima()andstatsmodelshandle the heavy lifting, so you can focus on interpreting results instead of hand-tuning formulas.
None of these ideas are actually complicated. They only sound complicated because of the acronyms. Strip away the jargon, and you’re left with three simple, intuitive techniques working together.
Watch the Full Video Walkthrough
This article covers the core ideas behind ARIMA and SARIMA models, but the video walks through every visual, diagram, and code example in real time. If you learn better by watching and listening, head over to Episode 87 on the Intelevo YouTube channel.
While you’re there, consider subscribing for new episodes every week. And if this article or the video helped something click for you, please leave a comment. Your feedback genuinely shapes which topics get covered next.
Coming up in Episode 88, we’re leaving classical statistics behind and stepping into machine learning. We’ll explore how models like Random Forests can forecast time series data in an entirely different way. See you there.
