Building ML APIs with FastAPI

Building ML APIs with FastAPI

You trained a model. It works well in your notebook. But right now, only you can use it.

That’s the gap this article closes. Building ML APIs with FastAPI turns a model that lives in a notebook into a service that anyone can call. And unlike some other frameworks, FastAPI does most of the hard work for you.

This article is the companion piece to EP99 of the Intelevo YouTube series. If you watched EP98, you already built a Flask API for your model. Today, we rebuild that same idea with FastAPI, and you’ll notice the difference immediately. Let’s get into it.

Why Move Beyond a Basic API

In EP98, we wrapped a trained model inside a Flask app. That app worked. However, it also came with hidden costs. We had to write our own input checks by hand and write our own error messages. We had to write our own documentation, too, and keep it updated every time the code changed.

That’s a lot of manual work for something that should feel automatic. Building ML APIs with FastAPI removes nearly all of it. FastAPI checks incoming data automatically. It generates documentation on its own. Plus, it handles many requests at once without any extra effort from you.

So, what exactly makes this possible? Let’s start with a simple picture.

The Analogy: A Self Check-In Kiosk

Picture an airport self check-in kiosk. You walk up, scan your passport, and the kiosk takes over from there. No agent manually checks your details. No one explains the process to you, either. The kiosk already knows the rules, and it applies them instantly.

Three things happen automatically at that kiosk. First, it checks your documents. A wrong passport format gets rejected before you ever reach the counter. Second, it prints your boarding pass without anyone writing instructions by hand. Third, many kiosks run at the same time. Dozens of travelers check in simultaneously, and no one blocks the line for anyone else.

Keep that picture in your mind. That’s exactly what this framework feels like once everything is set up.

Naming the Idea: What Is FastAPI

That kiosk has a real name in the Python world: FastAPI.

FastAPI is a modern Python framework for building web APIs. You describe what your data should look like using plain Python type hints. Then, FastAPI handles the checking, the errors, and the documentation on your behalf.

Four features make this framework stand out. Type hints act as your validation rulebook, so you don’t need a separate set of rules. Auto validation rejects bad requests immediately, with a clear message attached. Auto docs generate a live, interactive documentation page without any extra writing from you. Finally, the framework is async-ready, which means it can handle many requests at once right out of the box.

Together, these four features are why so many teams now reach for this framework instead of older tools.

FastAPI vs Flask: A Quick Comparison

Let’s compare the two frameworks directly, since this helps the difference click into place.

CapabilityFlaskFastAPI
Input validationYou write it by handBuilt in, from type hints
API documentationManual, plus an extra libraryGenerated automatically
Handling many requestsOne at a time by defaultAsync, many at once
Error messagesYou design and write themClear errors, out of the box

None of this means Flask is a bad choice. Flask remains simple and flexible, and plenty of production systems run on it happily. Still, when your priority is speed of development, automatic validation, and free documentation, FastAPI gets you there faster.

Pydantic: The Form-Checking Clerk

Here’s the engine behind FastAPI’s validation: a library called Pydantic.

Pydantic is the tool FastAPI uses to check incoming data. You describe the shape you expect, meaning the field names and their types. Then, Pydantic rejects anything that doesn’t match, and it does this before your own code ever sees the data.

Consider this small example:

from pydantic import BaseModel

class HouseFeatures(BaseModel):
    area_sqft: float
    bedrooms: int
    location_score: float

This class describes exactly what a valid request looks like. Send area_sqft="large", and Pydantic rejects it instantly, since that’s a string, not a number. Send area_sqft=1450, however, and it sails through without any extra code from you.

As a result, your endpoint function only ever sees clean, correctly typed data. You never need to write a manual if isinstance(...) check again.

From Request to Prediction: The Full Flow

So, how does a single request actually travel through your API? Here’s the journey in five simple steps.

First, the client sends a JSON request to your server. Second, Pydantic validates that request against your model definition. Third, your own function runs, using the now-clean data. Fourth, your ML model produces a prediction. Fifth, and finally, the response travels back to the client.

Here’s the part worth remembering: once you define your Pydantic model, steps one, two, and five happen automatically. You only ever write step three, which is your own prediction logic. Everything else is handled for you behind the scenes.

Hands-On: Your First FastAPI App

Theory is useful, but code makes it real. Let’s write an actual FastAPI app, starting with the simplest possible version.

from fastapi import FastAPI

app = FastAPI()

@app.get("/")
def home():
    return {"message": "Intelevo model API is live"}

Only four meaningful lines, and you already have a working web server. Let’s walk through what each one does.

First, you import FastAPI and create an app object. This object represents your entire web application. Next, the decorator @app.get("/") sits right above a plain Python function. That decorator is the piece of magic here: it turns an ordinary function into a live web endpoint, reachable at the root address of your server.

To run this app, open your terminal and type:

uvicorn main:app --reload

Uvicorn is the server that actually runs your FastAPI app behind the scenes. The --reload flag restarts your server automatically whenever you save a change, which saves you a lot of manual restarting during development. Once it’s running, open 127.0.0.1:8000 in your browser, and you’ll see your message appear immediately.

Hands-On: Serving Real Predictions

A “hello world” endpoint is a nice start, but it doesn’t solve any real problem. Let’s make it useful by serving actual predictions.

We’ll reuse the joblib model saved back in EP97, which covered model persistence with pickle and joblib. If you missed that episode, the short version is this: joblib saves a trained model to disk, so you don’t need to retrain it every time your app starts.

import joblib
from pydantic import BaseModel

model = joblib.load("house_price_model.pkl")

class HouseFeatures(BaseModel):
    area_sqft: float
    bedrooms: int
    location_score: float

@app.post("/predict")
def predict(data: HouseFeatures):
    x = [[data.area_sqft, data.bedrooms, data.location_score]]
    price = model.predict(x)[0]
    return {"predicted_price": round(price, 2)}

Let’s break this down piece by piece. First, we import joblib and Pydantic’s BaseModel. Then, we load our previously saved model with joblib.load. After that, we define HouseFeatures again, matching the three fields our model expects.

Next comes the interesting part. The decorator @app.post("/predict") creates a new endpoint that accepts POST requests. The function predict takes a single argument, data, typed as HouseFeatures. Because of that type hint, FastAPI automatically validates every incoming request against our Pydantic model before this function even runs.

Inside the function, we build a small list of values, feed it to our model’s predict method, and return the result as a rounded number. Notice something important here: we never wrote a single line checking whether area_sqft was actually a number. Pydantic already handled that for us, upstream of our own code.

This is the real payoff of building ML APIs with FastAPI: less validation code, and more time spent on the logic that actually matters.

Documentation You Never Had to Write

Here’s a feature that tends to surprise people the first time they see it.

Visit /docs in your browser, right after starting your server, and FastAPI shows you a live, interactive documentation page. It lists every endpoint you’ve built, generated straight from your code and your type hints. You never wrote a single word of that documentation, and it never goes out of date, since it updates automatically whenever your code changes.

On that page, you’ll see both routes we built. There’s a GET request at /, a simple health check confirming the API is running. There’s also a POST request at /predict, where you can send house features and receive a price prediction back. Better still, you can test both endpoints directly from that page, without writing any separate testing code at all.

For teams, this single feature saves hours. New teammates can explore your API and understand exactly how to call it, without ever needing to ask you directly.

Async: Why It Feels Fast

Let’s address one more concept, and don’t worry, there’s no heavy math involved here.

A normal function finishes one task completely before it starts the next one. An async function works differently. It can pause a slow task, like a database call, and let other requests move forward while it waits. In short, there’s no new math required, just a smarter queue underneath the hood.

Picture two lanes side by side. Without async, it’s one line and one counter. Five requests wait their turn, one after another, and each one blocks the next. With async, however, it’s many counters, all moving simultaneously. The same five requests get handled together, and nobody gets stuck waiting behind a slow one.

This is a major reason teams choose FastAPI for real production traffic, rather than just quick prototypes.

Why This Matters for Production

Let’s step back and connect these features to real business outcomes.

Reliability comes first. Bad input gets caught and rejected before it ever reaches your model, and long before it can crash your service. That single behavior prevents a whole category of production incidents.

Scale comes next. Async handling means your API can serve far more users at once, without falling over under moderate traffic. That matters the moment your project moves beyond a personal demo.

Team speed matters too. Auto-generated documentation means any teammate can call your API correctly, without ever needing to ask you how it works. That alone removes a common bottleneck in growing teams.

Together, these three benefits explain why building ML APIs with FastAPI has become such a popular choice among data scientists moving models into production for the first time.

Three Mistakes First-Timers Make

Even with all this automation, a few common mistakes still trip up beginners. Let’s cover them, along with their fixes.

The first mistake is placing blocking calls inside an async def function. Using slow, non-async code inside an async endpoint stalls every other request waiting behind it. The fix is simple: use a regular def endpoint instead, or reach for a proper async library when you truly need one.

The second mistake is skipping response_model. Without it, extra or sensitive fields can leak straight into your API response without your knowledge. The fix is to declare a response_model, so only the fields you actually intend to expose ever leave your server.

The third mistake is assuming validation guarantees correctness. Pydantic checks types and shape, but it doesn’t know whether a value makes real-world sense. A negative house price, for instance, still passes as a valid float. The fix is to add your own range checks for anything with genuine real-world limits, such as age or price.

Avoiding these three traps will save you real debugging time later.

Recap: What You Now Know

Let’s tie everything together.

FastAPI validates incoming data automatically, using Pydantic and Python type hints. A rejected request never reaches your model, since bad data gets caught right at the door. Every API you build gets free, live documentation at /docs, generated without any extra effort. Async lets your API serve many requests at once, and it does this with zero extra math required from you. Finally, you built a genuine, working prediction endpoint today, reusing the joblib model from EP97.

In short, you moved from a bare Flask endpoint to a self-documenting, validated, async-ready API, and you did it in a single episode. That’s the real value of building ML APIs with FastAPI: less boilerplate, fewer bugs, and a service that documents itself as you go.

Watch the Full Walkthrough

This article covers the core ideas, but the video walks through every line of code on screen, step by step, along with the live /docs page in action. If you learn better by watching someone build it in real time, the EP99 video on the Intelevo YouTube channel is the perfect next stop.

What’s Next: EP100

Your API now works beautifully on your own laptop. However, right now, only you can reach it. In EP100, Deploying ML Models to the Cloud, we take this exact API and put it online, so anyone, anywhere, can call your model. No more localhost. Just a real, working link you can share with the world.

If this article helped you, consider watching the full EP99 video for the live coding walkthrough, and subscribe to Intelevo so you don’t miss EP100. Building ML APIs with FastAPI is just the beginning; deploying them is where the real fun starts.

Leave a Comment

Your email address will not be published. Required fields are marked *