Type something to search...
Data Visualization in Python with Matplotlib and Seaborn

Data Visualization in Python with Matplotlib and Seaborn

Most Python plotting tutorials start with plt.plot(x, y) and plt.show(), and that works right up until you need two charts side by side, a legend that doesn't cover your data, or a figure you can drop into a report without fiddling. At that point the shortcuts stop helping and you need to understand how Matplotlib actually thinks about a chart.

Seaborn sits on top of Matplotlib and handles the statistical side: grouping by category, coloring by a column, drawing distributions and confidence intervals. It saves a lot of code, but every Seaborn chart is still a Matplotlib figure underneath, so the two skills go together.

This guide covers Matplotlib's figure and axes model, the core chart types, customization that matters, saving figures, and then Seaborn's axes-level and figure-level functions, with a clear picture of when to use which.

Installing the Libraries

Install both into a virtual environment. Seaborn pulls in pandas and NumPy, which you'll want anyway.

python -m pip install matplotlib seaborn

If you're new to installing packages, see what pip is and how to use it. The examples here were run with Matplotlib 3.11 and Seaborn 0.13, but everything shown uses long-stable APIs.

Matplotlib's Mental Model: Figure and Axes

A Matplotlib chart has two main objects:

  • Figure: the whole canvas. It holds one or more plots, a title, and controls size and resolution.
  • Axes: a single plot area with its own x-axis, y-axis, title, and data. Despite the name, an Axes is a plot, not an axis.

There are two ways to drive Matplotlib. The pyplot style (plt.plot, plt.title) acts on whatever the "current" axes happens to be. The object-oriented style creates the figure and axes explicitly and calls methods on them. The official docs recommend the object-oriented style for anything beyond a quick look, and it's the style used throughout this post because it stays predictable once you have more than one plot.

# line_chart.py
import matplotlib.pyplot as plt

months = ["Jan", "Feb", "Mar", "Apr", "May", "Jun"]
revenue = [12.4, 13.1, 15.8, 14.9, 17.2, 19.5]
costs = [9.8, 10.2, 11.0, 11.4, 12.1, 12.9]

fig, ax = plt.subplots(figsize=(8, 4.5), layout="constrained")
ax.plot(months, revenue, marker="o", label="Revenue")
ax.plot(months, costs, marker="s", linestyle="--", label="Costs")
ax.set_title("Revenue vs costs, H1 2026")
ax.set_xlabel("Month")
ax.set_ylabel("USD (thousands)")
ax.legend()
ax.grid(alpha=0.3)

fig.savefig("revenue.png", dpi=150)

plt.subplots() returns a Figure and an Axes. Everything about the plot itself goes through ax: ax.plot draws the lines, the set_* methods label things, and ax.legend() builds a legend from each line's label. figsize is in inches, and dpi on save controls pixel density, so (8, 4.5) at 150 dpi gives a 1200 by 675 pixel image.

layout="constrained" turns on Matplotlib's constrained layout engine, which adjusts spacing so titles, labels, and legends don't overlap or get clipped. It's the modern replacement for calling fig.tight_layout() at the end, and it's worth using by default.

Showing vs Saving

In a script, plt.show() opens a window and blocks until you close it. In Jupyter, figures render inline automatically. For anything automated (a report, a cron job, a server), skip show() and call fig.savefig(). On a headless machine you can force the non-interactive backend before importing pyplot:

import matplotlib

matplotlib.use("Agg")
import matplotlib.pyplot as plt

When a script generates many figures in a loop, close them with plt.close(fig) (or plt.close("all")). pyplot keeps references to every figure it created, and they'll pile up in memory otherwise.

The Core Chart Types

Most charts you'll need are a single method on an Axes. Here are four of the most common in one figure, which also shows how to work with a grid of subplots.

# chart_types.py
import matplotlib.pyplot as plt
import numpy as np

months = ["Jan", "Feb", "Mar", "Apr", "May", "Jun"]
revenue = [12.4, 13.1, 15.8, 14.9, 17.2, 19.5]
rng = np.random.default_rng(42)

fig, axes = plt.subplots(2, 2, figsize=(10, 7), layout="constrained")

axes[0, 0].bar(months, revenue, color="tab:blue")
axes[0, 0].set_title("Bar")

axes[0, 1].scatter(rng.normal(size=200), rng.normal(size=200), alpha=0.6, s=15)
axes[0, 1].set_title("Scatter")

axes[1, 0].hist(rng.normal(loc=50, scale=10, size=1000), bins=30, edgecolor="white")
axes[1, 0].set_title("Histogram")

axes[1, 1].barh(["A", "B", "C"], [5, 9, 3], color="tab:orange")
axes[1, 1].set_title("Horizontal bar")

fig.suptitle("Four common chart types")
fig.savefig("grid.png", dpi=150)

With plt.subplots(2, 2), axes is a 2 by 2 NumPy array of Axes objects, so you index it as axes[row, col]. With a single row or column (plt.subplots(1, 3)) it's a 1-D array, and with no arguments it's a single Axes. fig.suptitle() puts a title over the whole figure, while each set_title() labels one subplot.

A quick reference for picking a method:

You want to showMethodNotes
A trend over time or orderax.plotAdd marker= for few points
Comparison across categoriesax.bar / ax.barhUse barh for long labels
Relationship between two numbersax.scatteralpha helps with overlap
Distribution of one variableax.histTry a few bins values
Spread and outliers by groupax.boxplotSeaborn's version is easier
Part of a wholeax.pieA sorted bar chart is usually clearer
A 2-D grid of valuesax.imshowSeaborn's heatmap adds labels

s in scatter is marker area in points squared, and alpha (0 to 1) sets transparency. Both matter more than people expect: a scatter of 10,000 opaque points is just a blob.

Customizing Charts

Defaults get you a readable chart. A few targeted changes make it a good one.

Annotations, Spines, and Tick Formatting

# customize.py
import matplotlib.pyplot as plt

months = ["Jan", "Feb", "Mar", "Apr", "May", "Jun"]
revenue = [12.4, 13.1, 15.8, 14.9, 17.2, 19.5]

fig, ax = plt.subplots(layout="constrained")
ax.plot(months, revenue)

best = max(range(len(revenue)), key=revenue.__getitem__)
ax.annotate(
    f"Peak: {revenue[best]}k",
    xy=(best, revenue[best]),
    xytext=(best - 2, revenue[best] - 1),
    arrowprops={"arrowstyle": "->"},
)

ax.spines[["top", "right"]].set_visible(False)
ax.yaxis.set_major_formatter("${x:,.0f}k")

fig.savefig("annotated.png")

annotate places text at xytext with an arrow pointing to xy. When the x-values are category strings, Matplotlib plots them at positions 0, 1, 2 and so on, which is why the index best works as an x coordinate.

ax.spines gives you the four border lines. Hiding the top and right ones is a small change that makes most charts look less cluttered. set_major_formatter accepts a format string where x is the tick value, so "${x:,.0f}k" turns 15.0 into $15k. It uses the same mini-language as f-strings.

Two Y-Axes

When two series share an x-axis but have very different scales, twinx() creates a second Axes that shares the x-axis and gets its own y-axis on the right:

import matplotlib.pyplot as plt

months = ["Jan", "Feb", "Mar", "Apr", "May", "Jun"]
revenue = [12.4, 13.1, 15.8, 14.9, 17.2, 19.5]
signups = [3, 4, 2, 5, 6, 7]

fig, ax1 = plt.subplots(layout="constrained")
ax2 = ax1.twinx()

ax1.bar(months, revenue, alpha=0.5, label="Revenue")
ax2.plot(months, signups, color="tab:red", marker="o", label="Signups")
ax1.set_ylabel("Revenue (k)")
ax2.set_ylabel("Signups")

fig.savefig("twin.png")

Use this sparingly. Dual axes make it easy to imply a correlation just by choosing the scales, so label both axes clearly.

Styles

Matplotlib ships with style sheets that change colors, fonts, and grids in one go. plt.style.available lists them. Apply one globally with plt.style.use("ggplot"), or temporarily with a context manager:

import matplotlib.pyplot as plt

with plt.style.context("ggplot"):
    fig, ax = plt.subplots()
    ax.plot([1, 3, 2, 4])
    fig.savefig("styled.png")

Saving for Different Destinations

The file extension picks the format:

fig.savefig("chart.png", dpi=200)          # raster, for web and slides
fig.savefig("chart.svg")                   # vector, for web pages that scale
fig.savefig("chart.pdf")                   # vector, for print and LaTeX
fig.savefig("chart.png", transparent=True) # no background fill

Use PNG for screenshots and dashboards, SVG or PDF when the chart needs to stay sharp at any size. If parts of a label get cut off and you aren't using constrained layout, bbox_inches="tight" trims the saved image to its contents.

Seaborn: Statistical Plots from DataFrames

Matplotlib works with arrays. Seaborn works with tidy DataFrames: one row per observation, one column per variable. You name columns instead of passing arrays, and Seaborn handles the grouping, coloring, and legends.

The examples below use a synthetic usage dataset so they run without a network connection. (Seaborn's sns.load_dataset() is handy for practice, but it downloads data from GitHub on first use.)

# data.py
import numpy as np
import pandas as pd

rng = np.random.default_rng(7)
n = 300
df = pd.DataFrame(
    {
        "plan": rng.choice(["free", "pro", "team"], size=n, p=[0.5, 0.3, 0.2]),
        "region": rng.choice(["EU", "US"], size=n),
        "sessions": rng.poisson(12, size=n),
    }
)
df["minutes"] = (df["sessions"] * rng.uniform(2, 6, size=n)).round(1)
df.loc[df["plan"] == "team", "minutes"] *= 1.5

if __name__ == "__main__":
    print(df.groupby("plan")["minutes"].mean().round(1))

Running python data.py prints the average minutes per plan. The other scripts import df from this module, and the __main__ guard keeps them from printing it too.

plan
free    47.7
pro     50.3
team    72.8
Name: minutes, dtype: float64

If you need a refresher on building and grouping DataFrames like this, see using Python for data analysis.

Setting a Theme

sns.set_theme() applies Seaborn's styling to all Matplotlib figures in the session, including ones you draw with plain Matplotlib:

import seaborn as sns

sns.set_theme(style="whitegrid", palette="deep")

Styles include "darkgrid" (the default), "whitegrid", "dark", "white", and "ticks". You can also pass context="talk" or context="paper" to scale fonts and line widths for slides or print.

Axes-Level Functions

Most Seaborn functions draw onto a single Matplotlib Axes and accept an ax= argument. That means they slot straight into the figure and subplot code you already know.

# seaborn_scatter.py
import matplotlib.pyplot as plt
import seaborn as sns

from data import df

sns.set_theme(style="whitegrid")

fig, ax = plt.subplots(figsize=(7, 4), layout="constrained")
sns.scatterplot(data=df, x="sessions", y="minutes", hue="plan", style="region", ax=ax)
ax.set_title("Sessions vs minutes")
fig.savefig("seaborn_scatter.png", dpi=150)

Two lines of Seaborn did what would take a loop in Matplotlib: hue="plan" colors points by plan, style="region" changes marker shape by region, and the legend is built for you. After that, ax is a normal Axes, so ax.set_title() and every customization from earlier still apply.

Distribution and comparison charts are where Seaborn saves the most work:

# seaborn_distributions.py
import matplotlib.pyplot as plt
import seaborn as sns

from data import df

sns.set_theme(style="whitegrid")
fig, axes = plt.subplots(1, 3, figsize=(13, 4), layout="constrained")

sns.histplot(data=df, x="minutes", hue="plan", kde=True, element="step", ax=axes[0])
sns.boxplot(data=df, x="plan", y="minutes", order=["free", "pro", "team"], ax=axes[1])
sns.barplot(data=df, x="plan", y="minutes", hue="region", errorbar="sd", ax=axes[2])

fig.savefig("distributions.png")
  • histplot draws a histogram per plan, and kde=True overlays a smoothed density curve. element="step" draws outlines instead of filled bars so overlapping groups stay readable.
  • boxplot shows median, quartiles, and outliers per group. order= fixes the category order instead of using the order of first appearance.
  • barplot doesn't plot raw values. It aggregates (mean by default) and draws an error bar. errorbar="sd" shows one standard deviation; the default is a 95% bootstrapped confidence interval.

That last point trips people up: if your data is already aggregated, one row per bar, barplot still works, but the error bars disappear because there's nothing to estimate.

Here's a quick map of the most useful axes-level functions:

PurposeFunctions
Relationshipsscatterplot, lineplot, regplot
Distributionshistplot, kdeplot, ecdfplot
Categorical comparisonsboxplot, violinplot, stripplot, barplot, countplot
Matricesheatmap

lineplot deserves a mention: if your data has several y-values per x-value, it plots the mean and shades a confidence band automatically, which is exactly what you want for repeated measurements.

Figure-Level Functions and Faceting

Figure-level functions create their own figure and can split the data into a grid of subplots by column values. This is called faceting, and it's often clearer than cramming every group into one chart.

# seaborn_facets.py
import seaborn as sns

from data import df

sns.set_theme(style="whitegrid")

g = sns.relplot(
    data=df, x="sessions", y="minutes", hue="plan", col="region", kind="scatter", height=4
)
g.set_axis_labels("Sessions", "Minutes")
g.figure.suptitle("Usage by region", y=1.03)
g.savefig("by_region.png")

g = sns.catplot(data=df, x="plan", y="minutes", col="region", kind="violin", height=4)
g.savefig("violins.png")

col="region" makes one subplot per region, side by side; row= stacks them vertically, and you can use both. The return value is a FacetGrid, not an Axes, so you customize it through its own methods (set_axis_labels, set_titles) or reach into g.figure and g.axes for Matplotlib-level control.

The three figure-level functions map onto the axes-level families:

Figure-levelkind= options
relplot"scatter", "line"
displot"hist", "kde", "ecdf"
catplot"strip", "swarm", "box", "violin", "bar", "point", "count"

The rule of thumb: use an axes-level function when you're placing a chart into a layout you control, and a figure-level function when you want Seaborn to build the layout from your data. Don't pass ax= to a figure-level function; it creates its own figure.

Heatmaps and Pair Plots

A heatmap is a good fit for a pivot table or correlation matrix:

# seaborn_heatmap.py
import matplotlib.pyplot as plt
import seaborn as sns

from data import df

pivot = df.pivot_table(index="plan", columns="region", values="minutes", aggfunc="mean")

fig, ax = plt.subplots(figsize=(5, 4), layout="constrained")
sns.heatmap(pivot, annot=True, fmt=".1f", cmap="viridis", ax=ax)
ax.set_title("Mean minutes by plan and region")
fig.savefig("heatmap.png")

annot=True writes each value in its cell and fmt controls the number format. Seaborn uses the DataFrame's index and columns as tick labels, which is the main advantage over ax.imshow.

For a quick look at how several numeric columns relate, pairplot draws a scatter for every pair and a distribution on the diagonal:

import seaborn as sns

from data import df

g = sns.pairplot(df, hue="plan", vars=["sessions", "minutes"])
g.savefig("pairs.png")

It gets slow and unreadable with many columns, so pass vars= to choose a handful.

Plotting Straight from pandas

pandas has a .plot accessor that calls Matplotlib for you and returns an Axes. It's the fastest route from a groupby to a chart:

from data import df

ax = df.groupby("plan")["minutes"].mean().plot(kind="bar", rot=0, title="Mean minutes by plan")
ax.set_ylabel("Minutes")
ax.figure.savefig("pandas_bar.png")

Since you get an Axes back, you can keep customizing with the same methods. For exploratory work, pandas plotting and Seaborn cover most needs; drop to raw Matplotlib when you need precise control.

Choosing Between Matplotlib and Seaborn

You rarely choose one or the other. A practical split:

  • Reach for Seaborn when your data is in a DataFrame and you want to compare groups, show distributions, or facet by category. It handles colors, legends, and aggregation for you.
  • Reach for Matplotlib for layout (subplots, figure size, saving), fine-grained styling, annotations, unusual chart types, and anything that isn't a statistical summary.
  • Combine them by creating the figure with plt.subplots(), drawing with Seaborn using ax=, and finishing with Matplotlib methods on that same ax.

A few habits that consistently improve charts:

  • Label axes with units, and give each chart a title that states what it shows.
  • Start bar charts at zero. Line charts don't have to.
  • Sort categorical bars by value unless the categories have a natural order.
  • Use color to encode something. If every bar is a different color for no reason, it's noise.
  • Pick perceptually uniform colormaps like "viridis" for continuous values, and a diverging one like "vlag" when values are centered on zero.

Conclusion

Matplotlib's object-oriented API is the foundation: create a figure and axes with plt.subplots(), draw with methods on the axes, and save with fig.savefig(). Once that model clicks, Seaborn becomes a set of shortcuts that turn a tidy DataFrame into grouped, colored, statistically summarized charts in one call, while still handing you a normal Axes to polish.

Start with constrained layout and a Seaborn theme, use axes-level functions inside layouts you control, use figure-level functions when you want faceting, and close your figures when you generate them in bulk. That covers the large majority of real-world charting in Python.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading