Type something to search...
Data Validation in Python with Pydantic

Data Validation in Python with Pydantic

Every program eventually reads data it didn't create: JSON from an API, a request body, a config file, a CSV row, environment variables. That data is just strings and dictionaries until something checks it. Writing those checks by hand (is age present, is it an integer, is it positive, is email actually an email?) is tedious, easy to get wrong, and tends to get scattered across a codebase.

Pydantic turns that into a declaration. You describe the shape of your data with ordinary Python type hints, and Pydantic parses incoming data into that shape, converting what it safely can and raising a detailed error for everything else. It's the validation layer behind FastAPI and many other libraries, and it's just as useful on its own.

This guide covers Pydantic v2: models and type coercion, strict mode, field constraints, custom validators, serialization, aliases, discriminated unions, validating things that aren't models, and loading settings from the environment. The examples were run with Pydantic 2.13 on Python 3.13.

Installing Pydantic

python -m pip install pydantic

# Optional extras used later in this post
python -m pip install "pydantic[email]" pydantic-settings

The core validation engine, pydantic-core, is written in Rust, which is why v2 is dramatically faster than v1. If you're coming from v1, note that many method names changed (dict() became model_dump(), parse_obj() became model_validate(), @validator became @field_validator), and the old ones are deprecated.

Your First Model

A Pydantic model is a class that inherits from BaseModel, with fields declared as annotated class attributes:

from datetime import datetime

from pydantic import BaseModel, ValidationError


class User(BaseModel):
    id: int
    name: str
    email: str
    signed_up: datetime
    tags: list[str] = []
    nickname: str | None = None


user = User(id="42", name="Ada", email="ada@example.com", signed_up="2026-09-30T10:00:00Z")
print(user)
print(user.id, type(user.id).__name__)
print(user.signed_up)
id=42 name='Ada' email='ada@example.com' signed_up=datetime.datetime(2026, 9, 30, 10, 0, tzinfo=TzInfo(0)) tags=[] nickname=None
42 int
2026-09-30 10:00:00+00:00

Several things happened:

  • "42" became the integer 42, and the ISO 8601 string became a timezone-aware datetime. Pydantic parses data into the declared types, not just checks them.
  • Fields with defaults (tags, nickname) are optional. Fields without defaults are required.
  • The mutable default [] is safe here. Unlike a plain function default or a dataclass field, Pydantic copies defaults per instance.
  • str | None = None is the standard way to declare a field that may be missing or null.

When data doesn't fit, you get a ValidationError that reports every problem at once, not just the first:

try:
    User(id="forty-two", name="Ada", email="ada@example.com")
except ValidationError as exc:
    print(exc)
2 validation errors for User
id
  Input should be a valid integer, unable to parse string as an integer [type=int_parsing, input_value='forty-two', input_type=str]
    For further information visit https://errors.pydantic.dev/2.13/v/int_parsing
signed_up
  Field required [type=missing, input_value={'id': 'forty-two', 'name...ail': 'ada@example.com'}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing

For programmatic use (returning errors from an API, say), call exc.errors(), which gives a list of dictionaries with type, loc (the path to the bad field), msg, and input.

Three Ways In

You'll create models from three kinds of input:

User(id=1, name="Ada", ...)                      # keyword arguments
User.model_validate({"id": 1, "name": "Ada"})    # a dict (or other object)
User.model_validate_json('{"id": 1, ...}')       # a JSON string or bytes

Prefer model_validate_json() when you have raw JSON. It parses and validates in one pass inside the Rust core, which is faster than json.loads() followed by model_validate().

Coercion vs Strict Mode

By default Pydantic runs in lax mode: it accepts input that can be converted to the target type without losing information.

from pydantic import BaseModel, ValidationError


class Item(BaseModel):
    count: int
    price: float
    active: bool


print(Item(count="3", price="9.99", active="yes"))
print(Item(count=3.0, price=10, active=1))
count=3 price=9.99 active=True
count=3 price=10.0 active=True

"Without losing information" is the important part. Lax mode isn't the same as calling int() on everything:

for bad in [{"count": 3.5, "price": 1, "active": True}, {"count": 1, "price": 1, "active": "maybe"}]:
    try:
        Item(**bad)
    except ValidationError as e:
        for err in e.errors():
            print(err["loc"], err["type"], "-", err["msg"])
('count',) int_from_float - Input should be a valid integer, got a number with a fractional part
('active',) bool_parsing - Input should be a valid boolean, unable to interpret input

int(3.5) would silently truncate to 3; Pydantic refuses. Strings like "yes", "true", "1", and "on" are accepted for booleans, but "maybe" isn't.

Lax mode is ideal for data from forms, query strings, and environment variables, where everything arrives as a string. When you want exact types, you have three levels of strictness:

from typing import Annotated

from pydantic import BaseModel, Strict


class StrictItem(BaseModel, strict=True):  # whole model
    count: int


class Mixed(BaseModel):
    quantity: Annotated[int, Strict()]  # one field
    note: str


StrictItem(count="3")  # ValidationError: Input should be a valid integer
Item.model_validate({"count": "3", "price": 1, "active": True}, strict=True)  # one call

Strict mode is a little more forgiving with JSON input: StrictItem.model_validate_json('{"count": 3}') works as you'd expect, and types with no JSON equivalent, such as datetime or UUID, are still accepted from their standard string forms when parsing JSON.

Constraining Fields

Types alone don't say "positive integer" or "at most 500 characters". Field() adds constraints, and Annotated lets you package a type and its constraints into a reusable alias:

from decimal import Decimal
from typing import Annotated, Literal

from pydantic import BaseModel, EmailStr, Field, HttpUrl, ValidationError

Sku = Annotated[str, Field(pattern=r"^[A-Z]{3}-\d{4}$")]
Quantity = Annotated[int, Field(gt=0, le=1000)]


class LineItem(BaseModel):
    sku: Sku
    quantity: Quantity = 1
    unit_price: Annotated[Decimal, Field(ge=0, max_digits=10, decimal_places=2)]


class Order(BaseModel):
    customer_email: EmailStr
    status: Literal["pending", "paid", "shipped"] = "pending"
    items: list[LineItem] = Field(min_length=1)
    notes: str = Field(default="", max_length=500)
    website: HttpUrl | None = None
  • Numeric constraints: gt, ge, lt, le, multiple_of. Decimal fields also take max_digits and decimal_places, which matters for money.
  • String constraints: min_length, max_length, pattern.
  • Collection constraints: min_length and max_length on lists, sets and dicts.
  • Literal[...] restricts a value to a fixed set. An Enum works too, if you want the values as named constants.
  • EmailStr (needs the email extra) and HttpUrl are ready-made types for common formats.
  • Nested models just work: items is validated as a list of LineItem objects, dictionaries included.

Errors from nested data come with their full path:

try:
    Order(
        customer_email="not-an-email",
        status="lost",
        items=[{"sku": "mug-1", "quantity": 0, "unit_price": "1.999"}],
    )
except ValidationError as exc:
    for err in exc.errors():
        print(".".join(map(str, err["loc"])), "->", err["msg"])
customer_email -> value is not a valid email address: An email address must have an @-sign.
status -> Input should be 'pending', 'paid' or 'shipped'
items.0.sku -> String should match pattern '^[A-Z]{3}-\d{4}$'
items.0.quantity -> Input should be greater than 0
items.0.unit_price -> Decimal input should have no more than 2 decimal places

items.0.sku tells you exactly which list element failed. That's very helpful when you're returning errors to an API client who sent a large payload.

Custom Validators

When built-in constraints aren't enough, write a validator.

Field Validators

@field_validator runs custom logic for one or more fields. Raise ValueError to reject a value; return a value to accept (and possibly transform) it:

from datetime import date
from typing import Self

from pydantic import BaseModel, ValidationError, field_validator, model_validator


class Booking(BaseModel):
    guest: str
    check_in: date
    check_out: date
    guests: int = 1

    @field_validator("guest")
    @classmethod
    def normalize_guest(cls, value: str) -> str:
        value = " ".join(value.split())
        if not value:
            raise ValueError("guest name cannot be blank")
        return value.title()

    @model_validator(mode="after")
    def check_dates(self) -> Self:
        if self.check_out <= self.check_in:
            raise ValueError("check_out must be after check_in")
        return self


print(Booking(guest="  ada   lovelace ", check_in="2026-10-01", check_out="2026-10-04"))
guest='Ada Lovelace' check_in=datetime.date(2026, 10, 1) check_out=datetime.date(2026, 10, 4) guests=1

By default, a field validator is an "after" validator: it runs once Pydantic has already converted the input to str, so your function can rely on the type.

Model Validators

Rules that involve several fields belong in a @model_validator. With mode="after", it receives the fully validated instance (self), so you can compare fields with their real types. If it fails, the error is reported at the model level:

try:
    Booking(guest="Ada", check_in="2026-10-04", check_out="2026-10-01")
except ValidationError as exc:
    print(exc.errors()[0]["loc"], exc.errors()[0]["msg"])
() Value error, check_out must be after check_in

The empty loc tuple means "the whole model". A mode="before" model validator receives the raw input instead, which is useful for reshaping data (renaming legacy keys, unwrapping an envelope) before field validation runs.

Reusable Validators with Annotated

Validators attached to a model only help that model. To reuse logic, attach BeforeValidator and AfterValidator to a type alias:

from typing import Annotated, Any

from pydantic import AfterValidator, BaseModel, BeforeValidator


def split_csv(value: Any) -> Any:
    if isinstance(value, str):
        return [part.strip() for part in value.split(",") if part.strip()]
    return value


def dedupe(values: list[str]) -> list[str]:
    return sorted(set(v.lower() for v in values))


Tags = Annotated[list[str], BeforeValidator(split_csv), AfterValidator(dedupe)]


class Article(BaseModel):
    title: str
    tags: Tags = []


print(Article(title="Hello", tags="Python, pydantic,python"))
print(Article(title="Hello", tags=["B", "a", "b"]))
title='Hello' tags=['pydantic', 'python']
title='Hello' tags=['a', 'b']

The before validator turns a comma-separated string into a list, Pydantic validates it as list[str], and the after validator normalizes the result. Any model can now use Tags and get the same behavior.

Model Configuration

model_config controls model-wide behavior. Three settings are worth knowing early:

from pydantic import BaseModel, ConfigDict, ValidationError


class SignupRequest(BaseModel):
    model_config = ConfigDict(extra="forbid", str_strip_whitespace=True, frozen=True)

    username: str
    password: str


req = SignupRequest(username="  ada ", password="s3cret-pass")
print(repr(req.username))  # 'ada'

SignupRequest(username="ada", password="x", is_admin=True)
# ValidationError: extra_forbidden at ('is_admin',)

req.username = "mallory"
# ValidationError: frozen_instance
  • extra: by default, unknown fields are silently ignored ("ignore"). "forbid" rejects them, which is a good default for request bodies because it catches typos and attempts to set fields like is_admin. "allow" keeps them.
  • str_strip_whitespace: trims every string field.
  • frozen: makes instances immutable and hashable. To get a modified copy, use req.model_copy(update={"username": "grace"}), but note that model_copy doesn't re-run validation on the updated values.

Serialization

Validation is half the job. The other half is turning models back into dictionaries and JSON:

from datetime import datetime, timezone

from pydantic import BaseModel, ConfigDict, Field, computed_field


class Product(BaseModel):
    model_config = ConfigDict(validate_by_name=True)

    product_id: int = Field(alias="productId")
    name: str
    price_cents: int = Field(alias="priceCents")
    created_at: datetime = Field(default_factory=lambda: datetime(2026, 9, 30, tzinfo=timezone.utc))
    internal_note: str | None = Field(default=None, exclude=True)

    @computed_field
    @property
    def price(self) -> float:
        return self.price_cents / 100


p = Product.model_validate({"productId": 7, "name": "Mug", "priceCents": 1250, "internal_note": "fragile"})
print(p.model_dump())
print(p.model_dump(by_alias=True, exclude={"created_at"}))
print(p.model_dump(mode="json"))
print(p.model_dump_json(indent=2))
{'product_id': 7, 'name': 'Mug', 'price_cents': 1250, 'created_at': datetime.datetime(2026, 9, 30, 0, 0, tzinfo=datetime.timezone.utc), 'price': 12.5}
{'productId': 7, 'name': 'Mug', 'priceCents': 1250, 'price': 12.5}
{'product_id': 7, 'name': 'Mug', 'price_cents': 1250, 'created_at': '2026-09-30T00:00:00Z', 'price': 12.5}
{
  "product_id": 7,
  "name": "Mug",
  "price_cents": 1250,
  "created_at": "2026-09-30T00:00:00Z",
  "price": 12.5
}

What each piece does:

  • Aliases map external names to Python names. The input used productId and priceCents (camelCase, as a JavaScript client would send), while your code uses snake_case attributes. validate_by_name=True additionally accepts the Python names, so Product(product_id=8, name="Lamp", price_cents=6500) also works. Pass by_alias=True when dumping to send camelCase back out.
  • model_dump() returns Python objects (note the datetime). model_dump(mode="json") returns only JSON-compatible types, and model_dump_json() goes straight to a JSON string.
  • exclude=True on a field keeps it out of every dump, which is handy for internal or sensitive data. You can also pass include/exclude sets per call.
  • @computed_field adds a property to the serialized output. It's computed, not stored, and it also appears in the serialization-mode JSON schema.

Every model can also describe itself as JSON Schema with Model.model_json_schema(). That's how FastAPI produces its OpenAPI documentation.

Unions and Discriminated Unions

Real payloads often come in several shapes. A payment can be a card, a bank transfer, or PayPal, each with different fields. A plain union (CardPayment | BankTransfer | PayPal) makes Pydantic try each type, which is slower and produces confusing errors when nothing matches. If the shapes share a tag field, use a discriminated union:

from typing import Annotated, Literal

from pydantic import BaseModel, Field, ValidationError


class CardPayment(BaseModel):
    method: Literal["card"]
    last4: str = Field(pattern=r"^\d{4}$")


class BankTransfer(BaseModel):
    method: Literal["bank"]
    iban: str


class PayPal(BaseModel):
    method: Literal["paypal"]
    email: str


Payment = Annotated[CardPayment | BankTransfer | PayPal, Field(discriminator="method")]


class Checkout(BaseModel):
    order_id: int
    payment: Payment


c = Checkout.model_validate_json(
    '{"order_id": 1, "payment": {"method": "bank", "iban": "DE89370400440532013000"}}'
)
print(type(c.payment).__name__)  # BankTransfer

Pydantic reads method first and validates against exactly one model. Errors point at the right place:

('payment', 'card', 'last4') String should match pattern '^\d{4}$'
('payment',) Input tag 'cash' found using 'method' does not match any of the expected tags: 'card', 'bank', 'paypal'

The first error comes from a "card" payload with "last4": "12"; the second from a payload with "method": "cash". This pattern pairs naturally with structural pattern matching when you process the result.

Validating Without a Model

Sometimes the data isn't naturally a model: a list of events from an API, a dict of port numbers, a function's arguments. TypeAdapter validates against any type:

from datetime import date

from pydantic import BaseModel, TypeAdapter


class Event(BaseModel):
    name: str
    on: date


events_adapter = TypeAdapter(list[Event])
raw = '[{"name": "PyCon", "on": "2026-05-13"}, {"name": "Launch", "on": "2026-11-02"}]'
events = events_adapter.validate_json(raw)
print(events[1])
print(events_adapter.dump_json(events))

ports = TypeAdapter(dict[str, int]).validate_python({"http": "80", "https": 443})
print(ports)
name='Launch' on=datetime.date(2026, 11, 2)
b'[{"name":"PyCon","on":"2026-05-13"},{"name":"Launch","on":"2026-11-02"}]'
{'http': 80, 'https': 443}

Creating an adapter has a cost, so build it once at module level and reuse it, rather than inside a function that runs per request.

For functions, @validate_call validates arguments against the signature's type hints:

from pydantic import ValidationError, validate_call


@validate_call
def repeat(text: str, times: int = 2) -> str:
    return text * times


print(repeat("ab", "3"))  # ababab
repeat("ab", times="many")  # ValidationError at ('times',): int_parsing

It's useful at boundaries, such as functions called from a CLI or plugin system, but don't sprinkle it on every function; static type checking is the cheaper tool for code you control.

If you prefer dataclass syntax, pydantic.dataclasses.dataclass gives you a standard-looking dataclass with Pydantic validation.

Settings from the Environment

Configuration is untrusted input too: environment variables are all strings, and a typo in a variable name shouldn't surface as a crash an hour into production. The separate pydantic-settings package reads settings from the environment and .env files, with the same validation:

# settings.py
from pydantic import Field, PostgresDsn, SecretStr
from pydantic_settings import BaseSettings, SettingsConfigDict


class Settings(BaseSettings):
    model_config = SettingsConfigDict(env_prefix="APP_", env_file=".env")

    debug: bool = False
    database_url: PostgresDsn
    api_key: SecretStr
    allowed_hosts: list[str] = Field(default_factory=lambda: ["localhost"])
    workers: int = 2


settings = Settings()
print(settings.debug, settings.workers, settings.allowed_hosts)
print(settings.database_url)
print(settings.api_key)
print(settings.api_key.get_secret_value())
export APP_DEBUG=1 APP_WORKERS=4 APP_API_KEY=sk-123
export APP_DATABASE_URL=postgresql://app:pw@db:5432/shop
export APP_ALLOWED_HOSTS='["example.com","www.example.com"]'
python settings.py
True 4 ['example.com', 'www.example.com']
postgresql://app:pw@db:5432/shop
**********
sk-123
  • env_prefix="APP_" maps APP_DATABASE_URL to database_url (matching is case-insensitive by default).
  • Complex types like list[str] are read from the environment as JSON.
  • SecretStr hides its value in repr(), print() and logs; you must call get_secret_value() deliberately.
  • If APP_API_KEY is missing or APP_DATABASE_URL isn't a valid URL, Settings() raises a ValidationError immediately at startup, listing every problem.

Create a single settings instance at import time and import it wherever it's needed, so the app fails fast on bad configuration.

Where Pydantic Fits

Pydantic is most valuable at the edges of your program, where untrusted data enters:

  • HTTP request bodies and query parameters (FastAPI does this for you; see building a REST API with FastAPI)
  • Responses from third-party APIs
  • Config files and environment variables
  • Messages from queues, files from users, LLM output that should match a schema

Inside your program, once data has been validated, plain dataclasses or Pydantic models both work. Validating the same data over and over in internal code adds overhead without adding safety.

Conclusion

Pydantic lets you replace scattered if checks with a type-annotated declaration of what valid data looks like. Models coerce data sensibly in lax mode or exactly in strict mode, Field and Annotated add constraints, validators handle custom rules for single fields or whole models, and model_dump() and aliases handle the trip back out. Add discriminated unions for multi-shape payloads, TypeAdapter for anything that isn't a model, and pydantic-settings for configuration, and you have one consistent approach to every piece of data that crosses into your program.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading