
Data Validation in Python with Pydantic
Every program eventually reads data it didn't create: JSON from an API, a request body, a config file, a CSV row, environment variables. That data is just strings and dictionaries until something checks it. Writing those checks by hand (is age present, is it an integer, is it positive, is email actually an email?) is tedious, easy to get wrong, and tends to get scattered across a codebase.
Pydantic turns that into a declaration. You describe the shape of your data with ordinary Python type hints, and Pydantic parses incoming data into that shape, converting what it safely can and raising a detailed error for everything else. It's the validation layer behind FastAPI and many other libraries, and it's just as useful on its own.
This guide covers Pydantic v2: models and type coercion, strict mode, field constraints, custom validators, serialization, aliases, discriminated unions, validating things that aren't models, and loading settings from the environment. The examples were run with Pydantic 2.13 on Python 3.13.
Installing Pydantic
python -m pip install pydantic
# Optional extras used later in this post
python -m pip install "pydantic[email]" pydantic-settings
The core validation engine, pydantic-core, is written in Rust, which is why v2 is dramatically faster than v1. If you're coming from v1, note that many method names changed (dict() became model_dump(), parse_obj() became model_validate(), @validator became @field_validator), and the old ones are deprecated.
Your First Model
A Pydantic model is a class that inherits from BaseModel, with fields declared as annotated class attributes:
from datetime import datetime
from pydantic import BaseModel, ValidationError
class User(BaseModel):
id: int
name: str
email: str
signed_up: datetime
tags: list[str] = []
nickname: str | None = None
user = User(id="42", name="Ada", email="ada@example.com", signed_up="2026-09-30T10:00:00Z")
print(user)
print(user.id, type(user.id).__name__)
print(user.signed_up)
id=42 name='Ada' email='ada@example.com' signed_up=datetime.datetime(2026, 9, 30, 10, 0, tzinfo=TzInfo(0)) tags=[] nickname=None
42 int
2026-09-30 10:00:00+00:00
Several things happened:
"42"became the integer42, and the ISO 8601 string became a timezone-awaredatetime. Pydantic parses data into the declared types, not just checks them.- Fields with defaults (
tags,nickname) are optional. Fields without defaults are required. - The mutable default
[]is safe here. Unlike a plain function default or a dataclass field, Pydantic copies defaults per instance. str | None = Noneis the standard way to declare a field that may be missing or null.
When data doesn't fit, you get a ValidationError that reports every problem at once, not just the first:
try:
User(id="forty-two", name="Ada", email="ada@example.com")
except ValidationError as exc:
print(exc)
2 validation errors for User
id
Input should be a valid integer, unable to parse string as an integer [type=int_parsing, input_value='forty-two', input_type=str]
For further information visit https://errors.pydantic.dev/2.13/v/int_parsing
signed_up
Field required [type=missing, input_value={'id': 'forty-two', 'name...ail': 'ada@example.com'}, input_type=dict]
For further information visit https://errors.pydantic.dev/2.13/v/missing
For programmatic use (returning errors from an API, say), call exc.errors(), which gives a list of dictionaries with type, loc (the path to the bad field), msg, and input.
Three Ways In
You'll create models from three kinds of input:
User(id=1, name="Ada", ...) # keyword arguments
User.model_validate({"id": 1, "name": "Ada"}) # a dict (or other object)
User.model_validate_json('{"id": 1, ...}') # a JSON string or bytes
Prefer model_validate_json() when you have raw JSON. It parses and validates in one pass inside the Rust core, which is faster than json.loads() followed by model_validate().
Coercion vs Strict Mode
By default Pydantic runs in lax mode: it accepts input that can be converted to the target type without losing information.
from pydantic import BaseModel, ValidationError
class Item(BaseModel):
count: int
price: float
active: bool
print(Item(count="3", price="9.99", active="yes"))
print(Item(count=3.0, price=10, active=1))
count=3 price=9.99 active=True
count=3 price=10.0 active=True
"Without losing information" is the important part. Lax mode isn't the same as calling int() on everything:
for bad in [{"count": 3.5, "price": 1, "active": True}, {"count": 1, "price": 1, "active": "maybe"}]:
try:
Item(**bad)
except ValidationError as e:
for err in e.errors():
print(err["loc"], err["type"], "-", err["msg"])
('count',) int_from_float - Input should be a valid integer, got a number with a fractional part
('active',) bool_parsing - Input should be a valid boolean, unable to interpret input
int(3.5) would silently truncate to 3; Pydantic refuses. Strings like "yes", "true", "1", and "on" are accepted for booleans, but "maybe" isn't.
Lax mode is ideal for data from forms, query strings, and environment variables, where everything arrives as a string. When you want exact types, you have three levels of strictness:
from typing import Annotated
from pydantic import BaseModel, Strict
class StrictItem(BaseModel, strict=True): # whole model
count: int
class Mixed(BaseModel):
quantity: Annotated[int, Strict()] # one field
note: str
StrictItem(count="3") # ValidationError: Input should be a valid integer
Item.model_validate({"count": "3", "price": 1, "active": True}, strict=True) # one call
Strict mode is a little more forgiving with JSON input: StrictItem.model_validate_json('{"count": 3}') works as you'd expect, and types with no JSON equivalent, such as datetime or UUID, are still accepted from their standard string forms when parsing JSON.
Constraining Fields
Types alone don't say "positive integer" or "at most 500 characters". Field() adds constraints, and Annotated lets you package a type and its constraints into a reusable alias:
from decimal import Decimal
from typing import Annotated, Literal
from pydantic import BaseModel, EmailStr, Field, HttpUrl, ValidationError
Sku = Annotated[str, Field(pattern=r"^[A-Z]{3}-\d{4}$")]
Quantity = Annotated[int, Field(gt=0, le=1000)]
class LineItem(BaseModel):
sku: Sku
quantity: Quantity = 1
unit_price: Annotated[Decimal, Field(ge=0, max_digits=10, decimal_places=2)]
class Order(BaseModel):
customer_email: EmailStr
status: Literal["pending", "paid", "shipped"] = "pending"
items: list[LineItem] = Field(min_length=1)
notes: str = Field(default="", max_length=500)
website: HttpUrl | None = None
- Numeric constraints:
gt,ge,lt,le,multiple_of.Decimalfields also takemax_digitsanddecimal_places, which matters for money. - String constraints:
min_length,max_length,pattern. - Collection constraints:
min_lengthandmax_lengthon lists, sets and dicts. Literal[...]restricts a value to a fixed set. AnEnumworks too, if you want the values as named constants.EmailStr(needs theemailextra) andHttpUrlare ready-made types for common formats.- Nested models just work:
itemsis validated as a list ofLineItemobjects, dictionaries included.
Errors from nested data come with their full path:
try:
Order(
customer_email="not-an-email",
status="lost",
items=[{"sku": "mug-1", "quantity": 0, "unit_price": "1.999"}],
)
except ValidationError as exc:
for err in exc.errors():
print(".".join(map(str, err["loc"])), "->", err["msg"])
customer_email -> value is not a valid email address: An email address must have an @-sign.
status -> Input should be 'pending', 'paid' or 'shipped'
items.0.sku -> String should match pattern '^[A-Z]{3}-\d{4}$'
items.0.quantity -> Input should be greater than 0
items.0.unit_price -> Decimal input should have no more than 2 decimal places
items.0.sku tells you exactly which list element failed. That's very helpful when you're returning errors to an API client who sent a large payload.
Custom Validators
When built-in constraints aren't enough, write a validator.
Field Validators
@field_validator runs custom logic for one or more fields. Raise ValueError to reject a value; return a value to accept (and possibly transform) it:
from datetime import date
from typing import Self
from pydantic import BaseModel, ValidationError, field_validator, model_validator
class Booking(BaseModel):
guest: str
check_in: date
check_out: date
guests: int = 1
@field_validator("guest")
@classmethod
def normalize_guest(cls, value: str) -> str:
value = " ".join(value.split())
if not value:
raise ValueError("guest name cannot be blank")
return value.title()
@model_validator(mode="after")
def check_dates(self) -> Self:
if self.check_out <= self.check_in:
raise ValueError("check_out must be after check_in")
return self
print(Booking(guest=" ada lovelace ", check_in="2026-10-01", check_out="2026-10-04"))
guest='Ada Lovelace' check_in=datetime.date(2026, 10, 1) check_out=datetime.date(2026, 10, 4) guests=1
By default, a field validator is an "after" validator: it runs once Pydantic has already converted the input to str, so your function can rely on the type.
Model Validators
Rules that involve several fields belong in a @model_validator. With mode="after", it receives the fully validated instance (self), so you can compare fields with their real types. If it fails, the error is reported at the model level:
try:
Booking(guest="Ada", check_in="2026-10-04", check_out="2026-10-01")
except ValidationError as exc:
print(exc.errors()[0]["loc"], exc.errors()[0]["msg"])
() Value error, check_out must be after check_in
The empty loc tuple means "the whole model". A mode="before" model validator receives the raw input instead, which is useful for reshaping data (renaming legacy keys, unwrapping an envelope) before field validation runs.
Reusable Validators with Annotated
Validators attached to a model only help that model. To reuse logic, attach BeforeValidator and AfterValidator to a type alias:
from typing import Annotated, Any
from pydantic import AfterValidator, BaseModel, BeforeValidator
def split_csv(value: Any) -> Any:
if isinstance(value, str):
return [part.strip() for part in value.split(",") if part.strip()]
return value
def dedupe(values: list[str]) -> list[str]:
return sorted(set(v.lower() for v in values))
Tags = Annotated[list[str], BeforeValidator(split_csv), AfterValidator(dedupe)]
class Article(BaseModel):
title: str
tags: Tags = []
print(Article(title="Hello", tags="Python, pydantic,python"))
print(Article(title="Hello", tags=["B", "a", "b"]))
title='Hello' tags=['pydantic', 'python']
title='Hello' tags=['a', 'b']
The before validator turns a comma-separated string into a list, Pydantic validates it as list[str], and the after validator normalizes the result. Any model can now use Tags and get the same behavior.
Model Configuration
model_config controls model-wide behavior. Three settings are worth knowing early:
from pydantic import BaseModel, ConfigDict, ValidationError
class SignupRequest(BaseModel):
model_config = ConfigDict(extra="forbid", str_strip_whitespace=True, frozen=True)
username: str
password: str
req = SignupRequest(username=" ada ", password="s3cret-pass")
print(repr(req.username)) # 'ada'
SignupRequest(username="ada", password="x", is_admin=True)
# ValidationError: extra_forbidden at ('is_admin',)
req.username = "mallory"
# ValidationError: frozen_instance
extra: by default, unknown fields are silently ignored ("ignore")."forbid"rejects them, which is a good default for request bodies because it catches typos and attempts to set fields likeis_admin."allow"keeps them.str_strip_whitespace: trims every string field.frozen: makes instances immutable and hashable. To get a modified copy, usereq.model_copy(update={"username": "grace"}), but note thatmodel_copydoesn't re-run validation on the updated values.
Serialization
Validation is half the job. The other half is turning models back into dictionaries and JSON:
from datetime import datetime, timezone
from pydantic import BaseModel, ConfigDict, Field, computed_field
class Product(BaseModel):
model_config = ConfigDict(validate_by_name=True)
product_id: int = Field(alias="productId")
name: str
price_cents: int = Field(alias="priceCents")
created_at: datetime = Field(default_factory=lambda: datetime(2026, 9, 30, tzinfo=timezone.utc))
internal_note: str | None = Field(default=None, exclude=True)
@computed_field
@property
def price(self) -> float:
return self.price_cents / 100
p = Product.model_validate({"productId": 7, "name": "Mug", "priceCents": 1250, "internal_note": "fragile"})
print(p.model_dump())
print(p.model_dump(by_alias=True, exclude={"created_at"}))
print(p.model_dump(mode="json"))
print(p.model_dump_json(indent=2))
{'product_id': 7, 'name': 'Mug', 'price_cents': 1250, 'created_at': datetime.datetime(2026, 9, 30, 0, 0, tzinfo=datetime.timezone.utc), 'price': 12.5}
{'productId': 7, 'name': 'Mug', 'priceCents': 1250, 'price': 12.5}
{'product_id': 7, 'name': 'Mug', 'price_cents': 1250, 'created_at': '2026-09-30T00:00:00Z', 'price': 12.5}
{
"product_id": 7,
"name": "Mug",
"price_cents": 1250,
"created_at": "2026-09-30T00:00:00Z",
"price": 12.5
}
What each piece does:
- Aliases map external names to Python names. The input used
productIdandpriceCents(camelCase, as a JavaScript client would send), while your code uses snake_case attributes.validate_by_name=Trueadditionally accepts the Python names, soProduct(product_id=8, name="Lamp", price_cents=6500)also works. Passby_alias=Truewhen dumping to send camelCase back out. model_dump()returns Python objects (note thedatetime).model_dump(mode="json")returns only JSON-compatible types, andmodel_dump_json()goes straight to a JSON string.exclude=Trueon a field keeps it out of every dump, which is handy for internal or sensitive data. You can also passinclude/excludesets per call.@computed_fieldadds a property to the serialized output. It's computed, not stored, and it also appears in the serialization-mode JSON schema.
Every model can also describe itself as JSON Schema with Model.model_json_schema(). That's how FastAPI produces its OpenAPI documentation.
Unions and Discriminated Unions
Real payloads often come in several shapes. A payment can be a card, a bank transfer, or PayPal, each with different fields. A plain union (CardPayment | BankTransfer | PayPal) makes Pydantic try each type, which is slower and produces confusing errors when nothing matches. If the shapes share a tag field, use a discriminated union:
from typing import Annotated, Literal
from pydantic import BaseModel, Field, ValidationError
class CardPayment(BaseModel):
method: Literal["card"]
last4: str = Field(pattern=r"^\d{4}$")
class BankTransfer(BaseModel):
method: Literal["bank"]
iban: str
class PayPal(BaseModel):
method: Literal["paypal"]
email: str
Payment = Annotated[CardPayment | BankTransfer | PayPal, Field(discriminator="method")]
class Checkout(BaseModel):
order_id: int
payment: Payment
c = Checkout.model_validate_json(
'{"order_id": 1, "payment": {"method": "bank", "iban": "DE89370400440532013000"}}'
)
print(type(c.payment).__name__) # BankTransfer
Pydantic reads method first and validates against exactly one model. Errors point at the right place:
('payment', 'card', 'last4') String should match pattern '^\d{4}$'
('payment',) Input tag 'cash' found using 'method' does not match any of the expected tags: 'card', 'bank', 'paypal'
The first error comes from a "card" payload with "last4": "12"; the second from a payload with "method": "cash". This pattern pairs naturally with structural pattern matching when you process the result.
Validating Without a Model
Sometimes the data isn't naturally a model: a list of events from an API, a dict of port numbers, a function's arguments. TypeAdapter validates against any type:
from datetime import date
from pydantic import BaseModel, TypeAdapter
class Event(BaseModel):
name: str
on: date
events_adapter = TypeAdapter(list[Event])
raw = '[{"name": "PyCon", "on": "2026-05-13"}, {"name": "Launch", "on": "2026-11-02"}]'
events = events_adapter.validate_json(raw)
print(events[1])
print(events_adapter.dump_json(events))
ports = TypeAdapter(dict[str, int]).validate_python({"http": "80", "https": 443})
print(ports)
name='Launch' on=datetime.date(2026, 11, 2)
b'[{"name":"PyCon","on":"2026-05-13"},{"name":"Launch","on":"2026-11-02"}]'
{'http': 80, 'https': 443}
Creating an adapter has a cost, so build it once at module level and reuse it, rather than inside a function that runs per request.
For functions, @validate_call validates arguments against the signature's type hints:
from pydantic import ValidationError, validate_call
@validate_call
def repeat(text: str, times: int = 2) -> str:
return text * times
print(repeat("ab", "3")) # ababab
repeat("ab", times="many") # ValidationError at ('times',): int_parsing
It's useful at boundaries, such as functions called from a CLI or plugin system, but don't sprinkle it on every function; static type checking is the cheaper tool for code you control.
If you prefer dataclass syntax, pydantic.dataclasses.dataclass gives you a standard-looking dataclass with Pydantic validation.
Settings from the Environment
Configuration is untrusted input too: environment variables are all strings, and a typo in a variable name shouldn't surface as a crash an hour into production. The separate pydantic-settings package reads settings from the environment and .env files, with the same validation:
# settings.py
from pydantic import Field, PostgresDsn, SecretStr
from pydantic_settings import BaseSettings, SettingsConfigDict
class Settings(BaseSettings):
model_config = SettingsConfigDict(env_prefix="APP_", env_file=".env")
debug: bool = False
database_url: PostgresDsn
api_key: SecretStr
allowed_hosts: list[str] = Field(default_factory=lambda: ["localhost"])
workers: int = 2
settings = Settings()
print(settings.debug, settings.workers, settings.allowed_hosts)
print(settings.database_url)
print(settings.api_key)
print(settings.api_key.get_secret_value())
export APP_DEBUG=1 APP_WORKERS=4 APP_API_KEY=sk-123
export APP_DATABASE_URL=postgresql://app:pw@db:5432/shop
export APP_ALLOWED_HOSTS='["example.com","www.example.com"]'
python settings.py
True 4 ['example.com', 'www.example.com']
postgresql://app:pw@db:5432/shop
**********
sk-123
env_prefix="APP_"mapsAPP_DATABASE_URLtodatabase_url(matching is case-insensitive by default).- Complex types like
list[str]are read from the environment as JSON. SecretStrhides its value inrepr(),print()and logs; you must callget_secret_value()deliberately.- If
APP_API_KEYis missing orAPP_DATABASE_URLisn't a valid URL,Settings()raises aValidationErrorimmediately at startup, listing every problem.
Create a single settings instance at import time and import it wherever it's needed, so the app fails fast on bad configuration.
Where Pydantic Fits
Pydantic is most valuable at the edges of your program, where untrusted data enters:
- HTTP request bodies and query parameters (FastAPI does this for you; see building a REST API with FastAPI)
- Responses from third-party APIs
- Config files and environment variables
- Messages from queues, files from users, LLM output that should match a schema
Inside your program, once data has been validated, plain dataclasses or Pydantic models both work. Validating the same data over and over in internal code adds overhead without adding safety.
Conclusion
Pydantic lets you replace scattered if checks with a type-annotated declaration of what valid data looks like. Models coerce data sensibly in lax mode or exactly in strict mode, Field and Annotated add constraints, validators handle custom rules for single fields or whole models, and model_dump() and aliases handle the trip back out. Add discriminated unions for multi-shape payloads, TypeAdapter for anything that isn't a model, and pydantic-settings for configuration, and you have one consistent approach to every piece of data that crosses into your program.


