
Dataclasses in Python: Less Boilerplate, More Clarity
A lot of classes exist mainly to hold data: a product with a SKU and a price, a config object with a host and a port, a point with x and y. Writing them by hand means typing every field name three times in __init__, then writing a __repr__ so you can debug them, then an __eq__ so two equal objects compare equal, and maybe a __hash__ and some ordering methods. It's tedious, and every repeated name is a chance for a typo.
The dataclasses module, in the standard library since Python 3.7, does that work for you. You declare the fields with type annotations, add the @dataclass decorator, and Python generates the methods. The result is shorter, easier to read, and harder to get wrong.
This guide covers the basics and what gets generated, defaults and the mutable default rule, field() options, validation with __post_init__, frozen and ordered dataclasses, keyword-only fields, slots=True, converting to dicts, inheritance, and how dataclasses compare with NamedTuple, TypedDict, and Pydantic.
A Dataclass in Five Lines
from dataclasses import dataclass
@dataclass
class Product:
sku: str
name: str
price: float
in_stock: bool = True
mug = Product("MUG-01", "Mug", 12.0)
print(mug)
print(mug == Product("MUG-01", "Mug", 12.0))
print(Product(sku="LAMP-04", name="Lamp", price=65.0, in_stock=False))
Product(sku='MUG-01', name='Mug', price=12.0, in_stock=True)
True
Product(sku='LAMP-04', name='Lamp', price=65.0, in_stock=False)
Each annotated class attribute becomes a field. From those four lines, @dataclass generated:
__init__takingsku,name,price, and an optionalin_stock, in declaration order, positionally or by keyword.__repr__that shows every field, which is exactly what you want in logs and the debugger.__eq__that compares twoProductobjects field by field, as if they were tuples.
The equivalent hand-written class is about three times as long:
class Product:
def __init__(self, sku: str, name: str, price: float, in_stock: bool = True) -> None:
self.sku = sku
self.name = name
self.price = price
self.in_stock = in_stock
def __repr__(self) -> str:
return (
f"Product(sku={self.sku!r}, name={self.name!r}, "
f"price={self.price!r}, in_stock={self.in_stock!r})"
)
def __eq__(self, other: object) -> bool:
if not isinstance(other, Product):
return NotImplemented
return (self.sku, self.name, self.price, self.in_stock) == (
other.sku, other.name, other.price, other.in_stock
)
Add a field to that version and you have to remember to update three methods. With a dataclass you add one line. If you want to see how those generated methods work, Python dunder methods covers __repr__, __eq__, and friends in depth.
A dataclass is still a normal class. You can add methods, properties, class methods, and anything else you'd put in a class. The decorator only adds the generated methods; it doesn't change what else you can do.
Type Annotations Are Required, Not Enforced
A field is defined by its annotation. A class attribute without an annotation, like count = 0, is not a field and won't appear in __init__.
The types themselves aren't checked at runtime. Product("MUG-01", "Mug", "twelve") runs without complaint. The annotations are for readers and for static type checkers like mypy and Pyright, which will flag the mistake. If you need runtime validation, that's what __post_init__ (below) or Pydantic is for. For more on annotations in general, see type hints in Python.
Defaults and the Mutable Default Rule
Fields with defaults must come after fields without them, the same rule as function parameters:
@dataclass
class Bad:
a: int = 0
b: int
TypeError: non-default argument 'b' follows default argument 'a'
And a mutable default like a list is rejected outright:
@dataclass
class Bad:
tags: list[str] = []
ValueError: mutable default <class 'list'> for field tags is not allowed: use default_factory
This protects you from the classic bug where every instance would share one list (the same issue as the mutable default argument trap). The fix the error suggests is field(default_factory=...), which calls a function to create a fresh default for every instance.
field(): Per-Field Options
dataclasses.field() customizes individual fields. Here's an order model that uses several options:
# orders.py
from dataclasses import asdict, dataclass, field, replace
from datetime import UTC, datetime
from decimal import Decimal
from uuid import uuid4
@dataclass
class LineItem:
sku: str
quantity: int
unit_price: Decimal
def __post_init__(self) -> None:
if self.quantity < 1:
raise ValueError(f"quantity must be at least 1, got {self.quantity}")
@property
def subtotal(self) -> Decimal:
return self.quantity * self.unit_price
@dataclass
class Order:
customer: str
items: list[LineItem] = field(default_factory=list)
id: str = field(default_factory=lambda: uuid4().hex[:8], repr=False)
created_at: datetime = field(
default_factory=lambda: datetime.now(UTC), repr=False, compare=False
)
@property
def total(self) -> Decimal:
return sum((item.subtotal for item in self.items), Decimal("0"))
order = Order("ada", [LineItem("MUG-01", 2, Decimal("12.00"))])
order.items.append(LineItem("PEN-10", 10, Decimal("1.50")))
print(order)
print(order.total)
print(Order("grace").items)
Order(customer='ada', items=[LineItem(sku='MUG-01', quantity=2, unit_price=Decimal('12.00')), LineItem(sku='PEN-10', quantity=10, unit_price=Decimal('1.50'))])
39.00
[]
What each option does:
default_factory=listgives every order its own empty list.Order("grace").itemsis a fresh[], not shared with anyone.default_factory=lambda: ...works for any computed default, like a generated ID or the current time. The function is called once per instance, at creation.repr=Falsehides a field from__repr__. Use it for noisy fields, or for sensitive ones like tokens and passwords that shouldn't end up in logs.compare=Falseleaves a field out of__eq__(and ordering and hashing). Two orders created a millisecond apart but otherwise identical still compare equal.
Other field() options you'll occasionally need:
| Option | Effect |
|---|---|
default | A plain default value, same as x: int = 0 |
default_factory | A zero-argument callable that produces the default |
init=False | Leave the field out of __init__ (set it in __post_init__) |
repr=False | Leave it out of __repr__ |
compare=False | Leave it out of __eq__ and ordering |
hash | Override whether it's part of __hash__ (defaults to compare) |
kw_only=True | Make this field keyword-only in __init__ |
metadata | A read-only mapping for your own use, such as tools that read field info |
Notice total and subtotal are properties, not fields. They have no annotation at class level, so the dataclass ignores them, and they're computed from the real fields every time. That's the right way to add derived values; see @property in Python.
Validation with __post_init__
The generated __init__ just assigns fields. If you need to check or normalize values, define __post_init__, which the generated __init__ calls as its last step. LineItem above uses it to reject bad quantities:
LineItem("MUG-01", 0, Decimal("12.00"))
ValueError: quantity must be at least 1, got 0
Because the check runs inside __init__, an invalid LineItem can never exist.
Init-Only Values and Computed Fields
Sometimes a value is needed to build the object but shouldn't be stored on it. Annotate it as InitVar, and it's passed to __post_init__ instead of becoming a field. Combine it with field(init=False) for fields that are computed rather than passed in:
from dataclasses import InitVar, dataclass, field, fields
from typing import ClassVar
@dataclass
class Article:
title: str
body: InitVar[str]
word_count: int = field(init=False)
reading_time: int = field(init=False)
words_per_minute: ClassVar[int] = 200
def __post_init__(self, body: str) -> None:
self.word_count = len(body.split())
self.reading_time = max(1, round(self.word_count / self.words_per_minute))
post = Article("Dataclasses", "word " * 950)
print(post)
print([f.name for f in fields(Article)])
Article(title='Dataclasses', word_count=950, reading_time=5)
['title', 'word_count', 'reading_time']
Three different kinds of annotation are at work here:
body: InitVar[str]is an__init__parameter that's passed to__post_init__and then discarded. It's not a field and isn't stored.word_count: int = field(init=False)is a field that's not an__init__parameter.__post_init__sets it.words_per_minute: ClassVar[int]is a class attribute, explicitly excluded from fields. UseClassVarfor constants that belong to the class.
dataclasses.fields() lists the real fields, which is handy for introspection and generic code.
Frozen and Ordered Dataclasses
Two decorator arguments turn a dataclass into a proper value type:
import copy
from dataclasses import FrozenInstanceError, dataclass, field
@dataclass(frozen=True, order=True)
class Version:
major: int
minor: int
patch: int = 0
label: str = field(default="", compare=False)
def __str__(self) -> str:
return f"{self.major}.{self.minor}.{self.patch}"
@classmethod
def parse(cls, text: str) -> "Version":
return cls(*map(int, text.split(".")))
versions = [Version.parse(v) for v in ["3.13.2", "3.9.18", "3.14.0", "3.13.0"]]
print([str(v) for v in sorted(versions)])
print(max(versions), Version(3, 13) < Version(3, 13, 1))
print({Version(3, 13), Version(3, 13, 0, "stable")})
v = Version(3, 13, 2)
try:
v.minor = 14
except FrozenInstanceError as exc:
print(exc)
print(copy.replace(v, minor=14))
['3.9.18', '3.13.0', '3.13.2', '3.14.0']
3.14.0 True
{Version(major=3, minor=13, patch=0, label='')}
cannot assign to field 'minor'
3.14.2
order=True generates __lt__, __le__, __gt__, and __ge__, comparing fields in declaration order like a tuple. That's why 3.9.18 sorts before 3.13.0: it compares the integers 9 and 13, not strings. Field order therefore matters: put the most significant field first.
frozen=True makes instances immutable. Assigning to a field raises FrozenInstanceError. Because the fields can't change, the dataclass also generates a __hash__, so frozen instances work in sets and as dict keys. The set above has only one element because label has compare=False and is left out of both equality and the hash.
By default, a non-frozen dataclass with eq=True sets __hash__ to None, making instances unhashable. That's deliberate: hashing a mutable object whose equality can change would corrupt sets and dicts. If you need hashable dataclasses, make them frozen. (unsafe_hash=True exists, but the name says it all.)
Changing a Frozen Instance: replace()
Since you can't modify a frozen instance, you create a modified copy. dataclasses.replace(obj, **changes) has always done this. Python 3.13 added the general-purpose copy.replace(), which works with dataclasses and also named tuples, datetime objects, and other types that support it. Both run __init__ and __post_init__ again, so validation applies to the copy too.
Frozen Dataclasses and __post_init__
Assigning self.x = ... in __post_init__ raises FrozenInstanceError on a frozen dataclass. To set a computed field during initialization, go around the frozen __setattr__ with object.__setattr__:
from dataclasses import dataclass, field
@dataclass(frozen=True)
class Slug:
text: str
value: str = field(init=False)
def __post_init__(self) -> None:
object.__setattr__(self, "value", self.text.strip().lower().replace(" ", "-"))
print(Slug(" Hello World "))
Slug(text=' Hello World ', value='hello-world')
It's the officially documented workaround, and it's only appropriate inside __post_init__.
Keyword-Only Fields
For classes with many fields, especially config-style objects, positional arguments are hard to read: what does ServerConfig("example.com", 443, True) mean? kw_only=True forces every field to be passed by name:
from dataclasses import dataclass
@dataclass(kw_only=True)
class ServerConfig:
host: str = "localhost"
port: int = 8000
debug: bool = False
print(ServerConfig(port=9000, debug=True))
ServerConfig("example.com")
ServerConfig(host='localhost', port=9000, debug=True)
TypeError: ServerConfig.__init__() takes 1 positional argument but 2 were given
To make only some fields keyword-only, put a _: KW_ONLY sentinel line before them (import KW_ONLY from dataclasses); every field after it becomes keyword-only. Keyword-only fields also lift the "no non-default after default" rule, since ordering no longer matters for them, which is especially useful with inheritance.
Inheritance
A dataclass can inherit from another dataclass. The subclass's fields are added after the parent's:
from dataclasses import dataclass
@dataclass
class Base:
id: int
active: bool = True
@dataclass(kw_only=True)
class Customer(Base):
email: str
print(Customer(1, email="ada@example.com"))
Customer(id=1, active=True, email='ada@example.com')
Without kw_only=True, this would fail: email (no default) would follow active (which has a default), triggering the same "non-default argument follows default argument" error. Making the subclass's fields keyword-only is the clean fix. Keep dataclass hierarchies shallow, though; composition vs inheritance explains why that's good advice generally.
slots=True
By default, each instance stores its attributes in a __dict__. Passing slots=True (Python 3.10+) generates a class with __slots__ instead, which uses less memory per instance, makes attribute access slightly faster, and prevents typos from silently creating new attributes:
from dataclasses import dataclass
@dataclass(slots=True)
class Pixel:
x: int
y: int
p = Pixel(1, 2)
p.z = 3
AttributeError: 'Pixel' object has no attribute 'z' and no __dict__ for setting new attributes
It's a good default for small classes you create many instances of. The trade-offs, such as no functools.cached_property and some care with multiple inheritance, are covered in slots in Python.
Converting to Dicts and Tuples
asdict() and astuple() convert a dataclass into plain built-in types, recursing into nested dataclasses, lists, and dicts:
data = asdict(order)
print(data["items"][0])
print(list(data))
{'sku': 'MUG-01', 'quantity': 2, 'unit_price': Decimal('12.00')}
['customer', 'items', 'id', 'created_at']
That's the first step toward JSON, though values like Decimal and datetime still need converting; working with JSON in Python shows how. Note that asdict() deep-copies everything, which can be slow for large structures. For a shallow version, use {f.name: getattr(obj, f.name) for f in fields(obj)}.
There's no built-in way to go the other direction for nested data. Order(**data) would leave items as a list of dicts rather than LineItem objects. If you're loading nested data from JSON or an API, that's a strong sign you want Pydantic.
Pattern Matching
Dataclasses generate __match_args__ from their fields, so they work with positional patterns in match statements:
match Version(3, 13, 2):
case Version(3, minor) if minor >= 12:
print(f"modern Python 3.{minor}")
modern Python 3.13
See structural pattern matching in Python for more on class patterns.
Dataclass vs the Alternatives
Python has several ways to define a data-holding type. They overlap, but each has a sweet spot:
| Tool | Mutable | Runtime validation | Best for |
|---|---|---|---|
@dataclass | Yes (or frozen) | No (DIY in __post_init__) | Most internal data objects with behavior |
typing.NamedTuple | No | No | Lightweight immutable records that should also be tuples |
typing.TypedDict | Yes (it's a dict) | No | Typing dict-shaped data, like JSON you don't convert |
Pydantic BaseModel | Yes | Yes, with coercion | Untrusted input: APIs, config files, JSON parsing |
| attrs | Yes (or frozen) | Optional validators | Dataclass-like with more features (third-party) |
A reasonable default: use a dataclass for data your own code creates, and Pydantic at the boundaries where data comes in from outside. The Pydantic guide covers that side. A NamedTuple is worth choosing when tuple behavior (unpacking, indexing) is a feature; see the collections module for its sibling namedtuple.
Decorator Options at a Glance
| Argument | Default | Effect |
|---|---|---|
init | True | Generate __init__ |
repr | True | Generate __repr__ |
eq | True | Generate __eq__ |
order | False | Generate <, <=, >, >= |
frozen | False | Make instances immutable (and hashable) |
unsafe_hash | False | Force a __hash__ even if mutable |
kw_only | False | Make all fields keyword-only (3.10+) |
slots | False | Use __slots__ (3.10+) |
match_args | True | Generate __match_args__ (3.10+) |
weakref_slot | False | Add a weak reference slot when using slots (3.11+) |
If you define one of these methods yourself, the decorator leaves your version alone (for __init__, __repr__, and __eq__), so you can customize a single method without giving up the rest.
Conclusion
Dataclasses take the repetitive parts of writing a class, __init__, __repr__, __eq__, and optionally hashing and ordering, and generate them from a list of annotated fields. Use field(default_factory=...) for mutable defaults, __post_init__ for validation and derived values, frozen=True for immutable value types, kw_only=True for config-style classes and inheritance, and slots=True for lightweight objects you create in bulk. They don't validate types at runtime, so put Pydantic at the edges of your system where untrusted data comes in. For everything in between, a dataclass is usually the clearest way to say "this class holds this data".


