Type something to search...
Dataclasses in Python: Less Boilerplate, More Clarity

Dataclasses in Python: Less Boilerplate, More Clarity

A lot of classes exist mainly to hold data: a product with a SKU and a price, a config object with a host and a port, a point with x and y. Writing them by hand means typing every field name three times in __init__, then writing a __repr__ so you can debug them, then an __eq__ so two equal objects compare equal, and maybe a __hash__ and some ordering methods. It's tedious, and every repeated name is a chance for a typo.

The dataclasses module, in the standard library since Python 3.7, does that work for you. You declare the fields with type annotations, add the @dataclass decorator, and Python generates the methods. The result is shorter, easier to read, and harder to get wrong.

This guide covers the basics and what gets generated, defaults and the mutable default rule, field() options, validation with __post_init__, frozen and ordered dataclasses, keyword-only fields, slots=True, converting to dicts, inheritance, and how dataclasses compare with NamedTuple, TypedDict, and Pydantic.

A Dataclass in Five Lines

from dataclasses import dataclass


@dataclass
class Product:
    sku: str
    name: str
    price: float
    in_stock: bool = True


mug = Product("MUG-01", "Mug", 12.0)
print(mug)
print(mug == Product("MUG-01", "Mug", 12.0))
print(Product(sku="LAMP-04", name="Lamp", price=65.0, in_stock=False))
Product(sku='MUG-01', name='Mug', price=12.0, in_stock=True)
True
Product(sku='LAMP-04', name='Lamp', price=65.0, in_stock=False)

Each annotated class attribute becomes a field. From those four lines, @dataclass generated:

  • __init__ taking sku, name, price, and an optional in_stock, in declaration order, positionally or by keyword.
  • __repr__ that shows every field, which is exactly what you want in logs and the debugger.
  • __eq__ that compares two Product objects field by field, as if they were tuples.

The equivalent hand-written class is about three times as long:

class Product:
    def __init__(self, sku: str, name: str, price: float, in_stock: bool = True) -> None:
        self.sku = sku
        self.name = name
        self.price = price
        self.in_stock = in_stock

    def __repr__(self) -> str:
        return (
            f"Product(sku={self.sku!r}, name={self.name!r}, "
            f"price={self.price!r}, in_stock={self.in_stock!r})"
        )

    def __eq__(self, other: object) -> bool:
        if not isinstance(other, Product):
            return NotImplemented
        return (self.sku, self.name, self.price, self.in_stock) == (
            other.sku, other.name, other.price, other.in_stock
        )

Add a field to that version and you have to remember to update three methods. With a dataclass you add one line. If you want to see how those generated methods work, Python dunder methods covers __repr__, __eq__, and friends in depth.

A dataclass is still a normal class. You can add methods, properties, class methods, and anything else you'd put in a class. The decorator only adds the generated methods; it doesn't change what else you can do.

Type Annotations Are Required, Not Enforced

A field is defined by its annotation. A class attribute without an annotation, like count = 0, is not a field and won't appear in __init__.

The types themselves aren't checked at runtime. Product("MUG-01", "Mug", "twelve") runs without complaint. The annotations are for readers and for static type checkers like mypy and Pyright, which will flag the mistake. If you need runtime validation, that's what __post_init__ (below) or Pydantic is for. For more on annotations in general, see type hints in Python.

Defaults and the Mutable Default Rule

Fields with defaults must come after fields without them, the same rule as function parameters:

@dataclass
class Bad:
    a: int = 0
    b: int
TypeError: non-default argument 'b' follows default argument 'a'

And a mutable default like a list is rejected outright:

@dataclass
class Bad:
    tags: list[str] = []
ValueError: mutable default <class 'list'> for field tags is not allowed: use default_factory

This protects you from the classic bug where every instance would share one list (the same issue as the mutable default argument trap). The fix the error suggests is field(default_factory=...), which calls a function to create a fresh default for every instance.

field(): Per-Field Options

dataclasses.field() customizes individual fields. Here's an order model that uses several options:

# orders.py
from dataclasses import asdict, dataclass, field, replace
from datetime import UTC, datetime
from decimal import Decimal
from uuid import uuid4


@dataclass
class LineItem:
    sku: str
    quantity: int
    unit_price: Decimal

    def __post_init__(self) -> None:
        if self.quantity < 1:
            raise ValueError(f"quantity must be at least 1, got {self.quantity}")

    @property
    def subtotal(self) -> Decimal:
        return self.quantity * self.unit_price


@dataclass
class Order:
    customer: str
    items: list[LineItem] = field(default_factory=list)
    id: str = field(default_factory=lambda: uuid4().hex[:8], repr=False)
    created_at: datetime = field(
        default_factory=lambda: datetime.now(UTC), repr=False, compare=False
    )

    @property
    def total(self) -> Decimal:
        return sum((item.subtotal for item in self.items), Decimal("0"))


order = Order("ada", [LineItem("MUG-01", 2, Decimal("12.00"))])
order.items.append(LineItem("PEN-10", 10, Decimal("1.50")))
print(order)
print(order.total)
print(Order("grace").items)
Order(customer='ada', items=[LineItem(sku='MUG-01', quantity=2, unit_price=Decimal('12.00')), LineItem(sku='PEN-10', quantity=10, unit_price=Decimal('1.50'))])
39.00
[]

What each option does:

  • default_factory=list gives every order its own empty list. Order("grace").items is a fresh [], not shared with anyone.
  • default_factory=lambda: ... works for any computed default, like a generated ID or the current time. The function is called once per instance, at creation.
  • repr=False hides a field from __repr__. Use it for noisy fields, or for sensitive ones like tokens and passwords that shouldn't end up in logs.
  • compare=False leaves a field out of __eq__ (and ordering and hashing). Two orders created a millisecond apart but otherwise identical still compare equal.

Other field() options you'll occasionally need:

OptionEffect
defaultA plain default value, same as x: int = 0
default_factoryA zero-argument callable that produces the default
init=FalseLeave the field out of __init__ (set it in __post_init__)
repr=FalseLeave it out of __repr__
compare=FalseLeave it out of __eq__ and ordering
hashOverride whether it's part of __hash__ (defaults to compare)
kw_only=TrueMake this field keyword-only in __init__
metadataA read-only mapping for your own use, such as tools that read field info

Notice total and subtotal are properties, not fields. They have no annotation at class level, so the dataclass ignores them, and they're computed from the real fields every time. That's the right way to add derived values; see @property in Python.

Validation with __post_init__

The generated __init__ just assigns fields. If you need to check or normalize values, define __post_init__, which the generated __init__ calls as its last step. LineItem above uses it to reject bad quantities:

LineItem("MUG-01", 0, Decimal("12.00"))
ValueError: quantity must be at least 1, got 0

Because the check runs inside __init__, an invalid LineItem can never exist.

Init-Only Values and Computed Fields

Sometimes a value is needed to build the object but shouldn't be stored on it. Annotate it as InitVar, and it's passed to __post_init__ instead of becoming a field. Combine it with field(init=False) for fields that are computed rather than passed in:

from dataclasses import InitVar, dataclass, field, fields
from typing import ClassVar


@dataclass
class Article:
    title: str
    body: InitVar[str]
    word_count: int = field(init=False)
    reading_time: int = field(init=False)
    words_per_minute: ClassVar[int] = 200

    def __post_init__(self, body: str) -> None:
        self.word_count = len(body.split())
        self.reading_time = max(1, round(self.word_count / self.words_per_minute))


post = Article("Dataclasses", "word " * 950)
print(post)
print([f.name for f in fields(Article)])
Article(title='Dataclasses', word_count=950, reading_time=5)
['title', 'word_count', 'reading_time']

Three different kinds of annotation are at work here:

  • body: InitVar[str] is an __init__ parameter that's passed to __post_init__ and then discarded. It's not a field and isn't stored.
  • word_count: int = field(init=False) is a field that's not an __init__ parameter. __post_init__ sets it.
  • words_per_minute: ClassVar[int] is a class attribute, explicitly excluded from fields. Use ClassVar for constants that belong to the class.

dataclasses.fields() lists the real fields, which is handy for introspection and generic code.

Frozen and Ordered Dataclasses

Two decorator arguments turn a dataclass into a proper value type:

import copy
from dataclasses import FrozenInstanceError, dataclass, field


@dataclass(frozen=True, order=True)
class Version:
    major: int
    minor: int
    patch: int = 0
    label: str = field(default="", compare=False)

    def __str__(self) -> str:
        return f"{self.major}.{self.minor}.{self.patch}"

    @classmethod
    def parse(cls, text: str) -> "Version":
        return cls(*map(int, text.split(".")))


versions = [Version.parse(v) for v in ["3.13.2", "3.9.18", "3.14.0", "3.13.0"]]
print([str(v) for v in sorted(versions)])
print(max(versions), Version(3, 13) < Version(3, 13, 1))
print({Version(3, 13), Version(3, 13, 0, "stable")})

v = Version(3, 13, 2)
try:
    v.minor = 14
except FrozenInstanceError as exc:
    print(exc)
print(copy.replace(v, minor=14))
['3.9.18', '3.13.0', '3.13.2', '3.14.0']
3.14.0 True
{Version(major=3, minor=13, patch=0, label='')}
cannot assign to field 'minor'
3.14.2

order=True generates __lt__, __le__, __gt__, and __ge__, comparing fields in declaration order like a tuple. That's why 3.9.18 sorts before 3.13.0: it compares the integers 9 and 13, not strings. Field order therefore matters: put the most significant field first.

frozen=True makes instances immutable. Assigning to a field raises FrozenInstanceError. Because the fields can't change, the dataclass also generates a __hash__, so frozen instances work in sets and as dict keys. The set above has only one element because label has compare=False and is left out of both equality and the hash.

By default, a non-frozen dataclass with eq=True sets __hash__ to None, making instances unhashable. That's deliberate: hashing a mutable object whose equality can change would corrupt sets and dicts. If you need hashable dataclasses, make them frozen. (unsafe_hash=True exists, but the name says it all.)

Changing a Frozen Instance: replace()

Since you can't modify a frozen instance, you create a modified copy. dataclasses.replace(obj, **changes) has always done this. Python 3.13 added the general-purpose copy.replace(), which works with dataclasses and also named tuples, datetime objects, and other types that support it. Both run __init__ and __post_init__ again, so validation applies to the copy too.

Frozen Dataclasses and __post_init__

Assigning self.x = ... in __post_init__ raises FrozenInstanceError on a frozen dataclass. To set a computed field during initialization, go around the frozen __setattr__ with object.__setattr__:

from dataclasses import dataclass, field


@dataclass(frozen=True)
class Slug:
    text: str
    value: str = field(init=False)

    def __post_init__(self) -> None:
        object.__setattr__(self, "value", self.text.strip().lower().replace(" ", "-"))


print(Slug("  Hello World "))
Slug(text='  Hello World ', value='hello-world')

It's the officially documented workaround, and it's only appropriate inside __post_init__.

Keyword-Only Fields

For classes with many fields, especially config-style objects, positional arguments are hard to read: what does ServerConfig("example.com", 443, True) mean? kw_only=True forces every field to be passed by name:

from dataclasses import dataclass


@dataclass(kw_only=True)
class ServerConfig:
    host: str = "localhost"
    port: int = 8000
    debug: bool = False


print(ServerConfig(port=9000, debug=True))
ServerConfig("example.com")
ServerConfig(host='localhost', port=9000, debug=True)
TypeError: ServerConfig.__init__() takes 1 positional argument but 2 were given

To make only some fields keyword-only, put a _: KW_ONLY sentinel line before them (import KW_ONLY from dataclasses); every field after it becomes keyword-only. Keyword-only fields also lift the "no non-default after default" rule, since ordering no longer matters for them, which is especially useful with inheritance.

Inheritance

A dataclass can inherit from another dataclass. The subclass's fields are added after the parent's:

from dataclasses import dataclass


@dataclass
class Base:
    id: int
    active: bool = True


@dataclass(kw_only=True)
class Customer(Base):
    email: str


print(Customer(1, email="ada@example.com"))
Customer(id=1, active=True, email='ada@example.com')

Without kw_only=True, this would fail: email (no default) would follow active (which has a default), triggering the same "non-default argument follows default argument" error. Making the subclass's fields keyword-only is the clean fix. Keep dataclass hierarchies shallow, though; composition vs inheritance explains why that's good advice generally.

slots=True

By default, each instance stores its attributes in a __dict__. Passing slots=True (Python 3.10+) generates a class with __slots__ instead, which uses less memory per instance, makes attribute access slightly faster, and prevents typos from silently creating new attributes:

from dataclasses import dataclass


@dataclass(slots=True)
class Pixel:
    x: int
    y: int


p = Pixel(1, 2)
p.z = 3
AttributeError: 'Pixel' object has no attribute 'z' and no __dict__ for setting new attributes

It's a good default for small classes you create many instances of. The trade-offs, such as no functools.cached_property and some care with multiple inheritance, are covered in slots in Python.

Converting to Dicts and Tuples

asdict() and astuple() convert a dataclass into plain built-in types, recursing into nested dataclasses, lists, and dicts:

data = asdict(order)
print(data["items"][0])
print(list(data))
{'sku': 'MUG-01', 'quantity': 2, 'unit_price': Decimal('12.00')}
['customer', 'items', 'id', 'created_at']

That's the first step toward JSON, though values like Decimal and datetime still need converting; working with JSON in Python shows how. Note that asdict() deep-copies everything, which can be slow for large structures. For a shallow version, use {f.name: getattr(obj, f.name) for f in fields(obj)}.

There's no built-in way to go the other direction for nested data. Order(**data) would leave items as a list of dicts rather than LineItem objects. If you're loading nested data from JSON or an API, that's a strong sign you want Pydantic.

Pattern Matching

Dataclasses generate __match_args__ from their fields, so they work with positional patterns in match statements:

match Version(3, 13, 2):
    case Version(3, minor) if minor >= 12:
        print(f"modern Python 3.{minor}")
modern Python 3.13

See structural pattern matching in Python for more on class patterns.

Dataclass vs the Alternatives

Python has several ways to define a data-holding type. They overlap, but each has a sweet spot:

ToolMutableRuntime validationBest for
@dataclassYes (or frozen)No (DIY in __post_init__)Most internal data objects with behavior
typing.NamedTupleNoNoLightweight immutable records that should also be tuples
typing.TypedDictYes (it's a dict)NoTyping dict-shaped data, like JSON you don't convert
Pydantic BaseModelYesYes, with coercionUntrusted input: APIs, config files, JSON parsing
attrsYes (or frozen)Optional validatorsDataclass-like with more features (third-party)

A reasonable default: use a dataclass for data your own code creates, and Pydantic at the boundaries where data comes in from outside. The Pydantic guide covers that side. A NamedTuple is worth choosing when tuple behavior (unpacking, indexing) is a feature; see the collections module for its sibling namedtuple.

Decorator Options at a Glance

ArgumentDefaultEffect
initTrueGenerate __init__
reprTrueGenerate __repr__
eqTrueGenerate __eq__
orderFalseGenerate <, <=, >, >=
frozenFalseMake instances immutable (and hashable)
unsafe_hashFalseForce a __hash__ even if mutable
kw_onlyFalseMake all fields keyword-only (3.10+)
slotsFalseUse __slots__ (3.10+)
match_argsTrueGenerate __match_args__ (3.10+)
weakref_slotFalseAdd a weak reference slot when using slots (3.11+)

If you define one of these methods yourself, the decorator leaves your version alone (for __init__, __repr__, and __eq__), so you can customize a single method without giving up the rest.

Conclusion

Dataclasses take the repetitive parts of writing a class, __init__, __repr__, __eq__, and optionally hashing and ordering, and generate them from a list of annotated fields. Use field(default_factory=...) for mutable defaults, __post_init__ for validation and derived values, frozen=True for immutable value types, kw_only=True for config-style classes and inheritance, and slots=True for lightweight objects you create in bulk. They don't validate types at runtime, so put Pydantic at the edges of your system where untrusted data comes in. For everything in between, a dataclass is usually the clearest way to say "this class holds this data".

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading