
Property-Based Testing in Python with Hypothesis
Example-based tests check the cases you thought of. That's their strength and their limit: if you didn't think of an empty list, a string with an emoji, or a list whose length is one more than a multiple of the chunk size, your tests won't either. And the cases you don't think of are exactly where bugs live.
Property-based testing flips the approach. Instead of writing inputs and expected outputs, you describe what kind of inputs are valid and a property that should hold for all of them, such as "decoding an encoded value gives back the original". A library then generates hundreds of inputs, including nasty edge cases, and when one fails, it shrinks it down to the simplest input that still breaks your code.
In Python, that library is Hypothesis. This guide covers how to write your first property, the strategies that generate data, how to read Hypothesis's failure reports, the most useful kinds of properties, controlling the search with assume, @example, and settings, and stateful testing for objects that change over time. It all runs inside pytest.
Installing
python -m pip install hypothesis pytest
Hypothesis integrates with pytest automatically through a plugin that's installed with it. There's nothing to configure.
A Bug Example Tests Miss
Here's a function that splits a list into chunks:
# chunks.py
def chunk(items: list, size: int) -> list[list]:
"""Split items into consecutive lists of at most `size` elements."""
if size < 1:
raise ValueError("size must be at least 1")
return [items[i : i + size] for i in range(0, len(items) - 1, size)]
And some ordinary example tests:
def test_chunk_examples():
assert chunk([1, 2, 3, 4], 2) == [[1, 2], [3, 4]]
assert chunk([1, 2, 3, 4, 5, 6], 3) == [[1, 2, 3], [4, 5, 6]]
assert chunk([], 5) == []
They pass. They look thorough. But the - 1 in the range is a bug: whenever the last chunk would contain exactly one item, it's silently dropped.
Now the property-based version. Instead of specific lists, think about what's true of chunk for every input: joining the chunks back together gives you the original list, and every chunk has between 1 and size items.
# tests/test_chunks.py
from hypothesis import given
from hypothesis import strategies as st
from chunks import chunk
@given(st.lists(st.integers()), st.integers(min_value=1, max_value=10))
def test_chunks_rebuild_original(items, size):
result = chunk(items, size)
assert [x for part in result for x in part] == items
assert all(1 <= len(part) <= size for part in result)
@given tells Hypothesis to call the test many times (100 by default), each time with items drawn from "lists of integers" and size drawn from "integers from 1 to 10". Run it:
items = [0], size = 1
@given(st.lists(st.integers()), st.integers(min_value=1, max_value=10))
def test_chunks_rebuild_original(items, size):
result = chunk(items, size)
> assert [x for part in result for x in part] == items
E assert [] == [0]
E
E Right contains one more item: 0
E Use -v to get more diff
E Failing test case: test_chunks_rebuild_original(
E items=[0],
E size=1, # or any other generated value
E )
Hypothesis found the bug in a fraction of a second, and it didn't report the random 37-element list it probably stumbled on first. It shrank the failure to the simplest possible case: a one-item list. It even notes that size doesn't matter. That's the moment property-based testing earns its keep: the minimal example makes the bug obvious. Fixing the range to range(0, len(items), size) makes the test pass.
Strategies: Describing Your Inputs
The st.* functions are strategies, objects that know how to generate (and shrink) values of a certain shape. You'll use a handful constantly:
| Strategy | Generates |
|---|---|
st.integers(min_value=..., max_value=...) | int, optionally bounded |
st.floats(allow_nan=False, allow_infinity=False) | float; NaN and infinity are included unless excluded |
st.text(min_size=..., max_size=..., alphabet=...) | str, with full Unicode by default |
st.booleans() | True / False |
st.lists(elements, min_size=..., max_size=..., unique=...) | lists of another strategy's values |
st.dictionaries(keys, values) | dicts |
st.tuples(a, b, ...) | fixed-length tuples |
st.sampled_from(seq) | one element of a sequence or enum |
st.one_of(a, b) or a | b | a value from either strategy |
st.none(), st.just(x) | None, or always x |
st.dates(), st.datetimes(), st.uuids(), st.emails() | common value types |
You can peek at what a strategy produces from a REPL with .example():
>>> from hypothesis import strategies as st
>>> st.lists(st.integers(), max_size=5).example()
[-284, 11097, -241]
>>> st.text(max_size=8).example()
'÷/'
Your results will differ, since generation is random. Only use .example() interactively; inside tests, always use @given, which is what enables shrinking and replaying failures.
Transforming Strategies
Strategies compose. .map() transforms generated values and .filter() discards ones that don't fit:
even_numbers = st.integers().map(lambda n: n * 2)
non_blank = st.text().filter(lambda s: s.strip())
Prefer .map() and bounded strategies over .filter(). If a filter rejects most values, Hypothesis wastes time and eventually raises a health-check error telling you so.
Building Objects with st.builds
st.builds calls a class or function with arguments drawn from strategies, which is the easiest way to generate your own data types:
from dataclasses import dataclass
from datetime import date
from hypothesis import strategies as st
@dataclass
class Order:
order_id: int
customer: str
total_cents: int
placed_on: date
orders = st.builds(
Order,
order_id=st.integers(min_value=1),
customer=st.text(min_size=1, max_size=40),
total_cents=st.integers(min_value=0, max_value=10_000_000),
placed_on=st.dates(min_value=date(2020, 1, 1)),
)
If your class has type hints, st.builds(Order) with no keyword arguments infers strategies from the annotations, and st.from_type(Order) does the same. Spelling out the constraints is usually better, though, because real data has rules (no negative totals) that type hints don't express.
Dependent Values with @st.composite
Sometimes one value depends on another, like an index that must be valid for a generated list. @st.composite lets you write a strategy as a function that draws values step by step:
@st.composite
def list_and_index(draw):
items = draw(st.lists(st.integers(), min_size=1))
index = draw(st.integers(min_value=0, max_value=len(items) - 1))
return items, index
@given(list_and_index())
def test_pop_removes_one_item(data):
items, index = data
before = len(items)
value = items[index]
assert items.pop(index) == value
assert len(items) == before - 1
draw pulls a value from any strategy. Because Hypothesis controls every draw, shrinking still works on composite strategies.
Kinds of Properties Worth Testing
The hardest part of property-based testing is thinking of properties. "The output is correct" isn't one you can check without reimplementing the function. Fortunately, a few patterns cover most real code.
Round Trips
If you have an encoder and a decoder, serializer and parser, or save and load, decoding the encoded value must give back the original. This is the single most productive property:
import json
def order_to_json(order: Order) -> str:
return json.dumps({
"order_id": order.order_id,
"customer": order.customer,
"total_cents": order.total_cents,
"placed_on": order.placed_on.isoformat(),
})
def order_from_json(raw: str) -> Order:
data = json.loads(raw)
return Order(
order_id=data["order_id"],
customer=data["customer"],
total_cents=data["total_cents"],
placed_on=date.fromisoformat(data["placed_on"]),
)
@given(orders)
def test_order_json_round_trip(order):
assert order_from_json(order_to_json(order)) == order
Round-trip tests are especially good at finding problems with Unicode, escaping, and boundary values. Here's a run-length encoder with a compact string format, "aaab" becoming "3a1b":
# rle.py
def encode(text: str) -> list[tuple[str, int]]:
runs: list[tuple[str, int]] = []
for char in text:
if runs and runs[-1][0] == char:
runs[-1] = (char, runs[-1][1] + 1)
else:
runs.append((char, 1))
return runs
def to_string(runs: list[tuple[str, int]]) -> str:
return "".join(f"{count}{char}" for char, count in runs)
def from_string(data: str) -> list[tuple[str, int]]:
runs = []
count = ""
for ch in data:
if ch.isdigit():
count += ch
else:
runs.append((ch, int(count)))
count = ""
return runs
@given(st.text())
def test_string_form_round_trips(text):
runs = encode(text)
assert from_string(to_string(runs)) == runs
Hypothesis found two distinct bugs and reported both (as an ExceptionGroup, trimmed here):
ExceptionGroup: Hypothesis found 2 distinct failures. (2 sub-exceptions)
...
ValueError: invalid literal for int() with base 10: '1²1'
Failing test case: test_string_form_round_trips(
text='²:',
)
...
AssertionError: assert [] == [('0', 1)]
Failing test case: test_string_form_round_trips(
text='0',
)
The second failure is a design flaw: a text that contains digits can't be represented unambiguously in this format ("0" encodes to "10", which decodes as a run of ten of... nothing). The first is subtler. str.isdigit() returns True for characters like the superscript ², which int() refuses to parse. Neither is something you'd likely think to write an example for.
The fixes: document that the compact format doesn't support digit characters, check against an explicit set of ASCII digits instead of isdigit(), and tell Hypothesis about the precondition:
DIGITS = frozenset("0123456789")
no_digits = st.text(alphabet=st.characters(exclude_characters=DIGITS))
@given(no_digits)
def test_string_form_round_trips(text):
runs = encode(text)
assert from_string(to_string(runs)) == runs
Invariants
Some facts hold no matter what: sorting preserves length and produces ordered output, a discount never makes a total negative, a balanced tree stays balanced after an insert:
@given(st.lists(st.integers()))
def test_sorted_is_ordered_and_same_length(xs):
result = sorted(xs)
assert all(a <= b for a, b in zip(result, result[1:]))
assert len(result) == len(xs)
Oracles
If there's a simple, slow, obviously correct way to compute the answer, compare your clever, fast implementation against it. A custom sort can be checked against sorted, an optimized search against a linear scan, a new implementation against the legacy one you're replacing.
"Doesn't Crash"
The weakest property is still useful: for any valid input, the function returns without raising an unexpected exception. Pointing @given(st.text()) at a parser with no assertion at all often turns up crashes on the first run.
Controlling the Search
assume: Skipping Invalid Inputs
When a generated value doesn't satisfy a precondition, assume tells Hypothesis to discard it and try another:
from hypothesis import assume
@given(st.integers(), st.integers())
def test_divmod_identity(a, b):
assume(b != 0)
q, r = divmod(a, b)
assert q * b + r == a
Like .filter(), assume is fine for rare exclusions. If you're discarding most inputs, constrain the strategy instead.
@example: Always Test Specific Cases
@example adds explicit inputs that run every time, alongside the generated ones. Use it for known edge cases and for regressions you've fixed:
from hypothesis import example
@given(st.lists(st.integers()))
@example([])
@example([3, 1, 2])
def test_sorted_examples_and_properties(xs):
result = sorted(xs)
assert all(a <= b for a, b in zip(result, result[1:]))
@settings: How Hard to Look
@settings adjusts the search for a single test:
from hypothesis import settings
@settings(max_examples=500, deadline=None)
@given(st.lists(st.integers()), st.integers(min_value=1, max_value=10))
def test_chunk_thorough(items, size):
assert [x for part in chunk(items, size) for x in part] == items
max_examplesis how many inputs to try (default 100). More examples find rarer bugs but take longer.deadlinefails any single example that takes longer than the limit (200 ms by default), which catches accidental performance cliffs. Set it toNonefor tests that are legitimately slow, such as ones that touch a database.
Profiles for Local Runs and CI
Rather than tuning each test, register profiles once in conftest.py and choose one per environment:
# tests/conftest.py
import os
from hypothesis import settings
settings.register_profile("ci", max_examples=1000, deadline=None)
settings.register_profile("dev", max_examples=50)
settings.load_profile(os.getenv("HYPOTHESIS_PROFILE", "dev"))
Locally you get fast feedback with 50 examples per test; in CI you set HYPOTHESIS_PROFILE=ci and search much more thoroughly. The pytest plugin also accepts --hypothesis-profile=ci on the command line. To see what Hypothesis actually did for each test, run with --hypothesis-show-statistics.
The Example Database: Failures That Stay Fixed
When Hypothesis finds a failure, it saves the failing input in a .hypothesis/ directory in your project. The next run replays saved failures first, so a bug you're working on fails immediately and consistently instead of depending on random generation. Hypothesis adds a .gitignore inside the directory, so it stays out of version control by default, which is usually what you want.
Once you've fixed a bug, consider adding the shrunk input as an @example so it's checked forever, on every machine.
Stateful Testing
Some bugs only appear after a particular sequence of operations: push, push, pop, push on a full queue. Hypothesis can generate sequences too. You describe the operations as rules on a state machine, typically alongside a simple model of the expected state, and Hypothesis explores random sequences and shrinks failing ones.
Here's a stack with a capacity limit, tested against a plain list as the model:
# tests/test_stateful.py
from hypothesis import strategies as st
from hypothesis.stateful import RuleBasedStateMachine, invariant, rule
class BoundedStack:
def __init__(self, capacity: int) -> None:
self.capacity = capacity
self._items: list[int] = []
def push(self, item: int) -> bool:
if len(self._items) >= self.capacity:
return False
self._items.append(item)
return True
def pop(self) -> int | None:
return self._items.pop() if self._items else None
def __len__(self) -> int:
return len(self._items)
class BoundedStackMachine(RuleBasedStateMachine):
def __init__(self) -> None:
super().__init__()
self.stack = BoundedStack(capacity=3)
self.model: list[int] = []
@rule(item=st.integers())
def push(self, item):
accepted = self.stack.push(item)
assert accepted == (len(self.model) < 3)
if accepted:
self.model.append(item)
@rule()
def pop(self):
expected = self.model.pop() if self.model else None
assert self.stack.pop() == expected
@invariant()
def sizes_match(self):
assert len(self.stack) == len(self.model) <= 3
TestBoundedStack = BoundedStackMachine.TestCase
Each @rule is an operation Hypothesis may call, with arguments drawn from strategies. @invariant methods run after every step. The last line exposes the machine as a test case that pytest collects.
To see it work, introduce an off-by-one bug by changing >= to > in push. Hypothesis reports the shortest sequence of steps that breaks it:
Failing test case:
state = BoundedStackMachine()
state.sizes_match()
state.push(item=0)
state.sizes_match()
state.push(item=0)
state.sizes_match()
state.push(item=0)
state.sizes_match()
state.push(item=0)
state.teardown()
Four pushes onto a stack of capacity three: the fourth was accepted when it shouldn't have been. Stateful tests are especially valuable for caches, connection pools, and anything with an undo history, where bugs hide in sequences no one writes by hand.
Tips for Using Hypothesis Well
- Start with round trips and "doesn't crash". They're the easiest properties to write and find real bugs quickly.
- Keep example tests too. Properties check general rules; a few concrete examples document intent and are easier to read. They complement each other.
- Constrain strategies to the real domain. If your function only accepts positive amounts, generate positive amounts, and add a separate test for how it rejects negatives.
- Make tests deterministic. Hypothesis needs to re-run an input and get the same result in order to shrink it. Mock out clocks and randomness (see mocking with unittest.mock), and avoid shared state between examples. Function-scoped pytest fixtures are not reset between examples within one test, and Hypothesis raises a health-check error if you use one with
@given. - Read the shrunk example carefully. It's usually the clearest description of the bug you'll get.
Conclusion
Property-based testing changes the question from "does my code work for these inputs?" to "what must always be true, and can anything break it?". Hypothesis does the hard part: generating varied and adversarial inputs from simple strategy descriptions, shrinking failures to a minimal case, remembering them so they reproduce, and even exploring sequences of operations on stateful objects.
You don't need to rewrite your test suite to benefit. Pick one function with a natural round trip or invariant, write a single @given test for it, and see what turns up. Combined with ordinary example-based tests and coverage measurement, it's one of the most effective ways to find the bugs you didn't know to look for.


