
Mastering itertools in Python: Efficient Iteration Recipes
Most Python loops follow a handful of shapes: flatten some nested lists, take the first N items, walk through pairs of neighbors, split work into chunks, group consecutive rows by a key, try every combination of a few options. You can write each of these by hand with indexes and temporary lists, and plenty of codebases do. The itertools module gives you each one as a small, fast, lazy building block instead.
"Lazy" is the important part. Every function in itertools returns an iterator that produces values on demand. Nothing is computed until you ask for it, and nothing is stored that doesn't need to be. That means you can chain them over a 10 GB log file or an infinite stream and memory stays flat.
This post walks through the functions you'll actually use, grouped by what they do, then covers the gotchas and a set of practical recipes. Everything here runs on Python 3.13, and I'll flag the functions that were added recently.
Why Lazy Iteration Matters
A quick refresher. An iterator is an object you call next() on until it raises StopIteration. A for loop does that for you. Generators are one way to build iterators; itertools functions are another.
Compare these two:
squares_list = [n * n for n in range(10_000_000)] # builds 10 million ints up front
squares_iter = (n * n for n in range(10_000_000)) # builds nothing yet
The list takes hundreds of megabytes before you've used a single value. The generator expression costs a few hundred bytes and computes each square as the loop asks for it. itertools lets you keep working in that lazy style while still doing non-trivial things like slicing, grouping, and combining.
All examples below assume:
from itertools import (
accumulate, batched, chain, combinations, combinations_with_replacement,
compress, count, cycle, dropwhile, filterfalse, groupby, islice,
pairwise, permutations, product, repeat, starmap, takewhile, tee,
zip_longest,
)
Infinite Iterators: count, cycle, repeat
These never stop on their own, so you always pair them with something that does (zip, islice, takewhile, or a break).
print(list(islice(count(10, 5), 4)))
print(list(zip(count(1), ["a", "b", "c"])))
print(list(islice(cycle("RGB"), 7)))
print(list(repeat("x", 3)))
[10, 15, 20, 25]
[(1, 'a'), (2, 'b'), (3, 'c')]
['R', 'G', 'B', 'R', 'G', 'B', 'R']
['x', 'x', 'x']
count(start, step)is an endlessrange. Handy for generating IDs or numbering items from an unknown-length stream.cycle(iterable)loops over the input forever. Good for round-robin assignment (alternating row colors, rotating through a pool of workers).repeat(value, times)yields the same object repeatedly. Its most common use is feeding a constant intomap:map(pow, range(5), repeat(2))gives0, 1, 4, 9, 16.
Slicing and Joining: islice and chain
islice
You can't slice an iterator with [2:7], because it has no length and no random access. islice does the same job lazily:
print(list(islice("abcdefgh", 2, 7, 2))) # ['c', 'e', 'g']
It takes (iterable, stop) or (iterable, start, stop[, step]). Negative values aren't allowed, since that would require knowing where the end is.
The most useful form is islice(it, n): "give me the first n items". Combined with a file object, it reads only what it needs:
from itertools import islice
with open("access.log", encoding="utf-8") as f:
first_ten = list(islice(f, 10))
chain and chain.from_iterable
chain stitches iterables together end to end without building a combined list:
print(list(chain([1, 2], (3, 4), "ab")))
nested = [[1, 2], [3], [], [4, 5]]
print(list(chain.from_iterable(nested)))
[1, 2, 3, 4, 'a', 'b']
[1, 2, 3, 4, 5]
chain.from_iterable is the standard way to flatten one level of nesting. It accepts a single iterable of iterables, which can itself be lazy, so it works when you're reading many files and want one stream of lines:
from itertools import chain
from pathlib import Path
all_lines = chain.from_iterable(p.open(encoding="utf-8") for p in Path("logs").glob("*.log"))
(That example leaves the file handles for the garbage collector to close. In a long-running process, wrap it in a generator function that uses with.)
Filtering: takewhile, dropwhile, filterfalse, compress
data = [1, 3, 6, 2, 1]
print(list(takewhile(lambda x: x < 5, data)))
print(list(dropwhile(lambda x: x < 5, data)))
print(list(filterfalse(str.isdigit, ["a", "1", "b", "22"])))
print(list(compress("ABCDEF", [1, 0, 1, 0, 1, 1])))
[1, 3]
[6, 2, 1]
['a', 'b']
['A', 'C', 'E', 'F']
The key difference from filter: takewhile stops at the first failure and never looks at the rest, and dropwhile skips items only until the first failure and then yields everything after it, including the 2 and 1 that would have passed the test. That makes them ideal for sorted or time-ordered data, like "lines until the first blank line" (an HTTP header block) or "skip the comment header at the top of a file".
filterfalse is the mirror image of filter. compress keeps items where a parallel sequence of selectors is truthy, which is handy when the mask comes from somewhere else (a list of checkboxes, a column of booleans).
Running Totals and Neighbors: accumulate, pairwise, batched
accumulate
accumulate yields running results. By default it adds, but you can pass any two-argument function:
import operator
print(list(accumulate([1, 2, 3, 4, 5])))
print(list(accumulate([3, 1, 4, 1, 5, 9, 2], max)))
print(list(accumulate([1, 2, 3, 4], operator.mul)))
print(list(accumulate([100, -20, 50], initial=1000)))
[1, 3, 6, 10, 15]
[3, 3, 4, 4, 5, 9, 9]
[1, 2, 6, 24]
[1000, 1100, 1080, 1130]
Running totals, running maximums (the highest price so far), factorials, and account balances from a list of transactions all fall out of this one function. The initial value is yielded first, so the output is one item longer than the input.
pairwise (Python 3.10+)
pairwise yields overlapping pairs of neighbors. It replaces the zip(xs, xs[1:]) idiom, and unlike that idiom it works on any iterator, not just sequences:
readings = [10, 12, 11, 15, 20]
print([b - a for a, b in pairwise(readings)]) # [2, -1, 4, 5]
Use it for deltas between measurements, checking that a list is sorted (all(a <= b for a, b in pairwise(xs))), or building edges from a path of nodes.
batched (Python 3.12+)
batched splits an iterable into tuples of size n. The last batch may be shorter:
print(list(batched("ABCDEFG", 3)))
[('A', 'B', 'C'), ('D', 'E', 'F'), ('G',)]
This is the one I reach for most. Bulk inserts, API calls that accept 100 IDs at a time, and splitting work across processes are all batching problems:
from itertools import batched
rows = ({"id": i} for i in range(7))
for batch in batched(rows, 3):
print("insert", [r["id"] for r in batch])
insert [0, 1, 2]
insert [3, 4, 5]
insert [6]
Python 3.13 added a strict flag that raises if the final batch is incomplete, for cases where a short batch means bad input:
list(batched("ABCDEFG", 3, strict=True))
# ValueError: batched(): incomplete batch
Grouping: groupby
groupby groups consecutive items that share a key. That word "consecutive" is where almost everyone gets bitten the first time:
data = ["apple", "avocado", "banana", "blueberry", "cherry", "apricot"]
for key, group in groupby(data, key=lambda s: s[0]):
print(key, list(group))
a ['apple', 'avocado']
b ['banana', 'blueberry']
c ['cherry']
a ['apricot']
"apricot" ends up in its own group because it isn't next to the other a words. If you want one group per key, sort by the same key first:
for key, group in groupby(sorted(data), key=lambda s: s[0]):
print(key, list(group))
a ['apple', 'apricot', 'avocado']
b ['banana', 'blueberry']
c ['cherry']
That behavior isn't a bug. It's what makes groupby streamable: it never has to hold more than one group in memory, so it can process a sorted CSV with millions of rows. If your data isn't sorted and fits in memory, a defaultdict(list) (covered in the collections module post) is usually simpler.
Run-length encoding is a nice one-liner with groupby:
print([(k, len(list(g))) for k, g in groupby("aaabccdddd")])
# [('a', 3), ('b', 1), ('c', 2), ('d', 4)]
The Shared-Iterator Gotcha
Each group is a view into the same underlying iterator. As soon as groupby advances to the next key, the previous group is gone:
groups = dict(groupby("aaabbc"))
print({k: list(v) for k, v in groups.items()})
{'a': [], 'b': [], 'c': []}
Every group is empty because dict() consumed the whole groupby before anyone read the groups. Materialize each group as you go:
groups = {k: list(v) for k, v in groupby("aaabbc")}
# {'a': ['a', 'a', 'a'], 'b': ['b', 'b'], 'c': ['c']}
Combinatorics: product, permutations, combinations
These replace nested loops and hand-written recursion.
print(list(product("AB", [1, 2])))
print(list(permutations("ABC", 2)))
print(list(combinations("ABCD", 2)))
print(list(combinations_with_replacement("AB", 2)))
[('A', 1), ('A', 2), ('B', 1), ('B', 2)]
[('A', 'B'), ('A', 'C'), ('B', 'A'), ('B', 'C'), ('C', 'A'), ('C', 'B')]
[('A', 'B'), ('A', 'C'), ('A', 'D'), ('B', 'C'), ('B', 'D'), ('C', 'D')]
[('A', 'A'), ('A', 'B'), ('B', 'B')]
| Function | Order matters? | Repeats allowed? | Typical use |
|---|---|---|---|
product(a, b, ...) | Yes | One item from each input | Grids of options, nested loops |
product(xs, repeat=n) | Yes | Yes | All bit strings, PIN codes |
permutations(xs, r) | Yes | No | Orderings, seating plans |
combinations(xs, r) | No | No | Pairs to compare, team picks |
combinations_with_replacement(xs, r) | No | Yes | Coin combos, dice totals |
product is the one you'll use most in application code. Any time you have two or more nested for loops over independent lists, it flattens them:
sizes = ["S", "M"]
colors = ["red", "blue"]
print([f"{s}-{c}" for s, c in product(sizes, colors)])
# ['S-red', 'S-blue', 'M-red', 'M-blue']
It's also great for test matrices: product(browsers, viewports, locales) gives every configuration to check.
A warning about size: these grow explosively. permutations(range(12)) has 479 million results. They're lazy, so creating the iterator is free, but looping over it is not. If you only need the count, use math.comb(n, k) and math.perm(n, k) instead of generating anything. math.comb(52, 5) is 2598960, the number of five-card poker hands.
Smaller Tools: starmap, zip_longest, tee
print(list(starmap(pow, [(2, 5), (3, 2), (10, 3)])))
print(list(zip_longest("ABC", "x", fillvalue="-")))
[32, 9, 1000]
[('A', 'x'), ('B', '-'), ('C', '-')]
starmap(f, pairs)callsf(*pair)for each pair. Use it when your data is already tuples of arguments.zip_longestkeeps going until the longest input is exhausted, filling gaps withfillvalue. Regularzipstops at the shortest. (If the inputs should always be the same length, usezip(..., strict=True)so a mismatch raises instead.)tee(it, n)splits one iterator intonindependent ones. It buffers items that one copy has seen and another hasn't, so if one copy runs far ahead, memory grows. If you'll read all of one before starting the other,list(it)is simpler and faster.
Iterators Are Single-Use
This applies to everything in this post. An iterator can only be consumed once:
r = islice([1, 2, 3], 2)
print(list(r), list(r)) # [1, 2] []
And islice on a shared iterator advances that iterator:
it = iter([1, 2, 3, 4, 5])
print(list(islice(it, 2)), list(it)) # [1, 2] [3, 4, 5]
That second behavior is often exactly what you want, for example reading a header with islice(f, 3) and then looping over the rest of the same file. But if you call len() on something, index into it, or loop over it twice, convert it to a list first.
Recipes Worth Copying
The itertools documentation ends with a list of recipes built from these primitives. Several are worth keeping in a utils.py. Here are the ones I use most, with type hints added.
take and first_true
from collections.abc import Callable, Iterable
from itertools import islice
from typing import Any
def take[T](n: int, iterable: Iterable[T]) -> list[T]:
"""Return the first n items as a list."""
return list(islice(iterable, n))
def first_true(iterable: Iterable[Any], default: Any = None,
pred: Callable[[Any], bool] | None = None) -> Any:
"""Return the first item for which pred is true (or the first truthy item)."""
return next(filter(pred, iterable), default)
print(first_true([0, "", 7, 9])) # 7
print(first_true([1, 3, 8, 10], pred=lambda x: x % 2 == 0)) # 8
next(iterator, default) is the idiom underneath: it returns default instead of raising StopIteration when nothing matches. (The def take[T] syntax is the type parameter syntax from Python 3.12.)
sliding_window
pairwise gives windows of two; this generalizes it to any size by using a bounded deque:
from collections import deque
from collections.abc import Iterable, Iterator
from itertools import islice
def sliding_window[T](iterable: Iterable[T], n: int) -> Iterator[tuple[T, ...]]:
iterator = iter(iterable)
window = deque(islice(iterator, n - 1), maxlen=n)
for x in iterator:
window.append(x)
yield tuple(window)
print(list(sliding_window([1, 2, 3, 4, 5], 3)))
# [(1, 2, 3), (2, 3, 4), (3, 4, 5)]
unique_everseen
Remove duplicates while preserving order, optionally comparing by a key:
from collections.abc import Callable, Hashable, Iterable, Iterator
def unique_everseen[T](iterable: Iterable[T],
key: Callable[[T], Hashable] | None = None) -> Iterator[T]:
seen: set[Hashable] = set()
for element in iterable:
k = element if key is None else key(element)
if k not in seen:
seen.add(k)
yield element
print(list(unique_everseen("AAAABBBCCDAABBB"))) # ['A', 'B', 'C', 'D']
print(list(unique_everseen(["Apple", "apple", "Banana"], key=str.lower))) # ['Apple', 'Banana']
For hashable items without a key, list(dict.fromkeys(items)) does the same thing in one line, but the generator version works on streams and supports a key.
partition
Split one iterable into items that fail and pass a test, in a single pass over the source:
from itertools import filterfalse, tee
def partition(pred, iterable):
t1, t2 = tee(iterable)
return filterfalse(pred, t1), filter(pred, t2)
evens, odds = partition(lambda x: x % 2, range(10))
print(list(evens), list(odds))
# [0, 2, 4, 6, 8] [1, 3, 5, 7, 9]
roundrobin
Interleave several iterables, skipping ones that run out:
from itertools import cycle, islice
def roundrobin(*iterables):
iterators = map(iter, iterables)
for num_active in range(len(iterables), 0, -1):
iterators = cycle(islice(iterators, num_active))
yield from map(next, iterators)
print(list(roundrobin("ABC", "D", "EF")))
# ['A', 'D', 'E', 'B', 'F', 'C']
This one is clever enough that I wouldn't write it from memory. It's a good example of why the docs recipes are worth bookmarking. If you'd rather install them than copy them, the third-party more-itertools package ships all the official recipes and many more.
Building a Lazy Pipeline
The real payoff comes from composing these. Here's a pipeline that reads a large log, skips the comment header, keeps only error lines, and sends them to an API in batches of 50, without ever holding more than one batch in memory:
from collections.abc import Iterator
from itertools import batched, dropwhile
def error_lines(path: str) -> Iterator[str]:
with open(path, encoding="utf-8") as f:
body = dropwhile(lambda line: line.startswith("#"), f)
for line in body:
if " ERROR " in line:
yield line.rstrip("\n")
def ship(path: str) -> None:
for batch in batched(error_lines(path), 50):
send_to_api(list(batch)) # your HTTP call here
Each stage pulls one line at a time from the stage before it. The file can be any size. And each piece is a small, named, testable function rather than one long loop full of flags and counters.
Conclusion
itertools is a toolkit for describing what you want from a stream of data rather than how to index your way through it. The functions to remember are islice and chain for slicing and joining, pairwise and batched for neighbors and chunks, accumulate for running results, groupby for consecutive groups (sort first), and product for replacing nested loops.
Keep two rules in mind: iterators are single-use, and groupby only groups adjacent items. With those, you can replace a lot of index juggling with short pipelines that use constant memory and read clearly.


