
List, Dict, and Set Comprehensions in Python: Writing Cleaner Loops
A large share of the loops in any Python codebase do the same thing: start with an empty list or dict, loop over something, maybe skip a few items, and append a transformed value. That's four or five lines of ceremony for one idea. Comprehensions express that idea in a single line, and once you can read them fluently, they make code easier to scan, not harder.
They can also be overused. A comprehension with three nested loops and two conditions is harder to follow than the loop it replaced. The skill is knowing the shapes that work well and recognizing when to stop.
This post covers list, dict, and set comprehensions, filtering and conditional expressions, nested loops, generator expressions, scoping rules (including the class-body gotcha), and practical guidelines for when a plain for loop is the better choice.
From Loop to List Comprehension
Here's the pattern comprehensions are designed to replace:
nums = [1, 2, 3, 4, 5, 6]
squares = []
for n in nums:
squares.append(n * n)
print(squares)
[1, 4, 9, 16, 25, 36]
The same thing as a list comprehension:
squares = [n * n for n in nums]
Read it left to right as "a list of n * n for each n in nums". The general shape is:
[expression for item in iterable]
The expression comes first because it's what you care about: what each element of the new list looks like. The loop clause that follows says where the items come from.
Beyond being shorter, the comprehension makes the intent obvious. When you see [... for ... in ...], you know immediately that a new list is being built and nothing else is happening. A for loop could be doing anything, so you have to read the whole body to find out.
Filtering with if
Add an if clause at the end to keep only some items:
print([n for n in nums if n % 2 == 0])
print([n * n for n in nums if n % 2 == 0])
[2, 4, 6]
[4, 16, 36]
The filter runs first, then the expression is applied to items that pass.
Filter vs Conditional Expression
There are two places an if can go, and they do different things:
# Filter: drop items (if at the end, no else)
evens = [n for n in nums if n % 2 == 0]
# Conditional expression: transform every item (if/else at the front)
labels = ["even" if n % 2 == 0 else "odd" for n in nums]
print(labels)
['odd', 'even', 'odd', 'even', 'odd', 'even']
- An
ifafter the loop is a filter. It has noelse, and the output can be shorter than the input. - An
if ... elsebefore the loop is an ordinary conditional expression that produces a value for every item. The output has the same length as the input.
Mixing them up is the most common comprehension syntax error. [n for n in nums if n > 2 else 0] is invalid; if you want a value for every item, the if/else goes in front: [n if n > 2 else 0 for n in nums].
Dict Comprehensions
Dict comprehensions use braces and a key: value expression:
prices = {"mug": 12.0, "lamp": 65.5, "pen": 1.25}
with_tax = {name: round(p * 1.2, 2) for name, p in prices.items()}
print(with_tax)
{'mug': 14.4, 'lamp': 78.6, 'pen': 1.5}
Some patterns that come up constantly:
# Filter a dict
print({name: p for name, p in prices.items() if p > 5})
# Invert a dict (values must be unique and hashable)
print({p: name for name, p in prices.items()})
# Build a lookup table from a list
words = ["apple", "Avocado", "banana", "blueberry", "cherry"]
print({w: len(w) for w in words})
{'mug': 12.0, 'lamp': 65.5}
{12.0: 'mug', 65.5: 'lamp', 1.25: 'pen'}
{'apple': 5, 'Avocado': 7, 'banana': 6, 'blueberry': 9, 'cherry': 6}
Indexing records by a field is probably the single most useful dict comprehension in real code:
users = [
{"name": "ada", "active": True},
{"name": "bob", "active": False},
{"name": "cy", "active": True},
]
by_name = {u["name"]: u for u in users}
print(list(by_name))
['ada', 'bob', 'cy']
Now by_name["bob"] is a constant-time lookup instead of a search through the list. If two records share a key, the last one wins, so make sure the field is actually unique.
When you just want to pair two sequences, dict(zip(keys, values)) is simpler than a comprehension. Reach for the comprehension when you need to transform or filter along the way. Dictionaries themselves, including ordering and merging, are covered in Python Dictionaries Deep Dive.
Set Comprehensions
Set comprehensions also use braces, but with a single expression instead of key: value. Duplicates are removed automatically:
first_letters = {w[0].lower() for w in words}
print(sorted(first_letters))
['a', 'b', 'c']
I wrapped the result in sorted() because sets have no defined order; printing the set directly can show the letters in a different order from run to run.
One syntax note: {} on its own is an empty dict, not an empty set. Use set() for an empty set. A comprehension is unambiguous, though: {x for x in ...} is a set and {k: v for ...} is a dict.
Nested Loops
A comprehension can have more than one for clause. They run in the same order as nested for loops would, outermost first:
matrix = [[1, 2, 3], [4, 5, 6]]
flat = [x for row in matrix for x in row]
print(flat)
[1, 2, 3, 4, 5, 6]
The trick to reading this is to imagine the loops written out normally, keeping the clause order:
flat = []
for row in matrix:
for x in row:
flat.append(x)
The for clauses appear in the same order in both versions; only the expression moves to the front.
Multiple for clauses also produce combinations, and each can have its own filter:
print([(c, s) for c in "ab" for s in (1, 2)])
print([(x, y) for x in range(3) for y in range(3) if x < y])
[('a', 1), ('a', 2), ('b', 1), ('b', 2)]
[(0, 1), (0, 2), (1, 2)]
For combinations specifically, itertools.product() and itertools.combinations() often say it more clearly; see Mastering itertools in Python.
Nested Comprehensions
A comprehension can also appear inside the expression of another one. That builds nested structures rather than flattening them:
transposed = [[row[i] for row in matrix] for i in range(3)]
print(transposed)
[[1, 4], [2, 5], [3, 6]]
The outer comprehension produces one list per column index i; the inner one collects that column from each row. This is about the limit of what's comfortable to read in one line. (For this particular job, list(zip(*matrix)) gives you tuples with less effort.)
Generator Expressions
Replace the square brackets with parentheses and you get a generator expression. Instead of building the whole list in memory, it yields values one at a time as something consumes them:
gen = (n * n for n in nums)
print(type(gen).__name__)
generator
Generator expressions shine when you pass them straight into a function that consumes an iterable. You can drop the extra parentheses when the generator is the only argument:
print(sum(n * n for n in nums))
print(max(len(w) for w in words))
print(any(w.startswith("b") for w in words))
91
9
True
No intermediate list is ever created. With any() and all() there's a second benefit: they stop as soon as the answer is known, so the generator never computes the remaining values.
Use a list comprehension when you need the list itself (to index it, iterate it twice, or check its length). Use a generator expression when the result is consumed once, especially for large inputs. Note that there's no "tuple comprehension": (x for x in ...) is a generator. Write tuple(x for x in ...) if you want a tuple. Generators in general are covered in What Is a Generator in Python.
Scope: Comprehensions Don't Leak
The loop variable in a comprehension is local to the comprehension. It doesn't overwrite a variable of the same name outside:
x = "outer"
result = [x for x in range(3)]
print(x)
outer
That's different from a regular for loop, where the loop variable remains bound after the loop ends. (Python 2 list comprehensions did leak, which is one reason you may see older advice about avoiding name clashes.)
The Class Body Gotcha
There's one scoping rule that surprises nearly everyone. Inside a class body, a comprehension can't see other class-level names, except in its first (outermost) iterable:
class Config:
factor = 10
values = [1, 2, 3]
doubled = [n * 2 for n in values] # works: values is the first iterable
try:
scaled = [n * factor for n in range(3)] # fails: factor in the expression
except NameError as e:
err = str(e)
print(Config.doubled)
print(Config.err)
[2, 4, 6]
name 'factor' is not defined
Class bodies don't create an enclosing scope that nested functions or comprehensions can see. Only the first iterable is evaluated in the class body itself. Python 3.12 inlined comprehensions for speed (PEP 709) but deliberately kept this behavior. The fix is usually to compute the value outside the class, or in a method or classmethod. Python Scope Explained covers the general rules behind this.
Late Binding in Lambdas
A related trap: lambdas created inside a comprehension capture the variable, not its current value.
funcs = [lambda: i for i in range(3)]
print([f() for f in funcs])
funcs = [lambda i=i: i for i in range(3)]
print([f() for f in funcs])
[2, 2, 2]
[0, 1, 2]
All three lambdas in the first version look up i when they're called, after the comprehension has finished and i is 2. Binding it as a default argument captures the value at creation time.
Practical Patterns
A few everyday uses that show where comprehensions earn their place.
Cleaning input, keeping only non-blank lines:
raw = [" a ", " ", "b "]
print([s.strip() for s in raw if s.strip()])
['a', 'b']
This calls strip() twice per item. That's fine for short strings; the walrus operator can avoid the duplicate call, with trade-offs discussed in The Walrus Operator in Python.
Parsing simple config lines by feeding a generator into dict():
lines = ["# comment", "key=val", "", "x=1"]
cfg = dict(line.split("=", 1) for line in lines if line and not line.startswith("#"))
print(cfg)
{'key': 'val', 'x': '1'}
Extracting a field from a list of records:
active_names = [u["name"] for u in users if u["active"]]
print(active_names)
['ada', 'cy']
When Not to Use a Comprehension
Comprehensions are for building a collection. They're the wrong tool when:
- You want side effects.
[print(x) for x in items]builds a list ofNonevalues you throw away. Use aforloop when the point is to do something, not to produce something. - The logic needs statements. Comprehensions only hold expressions. If you need
try/except, multiple steps, logging, or early exit withbreak, write the loop. - It no longer fits in your head. A reasonable limit is two
forclauses and one condition. Beyond that, a loop with descriptive variable names, or a helper function used inside a simpler comprehension, is easier to read. - A built-in already does it.
list(range(10))beats[i for i in range(10)],dict(zip(a, b))beats the comprehension equivalent, andlist(map(str.upper, items))is fine when the function already exists.
Long comprehensions can also be split over several lines, which helps a lot:
report = [
f"{name}: {price:.2f}"
for name, price in prices.items()
if price > 5
]
Put the expression, each for, and each if on its own line, and a comprehension that would have been a wall of text becomes readable again.
Quick Reference
| Kind | Syntax | Result |
|---|---|---|
| List | [expr for x in it] | list |
| List with filter | [expr for x in it if cond] | list |
| Transform every item | [a if cond else b for x in it] | list |
| Dict | {k: v for x in it} | dict |
| Set | {expr for x in it} | set |
| Generator | (expr for x in it) | generator, lazy |
| Flatten | [y for x in it for y in x] | list |
Conclusion
Comprehensions replace the "empty collection, loop, append" pattern with a single expression that states what the result looks like. List comprehensions build lists, dict comprehensions build mappings, set comprehensions deduplicate, and generator expressions do the same work lazily when you only need to consume the values once.
Keep the filter (if at the end) and conditional expression (if/else at the front) straight, read multiple for clauses in the same order as nested loops, and watch for the class-body scoping rule. Most of all, stop when a comprehension gets crowded. The goal is cleaner loops, and sometimes the cleanest loop is a plain for.


