Type something to search...
List, Dict, and Set Comprehensions in Python: Writing Cleaner Loops

List, Dict, and Set Comprehensions in Python: Writing Cleaner Loops

A large share of the loops in any Python codebase do the same thing: start with an empty list or dict, loop over something, maybe skip a few items, and append a transformed value. That's four or five lines of ceremony for one idea. Comprehensions express that idea in a single line, and once you can read them fluently, they make code easier to scan, not harder.

They can also be overused. A comprehension with three nested loops and two conditions is harder to follow than the loop it replaced. The skill is knowing the shapes that work well and recognizing when to stop.

This post covers list, dict, and set comprehensions, filtering and conditional expressions, nested loops, generator expressions, scoping rules (including the class-body gotcha), and practical guidelines for when a plain for loop is the better choice.

From Loop to List Comprehension

Here's the pattern comprehensions are designed to replace:

nums = [1, 2, 3, 4, 5, 6]

squares = []
for n in nums:
    squares.append(n * n)

print(squares)
[1, 4, 9, 16, 25, 36]

The same thing as a list comprehension:

squares = [n * n for n in nums]

Read it left to right as "a list of n * n for each n in nums". The general shape is:

[expression for item in iterable]

The expression comes first because it's what you care about: what each element of the new list looks like. The loop clause that follows says where the items come from.

Beyond being shorter, the comprehension makes the intent obvious. When you see [... for ... in ...], you know immediately that a new list is being built and nothing else is happening. A for loop could be doing anything, so you have to read the whole body to find out.

Filtering with if

Add an if clause at the end to keep only some items:

print([n for n in nums if n % 2 == 0])
print([n * n for n in nums if n % 2 == 0])
[2, 4, 6]
[4, 16, 36]

The filter runs first, then the expression is applied to items that pass.

Filter vs Conditional Expression

There are two places an if can go, and they do different things:

# Filter: drop items (if at the end, no else)
evens = [n for n in nums if n % 2 == 0]

# Conditional expression: transform every item (if/else at the front)
labels = ["even" if n % 2 == 0 else "odd" for n in nums]
print(labels)
['odd', 'even', 'odd', 'even', 'odd', 'even']
  • An if after the loop is a filter. It has no else, and the output can be shorter than the input.
  • An if ... else before the loop is an ordinary conditional expression that produces a value for every item. The output has the same length as the input.

Mixing them up is the most common comprehension syntax error. [n for n in nums if n > 2 else 0] is invalid; if you want a value for every item, the if/else goes in front: [n if n > 2 else 0 for n in nums].

Dict Comprehensions

Dict comprehensions use braces and a key: value expression:

prices = {"mug": 12.0, "lamp": 65.5, "pen": 1.25}

with_tax = {name: round(p * 1.2, 2) for name, p in prices.items()}
print(with_tax)
{'mug': 14.4, 'lamp': 78.6, 'pen': 1.5}

Some patterns that come up constantly:

# Filter a dict
print({name: p for name, p in prices.items() if p > 5})

# Invert a dict (values must be unique and hashable)
print({p: name for name, p in prices.items()})

# Build a lookup table from a list
words = ["apple", "Avocado", "banana", "blueberry", "cherry"]
print({w: len(w) for w in words})
{'mug': 12.0, 'lamp': 65.5}
{12.0: 'mug', 65.5: 'lamp', 1.25: 'pen'}
{'apple': 5, 'Avocado': 7, 'banana': 6, 'blueberry': 9, 'cherry': 6}

Indexing records by a field is probably the single most useful dict comprehension in real code:

users = [
    {"name": "ada", "active": True},
    {"name": "bob", "active": False},
    {"name": "cy", "active": True},
]

by_name = {u["name"]: u for u in users}
print(list(by_name))
['ada', 'bob', 'cy']

Now by_name["bob"] is a constant-time lookup instead of a search through the list. If two records share a key, the last one wins, so make sure the field is actually unique.

When you just want to pair two sequences, dict(zip(keys, values)) is simpler than a comprehension. Reach for the comprehension when you need to transform or filter along the way. Dictionaries themselves, including ordering and merging, are covered in Python Dictionaries Deep Dive.

Set Comprehensions

Set comprehensions also use braces, but with a single expression instead of key: value. Duplicates are removed automatically:

first_letters = {w[0].lower() for w in words}
print(sorted(first_letters))
['a', 'b', 'c']

I wrapped the result in sorted() because sets have no defined order; printing the set directly can show the letters in a different order from run to run.

One syntax note: {} on its own is an empty dict, not an empty set. Use set() for an empty set. A comprehension is unambiguous, though: {x for x in ...} is a set and {k: v for ...} is a dict.

Nested Loops

A comprehension can have more than one for clause. They run in the same order as nested for loops would, outermost first:

matrix = [[1, 2, 3], [4, 5, 6]]

flat = [x for row in matrix for x in row]
print(flat)
[1, 2, 3, 4, 5, 6]

The trick to reading this is to imagine the loops written out normally, keeping the clause order:

flat = []
for row in matrix:
    for x in row:
        flat.append(x)

The for clauses appear in the same order in both versions; only the expression moves to the front.

Multiple for clauses also produce combinations, and each can have its own filter:

print([(c, s) for c in "ab" for s in (1, 2)])
print([(x, y) for x in range(3) for y in range(3) if x < y])
[('a', 1), ('a', 2), ('b', 1), ('b', 2)]
[(0, 1), (0, 2), (1, 2)]

For combinations specifically, itertools.product() and itertools.combinations() often say it more clearly; see Mastering itertools in Python.

Nested Comprehensions

A comprehension can also appear inside the expression of another one. That builds nested structures rather than flattening them:

transposed = [[row[i] for row in matrix] for i in range(3)]
print(transposed)
[[1, 4], [2, 5], [3, 6]]

The outer comprehension produces one list per column index i; the inner one collects that column from each row. This is about the limit of what's comfortable to read in one line. (For this particular job, list(zip(*matrix)) gives you tuples with less effort.)

Generator Expressions

Replace the square brackets with parentheses and you get a generator expression. Instead of building the whole list in memory, it yields values one at a time as something consumes them:

gen = (n * n for n in nums)
print(type(gen).__name__)
generator

Generator expressions shine when you pass them straight into a function that consumes an iterable. You can drop the extra parentheses when the generator is the only argument:

print(sum(n * n for n in nums))
print(max(len(w) for w in words))
print(any(w.startswith("b") for w in words))
91
9
True

No intermediate list is ever created. With any() and all() there's a second benefit: they stop as soon as the answer is known, so the generator never computes the remaining values.

Use a list comprehension when you need the list itself (to index it, iterate it twice, or check its length). Use a generator expression when the result is consumed once, especially for large inputs. Note that there's no "tuple comprehension": (x for x in ...) is a generator. Write tuple(x for x in ...) if you want a tuple. Generators in general are covered in What Is a Generator in Python.

Scope: Comprehensions Don't Leak

The loop variable in a comprehension is local to the comprehension. It doesn't overwrite a variable of the same name outside:

x = "outer"
result = [x for x in range(3)]
print(x)
outer

That's different from a regular for loop, where the loop variable remains bound after the loop ends. (Python 2 list comprehensions did leak, which is one reason you may see older advice about avoiding name clashes.)

The Class Body Gotcha

There's one scoping rule that surprises nearly everyone. Inside a class body, a comprehension can't see other class-level names, except in its first (outermost) iterable:

class Config:
    factor = 10
    values = [1, 2, 3]

    doubled = [n * 2 for n in values]  # works: values is the first iterable

    try:
        scaled = [n * factor for n in range(3)]  # fails: factor in the expression
    except NameError as e:
        err = str(e)


print(Config.doubled)
print(Config.err)
[2, 4, 6]
name 'factor' is not defined

Class bodies don't create an enclosing scope that nested functions or comprehensions can see. Only the first iterable is evaluated in the class body itself. Python 3.12 inlined comprehensions for speed (PEP 709) but deliberately kept this behavior. The fix is usually to compute the value outside the class, or in a method or classmethod. Python Scope Explained covers the general rules behind this.

Late Binding in Lambdas

A related trap: lambdas created inside a comprehension capture the variable, not its current value.

funcs = [lambda: i for i in range(3)]
print([f() for f in funcs])

funcs = [lambda i=i: i for i in range(3)]
print([f() for f in funcs])
[2, 2, 2]
[0, 1, 2]

All three lambdas in the first version look up i when they're called, after the comprehension has finished and i is 2. Binding it as a default argument captures the value at creation time.

Practical Patterns

A few everyday uses that show where comprehensions earn their place.

Cleaning input, keeping only non-blank lines:

raw = [" a ", "  ", "b  "]
print([s.strip() for s in raw if s.strip()])
['a', 'b']

This calls strip() twice per item. That's fine for short strings; the walrus operator can avoid the duplicate call, with trade-offs discussed in The Walrus Operator in Python.

Parsing simple config lines by feeding a generator into dict():

lines = ["# comment", "key=val", "", "x=1"]

cfg = dict(line.split("=", 1) for line in lines if line and not line.startswith("#"))
print(cfg)
{'key': 'val', 'x': '1'}

Extracting a field from a list of records:

active_names = [u["name"] for u in users if u["active"]]
print(active_names)
['ada', 'cy']

When Not to Use a Comprehension

Comprehensions are for building a collection. They're the wrong tool when:

  • You want side effects. [print(x) for x in items] builds a list of None values you throw away. Use a for loop when the point is to do something, not to produce something.
  • The logic needs statements. Comprehensions only hold expressions. If you need try/except, multiple steps, logging, or early exit with break, write the loop.
  • It no longer fits in your head. A reasonable limit is two for clauses and one condition. Beyond that, a loop with descriptive variable names, or a helper function used inside a simpler comprehension, is easier to read.
  • A built-in already does it. list(range(10)) beats [i for i in range(10)], dict(zip(a, b)) beats the comprehension equivalent, and list(map(str.upper, items)) is fine when the function already exists.

Long comprehensions can also be split over several lines, which helps a lot:

report = [
    f"{name}: {price:.2f}"
    for name, price in prices.items()
    if price > 5
]

Put the expression, each for, and each if on its own line, and a comprehension that would have been a wall of text becomes readable again.

Quick Reference

KindSyntaxResult
List[expr for x in it]list
List with filter[expr for x in it if cond]list
Transform every item[a if cond else b for x in it]list
Dict{k: v for x in it}dict
Set{expr for x in it}set
Generator(expr for x in it)generator, lazy
Flatten[y for x in it for y in x]list

Conclusion

Comprehensions replace the "empty collection, loop, append" pattern with a single expression that states what the result looks like. List comprehensions build lists, dict comprehensions build mappings, set comprehensions deduplicate, and generator expressions do the same work lazily when you only need to consume the values once.

Keep the filter (if at the end) and conditional expression (if/else at the front) straight, read multiple for clauses in the same order as nested loops, and watch for the class-body scoping rule. Most of all, stop when a comprehension gets crowded. The goal is cleaner loops, and sometimes the cleanest loop is a plain for.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading