Type something to search...
Iterators and the Iterator Protocol in Python

Iterators and the Iterator Protocol in Python

A for loop in Python works on lists, strings, dictionaries, files, database cursors, range objects, and anything else that knows how to hand out its items one at a time. It doesn't index into them, and it doesn't need to know how many items there are. All it relies on is a two-method contract called the iterator protocol.

Knowing that protocol pays off in two ways. It explains behaviour that otherwise looks like a bug, such as a map() result that's empty the second time you loop over it. And it lets you make your own classes work with for, list(), sum(), unpacking, and every other tool that consumes iterables.

This post covers the difference between an iterable and an iterator, what iter() and next() do, how a for loop is built from them, how to write iterator and iterable classes, the two-argument form of iter(), and the exhaustion pitfalls to watch for.

Iterables vs Iterators

These two words sound interchangeable, but they mean different things:

  • An iterable is anything you can loop over. It has an __iter__ method that returns an iterator. Lists, tuples, strings, dicts, sets, and range objects are iterables.
  • An iterator is the object that does the looping. It has a __next__ method that returns the next item, or raises StopIteration when there are none left. It also has an __iter__ method that returns itself.

The built-in iter() asks an iterable for an iterator, and next() asks an iterator for its next item:

nums = [10, 20, 30]
it = iter(nums)
print(it)
print(next(it), next(it), next(it))
next(it)
<list_iterator object at 0x1048348b0>
10 20 30
StopIteration

The list itself doesn't track a position. The list_iterator does. That's the key separation: an iterable is a source of items, an iterator is a cursor over that source.

You can't call next() on a list directly:

next(nums)
# TypeError: 'list' object is not an iterator

And collections.abc confirms the distinction:

from collections.abc import Iterable, Iterator

print(isinstance(nums, Iterable), isinstance(nums, Iterator), isinstance(it, Iterator))
# True False True

Every iterator is also an iterable (its __iter__ returns itself), but not every iterable is an iterator.

A Default for next()

next() accepts a second argument that's returned instead of raising StopIteration:

print(next(it, "done"))  # done

That makes a clean "first item or default" helper:

def first(iterable, default=None):
    return next(iter(iterable), default)


print(first([]), first("xyz"), first(x for x in range(10) if x > 6))
# None x 7

How a for Loop Actually Works

A for loop is the protocol plus exception handling. This loop:

for item in nums:
    print(item)

does essentially this:

iterator = iter(nums)
while True:
    try:
        item = next(iterator)
    except StopIteration:
        break
    print(item)

Python calls iter() once at the start, then next() repeatedly until StopIteration signals the end. StopIteration isn't an error here; it's the normal way an iterator says "I'm finished", and the loop catches it for you.

The same protocol powers far more than for: list(), tuple(), set(), dict(), sum(), min(), max(), sorted(), any(), all(), str.join(), zip(), enumerate(), unpacking like a, b, *rest = thing, comprehensions, and in (when the object has no faster __contains__). Implement the protocol once and your object works with all of them.

Which Built-ins Are Iterators?

Many built-ins return iterators rather than lists. A quick way to tell: an iterator's iter() returns the object itself.

for obj in (range(3), {"a": 1}, "hi", open(__file__), enumerate([]), reversed([1])):
    print(type(obj).__name__, iter(obj) is obj)
range False
dict False
str False
TextIOWrapper True
enumerate True
list_reverseiterator True

range, dict, and str are reusable iterables. File objects, enumerate, reversed, zip, map, filter, generators, and most of itertools are iterators. That distinction matters, as the next section shows.

Iterators Are Single-Use

An iterator moves forward only and can't be reset. Once it's exhausted, it stays exhausted:

squares = map(lambda n: n * n, [1, 2, 3])
print(list(squares), list(squares))
[1, 4, 9] []

The second list() gets nothing, with no error. This is the source of a whole family of bugs. Here's a function that works with a list and silently breaks with an iterator:

def total_and_count(numbers):
    total = sum(numbers)
    count = len(list(numbers))
    return total, count


print(total_and_count([1, 2, 3]))        # (6, 3)
print(total_and_count(iter([1, 2, 3])))  # (6, 0)

sum() consumed the iterator, so list(numbers) saw nothing. If a function needs to pass over its input more than once, either convert it once at the top (numbers = list(numbers)) or document that it requires a re-iterable collection, such as a collections.abc.Collection.

Membership tests on an iterator also consume it, up to and including the match:

it = iter([10, 20, 30])
print(20 in it, list(it))  # True [30]

Writing an Iterator Class

To make your own iterator, implement __next__ (return the next item or raise StopIteration) and __iter__ (return self):

class Countdown:
    """An iterator that counts down from start to 1."""

    def __init__(self, start: int) -> None:
        self.current = start

    def __iter__(self) -> "Countdown":
        return self

    def __next__(self) -> int:
        if self.current <= 0:
            raise StopIteration
        value = self.current
        self.current -= 1
        return value


for n in Countdown(2):
    print(n)

cd = Countdown(3)
print(list(cd))
print(list(cd))
2
1
[3, 2, 1]
[]

It works with for and list(), and like every iterator it's exhausted after one pass. The state (self.current) lives on the iterator, so there's no way to restart it except creating a new one.

Two rules to follow:

  • Once __next__ has raised StopIteration, it should keep raising it on every later call. Code that loops over an iterator expects "finished" to be permanent.
  • __iter__ must return self. Without it, your iterator can't be used directly in a for loop, because the loop calls iter() first.

Iterables That Can Be Looped Over Many Times

Usually you don't want your collection to be single-use. A list can be looped over any number of times because each for gets a new list_iterator. To get the same behaviour, separate the iterable from its iterator: __iter__ returns a fresh iterator object each time.

class NumberRange:
    """An iterable: each for-loop gets a fresh iterator."""

    def __init__(self, start: int, stop: int) -> None:
        self.start = start
        self.stop = stop

    def __iter__(self) -> "NumberRangeIterator":
        return NumberRangeIterator(self.start, self.stop)


class NumberRangeIterator:
    def __init__(self, current: int, stop: int) -> None:
        self.current = current
        self.stop = stop

    def __iter__(self) -> "NumberRangeIterator":
        return self

    def __next__(self) -> int:
        if self.current >= self.stop:
            raise StopIteration
        value = self.current
        self.current += 1
        return value


r = NumberRange(1, 4)
print(list(r), list(r))
print([(a, b) for a in r for b in r if a < b])
[1, 2, 3] [1, 2, 3]
[(1, 2), (1, 3), (2, 3)]

The nested comprehension loops over r twice at the same time. That works because each loop has its own independent iterator with its own position. If NumberRange were its own iterator, the inner loop would exhaust the shared state and the result would be wrong.

The Shortcut: Make __iter__ a Generator

Writing a separate iterator class is educational, but in practice almost nobody does it. If __iter__ is a generator function (it contains yield), Python builds the iterator for you, and you get a fresh one on every call:

from collections.abc import Iterator
from datetime import date, timedelta


class DaySpan:
    def __init__(self, start: date, end: date) -> None:
        self.start = start
        self.end = end

    def __iter__(self) -> Iterator[date]:
        day = self.start
        while day <= self.end:
            yield day
            day += timedelta(days=1)

    def __len__(self) -> int:
        return (self.end - self.start).days + 1


span = DaySpan(date(2026, 9, 28), date(2026, 10, 1))
print([d.isoformat() for d in span], len(span))
print([d.day for d in span])
['2026-09-28', '2026-09-29', '2026-09-30', '2026-10-01'] 4
[28, 29, 30, 1]

Calling iter(span) returns a generator object, which is a full iterator: it has __next__, its __iter__ returns itself, and it raises StopIteration automatically when the function body finishes. This is the pattern I'd recommend for most custom iterables. Generators themselves (yield, generator expressions, yield from) are covered in What Is a Generator in Python.

Adding __len__ is optional, but it's useful when the length is cheap to compute, and list() uses it as a size hint.

The Two-Argument Form of iter()

iter() has a second, lesser-known form: iter(callable, sentinel). It calls callable with no arguments repeatedly and yields each result until one equals sentinel.

That's ideal for reading fixed-size chunks:

import io
from functools import partial

stream = io.BytesIO(b"abcdefghij")
for chunk in iter(partial(stream.read, 4), b""):
    print(chunk)
b'abcd'
b'efgh'
b'ij'

stream.read(4) returns b"" at end of file, so that's the sentinel. The same pattern works for sockets, binary files, and paginated reads. It also works with any zero-argument function, like "roll a die until you get a 6":

import random

random.seed(3)
rolls = list(iter(lambda: random.randint(1, 6), 6))
print(rolls)  # [2, 5, 5, 2, 3, 5, 4]

Or reading lines from a file until a blank line, a common format for headers:

# notes.txt contains: alpha, beta, a blank line, gamma
with open("notes.txt", encoding="utf-8") as f:
    for line in iter(f.readline, "\n"):
        print(repr(line))
'alpha\n'
'beta\n'

The Old Sequence Protocol

There's one more way to be iterable. If a class has no __iter__ but does have a __getitem__ that accepts integers starting from 0, iter() falls back to calling it with 0, 1, 2, ... until it raises IndexError:

class Sheet:
    def __init__(self, rows: list[str]) -> None:
        self.rows = rows

    def __getitem__(self, index: int) -> str:
        return self.rows[index]


print(list(Sheet(["a", "b"])), "b" in Sheet(["a", "b"]))
# ['a', 'b'] True

It works, but it's a legacy path, and isinstance(Sheet([]), Iterable) returns False because the ABC only checks for __iter__. Define __iter__ explicitly in new code.

Common Pitfalls

Modifying a Collection While Iterating

Changing a dict or set's size during iteration raises an error, because the iterator's internal position becomes meaningless:

d = {"a": 1, "b": 2}
for k in d:
    d[k + k] = 0
# RuntimeError: dictionary changed size during iteration

Iterate over a snapshot instead: for k in list(d):. Lists don't raise in this situation, which is worse: removing items while iterating silently skips elements. Build a new list with a comprehension rather than mutating the one you're looping over.

Mismatched Lengths in zip

zip() stops at the shortest input, which can hide data problems. Since Python 3.10, strict=True makes it raise instead:

list(zip([1, 2, 3], "ab", strict=True))
# ValueError: zip() argument 2 is shorter than argument 1

Consuming Part of an Iterator on Purpose

Single-use behaviour isn't always a problem. Sometimes it's exactly what you want, like skipping a header line and then processing the rest:

lines = iter(["# header", "a", "b"])
header = next(lines)
print(header, list(lines))  # # header ['a', 'b']

The same idea works with a file object, which is an iterator over its lines. For slicing, chaining, and grouping iterators without building lists, see Mastering itertools in Python.

Why Iterators Matter: Laziness

Iterators produce items on demand, so they don't need to hold everything in memory. A range of a billion numbers and its iterator are each a few dozen bytes:

import sys

print(sys.getsizeof(range(10**9)), sys.getsizeof(iter(range(10**9))))  # 48 40

That's why you can loop over a multi-gigabyte log file line by line, stream rows from a database cursor, or chain map() and filter() over huge inputs without running out of memory. Each step pulls one item from the step before it, only when needed.

The protocol also has an async version for code using asyncio: __aiter__ and __anext__, consumed with async for. The ideas are the same; each step just awaits.

Conclusion

The iterator protocol is two methods. An iterable's __iter__ returns an iterator; an iterator's __next__ returns items until it raises StopIteration, and its __iter__ returns itself. A for loop is iter() once and next() until StopIteration, and nearly every Python tool that takes "a bunch of things" uses the same contract.

Keep the distinction straight and the surprises go away: lists, dicts, strings, and range can be looped over repeatedly, while map, zip, enumerate, files, and generators are single-use. When you write your own iterable, make __iter__ a generator so every loop gets a fresh iterator, and reach for iter(callable, sentinel) when you're reading until some end marker.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading