
Iterators and the Iterator Protocol in Python
A for loop in Python works on lists, strings, dictionaries, files, database cursors, range objects, and anything else that knows how to hand out its items one at a time. It doesn't index into them, and it doesn't need to know how many items there are. All it relies on is a two-method contract called the iterator protocol.
Knowing that protocol pays off in two ways. It explains behaviour that otherwise looks like a bug, such as a map() result that's empty the second time you loop over it. And it lets you make your own classes work with for, list(), sum(), unpacking, and every other tool that consumes iterables.
This post covers the difference between an iterable and an iterator, what iter() and next() do, how a for loop is built from them, how to write iterator and iterable classes, the two-argument form of iter(), and the exhaustion pitfalls to watch for.
Iterables vs Iterators
These two words sound interchangeable, but they mean different things:
- An iterable is anything you can loop over. It has an
__iter__method that returns an iterator. Lists, tuples, strings, dicts, sets, andrangeobjects are iterables. - An iterator is the object that does the looping. It has a
__next__method that returns the next item, or raisesStopIterationwhen there are none left. It also has an__iter__method that returns itself.
The built-in iter() asks an iterable for an iterator, and next() asks an iterator for its next item:
nums = [10, 20, 30]
it = iter(nums)
print(it)
print(next(it), next(it), next(it))
next(it)
<list_iterator object at 0x1048348b0>
10 20 30
StopIteration
The list itself doesn't track a position. The list_iterator does. That's the key separation: an iterable is a source of items, an iterator is a cursor over that source.
You can't call next() on a list directly:
next(nums)
# TypeError: 'list' object is not an iterator
And collections.abc confirms the distinction:
from collections.abc import Iterable, Iterator
print(isinstance(nums, Iterable), isinstance(nums, Iterator), isinstance(it, Iterator))
# True False True
Every iterator is also an iterable (its __iter__ returns itself), but not every iterable is an iterator.
A Default for next()
next() accepts a second argument that's returned instead of raising StopIteration:
print(next(it, "done")) # done
That makes a clean "first item or default" helper:
def first(iterable, default=None):
return next(iter(iterable), default)
print(first([]), first("xyz"), first(x for x in range(10) if x > 6))
# None x 7
How a for Loop Actually Works
A for loop is the protocol plus exception handling. This loop:
for item in nums:
print(item)
does essentially this:
iterator = iter(nums)
while True:
try:
item = next(iterator)
except StopIteration:
break
print(item)
Python calls iter() once at the start, then next() repeatedly until StopIteration signals the end. StopIteration isn't an error here; it's the normal way an iterator says "I'm finished", and the loop catches it for you.
The same protocol powers far more than for: list(), tuple(), set(), dict(), sum(), min(), max(), sorted(), any(), all(), str.join(), zip(), enumerate(), unpacking like a, b, *rest = thing, comprehensions, and in (when the object has no faster __contains__). Implement the protocol once and your object works with all of them.
Which Built-ins Are Iterators?
Many built-ins return iterators rather than lists. A quick way to tell: an iterator's iter() returns the object itself.
for obj in (range(3), {"a": 1}, "hi", open(__file__), enumerate([]), reversed([1])):
print(type(obj).__name__, iter(obj) is obj)
range False
dict False
str False
TextIOWrapper True
enumerate True
list_reverseiterator True
range, dict, and str are reusable iterables. File objects, enumerate, reversed, zip, map, filter, generators, and most of itertools are iterators. That distinction matters, as the next section shows.
Iterators Are Single-Use
An iterator moves forward only and can't be reset. Once it's exhausted, it stays exhausted:
squares = map(lambda n: n * n, [1, 2, 3])
print(list(squares), list(squares))
[1, 4, 9] []
The second list() gets nothing, with no error. This is the source of a whole family of bugs. Here's a function that works with a list and silently breaks with an iterator:
def total_and_count(numbers):
total = sum(numbers)
count = len(list(numbers))
return total, count
print(total_and_count([1, 2, 3])) # (6, 3)
print(total_and_count(iter([1, 2, 3]))) # (6, 0)
sum() consumed the iterator, so list(numbers) saw nothing. If a function needs to pass over its input more than once, either convert it once at the top (numbers = list(numbers)) or document that it requires a re-iterable collection, such as a collections.abc.Collection.
Membership tests on an iterator also consume it, up to and including the match:
it = iter([10, 20, 30])
print(20 in it, list(it)) # True [30]
Writing an Iterator Class
To make your own iterator, implement __next__ (return the next item or raise StopIteration) and __iter__ (return self):
class Countdown:
"""An iterator that counts down from start to 1."""
def __init__(self, start: int) -> None:
self.current = start
def __iter__(self) -> "Countdown":
return self
def __next__(self) -> int:
if self.current <= 0:
raise StopIteration
value = self.current
self.current -= 1
return value
for n in Countdown(2):
print(n)
cd = Countdown(3)
print(list(cd))
print(list(cd))
2
1
[3, 2, 1]
[]
It works with for and list(), and like every iterator it's exhausted after one pass. The state (self.current) lives on the iterator, so there's no way to restart it except creating a new one.
Two rules to follow:
- Once
__next__has raisedStopIteration, it should keep raising it on every later call. Code that loops over an iterator expects "finished" to be permanent. __iter__must returnself. Without it, your iterator can't be used directly in aforloop, because the loop callsiter()first.
Iterables That Can Be Looped Over Many Times
Usually you don't want your collection to be single-use. A list can be looped over any number of times because each for gets a new list_iterator. To get the same behaviour, separate the iterable from its iterator: __iter__ returns a fresh iterator object each time.
class NumberRange:
"""An iterable: each for-loop gets a fresh iterator."""
def __init__(self, start: int, stop: int) -> None:
self.start = start
self.stop = stop
def __iter__(self) -> "NumberRangeIterator":
return NumberRangeIterator(self.start, self.stop)
class NumberRangeIterator:
def __init__(self, current: int, stop: int) -> None:
self.current = current
self.stop = stop
def __iter__(self) -> "NumberRangeIterator":
return self
def __next__(self) -> int:
if self.current >= self.stop:
raise StopIteration
value = self.current
self.current += 1
return value
r = NumberRange(1, 4)
print(list(r), list(r))
print([(a, b) for a in r for b in r if a < b])
[1, 2, 3] [1, 2, 3]
[(1, 2), (1, 3), (2, 3)]
The nested comprehension loops over r twice at the same time. That works because each loop has its own independent iterator with its own position. If NumberRange were its own iterator, the inner loop would exhaust the shared state and the result would be wrong.
The Shortcut: Make __iter__ a Generator
Writing a separate iterator class is educational, but in practice almost nobody does it. If __iter__ is a generator function (it contains yield), Python builds the iterator for you, and you get a fresh one on every call:
from collections.abc import Iterator
from datetime import date, timedelta
class DaySpan:
def __init__(self, start: date, end: date) -> None:
self.start = start
self.end = end
def __iter__(self) -> Iterator[date]:
day = self.start
while day <= self.end:
yield day
day += timedelta(days=1)
def __len__(self) -> int:
return (self.end - self.start).days + 1
span = DaySpan(date(2026, 9, 28), date(2026, 10, 1))
print([d.isoformat() for d in span], len(span))
print([d.day for d in span])
['2026-09-28', '2026-09-29', '2026-09-30', '2026-10-01'] 4
[28, 29, 30, 1]
Calling iter(span) returns a generator object, which is a full iterator: it has __next__, its __iter__ returns itself, and it raises StopIteration automatically when the function body finishes. This is the pattern I'd recommend for most custom iterables. Generators themselves (yield, generator expressions, yield from) are covered in What Is a Generator in Python.
Adding __len__ is optional, but it's useful when the length is cheap to compute, and list() uses it as a size hint.
The Two-Argument Form of iter()
iter() has a second, lesser-known form: iter(callable, sentinel). It calls callable with no arguments repeatedly and yields each result until one equals sentinel.
That's ideal for reading fixed-size chunks:
import io
from functools import partial
stream = io.BytesIO(b"abcdefghij")
for chunk in iter(partial(stream.read, 4), b""):
print(chunk)
b'abcd'
b'efgh'
b'ij'
stream.read(4) returns b"" at end of file, so that's the sentinel. The same pattern works for sockets, binary files, and paginated reads. It also works with any zero-argument function, like "roll a die until you get a 6":
import random
random.seed(3)
rolls = list(iter(lambda: random.randint(1, 6), 6))
print(rolls) # [2, 5, 5, 2, 3, 5, 4]
Or reading lines from a file until a blank line, a common format for headers:
# notes.txt contains: alpha, beta, a blank line, gamma
with open("notes.txt", encoding="utf-8") as f:
for line in iter(f.readline, "\n"):
print(repr(line))
'alpha\n'
'beta\n'
The Old Sequence Protocol
There's one more way to be iterable. If a class has no __iter__ but does have a __getitem__ that accepts integers starting from 0, iter() falls back to calling it with 0, 1, 2, ... until it raises IndexError:
class Sheet:
def __init__(self, rows: list[str]) -> None:
self.rows = rows
def __getitem__(self, index: int) -> str:
return self.rows[index]
print(list(Sheet(["a", "b"])), "b" in Sheet(["a", "b"]))
# ['a', 'b'] True
It works, but it's a legacy path, and isinstance(Sheet([]), Iterable) returns False because the ABC only checks for __iter__. Define __iter__ explicitly in new code.
Common Pitfalls
Modifying a Collection While Iterating
Changing a dict or set's size during iteration raises an error, because the iterator's internal position becomes meaningless:
d = {"a": 1, "b": 2}
for k in d:
d[k + k] = 0
# RuntimeError: dictionary changed size during iteration
Iterate over a snapshot instead: for k in list(d):. Lists don't raise in this situation, which is worse: removing items while iterating silently skips elements. Build a new list with a comprehension rather than mutating the one you're looping over.
Mismatched Lengths in zip
zip() stops at the shortest input, which can hide data problems. Since Python 3.10, strict=True makes it raise instead:
list(zip([1, 2, 3], "ab", strict=True))
# ValueError: zip() argument 2 is shorter than argument 1
Consuming Part of an Iterator on Purpose
Single-use behaviour isn't always a problem. Sometimes it's exactly what you want, like skipping a header line and then processing the rest:
lines = iter(["# header", "a", "b"])
header = next(lines)
print(header, list(lines)) # # header ['a', 'b']
The same idea works with a file object, which is an iterator over its lines. For slicing, chaining, and grouping iterators without building lists, see Mastering itertools in Python.
Why Iterators Matter: Laziness
Iterators produce items on demand, so they don't need to hold everything in memory. A range of a billion numbers and its iterator are each a few dozen bytes:
import sys
print(sys.getsizeof(range(10**9)), sys.getsizeof(iter(range(10**9)))) # 48 40
That's why you can loop over a multi-gigabyte log file line by line, stream rows from a database cursor, or chain map() and filter() over huge inputs without running out of memory. Each step pulls one item from the step before it, only when needed.
The protocol also has an async version for code using asyncio: __aiter__ and __anext__, consumed with async for. The ideas are the same; each step just awaits.
Conclusion
The iterator protocol is two methods. An iterable's __iter__ returns an iterator; an iterator's __next__ returns items until it raises StopIteration, and its __iter__ returns itself. A for loop is iter() once and next() until StopIteration, and nearly every Python tool that takes "a bunch of things" uses the same contract.
Keep the distinction straight and the surprises go away: lists, dicts, strings, and range can be looped over repeatedly, while map, zip, enumerate, files, and generators are single-use. When you write your own iterable, make __iter__ a generator so every loop gets a fresh iterator, and reach for iter(callable, sentinel) when you're reading until some end marker.


