Type something to search...
Regular Expressions in Python with the re Module

Regular Expressions in Python with the re Module

Regular expressions are a small language for describing patterns in text: "four digits, a dash, two digits", "a word that starts with a capital letter", "everything between quotes". When string methods like split(), startswith(), and in aren't enough, a regex can usually find, extract, validate, or rewrite the text you care about in a single line.

They also have a reputation for being write-only. A pattern that made perfect sense when you wrote it can look like line noise a month later. The trick is to learn the handful of pieces you'll use 95% of the time, and to use the tools Python gives you for keeping patterns readable.

This post covers the re module's functions, the core syntax, groups and named groups, substitution, flags, lookarounds, greedy vs lazy matching, and the pitfalls that catch people most often. Examples run on Python 3.13.

Your First Match

import re

text = "Order #1042 shipped on 2026-09-18, order #1043 on 2026-09-19."

m = re.search(r"\d{4}-\d{2}-\d{2}", text)
print(m)
print(m.group(), m.span())
<re.Match object; span=(23, 33), match='2026-09-18'>
2026-09-18 (23, 33)

re.search() scans the string for the first place the pattern matches and returns a Match object, or None if nothing matches. The pattern \d{4}-\d{2}-\d{2} reads as "four digits, a dash, two digits, a dash, two digits".

Always Use Raw Strings

Notice the r before the pattern. Regex syntax is full of backslashes (\d, \w, \b), and so are Python's own string escapes (\n, \t). A raw string tells Python to leave backslashes alone so the regex engine sees them exactly as written.

Without the r, some patterns happen to work by accident, while others break silently. "\b" in a normal string is a backspace character, not a word boundary. Python 3.12+ also emits a SyntaxWarning for unknown escapes like "\d". Make r"..." a habit for every pattern.

Check for None Before Using a Match

When there's no match, search() returns None, and calling .group() on it raises AttributeError: 'NoneType' object has no attribute 'group'. The walrus operator makes the check concise:

if m := re.search(r"(\d+) items", "Cart: 3 items"):
    print(int(m.group(1)))   # 3

The re Functions

FunctionWhat it doesReturns
re.search(p, s)First match anywhere in the stringMatch or None
re.match(p, s)Match only at the start of the stringMatch or None
re.fullmatch(p, s)The whole string must matchMatch or None
re.findall(p, s)All non-overlapping matcheslist
re.finditer(p, s)All matches, lazilyiterator of Match
re.sub(p, repl, s)Replace matchesnew string
re.subn(p, repl, s)Replace and count(string, count)
re.split(p, s)Split on matcheslist
re.compile(p)Pre-compile a patternPattern

search vs match vs fullmatch

This is the most common source of confusion:

print(re.match(r"\d+", text))            # None: text starts with "Order"
print(re.match(r"Order", text))          # matches at position 0
print(re.fullmatch(r"\d{4}", "2026"))    # matches
print(re.fullmatch(r"\d{4}", "20261"))   # None: extra character

match() only checks the beginning but doesn't care what comes after. For validation, where the entire input must fit the pattern, use fullmatch(). A validator built with match(r"\d{4}") would happily accept "2026abc".

findall and finditer

findall() returns every match as a list of strings:

print(re.findall(r"#(\d+)", text))
print(re.findall(r"(\d{4})-(\d{2})-(\d{2})", text))
['1042', '1043']
[('2026', '09', '18'), ('2026', '09', '19')]

Its return type depends on the groups in the pattern, which trips people up:

  • No groups: a list of the full matches.
  • One group: a list of that group's text (not the full match). That's why #(\d+) returned only the numbers, without the #.
  • Several groups: a list of tuples.

When you need positions, named groups, or more than one piece of each match, use finditer(), which yields full Match objects:

for m in re.finditer(r"#(\d+)", text):
    print(m.group(1), m.span())
1042 (6, 11)
1043 (41, 46)

Core Syntax

You don't need every feature of regex syntax. These cover the vast majority of real patterns.

Characters and Classes

PatternMatches
.Any character except a newline
\d / \DA digit / a non-digit
\w / \WA word character (letter, digit, underscore) / anything else
\s / \SWhitespace / non-whitespace
[abc]One of a, b, or c
[a-z0-9]A character in a range
[^abc]Any character except those
\., \(, \$A literal special character, escaped

Quantifiers

PatternMeaning
*0 or more
+1 or more
?0 or 1 (optional)
{3}Exactly 3
{2,5}Between 2 and 5
{2,}2 or more

So colou?r matches both color and colour, and \d{2,4} matches two to four digits.

Anchors and Boundaries

  • ^ matches at the start of the string, $ at the end.
  • \b matches a word boundary: the position between a word character and a non-word character.

Word boundaries are how you match whole words:

print(re.findall(r"\bcat\b", "cat concat cat's catalog"))   # ['cat', 'cat']

It skips concat and catalog but matches the cat in cat's, because the apostrophe isn't a word character.

Alternation and Grouping

A vertical bar means "or": cat|dog matches either word. Parentheses group parts of a pattern so quantifiers and alternation apply to the whole group: (ab)+ matches ab, abab, and so on, and gr(a|e)y matches gray and grey.

Groups: Extracting Pieces of a Match

Parentheses also capture the text they match, so you can pull it out afterward. Groups are numbered from 1 by their opening parenthesis; group 0 is the whole match.

Named Groups

Numbered groups get hard to follow in longer patterns. Named groups, written (?P<name>...), make both the pattern and the code that uses it self-documenting:

m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})", text)

print(m.group("year"), m["month"])
print(m.groupdict())
print(m.groups())
2026 09
{'year': '2026', 'month': '09', 'day': '18'}
('2026', '09', '18')

m["month"] is shorthand for m.group("month"), and groupdict() gives you everything at once, ready to pass into a dataclass or a dict of parsed fields.

An optional group that didn't participate in the match returns None, not an empty string:

m = re.search(r"(\d+)-(\d+)?", "10-")
print(m.groups())   # ('10', None)

Non-Capturing Groups

Sometimes you need parentheses just for grouping, not capturing. Write (?:...). This matters most with findall(), because capturing groups change what it returns:

print(re.findall(r"(ab)+", "ababab abab"))     # ['ab', 'ab']
print(re.findall(r"(?:ab)+", "ababab abab"))   # ['ababab', 'abab']

With a capturing group, findall returns only the last repetition of the group. Non-capturing groups keep the full match.

Backreferences

\1 refers back to whatever group 1 matched, which is useful for finding repeated words:

print(re.findall(r"\b(\w+) \1\b", "this is is a test test"))   # ['is', 'test']

Substitution with re.sub

re.sub(pattern, replacement, string) replaces every match. The simplest use is normalizing text:

print(re.sub(r"\s+", " ", "too    many   spaces"))   # too many spaces

The replacement string can refer to groups with \1 or, for named groups, \g<name>:

print(re.sub(r"(\d{4})-(\d{2})-(\d{2})", r"\3/\2/\1", text))
print(re.sub(r"(?P<y>\d{4})-(?P<m>\d{2})-(?P<d>\d{2})", r"\g<d>.\g<m>.\g<y>", text))
Order #1042 shipped on 18/09/2026, order #1043 on 19/09/2026.
Order #1042 shipped on 18.09.2026, order #1043 on 19.09.2026.

Replacing with a Function

When the replacement depends on the matched text, pass a function. It receives each Match and returns the replacement string:

print(re.sub(r"#(\d+)", lambda m: f"#{int(m.group(1)) + 1000}", text))
Order #2042 shipped on 2026-09-18, order #2043 on 2026-09-19.

That's how you do things like converting units, looking up values in a dict, or masking all but the last four digits of a card number.

re.subn() does the same thing but also returns how many replacements were made. Both accept count= to limit the number of replacements. As of Python 3.13, pass count, maxsplit, and flags as keyword arguments; passing them positionally is deprecated.

Two Small, Useful Recipes

A slug generator and a camelCase to snake_case converter:

import re


def slugify(title: str) -> str:
    return re.sub(r"[^a-z0-9]+", "-", title.lower()).strip("-")


def camel_to_snake(name: str) -> str:
    name = re.sub(r"(.)([A-Z][a-z]+)", r"\1_\2", name)
    return re.sub(r"([a-z0-9])([A-Z])", r"\1_\2", name).lower()


print(slugify("  Hello, World! 2026 "))            # hello-world-2026
print(camel_to_snake("parseHTTPResponseCode"))     # parse_http_response_code
print(camel_to_snake("userId"))                    # user_id

Splitting with re.split

str.split() only splits on one fixed separator. re.split() splits on a pattern:

print(re.split(r"[,;]\s*", "a, b;c ,d"))
print(re.split(r"(\d)", "a1b2c"))
['a', 'b', 'c ', 'd']
['a', '1', 'b', '2', 'c']

If the pattern contains a capturing group, the separators are included in the result, as in the second example. Use (?:...) if you don't want them.

Flags

Flags change how a pattern is interpreted. Pass them as flags= and combine them with |:

FlagEffect
re.IGNORECASE (re.I)Case-insensitive matching
re.MULTILINE (re.M)^ and $ match at the start/end of each line
re.DOTALL (re.S). also matches newlines
re.VERBOSE (re.X)Ignore whitespace and allow comments in the pattern
re.ASCII (re.A)\w, \d, \s match ASCII only

MULTILINE is the one people forget when processing a whole file at once:

s = "first line\nsecond line"
print(re.findall(r"^\w+", s))                       # ['first']
print(re.findall(r"^\w+", s, flags=re.MULTILINE))   # ['first', 'second']

You can also put flags inline at the very start of a pattern, like (?i)hello, which is handy when the pattern lives in a config file.

re.VERBOSE: Readable Patterns

re.VERBOSE is the best tool for keeping regexes maintainable. Whitespace in the pattern is ignored (unless escaped or inside a character class), and # starts a comment:

import re

EMAIL_RE = re.compile(
    r"""
    ^(?P<user>[\w.+-]+)                 # local part
    @
    (?P<domain>[\w-]+(?:\.[\w-]+)+)$    # domain with at least one dot
    """,
    re.VERBOSE,
)

print(EMAIL_RE.match("ada.lovelace+news@example.co.uk").groupdict())
print(EMAIL_RE.match("not-an-email"))
{'user': 'ada.lovelace+news', 'domain': 'example.co.uk'}
None

(This is a reasonable sanity check, not a full email validator. Real-world email validation is better done by sending a confirmation email.)

Unicode by Default

In Python 3, str patterns are Unicode-aware. \w matches letters from any script and \d matches digits from any script:

print(re.findall(r"\w+", "café naïve 東京"))   # ['café', 'naïve', '東京']
print(re.findall(r"\d", "١٢٣ 123"))           # ['١', '٢', '٣', '1', '2', '3']
print(re.findall(r"[0-9]", "١٢٣ 123"))        # ['1', '2', '3']

That's usually what you want for text, but not for validating something like an ID or a port number, where int() would later choke on Arabic-Indic digits. Use [0-9] or the re.ASCII flag when you mean ASCII digits only.

Greedy vs Lazy Quantifiers

Quantifiers are greedy: they match as much as possible and only back off if the rest of the pattern fails. That causes the classic HTML-tag mistake:

html = "<b>bold</b> and <i>italic</i>"

print(re.findall(r"<.+>", html))     # ['<b>bold</b> and <i>italic</i>']
print(re.findall(r"<.+?>", html))    # ['<b>', '</b>', '<i>', '</i>']
print(re.findall(r"<[^>]+>", html))  # ['<b>', '</b>', '<i>', '</i>']

.+ grabbed everything up to the last >. Adding ? after a quantifier (+?, *?, ??) makes it lazy, matching as little as possible. A negated character class like [^>]+ is often even better: it states exactly what's allowed and avoids backtracking.

(And for real HTML, use a parser like Beautiful Soup, not regular expressions.)

Lookarounds

Lookarounds check what's before or after a position without including it in the match:

SyntaxNameMeaning
(?=...)LookaheadFollowed by ...
(?!...)Negative lookaheadNot followed by ...
(?<=...)LookbehindPreceded by ...
(?<!...)Negative lookbehindNot preceded by ...
print(re.findall(r"\d+(?= USD)", "10 USD, 20 EUR, 30 USD"))           # ['10', '30']
print(re.findall(r"(?<=\$)\d+(?:\.\d{2})?", "Total: $42.50, tax $3"))  # ['42.50', '3']
print(re.findall(r"(?<!-)\b\d+", "5 -3 12"))                            # ['5', '12']

The amounts come back without the currency markers because lookarounds are zero-width. One limitation in the standard re module: lookbehinds must have a fixed width, so (?<=\$+) isn't allowed.

Compiling Patterns

re.compile() turns a pattern into a Pattern object with the same methods (search, findall, sub, and so on):

import re

LOG_RE = re.compile(
    r'(?P<ip>\S+) \S+ \S+ \[(?P<time>[^\]]+)\] '
    r'"(?P<method>[A-Z]+) (?P<path>\S+) [^"]*" '
    r'(?P<status>\d{3}) (?P<size>\d+|-)'
)

line = '127.0.0.1 - - [18/Sep/2026:10:15:32 +0000] "GET /blog/python HTTP/1.1" 200 5123'
print(LOG_RE.match(line).groupdict())
{'ip': '127.0.0.1', 'time': '18/Sep/2026:10:15:32 +0000', 'method': 'GET', 'path': '/blog/python', 'status': '200', 'size': '5123'}

Is compiling faster? A little, but less than people think: the module-level functions cache recently used compiled patterns internally. The real benefit is organization. A compiled pattern at module level, with a descriptive name, documents what it's for and keeps the regex out of your loop body.

Escaping User Input

If you build a pattern from a variable, special characters in it will be interpreted as regex syntax. re.escape() neutralizes them:

term = "c++"
print(re.findall(re.escape(term), "I like c++ and c"))   # ['c++']
print(re.escape("price (USD): $5.00"))
['c++']
price\ \(USD\):\ \$5\.00

Without re.escape, "c++" is read as "one or more c" with a possessive quantifier (supported since Python 3.11), so it returns ['c', 'c'] instead of ['c++']. Similarly, "a.b" would match "axb", and a term with an unbalanced ( would raise an error. Escape any untrusted input before putting it into a pattern.

Common Pitfalls

  • Forgetting the r prefix. Use raw strings for every pattern.
  • Using match() for validation. It only anchors at the start. Use fullmatch().
  • Surprising findall() output. Capturing groups change the return value. Use (?:...) or finditer().
  • Greedy .* eating too much. Use a lazy quantifier or a negated character class.
  • Unescaped dots. . matches any character, so example.com also matches exampleXcom. Write example\.com.
  • Catastrophic backtracking. Nested quantifiers like (a+)+$ can take exponential time on inputs that almost match. If a regex runs on untrusted input, keep it simple, prefer specific character classes over .*, and test it against long, non-matching strings.
  • Invalid patterns. A malformed pattern raises re.PatternError (the new name in Python 3.13; re.error still works as an alias), for example missing ), unterminated subpattern at position 0.
  • Reaching for regex when a string method works. s.startswith("http"), "@" in s, and s.split(",") are faster and clearer. See the Python strings deep dive for what's available before you write a pattern.

Conclusion

The re module comes down to a few decisions: search() to find, fullmatch() to validate, finditer() or findall() to extract everything, sub() to rewrite, and split() to break text apart. Learn the core syntax (character classes, quantifiers, anchors, groups), lean on named groups and re.VERBOSE to keep patterns readable, and remember that quantifiers are greedy unless you add ?.

Write every pattern as a raw string, escape anything that comes from users, and test your regex against inputs that shouldn't match as carefully as the ones that should. With those habits, regular expressions stop being line noise and become one of the most useful tools you have for working with text.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading