
Regular Expressions in Python with the re Module
Regular expressions are a small language for describing patterns in text: "four digits, a dash, two digits", "a word that starts with a capital letter", "everything between quotes". When string methods like split(), startswith(), and in aren't enough, a regex can usually find, extract, validate, or rewrite the text you care about in a single line.
They also have a reputation for being write-only. A pattern that made perfect sense when you wrote it can look like line noise a month later. The trick is to learn the handful of pieces you'll use 95% of the time, and to use the tools Python gives you for keeping patterns readable.
This post covers the re module's functions, the core syntax, groups and named groups, substitution, flags, lookarounds, greedy vs lazy matching, and the pitfalls that catch people most often. Examples run on Python 3.13.
Your First Match
import re
text = "Order #1042 shipped on 2026-09-18, order #1043 on 2026-09-19."
m = re.search(r"\d{4}-\d{2}-\d{2}", text)
print(m)
print(m.group(), m.span())
<re.Match object; span=(23, 33), match='2026-09-18'>
2026-09-18 (23, 33)
re.search() scans the string for the first place the pattern matches and returns a Match object, or None if nothing matches. The pattern \d{4}-\d{2}-\d{2} reads as "four digits, a dash, two digits, a dash, two digits".
Always Use Raw Strings
Notice the r before the pattern. Regex syntax is full of backslashes (\d, \w, \b), and so are Python's own string escapes (\n, \t). A raw string tells Python to leave backslashes alone so the regex engine sees them exactly as written.
Without the r, some patterns happen to work by accident, while others break silently. "\b" in a normal string is a backspace character, not a word boundary. Python 3.12+ also emits a SyntaxWarning for unknown escapes like "\d". Make r"..." a habit for every pattern.
Check for None Before Using a Match
When there's no match, search() returns None, and calling .group() on it raises AttributeError: 'NoneType' object has no attribute 'group'. The walrus operator makes the check concise:
if m := re.search(r"(\d+) items", "Cart: 3 items"):
print(int(m.group(1))) # 3
The re Functions
| Function | What it does | Returns |
|---|---|---|
re.search(p, s) | First match anywhere in the string | Match or None |
re.match(p, s) | Match only at the start of the string | Match or None |
re.fullmatch(p, s) | The whole string must match | Match or None |
re.findall(p, s) | All non-overlapping matches | list |
re.finditer(p, s) | All matches, lazily | iterator of Match |
re.sub(p, repl, s) | Replace matches | new string |
re.subn(p, repl, s) | Replace and count | (string, count) |
re.split(p, s) | Split on matches | list |
re.compile(p) | Pre-compile a pattern | Pattern |
search vs match vs fullmatch
This is the most common source of confusion:
print(re.match(r"\d+", text)) # None: text starts with "Order"
print(re.match(r"Order", text)) # matches at position 0
print(re.fullmatch(r"\d{4}", "2026")) # matches
print(re.fullmatch(r"\d{4}", "20261")) # None: extra character
match() only checks the beginning but doesn't care what comes after. For validation, where the entire input must fit the pattern, use fullmatch(). A validator built with match(r"\d{4}") would happily accept "2026abc".
findall and finditer
findall() returns every match as a list of strings:
print(re.findall(r"#(\d+)", text))
print(re.findall(r"(\d{4})-(\d{2})-(\d{2})", text))
['1042', '1043']
[('2026', '09', '18'), ('2026', '09', '19')]
Its return type depends on the groups in the pattern, which trips people up:
- No groups: a list of the full matches.
- One group: a list of that group's text (not the full match). That's why
#(\d+)returned only the numbers, without the#. - Several groups: a list of tuples.
When you need positions, named groups, or more than one piece of each match, use finditer(), which yields full Match objects:
for m in re.finditer(r"#(\d+)", text):
print(m.group(1), m.span())
1042 (6, 11)
1043 (41, 46)
Core Syntax
You don't need every feature of regex syntax. These cover the vast majority of real patterns.
Characters and Classes
| Pattern | Matches |
|---|---|
. | Any character except a newline |
\d / \D | A digit / a non-digit |
\w / \W | A word character (letter, digit, underscore) / anything else |
\s / \S | Whitespace / non-whitespace |
[abc] | One of a, b, or c |
[a-z0-9] | A character in a range |
[^abc] | Any character except those |
\., \(, \$ | A literal special character, escaped |
Quantifiers
| Pattern | Meaning |
|---|---|
* | 0 or more |
+ | 1 or more |
? | 0 or 1 (optional) |
{3} | Exactly 3 |
{2,5} | Between 2 and 5 |
{2,} | 2 or more |
So colou?r matches both color and colour, and \d{2,4} matches two to four digits.
Anchors and Boundaries
^matches at the start of the string,$at the end.\bmatches a word boundary: the position between a word character and a non-word character.
Word boundaries are how you match whole words:
print(re.findall(r"\bcat\b", "cat concat cat's catalog")) # ['cat', 'cat']
It skips concat and catalog but matches the cat in cat's, because the apostrophe isn't a word character.
Alternation and Grouping
A vertical bar means "or": cat|dog matches either word. Parentheses group parts of a pattern so quantifiers and alternation apply to the whole group: (ab)+ matches ab, abab, and so on, and gr(a|e)y matches gray and grey.
Groups: Extracting Pieces of a Match
Parentheses also capture the text they match, so you can pull it out afterward. Groups are numbered from 1 by their opening parenthesis; group 0 is the whole match.
Named Groups
Numbered groups get hard to follow in longer patterns. Named groups, written (?P<name>...), make both the pattern and the code that uses it self-documenting:
m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})", text)
print(m.group("year"), m["month"])
print(m.groupdict())
print(m.groups())
2026 09
{'year': '2026', 'month': '09', 'day': '18'}
('2026', '09', '18')
m["month"] is shorthand for m.group("month"), and groupdict() gives you everything at once, ready to pass into a dataclass or a dict of parsed fields.
An optional group that didn't participate in the match returns None, not an empty string:
m = re.search(r"(\d+)-(\d+)?", "10-")
print(m.groups()) # ('10', None)
Non-Capturing Groups
Sometimes you need parentheses just for grouping, not capturing. Write (?:...). This matters most with findall(), because capturing groups change what it returns:
print(re.findall(r"(ab)+", "ababab abab")) # ['ab', 'ab']
print(re.findall(r"(?:ab)+", "ababab abab")) # ['ababab', 'abab']
With a capturing group, findall returns only the last repetition of the group. Non-capturing groups keep the full match.
Backreferences
\1 refers back to whatever group 1 matched, which is useful for finding repeated words:
print(re.findall(r"\b(\w+) \1\b", "this is is a test test")) # ['is', 'test']
Substitution with re.sub
re.sub(pattern, replacement, string) replaces every match. The simplest use is normalizing text:
print(re.sub(r"\s+", " ", "too many spaces")) # too many spaces
The replacement string can refer to groups with \1 or, for named groups, \g<name>:
print(re.sub(r"(\d{4})-(\d{2})-(\d{2})", r"\3/\2/\1", text))
print(re.sub(r"(?P<y>\d{4})-(?P<m>\d{2})-(?P<d>\d{2})", r"\g<d>.\g<m>.\g<y>", text))
Order #1042 shipped on 18/09/2026, order #1043 on 19/09/2026.
Order #1042 shipped on 18.09.2026, order #1043 on 19.09.2026.
Replacing with a Function
When the replacement depends on the matched text, pass a function. It receives each Match and returns the replacement string:
print(re.sub(r"#(\d+)", lambda m: f"#{int(m.group(1)) + 1000}", text))
Order #2042 shipped on 2026-09-18, order #2043 on 2026-09-19.
That's how you do things like converting units, looking up values in a dict, or masking all but the last four digits of a card number.
re.subn() does the same thing but also returns how many replacements were made. Both accept count= to limit the number of replacements. As of Python 3.13, pass count, maxsplit, and flags as keyword arguments; passing them positionally is deprecated.
Two Small, Useful Recipes
A slug generator and a camelCase to snake_case converter:
import re
def slugify(title: str) -> str:
return re.sub(r"[^a-z0-9]+", "-", title.lower()).strip("-")
def camel_to_snake(name: str) -> str:
name = re.sub(r"(.)([A-Z][a-z]+)", r"\1_\2", name)
return re.sub(r"([a-z0-9])([A-Z])", r"\1_\2", name).lower()
print(slugify(" Hello, World! 2026 ")) # hello-world-2026
print(camel_to_snake("parseHTTPResponseCode")) # parse_http_response_code
print(camel_to_snake("userId")) # user_id
Splitting with re.split
str.split() only splits on one fixed separator. re.split() splits on a pattern:
print(re.split(r"[,;]\s*", "a, b;c ,d"))
print(re.split(r"(\d)", "a1b2c"))
['a', 'b', 'c ', 'd']
['a', '1', 'b', '2', 'c']
If the pattern contains a capturing group, the separators are included in the result, as in the second example. Use (?:...) if you don't want them.
Flags
Flags change how a pattern is interpreted. Pass them as flags= and combine them with |:
| Flag | Effect |
|---|---|
re.IGNORECASE (re.I) | Case-insensitive matching |
re.MULTILINE (re.M) | ^ and $ match at the start/end of each line |
re.DOTALL (re.S) | . also matches newlines |
re.VERBOSE (re.X) | Ignore whitespace and allow comments in the pattern |
re.ASCII (re.A) | \w, \d, \s match ASCII only |
MULTILINE is the one people forget when processing a whole file at once:
s = "first line\nsecond line"
print(re.findall(r"^\w+", s)) # ['first']
print(re.findall(r"^\w+", s, flags=re.MULTILINE)) # ['first', 'second']
You can also put flags inline at the very start of a pattern, like (?i)hello, which is handy when the pattern lives in a config file.
re.VERBOSE: Readable Patterns
re.VERBOSE is the best tool for keeping regexes maintainable. Whitespace in the pattern is ignored (unless escaped or inside a character class), and # starts a comment:
import re
EMAIL_RE = re.compile(
r"""
^(?P<user>[\w.+-]+) # local part
@
(?P<domain>[\w-]+(?:\.[\w-]+)+)$ # domain with at least one dot
""",
re.VERBOSE,
)
print(EMAIL_RE.match("ada.lovelace+news@example.co.uk").groupdict())
print(EMAIL_RE.match("not-an-email"))
{'user': 'ada.lovelace+news', 'domain': 'example.co.uk'}
None
(This is a reasonable sanity check, not a full email validator. Real-world email validation is better done by sending a confirmation email.)
Unicode by Default
In Python 3, str patterns are Unicode-aware. \w matches letters from any script and \d matches digits from any script:
print(re.findall(r"\w+", "café naïve 東京")) # ['café', 'naïve', '東京']
print(re.findall(r"\d", "١٢٣ 123")) # ['١', '٢', '٣', '1', '2', '3']
print(re.findall(r"[0-9]", "١٢٣ 123")) # ['1', '2', '3']
That's usually what you want for text, but not for validating something like an ID or a port number, where int() would later choke on Arabic-Indic digits. Use [0-9] or the re.ASCII flag when you mean ASCII digits only.
Greedy vs Lazy Quantifiers
Quantifiers are greedy: they match as much as possible and only back off if the rest of the pattern fails. That causes the classic HTML-tag mistake:
html = "<b>bold</b> and <i>italic</i>"
print(re.findall(r"<.+>", html)) # ['<b>bold</b> and <i>italic</i>']
print(re.findall(r"<.+?>", html)) # ['<b>', '</b>', '<i>', '</i>']
print(re.findall(r"<[^>]+>", html)) # ['<b>', '</b>', '<i>', '</i>']
.+ grabbed everything up to the last >. Adding ? after a quantifier (+?, *?, ??) makes it lazy, matching as little as possible. A negated character class like [^>]+ is often even better: it states exactly what's allowed and avoids backtracking.
(And for real HTML, use a parser like Beautiful Soup, not regular expressions.)
Lookarounds
Lookarounds check what's before or after a position without including it in the match:
| Syntax | Name | Meaning |
|---|---|---|
(?=...) | Lookahead | Followed by ... |
(?!...) | Negative lookahead | Not followed by ... |
(?<=...) | Lookbehind | Preceded by ... |
(?<!...) | Negative lookbehind | Not preceded by ... |
print(re.findall(r"\d+(?= USD)", "10 USD, 20 EUR, 30 USD")) # ['10', '30']
print(re.findall(r"(?<=\$)\d+(?:\.\d{2})?", "Total: $42.50, tax $3")) # ['42.50', '3']
print(re.findall(r"(?<!-)\b\d+", "5 -3 12")) # ['5', '12']
The amounts come back without the currency markers because lookarounds are zero-width. One limitation in the standard re module: lookbehinds must have a fixed width, so (?<=\$+) isn't allowed.
Compiling Patterns
re.compile() turns a pattern into a Pattern object with the same methods (search, findall, sub, and so on):
import re
LOG_RE = re.compile(
r'(?P<ip>\S+) \S+ \S+ \[(?P<time>[^\]]+)\] '
r'"(?P<method>[A-Z]+) (?P<path>\S+) [^"]*" '
r'(?P<status>\d{3}) (?P<size>\d+|-)'
)
line = '127.0.0.1 - - [18/Sep/2026:10:15:32 +0000] "GET /blog/python HTTP/1.1" 200 5123'
print(LOG_RE.match(line).groupdict())
{'ip': '127.0.0.1', 'time': '18/Sep/2026:10:15:32 +0000', 'method': 'GET', 'path': '/blog/python', 'status': '200', 'size': '5123'}
Is compiling faster? A little, but less than people think: the module-level functions cache recently used compiled patterns internally. The real benefit is organization. A compiled pattern at module level, with a descriptive name, documents what it's for and keeps the regex out of your loop body.
Escaping User Input
If you build a pattern from a variable, special characters in it will be interpreted as regex syntax. re.escape() neutralizes them:
term = "c++"
print(re.findall(re.escape(term), "I like c++ and c")) # ['c++']
print(re.escape("price (USD): $5.00"))
['c++']
price\ \(USD\):\ \$5\.00
Without re.escape, "c++" is read as "one or more c" with a possessive quantifier (supported since Python 3.11), so it returns ['c', 'c'] instead of ['c++']. Similarly, "a.b" would match "axb", and a term with an unbalanced ( would raise an error. Escape any untrusted input before putting it into a pattern.
Common Pitfalls
- Forgetting the
rprefix. Use raw strings for every pattern. - Using
match()for validation. It only anchors at the start. Usefullmatch(). - Surprising
findall()output. Capturing groups change the return value. Use(?:...)orfinditer(). - Greedy
.*eating too much. Use a lazy quantifier or a negated character class. - Unescaped dots.
.matches any character, soexample.comalso matchesexampleXcom. Writeexample\.com. - Catastrophic backtracking. Nested quantifiers like
(a+)+$can take exponential time on inputs that almost match. If a regex runs on untrusted input, keep it simple, prefer specific character classes over.*, and test it against long, non-matching strings. - Invalid patterns. A malformed pattern raises
re.PatternError(the new name in Python 3.13;re.errorstill works as an alias), for examplemissing ), unterminated subpattern at position 0. - Reaching for regex when a string method works.
s.startswith("http"),"@" in s, ands.split(",")are faster and clearer. See the Python strings deep dive for what's available before you write a pattern.
Conclusion
The re module comes down to a few decisions: search() to find, fullmatch() to validate, finditer() or findall() to extract everything, sub() to rewrite, and split() to break text apart. Learn the core syntax (character classes, quantifiers, anchors, groups), lean on named groups and re.VERBOSE to keep patterns readable, and remember that quantifiers are greedy unless you add ?.
Write every pattern as a raw string, escape anything that comes from users, and test your regex against inputs that shouldn't match as carefully as the ones that should. With those habits, regular expressions stop being line noise and become one of the most useful tools you have for working with text.


