
Python Strings Deep Dive: Methods Every Developer Should Know
Python's str type has around 47 methods. Most developers use five or six of them regularly and reach for regular expressions or hand-written loops for everything else. That's a shame, because the built-in methods are fast, readable, and handle edge cases you'd probably forget, like Unicode case folding or splitting on any run of whitespace.
This post walks through the string methods worth knowing, grouped by the job they do: changing case, trimming, splitting and joining, searching, validating, padding and aligning, and character-level replacement. Along the way I'll point out the traps, such as why rstrip(".txt") doesn't do what it looks like and why isdigit() accepts characters int() rejects. String formatting gets its own treatment in the f-strings post, so I'll only touch on it here.
Strings Are Immutable
One fact shapes everything else: strings can't be changed in place.
s = "hello"
s[0] = "H"
TypeError: 'str' object does not support item assignment
Every string method returns a new string and leaves the original untouched:
t = s.upper()
print(s, t)
hello HELLO
The most common beginner bug with strings is calling name.strip() on its own line and expecting name to change. You always need to use the return value: name = name.strip().
Changing Case
The basic case methods are lower(), upper(), capitalize() (first character uppercase, rest lowercase), swapcase(), and title():
print("hELLO wORLD".capitalize(), "Hello".swapcase())
print("hello world".title(), "they're bill's".title())
Hello world hELLO
Hello World They'Re Bill'S
title() capitalizes the first letter after any non-letter, so apostrophes produce They'Re. For human-readable titles, string.capwords() splits on whitespace instead, which is usually what you want:
import string
print(string.capwords("they're bill's friends"))
They're Bill's Friends
casefold() for Case-Insensitive Comparison
When you compare strings without regard to case, use casefold() rather than lower(). It's a more aggressive normalization designed exactly for caseless matching:
print("Straße".lower(), "Straße".casefold(), "Straße".upper())
print("straße".lower() == "STRASSE".lower())
print("straße".casefold() == "STRASSE".casefold())
straße strasse STRASSE
False
True
The German ß uppercases to SS, so lower() can't match the two spellings, but casefold() maps ß to ss and they compare equal. For English-only ASCII text the two behave the same, so casefold() costs you nothing and protects you when the data turns out to be international.
Trimming: strip, lstrip, rstrip
With no arguments, strip() removes leading and trailing whitespace, which includes spaces, tabs, and newlines. lstrip() and rstrip() handle one side only:
raw = " \t Hello, World! \n"
print(repr(raw.strip()), repr(raw.lstrip()), repr(raw.rstrip()))
'Hello, World!' 'Hello, World! \n' ' \t Hello, World!'
rstrip() is the standard way to drop the trailing newline from lines read from a file.
The rstrip(".txt") Trap
When you pass an argument, it's treated as a set of characters, not a substring. Python strips any combination of those characters from the end:
print("xxhixyx".strip("xy"))
print("report.txt".rstrip(".txt"))
hi
repor
"report.txt".rstrip(".txt") removes t, x, t, ., and then keeps going because the t at the end of report is in the set too. The bug is nasty because it works fine for filenames like readme.txt (which ends in e before the extension) and only breaks on some inputs.
removeprefix and removesuffix
Python 3.9 added the methods people actually wanted. They remove an exact substring, and only if it's present:
print("report.txt".removesuffix(".txt"), "report.txt".removeprefix("draft-"))
report report.txt
The second call returns the string unchanged because it doesn't start with "draft-". Use these any time you're removing a known prefix or suffix. For file extensions specifically, pathlib.Path("report.txt").stem is even clearer.
Splitting
split() With and Without an Argument
split() behaves very differently depending on whether you give it a separator:
print("a,b,,c".split(","))
print(" many spaces here ".split())
print(" many spaces here ".split(" "))
['a', 'b', '', 'c']
['many', 'spaces', 'here']
['', '', 'many', '', '', 'spaces', 'here', '']
- With an explicit separator, every occurrence splits, and consecutive separators produce empty strings. That's correct for CSV-like data where an empty field is meaningful.
- With no argument (or
None), Python splits on runs of any whitespace and ignores leading and trailing whitespace entirely. That's what you want for splitting words or command-line-style input.
Passing " " explicitly almost never does what people expect. Default to split() for whitespace.
Limiting Splits
maxsplit caps the number of splits, and rsplit() counts from the right:
print("k=v=w".split("=", 1), "a/b/c/d.txt".rsplit("/", 1))
['k', 'v=w'] ['a/b/c', 'd.txt']
split("=", 1) is the right way to parse key=value pairs where the value might itself contain =.
splitlines()
To break text into lines, use splitlines() rather than split("\n"):
print("line1\nline2\r\nline3\n".splitlines())
print("line1\nline2\n".split("\n"))
print("line1\nline2\n".splitlines(keepends=True))
['line1', 'line2', 'line3']
['line1', 'line2', '']
['line1\n', 'line2\n']
splitlines() handles \r\n Windows line endings (and several other Unicode line boundaries), and it doesn't produce a trailing empty string when the text ends with a newline.
partition() and rpartition()
partition(sep) splits on the first occurrence and always returns a 3-tuple: the part before, the separator, and the part after. If the separator is missing, you get the whole string and two empty strings:
print("host:8080".partition(":"))
print("localhost".partition(":"))
print("a.b.c".rpartition("."))
('host', ':', '8080')
('localhost', '', '')
('a.b', '.', 'c')
Because the result always has three items, you can unpack it without checking the length first:
host, _, port = address.partition(":")
port = port or "80"
That's cleaner and safer than split(":") followed by an if len(parts) == 2 check.
Joining
str.join() is the inverse of split(). You call it on the separator and pass an iterable of strings:
print(", ".join(["red", "green", "blue"]))
red, green, blue
Every item must already be a string. Joining numbers fails:
", ".join([1, 2, 3])
TypeError: sequence item 0: expected str instance, int found
Convert them first with map(str, ...) or a generator expression:
print(", ".join(map(str, [1, 2, 3])))
1, 2, 3
Build Lists, Then Join
Because strings are immutable, result += piece in a loop creates a new string each time. CPython has an optimization that often makes this fast in practice, but it isn't guaranteed (other implementations don't have it, and it doesn't apply in every situation). The reliably efficient pattern is to collect pieces in a list and join once at the end:
parts = []
for i in range(3):
parts.append(f"item{i}")
print("".join(parts))
item0item1item2
For larger text-building jobs, io.StringIO works too.
Searching
in, find(), and index()
For "does it contain this?", use the in operator. For "where is it?", use find() or index():
text = "the cat sat on the mat"
print("cat" in text)
print(text.find("at"), text.find("dog"), text.rfind("at"))
True
5 -1 20
find() returns -1 when there's no match, while index() raises ValueError: substring not found. Be careful with find(): -1 is a valid index in Python (the last character), so code like text[text.find("x"):] silently does the wrong thing on a miss. Prefer index() inside a try when a miss is an error, or check in first. rfind() and rindex() search from the right.
count()
print(text.count("at"), text.count("the"))
3 2
count() counts non-overlapping occurrences.
startswith() and endswith() Take Tuples
Both methods accept a tuple of options, which replaces chains of or:
print("photo.JPG".lower().endswith((".jpg", ".jpeg", ".png")))
print(text.startswith("cat", 4))
True
True
The optional second argument is a start position, so text.startswith("cat", 4) checks for "cat" at index 4 without slicing.
replace()
replace(old, new) replaces every occurrence. An optional count limits how many:
print(text.replace("the", "a"))
print(text.replace("the", "a", 1))
a cat sat on a mat
a cat sat on the mat
When you need patterns rather than fixed text, that's the point to switch to the re module, covered in the regular expressions post.
Validation: the is* Methods
These methods return True only for non-empty strings where every character passes the test:
print("abc".isalpha(), "abc123".isalnum(), " ".isspace(), "".isspace())
print("Hello World".istitle(), "my_var".isidentifier())
print("ascii".isascii(), "café".isascii())
True True True False
True True
True False
Note that "".isspace() is False: empty strings fail every is* check except isascii(), which returns True for an empty string. Also note that "class".isidentifier() returns True; to rule out keywords, combine it with keyword.iskeyword().
isdecimal() vs isdigit() vs isnumeric()
These three look interchangeable and aren't:
for v in ["42", "4.2", "-3", "²", "½", "٣", " 7"]:
print(repr(v), v.isdecimal(), v.isdigit(), v.isnumeric())
'42' True True True
'4.2' False False False
'-3' False False False
'²' False True True
'½' False False True
'٣' True True True
' 7' False False False
isdecimal()is the strictest: characters usable to write base-10 numbers. That includes Arabic-Indic digits like٣, whichint()happily accepts.isdigit()also accepts superscripts like², whichint()rejects.isnumeric()adds fractions like½and other numeric characters.
None of them accept signs, decimal points, or surrounding whitespace. If you want to know whether int() will succeed, isdecimal() is the closest match, but the most reliable check is to just call int() and catch ValueError. The post on converting between data types covers that pattern.
Padding and Alignment
center(), ljust(), and rjust() pad a string to a given width, with an optional fill character. zfill() pads numbers with zeros and respects the sign:
print("|" + "hi".center(10) + "|", "|" + "hi".ljust(6, ".") + "|", "|" + "hi".rjust(6) + "|")
print("42".zfill(5), "-42".zfill(5))
| hi | |hi....| | hi|
00042 -0042
They're handy for simple text reports:
for name, price in [("Coffee", 3.5), ("Bagel", 12.25)]:
print(name.ljust(10, ".") + f"{price:>8.2f}")
Coffee.... 3.50
Bagel..... 12.25
The f-string format spec can do the same alignment (f"{name:.<10}"), so pick whichever reads better in context.
Character-Level Replacement with translate()
To replace or delete many individual characters at once, build a translation table with str.maketrans() and apply it with translate(). One pass handles every character:
table = str.maketrans({"-": "_", " ": "_", "!": None})
print("hello world-again!".translate(table))
print("hello".translate(str.maketrans("el", "ip")))
hello_world_again
hippo
The dictionary form maps characters to replacement strings, or to None to delete them. The two-string form maps each character in the first string to the character at the same position in the second. This is much faster and clearer than chaining several replace() calls.
Other Useful Odds and Ends
expandtabs(n)converts tab characters to spaces aligned to tab stops of widthn.format_map(mapping)works likestr.format(**mapping)but takes the mapping directly, which lets you pass adictsubclass with a__missing__method for default values.- The
textwrapmodule (not a method, but closely related) hasfill()for wrapping paragraphs,shorten()for truncating with a placeholder, anddedent()for removing common indentation from triple-quoted strings.
import textwrap
print(textwrap.shorten("The quick brown fox jumps over the lazy dog", width=25))
The quick brown fox [...]
Putting It Together
Here's a small function that combines several methods into a URL slug generator:
def slugify(title: str) -> str:
keep = "".join(ch if ch.isalnum() else " " for ch in title.casefold())
return "-".join(keep.split())
print(slugify(" Python Strings: Deep Dive! "))
python-strings-deep-dive
casefold() normalizes case, the generator replaces punctuation with spaces, split() with no argument collapses all the whitespace runs, and "-".join() reassembles the words. No regex required. (Accented letters pass isalnum() and are kept as-is; normalizing them to ASCII involves unicodedata, which the bytes, encoding, and Unicode post touches on.)
Quick Reference
| Task | Method |
|---|---|
| Case-insensitive compare | a.casefold() == b.casefold() |
| Trim whitespace | strip(), lstrip(), rstrip() |
| Remove an exact prefix/suffix | removeprefix(), removesuffix() |
| Split into words | split() (no argument) |
| Split into lines | splitlines() |
| Split once on a separator | partition(), split(sep, 1) |
| Join pieces | sep.join(iterable) |
| Contains / position | in, find(), index() |
| Multiple prefixes/suffixes | startswith((...)), endswith((...)) |
| Many single-character replacements | translate(str.maketrans(...)) |
| Check for decimal digits | isdecimal() |
| Zero-pad a number | zfill() |
Conclusion
Most string-processing code gets shorter and more correct once you know what's built in. Use casefold() for comparisons, removeprefix() and removesuffix() instead of strip() with an argument, split() with no argument for whitespace, partition() when you need exactly one split, join() for building strings, and translate() for bulk character changes.
And remember that every method returns a new string. If a string "isn't changing", the fix is almost always to assign the result.


