Type something to search...
Python Strings Deep Dive: Methods Every Developer Should Know

Python Strings Deep Dive: Methods Every Developer Should Know

Python's str type has around 47 methods. Most developers use five or six of them regularly and reach for regular expressions or hand-written loops for everything else. That's a shame, because the built-in methods are fast, readable, and handle edge cases you'd probably forget, like Unicode case folding or splitting on any run of whitespace.

This post walks through the string methods worth knowing, grouped by the job they do: changing case, trimming, splitting and joining, searching, validating, padding and aligning, and character-level replacement. Along the way I'll point out the traps, such as why rstrip(".txt") doesn't do what it looks like and why isdigit() accepts characters int() rejects. String formatting gets its own treatment in the f-strings post, so I'll only touch on it here.

Strings Are Immutable

One fact shapes everything else: strings can't be changed in place.

s = "hello"
s[0] = "H"
TypeError: 'str' object does not support item assignment

Every string method returns a new string and leaves the original untouched:

t = s.upper()
print(s, t)
hello HELLO

The most common beginner bug with strings is calling name.strip() on its own line and expecting name to change. You always need to use the return value: name = name.strip().

Changing Case

The basic case methods are lower(), upper(), capitalize() (first character uppercase, rest lowercase), swapcase(), and title():

print("hELLO wORLD".capitalize(), "Hello".swapcase())
print("hello world".title(), "they're bill's".title())
Hello world hELLO
Hello World They'Re Bill'S

title() capitalizes the first letter after any non-letter, so apostrophes produce They'Re. For human-readable titles, string.capwords() splits on whitespace instead, which is usually what you want:

import string

print(string.capwords("they're bill's friends"))
They're Bill's Friends

casefold() for Case-Insensitive Comparison

When you compare strings without regard to case, use casefold() rather than lower(). It's a more aggressive normalization designed exactly for caseless matching:

print("Straße".lower(), "Straße".casefold(), "Straße".upper())
print("straße".lower() == "STRASSE".lower())
print("straße".casefold() == "STRASSE".casefold())
straße strasse STRASSE
False
True

The German ß uppercases to SS, so lower() can't match the two spellings, but casefold() maps ß to ss and they compare equal. For English-only ASCII text the two behave the same, so casefold() costs you nothing and protects you when the data turns out to be international.

Trimming: strip, lstrip, rstrip

With no arguments, strip() removes leading and trailing whitespace, which includes spaces, tabs, and newlines. lstrip() and rstrip() handle one side only:

raw = "  \t Hello, World! \n"
print(repr(raw.strip()), repr(raw.lstrip()), repr(raw.rstrip()))
'Hello, World!' 'Hello, World! \n' '  \t Hello, World!'

rstrip() is the standard way to drop the trailing newline from lines read from a file.

The rstrip(".txt") Trap

When you pass an argument, it's treated as a set of characters, not a substring. Python strips any combination of those characters from the end:

print("xxhixyx".strip("xy"))
print("report.txt".rstrip(".txt"))
hi
repor

"report.txt".rstrip(".txt") removes t, x, t, ., and then keeps going because the t at the end of report is in the set too. The bug is nasty because it works fine for filenames like readme.txt (which ends in e before the extension) and only breaks on some inputs.

removeprefix and removesuffix

Python 3.9 added the methods people actually wanted. They remove an exact substring, and only if it's present:

print("report.txt".removesuffix(".txt"), "report.txt".removeprefix("draft-"))
report report.txt

The second call returns the string unchanged because it doesn't start with "draft-". Use these any time you're removing a known prefix or suffix. For file extensions specifically, pathlib.Path("report.txt").stem is even clearer.

Splitting

split() With and Without an Argument

split() behaves very differently depending on whether you give it a separator:

print("a,b,,c".split(","))
print("  many   spaces here ".split())
print("  many   spaces here ".split(" "))
['a', 'b', '', 'c']
['many', 'spaces', 'here']
['', '', 'many', '', '', 'spaces', 'here', '']
  • With an explicit separator, every occurrence splits, and consecutive separators produce empty strings. That's correct for CSV-like data where an empty field is meaningful.
  • With no argument (or None), Python splits on runs of any whitespace and ignores leading and trailing whitespace entirely. That's what you want for splitting words or command-line-style input.

Passing " " explicitly almost never does what people expect. Default to split() for whitespace.

Limiting Splits

maxsplit caps the number of splits, and rsplit() counts from the right:

print("k=v=w".split("=", 1), "a/b/c/d.txt".rsplit("/", 1))
['k', 'v=w'] ['a/b/c', 'd.txt']

split("=", 1) is the right way to parse key=value pairs where the value might itself contain =.

splitlines()

To break text into lines, use splitlines() rather than split("\n"):

print("line1\nline2\r\nline3\n".splitlines())
print("line1\nline2\n".split("\n"))
print("line1\nline2\n".splitlines(keepends=True))
['line1', 'line2', 'line3']
['line1', 'line2', '']
['line1\n', 'line2\n']

splitlines() handles \r\n Windows line endings (and several other Unicode line boundaries), and it doesn't produce a trailing empty string when the text ends with a newline.

partition() and rpartition()

partition(sep) splits on the first occurrence and always returns a 3-tuple: the part before, the separator, and the part after. If the separator is missing, you get the whole string and two empty strings:

print("host:8080".partition(":"))
print("localhost".partition(":"))
print("a.b.c".rpartition("."))
('host', ':', '8080')
('localhost', '', '')
('a.b', '.', 'c')

Because the result always has three items, you can unpack it without checking the length first:

host, _, port = address.partition(":")
port = port or "80"

That's cleaner and safer than split(":") followed by an if len(parts) == 2 check.

Joining

str.join() is the inverse of split(). You call it on the separator and pass an iterable of strings:

print(", ".join(["red", "green", "blue"]))
red, green, blue

Every item must already be a string. Joining numbers fails:

", ".join([1, 2, 3])
TypeError: sequence item 0: expected str instance, int found

Convert them first with map(str, ...) or a generator expression:

print(", ".join(map(str, [1, 2, 3])))
1, 2, 3

Build Lists, Then Join

Because strings are immutable, result += piece in a loop creates a new string each time. CPython has an optimization that often makes this fast in practice, but it isn't guaranteed (other implementations don't have it, and it doesn't apply in every situation). The reliably efficient pattern is to collect pieces in a list and join once at the end:

parts = []
for i in range(3):
    parts.append(f"item{i}")

print("".join(parts))
item0item1item2

For larger text-building jobs, io.StringIO works too.

Searching

in, find(), and index()

For "does it contain this?", use the in operator. For "where is it?", use find() or index():

text = "the cat sat on the mat"

print("cat" in text)
print(text.find("at"), text.find("dog"), text.rfind("at"))
True
5 -1 20

find() returns -1 when there's no match, while index() raises ValueError: substring not found. Be careful with find(): -1 is a valid index in Python (the last character), so code like text[text.find("x"):] silently does the wrong thing on a miss. Prefer index() inside a try when a miss is an error, or check in first. rfind() and rindex() search from the right.

count()

print(text.count("at"), text.count("the"))
3 2

count() counts non-overlapping occurrences.

startswith() and endswith() Take Tuples

Both methods accept a tuple of options, which replaces chains of or:

print("photo.JPG".lower().endswith((".jpg", ".jpeg", ".png")))
print(text.startswith("cat", 4))
True
True

The optional second argument is a start position, so text.startswith("cat", 4) checks for "cat" at index 4 without slicing.

replace()

replace(old, new) replaces every occurrence. An optional count limits how many:

print(text.replace("the", "a"))
print(text.replace("the", "a", 1))
a cat sat on a mat
a cat sat on the mat

When you need patterns rather than fixed text, that's the point to switch to the re module, covered in the regular expressions post.

Validation: the is* Methods

These methods return True only for non-empty strings where every character passes the test:

print("abc".isalpha(), "abc123".isalnum(), "  ".isspace(), "".isspace())
print("Hello World".istitle(), "my_var".isidentifier())
print("ascii".isascii(), "café".isascii())
True True True False
True True
True False

Note that "".isspace() is False: empty strings fail every is* check except isascii(), which returns True for an empty string. Also note that "class".isidentifier() returns True; to rule out keywords, combine it with keyword.iskeyword().

isdecimal() vs isdigit() vs isnumeric()

These three look interchangeable and aren't:

for v in ["42", "4.2", "-3", "²", "½", "٣", " 7"]:
    print(repr(v), v.isdecimal(), v.isdigit(), v.isnumeric())
'42' True True True
'4.2' False False False
'-3' False False False
'²' False True True
'½' False False True
'٣' True True True
' 7' False False False
  • isdecimal() is the strictest: characters usable to write base-10 numbers. That includes Arabic-Indic digits like ٣, which int() happily accepts.
  • isdigit() also accepts superscripts like ², which int() rejects.
  • isnumeric() adds fractions like ½ and other numeric characters.

None of them accept signs, decimal points, or surrounding whitespace. If you want to know whether int() will succeed, isdecimal() is the closest match, but the most reliable check is to just call int() and catch ValueError. The post on converting between data types covers that pattern.

Padding and Alignment

center(), ljust(), and rjust() pad a string to a given width, with an optional fill character. zfill() pads numbers with zeros and respects the sign:

print("|" + "hi".center(10) + "|", "|" + "hi".ljust(6, ".") + "|", "|" + "hi".rjust(6) + "|")
print("42".zfill(5), "-42".zfill(5))
|    hi    | |hi....| |    hi|
00042 -0042

They're handy for simple text reports:

for name, price in [("Coffee", 3.5), ("Bagel", 12.25)]:
    print(name.ljust(10, ".") + f"{price:>8.2f}")
Coffee....    3.50
Bagel.....   12.25

The f-string format spec can do the same alignment (f"{name:.<10}"), so pick whichever reads better in context.

Character-Level Replacement with translate()

To replace or delete many individual characters at once, build a translation table with str.maketrans() and apply it with translate(). One pass handles every character:

table = str.maketrans({"-": "_", " ": "_", "!": None})
print("hello world-again!".translate(table))

print("hello".translate(str.maketrans("el", "ip")))
hello_world_again
hippo

The dictionary form maps characters to replacement strings, or to None to delete them. The two-string form maps each character in the first string to the character at the same position in the second. This is much faster and clearer than chaining several replace() calls.

Other Useful Odds and Ends

  • expandtabs(n) converts tab characters to spaces aligned to tab stops of width n.
  • format_map(mapping) works like str.format(**mapping) but takes the mapping directly, which lets you pass a dict subclass with a __missing__ method for default values.
  • The textwrap module (not a method, but closely related) has fill() for wrapping paragraphs, shorten() for truncating with a placeholder, and dedent() for removing common indentation from triple-quoted strings.
import textwrap

print(textwrap.shorten("The quick brown fox jumps over the lazy dog", width=25))
The quick brown fox [...]

Putting It Together

Here's a small function that combines several methods into a URL slug generator:

def slugify(title: str) -> str:
    keep = "".join(ch if ch.isalnum() else " " for ch in title.casefold())
    return "-".join(keep.split())


print(slugify("  Python Strings: Deep Dive!  "))
python-strings-deep-dive

casefold() normalizes case, the generator replaces punctuation with spaces, split() with no argument collapses all the whitespace runs, and "-".join() reassembles the words. No regex required. (Accented letters pass isalnum() and are kept as-is; normalizing them to ASCII involves unicodedata, which the bytes, encoding, and Unicode post touches on.)

Quick Reference

TaskMethod
Case-insensitive comparea.casefold() == b.casefold()
Trim whitespacestrip(), lstrip(), rstrip()
Remove an exact prefix/suffixremoveprefix(), removesuffix()
Split into wordssplit() (no argument)
Split into linessplitlines()
Split once on a separatorpartition(), split(sep, 1)
Join piecessep.join(iterable)
Contains / positionin, find(), index()
Multiple prefixes/suffixesstartswith((...)), endswith((...))
Many single-character replacementstranslate(str.maketrans(...))
Check for decimal digitsisdecimal()
Zero-pad a numberzfill()

Conclusion

Most string-processing code gets shorter and more correct once you know what's built in. Use casefold() for comparisons, removeprefix() and removesuffix() instead of strip() with an argument, split() with no argument for whitespace, partition() when you need exactly one split, join() for building strings, and translate() for bulk character changes.

And remember that every method returns a new string. If a string "isn't changing", the fix is almost always to assign the result.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading