
Web Scraping in Python with Requests and Beautiful Soup
Sometimes the data you need is sitting on a web page with no API in sight: a product catalog, a table of results, a list of events. Web scraping is how you get it out. In Python, the classic pairing is Requests to download the HTML and Beautiful Soup to parse it, and for server-rendered pages it's still the simplest tool for the job.
The hard part of scraping isn't the first request. It's writing a scraper that handles encodings, relative links, pagination, missing elements, and flaky responses, without hammering the site you're scraping. This post builds one step by step against Books to Scrape, a sandbox site made specifically for practicing scraping, so you can run every example yourself.
Before You Scrape
A few ground rules worth taking seriously:
- Check for an API first. If the site offers one, it's more stable and usually allowed by the terms.
- Read the terms of service and
robots.txt. Some sites prohibit scraping outright.robots.txtisn't legally binding everywhere, but it tells you what the owner wants crawled. - Be gentle. Add delays between requests, identify yourself with a
User-Agent, and cache pages during development so you aren't refetching them on every run. - Be careful with personal data. Scraping personal information can bring privacy laws like GDPR into play regardless of whether the page is public.
Installing the Libraries
python -m pip install requests beautifulsoup4
Note the package name is beautifulsoup4, but you import it as bs4. If you need a refresher on installing packages, see what pip is and how to use it.
Fetching a Page with Requests
# first_request.py
import requests
response = requests.get("https://books.toscrape.com/", timeout=10)
print(response.status_code)
print(response.headers["Content-Type"])
print(response.encoding)
200
text/html
ISO-8859-1
Always pass timeout. Requests has no default timeout, so without one a stalled server can hang your script forever. timeout=10 means "give up if the connection or a read stalls for 10 seconds".
That third line hides the most common scraping bug. The server sent Content-Type: text/html with no charset, and in that case Requests falls back to ISO-8859-1 as the HTTP spec historically required. The page is actually UTF-8, so if you decode it with response.text, the pound sign in the prices turns into £.
The fix is to give Beautiful Soup the raw bytes, response.content, instead of response.text. Beautiful Soup reads the <meta charset> tag inside the HTML and decodes it correctly:
from bs4 import BeautifulSoup
soup = BeautifulSoup(response.content, "html.parser") # bytes, not text
Checking for Errors
requests.get doesn't raise on a 404 or 500. It returns a response with that status code. Call raise_for_status() to turn error statuses into an HTTPError exception:
response = requests.get("https://books.toscrape.com/no-such-page", timeout=10)
response.raise_for_status() # raises requests.HTTPError: 404 Client Error: Not Found
Parsing HTML with Beautiful Soup
Beautiful Soup turns HTML into a tree you can search. Open the page in your browser, right-click a product, and choose "Inspect" to see the structure. Each book on the listing page looks roughly like this:
<article class="product_pod">
<div class="image_container">...</div>
<p class="star-rating Three">...</p>
<h3>
<a
href="catalogue/a-light-in-the-attic_1000/index.html"
title="A Light in the Attic"
>A Light in the ...</a
>
</h3>
<div class="product_price">
<p class="price_color">£51.77</p>
<p class="instock availability">In stock</p>
</div>
</article>
Notice that the visible link text is truncated ("A Light in the ..."), while the title attribute has the full title. Inspecting the real HTML instead of trusting what you see on screen saves a lot of confusion.
Here's how to pull data out of that structure:
# parse_listing.py
import requests
from bs4 import BeautifulSoup
response = requests.get("https://books.toscrape.com/", timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")
print(soup.title.get_text(strip=True))
first = soup.select_one("article.product_pod")
print(first.h3.a["title"])
print(first.select_one("p.price_color").get_text())
print(first.select_one("p.star-rating")["class"])
print(first.h3.a.get("href"))
print(len(soup.select("article.product_pod")))
All products | Books to Scrape - Sandbox
A Light in the Attic
£51.77
['star-rating', 'Three']
catalogue/a-light-in-the-attic_1000/index.html
20
The pieces you'll use constantly:
soup.select_one(css)returns the first element matching a CSS selector, orNone.soup.select(css)returns a list of all matches.element["attr"]reads an attribute and raisesKeyErrorif it's missing;element.get("attr")returnsNoneinstead.element.get_text(strip=True)returns the text content with surrounding whitespace removed.element.h3.ais a shortcut for "the firsth3inside, then the firstainside that".classis a multi-valued attribute, so it comes back as a list.
CSS Selectors vs find and find_all
Beautiful Soup has two search styles. The older one uses find and find_all with keyword filters:
next_item = soup.find("li", class_="next")
links = soup.find_all("a", limit=3)
class_ has a trailing underscore because class is a Python keyword. The newer style uses CSS selectors through select and select_one, powered by the Soup Sieve library that's installed alongside Beautiful Soup. Both work; selectors are usually shorter and match what you see in browser dev tools, so I'll use them from here on.
Some selectors that come up a lot in scraping:
| Selector | Matches |
|---|---|
article.product_pod | article elements with class product_pod |
#product_description | The element with that id |
div.product_price > p | p elements that are direct children |
a[href^="https"] | Links whose href starts with https |
tr:nth-of-type(2) td | Cells in the second table row |
#product_description + p | The p immediately after that element |
Parsers
The second argument to BeautifulSoup picks the parser. "html.parser" is built into Python and needs nothing extra. "lxml" is faster and more lenient with broken markup but requires pip install lxml. For most scraping jobs, html.parser is fine; switch to lxml if you're parsing thousands of pages and speed matters.
Scraping a Detail Page
Listing pages give you a summary. Detail pages often hold the data you really want, frequently in a table:
# detail_page.py
import requests
from bs4 import BeautifulSoup
url = "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
response = requests.get(url, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")
info = {
row.th.get_text(strip=True): row.td.get_text(strip=True)
for row in soup.select("table.table-striped tr")
}
print(info["UPC"], info["Availability"])
description = soup.select_one("#product_description + p")
print(description.get_text(strip=True)[:60])
a897fe39b1053632 In stock (22 available)
It's hard to imagine a world without A Light in the Attic. T
The dict comprehension turns a two-column table of th/td rows into a lookup by label, which is much more robust than relying on row positions. The description has no class of its own, so the adjacent-sibling selector #product_description + p grabs the paragraph that follows the heading with that id.
Building a Robust Scraper
Now let's put it together into a scraper that follows pagination, retries transient errors, respects robots.txt, and writes a CSV.
A Session with Retries and a User-Agent
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
USER_AGENT = "tidewave-book-scraper/1.0 (+mailto:you@example.com)"
def make_session() -> requests.Session:
retry = Retry(
total=3,
backoff_factor=1,
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET"],
)
session = requests.Session()
session.headers["User-Agent"] = USER_AGENT
session.mount("https://", HTTPAdapter(max_retries=retry))
return session
A Session reuses the underlying TCP connection across requests, which is faster and kinder to the server than opening a new connection every time. It also stores headers and cookies, so the User-Agent applies to every request.
Mounting an HTTPAdapter with a urllib3 Retry gives you automatic retries with exponential backoff for rate limiting (429) and server errors (5xx). With backoff_factor=1 and urllib3 2.x, the first retry happens immediately and later ones wait 2s, then 4s. urllib3 also honors a Retry-After header if the server sends one.
Respecting robots.txt
The standard library can parse robots.txt for you:
from urllib.parse import urljoin
from urllib.robotparser import RobotFileParser
BASE_URL = "https://books.toscrape.com/"
def allowed_by_robots(url: str) -> bool:
parser = RobotFileParser(urljoin(BASE_URL, "/robots.txt"))
parser.read()
return parser.can_fetch(USER_AGENT, url)
RobotFileParser treats a missing robots.txt (a 404) as "everything allowed" and a 401 or 403 as "nothing allowed". Books to Scrape doesn't have one, so this returns True. On a real site, check once at startup and also read the site's terms.
Parsing Each Item into a Dataclass
from dataclasses import dataclass
from decimal import Decimal
RATINGS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}
@dataclass
class Book:
title: str
price: Decimal
rating: int
in_stock: bool
url: str
def parse_book(card, page_url: str) -> Book:
link = card.select_one("h3 a")
rating_class = card.select_one("p.star-rating")["class"][1]
price_text = card.select_one("p.price_color").get_text(strip=True)
stock_text = card.select_one("p.availability").get_text(strip=True)
return Book(
title=link["title"],
price=Decimal(price_text.lstrip("£")),
rating=RATINGS.get(rating_class, 0),
in_stock="In stock" in stock_text,
url=urljoin(page_url, link["href"]),
)
This is where scraped strings become real data. The star rating lives in a class name (star-rating Three), so it's mapped to an integer. Prices become Decimal rather than float to avoid rounding surprises with money.
The urljoin(page_url, link["href"]) call matters more than it looks. The href values are relative (catalogue/...), and what they're relative to changes depending on which page you're on. Page 1 is at the site root, but page 2 lives under /catalogue/, where the links drop the catalogue/ prefix. Joining against the URL of the page you actually fetched always produces the right absolute URL.
Following Pagination
import time
def fetch(session: requests.Session, url: str) -> BeautifulSoup:
response = session.get(url, timeout=10)
response.raise_for_status()
return BeautifulSoup(response.content, "html.parser")
def scrape(max_pages: int = 3, delay: float = 1.0) -> list[Book]:
session = make_session()
url: str | None = BASE_URL
books: list[Book] = []
for _ in range(max_pages):
if url is None:
break
soup = fetch(session, url)
books.extend(parse_book(card, url) for card in soup.select("article.product_pod"))
next_link = soup.select_one("li.next a")
url = urljoin(url, next_link["href"]) if next_link else None
time.sleep(delay)
return books
Rather than guessing URLs like page-1.html, page-2.html, the scraper follows the "next" link until there isn't one. That keeps working if the site changes its URL scheme, and it stops naturally on the last page. max_pages is a safety limit so a bug can't send you crawling forever, and time.sleep(delay) spaces out requests.
The Complete Script
# scraper.py
import csv
import time
from dataclasses import asdict, dataclass, fields
from decimal import Decimal
from urllib.parse import urljoin
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
BASE_URL = "https://books.toscrape.com/"
USER_AGENT = "tidewave-book-scraper/1.0 (+mailto:you@example.com)"
RATINGS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}
@dataclass
class Book:
title: str
price: Decimal
rating: int
in_stock: bool
url: str
def make_session() -> requests.Session:
retry = Retry(
total=3,
backoff_factor=1,
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET"],
)
session = requests.Session()
session.headers["User-Agent"] = USER_AGENT
session.mount("https://", HTTPAdapter(max_retries=retry))
return session
def allowed_by_robots(url: str) -> bool:
parser = RobotFileParser(urljoin(BASE_URL, "/robots.txt"))
parser.read()
return parser.can_fetch(USER_AGENT, url)
def fetch(session: requests.Session, url: str) -> BeautifulSoup:
response = session.get(url, timeout=10)
response.raise_for_status()
return BeautifulSoup(response.content, "html.parser")
def parse_book(card, page_url: str) -> Book:
link = card.select_one("h3 a")
rating_class = card.select_one("p.star-rating")["class"][1]
price_text = card.select_one("p.price_color").get_text(strip=True)
stock_text = card.select_one("p.availability").get_text(strip=True)
return Book(
title=link["title"],
price=Decimal(price_text.lstrip("£")),
rating=RATINGS.get(rating_class, 0),
in_stock="In stock" in stock_text,
url=urljoin(page_url, link["href"]),
)
def scrape(max_pages: int = 3, delay: float = 1.0) -> list[Book]:
session = make_session()
url: str | None = BASE_URL
books: list[Book] = []
for _ in range(max_pages):
if url is None:
break
soup = fetch(session, url)
books.extend(parse_book(card, url) for card in soup.select("article.product_pod"))
next_link = soup.select_one("li.next a")
url = urljoin(url, next_link["href"]) if next_link else None
time.sleep(delay)
return books
def save_csv(books: list[Book], path: str) -> None:
with open(path, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=[f.name for f in fields(Book)])
writer.writeheader()
writer.writerows(asdict(book) for book in books)
if __name__ == "__main__":
if not allowed_by_robots(BASE_URL):
raise SystemExit("robots.txt disallows scraping this site")
books = scrape(max_pages=3)
save_csv(books, "books.csv")
print(f"Saved {len(books)} books")
for book in books[:3]:
print(f"{book.rating}/5 £{book.price:>6} {book.title}")
python scraper.py
Saved 60 books
3/5 £ 51.77 A Light in the Attic
1/5 £ 53.74 Tipping the Velvet
1/5 £ 50.10 Soumission
save_csv uses csv.DictWriter with field names taken from the dataclass, and asdict converts each Book to a dict. Opening the file with newline="" is what the csv module docs require to avoid blank lines on Windows. The result loads straight into a spreadsheet or into pandas with pd.read_csv("books.csv").
Handling Missing Elements
Real pages are inconsistent. A product might not have a rating, or a sale price might replace the regular one. select_one returns None when nothing matches, and calling .get_text() on None raises AttributeError. A small helper keeps parsing code clean:
def text_or_none(parent, selector: str) -> str | None:
element = parent.select_one(selector)
return element.get_text(strip=True) if element else None
Decide per field whether a missing value should be None, a default, or a hard error. For required fields, failing loudly is usually better than silently writing empty rows, because it tells you when the site's markup has changed.
When Requests and Beautiful Soup Aren't Enough
This stack downloads the HTML the server sends. It doesn't run JavaScript. If you fetch a page and the data you see in your browser isn't in response.content, the page is rendering it client-side. Before reaching for a browser automation tool, open your browser's dev tools, go to the Network tab, filter by Fetch/XHR, and reload. Very often the page loads its data from a JSON endpoint you can call directly with Requests, which is faster and more reliable than parsing HTML at all.
If there's no such endpoint, a headless browser like Playwright can render the page for you, and you can still hand the rendered HTML to Beautiful Soup for parsing.
For large crawls across many pages and domains, a framework like Scrapy adds scheduling, concurrency, throttling, and pipelines that you'd otherwise build yourself.
Practical Tips
- Cache while developing. Save fetched HTML to disk and parse from the file while you work out selectors, so you only hit the site once.
- Prefer stable hooks. IDs,
data-*attributes, and semantic class names change less often than deep positional selectors likediv > div:nth-child(3) span. - Log what you skip. When a card fails to parse, log the URL so you can inspect it rather than losing data quietly.
- Validate the output. Check row counts and spot-check values. Scrapers fail silently when markup changes.
Conclusion
Requests and Beautiful Soup cover a large share of scraping work: fetch the page with a timeout, parse response.content so the encoding comes out right, select elements with CSS selectors, and convert strings into typed data. Wrap that in a session with retries, follow "next" links with urljoin, pause between requests, and you have a scraper that's both robust and polite.
When the data isn't in the HTML, look for the JSON endpoint behind the page before reaching for a headless browser. And whatever you build, check the site's rules first and scrape at a pace the site won't notice.


