Type something to search...
Web Scraping in Python with Requests and Beautiful Soup

Web Scraping in Python with Requests and Beautiful Soup

Sometimes the data you need is sitting on a web page with no API in sight: a product catalog, a table of results, a list of events. Web scraping is how you get it out. In Python, the classic pairing is Requests to download the HTML and Beautiful Soup to parse it, and for server-rendered pages it's still the simplest tool for the job.

The hard part of scraping isn't the first request. It's writing a scraper that handles encodings, relative links, pagination, missing elements, and flaky responses, without hammering the site you're scraping. This post builds one step by step against Books to Scrape, a sandbox site made specifically for practicing scraping, so you can run every example yourself.

Before You Scrape

A few ground rules worth taking seriously:

  • Check for an API first. If the site offers one, it's more stable and usually allowed by the terms.
  • Read the terms of service and robots.txt. Some sites prohibit scraping outright. robots.txt isn't legally binding everywhere, but it tells you what the owner wants crawled.
  • Be gentle. Add delays between requests, identify yourself with a User-Agent, and cache pages during development so you aren't refetching them on every run.
  • Be careful with personal data. Scraping personal information can bring privacy laws like GDPR into play regardless of whether the page is public.

Installing the Libraries

python -m pip install requests beautifulsoup4

Note the package name is beautifulsoup4, but you import it as bs4. If you need a refresher on installing packages, see what pip is and how to use it.

Fetching a Page with Requests

# first_request.py
import requests

response = requests.get("https://books.toscrape.com/", timeout=10)
print(response.status_code)
print(response.headers["Content-Type"])
print(response.encoding)
200
text/html
ISO-8859-1

Always pass timeout. Requests has no default timeout, so without one a stalled server can hang your script forever. timeout=10 means "give up if the connection or a read stalls for 10 seconds".

That third line hides the most common scraping bug. The server sent Content-Type: text/html with no charset, and in that case Requests falls back to ISO-8859-1 as the HTTP spec historically required. The page is actually UTF-8, so if you decode it with response.text, the pound sign in the prices turns into £.

The fix is to give Beautiful Soup the raw bytes, response.content, instead of response.text. Beautiful Soup reads the <meta charset> tag inside the HTML and decodes it correctly:

from bs4 import BeautifulSoup

soup = BeautifulSoup(response.content, "html.parser")  # bytes, not text

Checking for Errors

requests.get doesn't raise on a 404 or 500. It returns a response with that status code. Call raise_for_status() to turn error statuses into an HTTPError exception:

response = requests.get("https://books.toscrape.com/no-such-page", timeout=10)
response.raise_for_status()  # raises requests.HTTPError: 404 Client Error: Not Found

Parsing HTML with Beautiful Soup

Beautiful Soup turns HTML into a tree you can search. Open the page in your browser, right-click a product, and choose "Inspect" to see the structure. Each book on the listing page looks roughly like this:

<article class="product_pod">
  <div class="image_container">...</div>
  <p class="star-rating Three">...</p>
  <h3>
    <a
      href="catalogue/a-light-in-the-attic_1000/index.html"
      title="A Light in the Attic"
      >A Light in the ...</a
    >
  </h3>
  <div class="product_price">
    <p class="price_color">£51.77</p>
    <p class="instock availability">In stock</p>
  </div>
</article>

Notice that the visible link text is truncated ("A Light in the ..."), while the title attribute has the full title. Inspecting the real HTML instead of trusting what you see on screen saves a lot of confusion.

Here's how to pull data out of that structure:

# parse_listing.py
import requests
from bs4 import BeautifulSoup

response = requests.get("https://books.toscrape.com/", timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")

print(soup.title.get_text(strip=True))

first = soup.select_one("article.product_pod")
print(first.h3.a["title"])
print(first.select_one("p.price_color").get_text())
print(first.select_one("p.star-rating")["class"])
print(first.h3.a.get("href"))
print(len(soup.select("article.product_pod")))
All products | Books to Scrape - Sandbox
A Light in the Attic
£51.77
['star-rating', 'Three']
catalogue/a-light-in-the-attic_1000/index.html
20

The pieces you'll use constantly:

  • soup.select_one(css) returns the first element matching a CSS selector, or None.
  • soup.select(css) returns a list of all matches.
  • element["attr"] reads an attribute and raises KeyError if it's missing; element.get("attr") returns None instead.
  • element.get_text(strip=True) returns the text content with surrounding whitespace removed.
  • element.h3.a is a shortcut for "the first h3 inside, then the first a inside that".
  • class is a multi-valued attribute, so it comes back as a list.

CSS Selectors vs find and find_all

Beautiful Soup has two search styles. The older one uses find and find_all with keyword filters:

next_item = soup.find("li", class_="next")
links = soup.find_all("a", limit=3)

class_ has a trailing underscore because class is a Python keyword. The newer style uses CSS selectors through select and select_one, powered by the Soup Sieve library that's installed alongside Beautiful Soup. Both work; selectors are usually shorter and match what you see in browser dev tools, so I'll use them from here on.

Some selectors that come up a lot in scraping:

SelectorMatches
article.product_podarticle elements with class product_pod
#product_descriptionThe element with that id
div.product_price > pp elements that are direct children
a[href^="https"]Links whose href starts with https
tr:nth-of-type(2) tdCells in the second table row
#product_description + pThe p immediately after that element

Parsers

The second argument to BeautifulSoup picks the parser. "html.parser" is built into Python and needs nothing extra. "lxml" is faster and more lenient with broken markup but requires pip install lxml. For most scraping jobs, html.parser is fine; switch to lxml if you're parsing thousands of pages and speed matters.

Scraping a Detail Page

Listing pages give you a summary. Detail pages often hold the data you really want, frequently in a table:

# detail_page.py
import requests
from bs4 import BeautifulSoup

url = "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
response = requests.get(url, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")

info = {
    row.th.get_text(strip=True): row.td.get_text(strip=True)
    for row in soup.select("table.table-striped tr")
}
print(info["UPC"], info["Availability"])

description = soup.select_one("#product_description + p")
print(description.get_text(strip=True)[:60])
a897fe39b1053632 In stock (22 available)
It's hard to imagine a world without A Light in the Attic. T

The dict comprehension turns a two-column table of th/td rows into a lookup by label, which is much more robust than relying on row positions. The description has no class of its own, so the adjacent-sibling selector #product_description + p grabs the paragraph that follows the heading with that id.

Building a Robust Scraper

Now let's put it together into a scraper that follows pagination, retries transient errors, respects robots.txt, and writes a CSV.

A Session with Retries and a User-Agent

from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

USER_AGENT = "tidewave-book-scraper/1.0 (+mailto:you@example.com)"


def make_session() -> requests.Session:
    retry = Retry(
        total=3,
        backoff_factor=1,
        status_forcelist=[429, 500, 502, 503, 504],
        allowed_methods=["GET"],
    )
    session = requests.Session()
    session.headers["User-Agent"] = USER_AGENT
    session.mount("https://", HTTPAdapter(max_retries=retry))
    return session

A Session reuses the underlying TCP connection across requests, which is faster and kinder to the server than opening a new connection every time. It also stores headers and cookies, so the User-Agent applies to every request.

Mounting an HTTPAdapter with a urllib3 Retry gives you automatic retries with exponential backoff for rate limiting (429) and server errors (5xx). With backoff_factor=1 and urllib3 2.x, the first retry happens immediately and later ones wait 2s, then 4s. urllib3 also honors a Retry-After header if the server sends one.

Respecting robots.txt

The standard library can parse robots.txt for you:

from urllib.parse import urljoin
from urllib.robotparser import RobotFileParser

BASE_URL = "https://books.toscrape.com/"


def allowed_by_robots(url: str) -> bool:
    parser = RobotFileParser(urljoin(BASE_URL, "/robots.txt"))
    parser.read()
    return parser.can_fetch(USER_AGENT, url)

RobotFileParser treats a missing robots.txt (a 404) as "everything allowed" and a 401 or 403 as "nothing allowed". Books to Scrape doesn't have one, so this returns True. On a real site, check once at startup and also read the site's terms.

Parsing Each Item into a Dataclass

from dataclasses import dataclass
from decimal import Decimal

RATINGS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}


@dataclass
class Book:
    title: str
    price: Decimal
    rating: int
    in_stock: bool
    url: str


def parse_book(card, page_url: str) -> Book:
    link = card.select_one("h3 a")
    rating_class = card.select_one("p.star-rating")["class"][1]
    price_text = card.select_one("p.price_color").get_text(strip=True)
    stock_text = card.select_one("p.availability").get_text(strip=True)
    return Book(
        title=link["title"],
        price=Decimal(price_text.lstrip("£")),
        rating=RATINGS.get(rating_class, 0),
        in_stock="In stock" in stock_text,
        url=urljoin(page_url, link["href"]),
    )

This is where scraped strings become real data. The star rating lives in a class name (star-rating Three), so it's mapped to an integer. Prices become Decimal rather than float to avoid rounding surprises with money.

The urljoin(page_url, link["href"]) call matters more than it looks. The href values are relative (catalogue/...), and what they're relative to changes depending on which page you're on. Page 1 is at the site root, but page 2 lives under /catalogue/, where the links drop the catalogue/ prefix. Joining against the URL of the page you actually fetched always produces the right absolute URL.

Following Pagination

import time


def fetch(session: requests.Session, url: str) -> BeautifulSoup:
    response = session.get(url, timeout=10)
    response.raise_for_status()
    return BeautifulSoup(response.content, "html.parser")


def scrape(max_pages: int = 3, delay: float = 1.0) -> list[Book]:
    session = make_session()
    url: str | None = BASE_URL
    books: list[Book] = []

    for _ in range(max_pages):
        if url is None:
            break
        soup = fetch(session, url)
        books.extend(parse_book(card, url) for card in soup.select("article.product_pod"))

        next_link = soup.select_one("li.next a")
        url = urljoin(url, next_link["href"]) if next_link else None
        time.sleep(delay)

    return books

Rather than guessing URLs like page-1.html, page-2.html, the scraper follows the "next" link until there isn't one. That keeps working if the site changes its URL scheme, and it stops naturally on the last page. max_pages is a safety limit so a bug can't send you crawling forever, and time.sleep(delay) spaces out requests.

The Complete Script

# scraper.py
import csv
import time
from dataclasses import asdict, dataclass, fields
from decimal import Decimal
from urllib.parse import urljoin
from urllib.robotparser import RobotFileParser

import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

BASE_URL = "https://books.toscrape.com/"
USER_AGENT = "tidewave-book-scraper/1.0 (+mailto:you@example.com)"
RATINGS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}


@dataclass
class Book:
    title: str
    price: Decimal
    rating: int
    in_stock: bool
    url: str


def make_session() -> requests.Session:
    retry = Retry(
        total=3,
        backoff_factor=1,
        status_forcelist=[429, 500, 502, 503, 504],
        allowed_methods=["GET"],
    )
    session = requests.Session()
    session.headers["User-Agent"] = USER_AGENT
    session.mount("https://", HTTPAdapter(max_retries=retry))
    return session


def allowed_by_robots(url: str) -> bool:
    parser = RobotFileParser(urljoin(BASE_URL, "/robots.txt"))
    parser.read()
    return parser.can_fetch(USER_AGENT, url)


def fetch(session: requests.Session, url: str) -> BeautifulSoup:
    response = session.get(url, timeout=10)
    response.raise_for_status()
    return BeautifulSoup(response.content, "html.parser")


def parse_book(card, page_url: str) -> Book:
    link = card.select_one("h3 a")
    rating_class = card.select_one("p.star-rating")["class"][1]
    price_text = card.select_one("p.price_color").get_text(strip=True)
    stock_text = card.select_one("p.availability").get_text(strip=True)
    return Book(
        title=link["title"],
        price=Decimal(price_text.lstrip("£")),
        rating=RATINGS.get(rating_class, 0),
        in_stock="In stock" in stock_text,
        url=urljoin(page_url, link["href"]),
    )


def scrape(max_pages: int = 3, delay: float = 1.0) -> list[Book]:
    session = make_session()
    url: str | None = BASE_URL
    books: list[Book] = []

    for _ in range(max_pages):
        if url is None:
            break
        soup = fetch(session, url)
        books.extend(parse_book(card, url) for card in soup.select("article.product_pod"))

        next_link = soup.select_one("li.next a")
        url = urljoin(url, next_link["href"]) if next_link else None
        time.sleep(delay)

    return books


def save_csv(books: list[Book], path: str) -> None:
    with open(path, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=[f.name for f in fields(Book)])
        writer.writeheader()
        writer.writerows(asdict(book) for book in books)


if __name__ == "__main__":
    if not allowed_by_robots(BASE_URL):
        raise SystemExit("robots.txt disallows scraping this site")
    books = scrape(max_pages=3)
    save_csv(books, "books.csv")
    print(f"Saved {len(books)} books")
    for book in books[:3]:
        print(f"{book.rating}/5  £{book.price:>6}  {book.title}")
python scraper.py
Saved 60 books
3/5  £ 51.77  A Light in the Attic
1/5  £ 53.74  Tipping the Velvet
1/5  £ 50.10  Soumission

save_csv uses csv.DictWriter with field names taken from the dataclass, and asdict converts each Book to a dict. Opening the file with newline="" is what the csv module docs require to avoid blank lines on Windows. The result loads straight into a spreadsheet or into pandas with pd.read_csv("books.csv").

Handling Missing Elements

Real pages are inconsistent. A product might not have a rating, or a sale price might replace the regular one. select_one returns None when nothing matches, and calling .get_text() on None raises AttributeError. A small helper keeps parsing code clean:

def text_or_none(parent, selector: str) -> str | None:
    element = parent.select_one(selector)
    return element.get_text(strip=True) if element else None

Decide per field whether a missing value should be None, a default, or a hard error. For required fields, failing loudly is usually better than silently writing empty rows, because it tells you when the site's markup has changed.

When Requests and Beautiful Soup Aren't Enough

This stack downloads the HTML the server sends. It doesn't run JavaScript. If you fetch a page and the data you see in your browser isn't in response.content, the page is rendering it client-side. Before reaching for a browser automation tool, open your browser's dev tools, go to the Network tab, filter by Fetch/XHR, and reload. Very often the page loads its data from a JSON endpoint you can call directly with Requests, which is faster and more reliable than parsing HTML at all.

If there's no such endpoint, a headless browser like Playwright can render the page for you, and you can still hand the rendered HTML to Beautiful Soup for parsing.

For large crawls across many pages and domains, a framework like Scrapy adds scheduling, concurrency, throttling, and pipelines that you'd otherwise build yourself.

Practical Tips

  • Cache while developing. Save fetched HTML to disk and parse from the file while you work out selectors, so you only hit the site once.
  • Prefer stable hooks. IDs, data-* attributes, and semantic class names change less often than deep positional selectors like div > div:nth-child(3) span.
  • Log what you skip. When a card fails to parse, log the URL so you can inspect it rather than losing data quietly.
  • Validate the output. Check row counts and spot-check values. Scrapers fail silently when markup changes.

Conclusion

Requests and Beautiful Soup cover a large share of scraping work: fetch the page with a timeout, parse response.content so the encoding comes out right, select elements with CSS selectors, and convert strings into typed data. Wrap that in a session with retries, follow "next" links with urljoin, pause between requests, and you have a scraper that's both robust and polite.

When the data isn't in the HTML, look for the JSON endpoint behind the page before reaching for a headless browser. And whatever you build, check the site's rules first and scrape at a pace the site won't notice.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading