Type something to search...
Threading vs Multiprocessing vs Asyncio in Python: Which Should You Use?

Threading vs Multiprocessing vs Asyncio in Python: Which Should You Use?

Python gives you three built-in ways to do more than one thing at a time: threads, processes, and asyncio. They look similar from a distance, since all three let you "run things concurrently", but they solve different problems, and picking the wrong one can leave you with code that's more complicated and no faster than a plain loop.

The deciding question is almost always the same: is your program slow because it's waiting, or because it's computing? Once you know that, the choice mostly makes itself.

In this post I'll explain how each model works, run the same workloads through all three so you can see the actual numbers, cover the pitfalls specific to each one, and finish with a short decision guide. The examples use Python 3.13 and the standard library, with notes on how free-threaded Python is changing the picture.

Concurrency vs Parallelism

Two terms that get mixed up constantly:

  • Concurrency means several tasks are in progress at the same time. They might take turns on a single CPU core. While one waits for a network response, another runs.
  • Parallelism means several tasks are executing at literally the same instant, on different CPU cores.

Waiting doesn't need a CPU core, so I/O-bound work only needs concurrency. Computation does need cores, so CPU-bound work needs parallelism to go faster. Keep that distinction in mind; it's the whole story in two sentences.

The Three Models

Threading

The threading module runs functions in OS threads that share the same memory. Threads are cheap to start and can read and write the same objects, which makes passing data around easy and also makes it easy to corrupt.

In the standard CPython build, the Global Interpreter Lock (GIL) allows only one thread to execute Python bytecode at a time. A thread releases the GIL while it waits on I/O (sockets, files, time.sleep()) and during many C-level operations, so other threads can run then. The result: threads are excellent for overlapping waits, but they don't make pure-Python computation faster. The OS decides when to switch threads, which is called preemptive multitasking, so a switch can happen between almost any two bytecode instructions.

Multiprocessing

The multiprocessing module (and concurrent.futures.ProcessPoolExecutor) runs work in separate Python processes. Each process has its own interpreter, its own memory, and its own GIL, so they genuinely run in parallel on multiple cores.

The cost is isolation. Processes don't share memory, so arguments and return values have to be serialized (pickled), sent across, and deserialized. Starting a process is much slower than starting a thread, and each one uses tens of megabytes of memory.

Asyncio

asyncio runs many coroutines in a single thread, using an event loop. A coroutine runs until it reaches an await on something that isn't ready, then hands control back to the loop, which runs another coroutine. Switches only happen at await points, which is called cooperative multitasking.

Because there's one thread and switching is explicit, asyncio can handle thousands of concurrent connections with very little overhead. But it needs async-aware libraries for all I/O, and any blocking call freezes everything. Asyncio in Python: A Beginner's Guide to Asynchronous Programming covers it in depth.

Side by Side

ThreadingMultiprocessingAsyncio
Runs onMultiple OS threads, one processMultiple processesOne thread, event loop
SwitchingPreemptive (OS decides)Truly parallelCooperative (at await)
Speeds up I/O-bound workYesYes, with high overheadYes
Speeds up CPU-bound workNo (standard build)YesNo
Shared memoryYes, needs locksNo, data is pickledYes, but only one coroutine runs at a time
Overhead per unitModerate (OS thread)High (whole process)Very low (coroutine object)
Practical scaleTens to hundredsRoughly one per CPU coreThousands or more
Works with blocking librariesYesYesOnly via threads

Benchmark 1: I/O-Bound Work

Let's make 20 "requests" that each wait half a second, using time.sleep() and asyncio.sleep() as stand-ins for network latency. Each approach uses the simplest standard-library tool for the job:

# io_bound.py
import asyncio
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor


def fetch_sync(i: int) -> int:
    time.sleep(0.5)  # simulates waiting on a network response
    return i


async def fetch_async(i: int) -> int:
    await asyncio.sleep(0.5)
    return i


def timed(label: str, func) -> None:
    start = time.perf_counter()
    func()
    print(f"{label:<16}{time.perf_counter() - start:5.2f}s")


def sequential() -> None:
    [fetch_sync(i) for i in range(20)]


def threads() -> None:
    with ThreadPoolExecutor(max_workers=20) as pool:
        list(pool.map(fetch_sync, range(20)))


def processes() -> None:
    with ProcessPoolExecutor(max_workers=20) as pool:
        list(pool.map(fetch_sync, range(20)))


def with_asyncio() -> None:
    async def main() -> None:
        async with asyncio.TaskGroup() as tg:
            for i in range(20):
                tg.create_task(fetch_async(i))
    asyncio.run(main())


if __name__ == "__main__":
    timed("sequential", sequential)
    timed("threads", threads)
    timed("processes", processes)
    timed("asyncio", with_asyncio)

On my 8-core machine:

sequential      10.13s
threads          0.51s
processes        0.98s
asyncio          0.50s

The sequential version waits 20 times in a row. Threads and asyncio overlap all the waits and finish in about one wait's worth of time. Processes also overlap the waits, but spinning up 20 Python processes added roughly half a second of pure overhead, doubling the runtime. For waiting, processes are the wrong tool.

Between threads and asyncio, the numbers are a tie at this scale. The differences show up at larger scale and in how the code is written, which we'll get to.

Benchmark 2: CPU-Bound Work

Now the opposite: pure computation, with no waiting at all. Eight jobs, each counting primes the slow way:

# cpu_bound.py
import os
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor


def count_primes(limit: int) -> int:
    count = 0
    for n in range(2, limit):
        if all(n % d for d in range(2, int(n ** 0.5) + 1)):
            count += 1
    return count


JOBS = [200_000] * 8


def timed(label: str, func) -> None:
    start = time.perf_counter()
    func()
    print(f"{label:<16}{time.perf_counter() - start:5.2f}s")


def sequential() -> None:
    [count_primes(n) for n in JOBS]


def threads() -> None:
    with ThreadPoolExecutor(max_workers=8) as pool:
        list(pool.map(count_primes, JOBS))


def processes() -> None:
    with ProcessPoolExecutor() as pool:
        list(pool.map(count_primes, JOBS))


if __name__ == "__main__":
    print(f"CPU cores: {os.cpu_count()}")
    timed("sequential", sequential)
    timed("threads", threads)
    timed("processes", processes)

Output:

CPU cores: 8
sequential       2.11s
threads          2.09s
processes        0.55s

Eight threads were no faster than one, because the GIL let only one of them run Python code at any moment. Processes spread the work across cores and finished almost four times faster. (Not eight times: process startup, pickling, and (on many laptops) a mix of performance and efficiency cores all eat into the ideal speedup.)

I didn't include asyncio here because there's nothing for it to do. With no await on anything slow, coroutines would just run one after another on a single thread, exactly like the sequential version.

Your numbers will differ by machine, but the shape of the results won't.

Pitfalls of Each Model

Threads: Race Conditions

Threads share memory and can be switched at almost any moment, so any "read, modify, write" sequence on shared data can interleave badly. Here four threads each deposit 1 into a shared balance 1,000 times:

# race.py
import threading
import time

balance = 0
lock = threading.Lock()


def deposit_unsafe(times: int) -> None:
    global balance
    for _ in range(times):
        current = balance
        time.sleep(0)  # give other threads a chance to run here
        balance = current + 1


def deposit_safe(times: int) -> None:
    global balance
    for _ in range(times):
        with lock:
            current = balance
            time.sleep(0)
            balance = current + 1


def run(target) -> int:
    global balance
    balance = 0
    threads = [threading.Thread(target=target, args=(1_000,)) for _ in range(4)]
    for t in threads:
        t.start()
    for t in threads:
        t.join()
    return balance


print("unsafe:", run(deposit_unsafe))
print("safe:  ", run(deposit_safe))

Output (the unsafe number changes every run):

unsafe: 1022
safe:   4000

The correct answer is 4,000. Without the lock, threads read the same balance, each add one, and overwrite each other's updates. The time.sleep(0) just makes the problem show up reliably; in real code the same bug appears rarely and unpredictably, which is worse. The GIL does not protect you here; it keeps the interpreter's internals consistent, not your data.

Ways to stay safe with threads:

  • Protect shared mutable state with a threading.Lock, used as a context manager.
  • Better, avoid sharing: pass work in and get results out through queue.Queue or the futures returned by ThreadPoolExecutor.
  • Keep threads to I/O work, where they shine, and keep the shared parts small.

Multiprocessing: Pickling and the Main Guard

Because work crosses a process boundary, every function and argument you send must be picklable. That rules out lambdas, nested functions, open files, database connections, and locks:

_pickle.PicklingError: Can't pickle <function <lambda> at 0x1057f80e0>: attribute lookup <lambda> on __main__ failed

Define worker functions at the top level of a module, and pass plain data (numbers, strings, lists, dicts, dataclasses) as arguments.

The other classic mistake is forgetting the if __name__ == "__main__": guard. On macOS and Windows, new processes are started with the spawn method, which imports your main module fresh in each child. Without the guard, each child tries to create its own pool, and Python stops it with a RuntimeError about "the bootstrapping phase". On Linux, Python 3.14 also moved away from the old fork default to forkserver, so always write the guard regardless of platform:

# word_counts.py
from collections import Counter
from multiprocessing import Pool


def count_words(text: str) -> Counter[str]:
    return Counter(text.lower().split())


if __name__ == "__main__":
    documents = [
        "the quick brown fox",
        "the lazy dog",
        "the fox jumps over the dog",
    ]
    with Pool(processes=3) as pool:
        partials = pool.map(count_words, documents)
    total = sum(partials, Counter())
    print(total.most_common(3))

Output:

[('the', 4), ('fox', 2), ('dog', 2)]

Also watch the data volume. If each task sends a large object to a worker and gets a large object back, pickling can cost more than the computation saves. Process pools work best for chunky tasks: lots of computation per byte transferred.

Asyncio: Blocking the Loop and the Async Ecosystem

asyncio's pitfalls are different:

  • One blocking call stalls everything. A time.sleep(), a requests.get(), or a heavy loop inside a coroutine freezes every other task. Use async libraries (httpx, aiohttp, asyncpg) and push unavoidable blocking calls to a thread with asyncio.to_thread().
  • It's contagious. await only works inside async def, so async tends to spread up the call stack. That's fine in a new async service; retrofitting it into a large synchronous codebase is a lot of work.
  • It needs library support. If the library you depend on is synchronous only, you'll be running it in threads anyway.

On the plus side, data races are much rarer in asyncio. Switches only happen at await, so code between two awaits runs without interruption. A read-modify-write with no await in the middle is safe.

Combining Models

These models aren't mutually exclusive. Some combinations are very common:

  • asyncio + threads: an async web service calls a blocking SDK through asyncio.to_thread().
  • asyncio + processes: an async service offloads CPU-heavy work (image processing, report generation) with loop.run_in_executor() and a ProcessPoolExecutor, so the event loop stays responsive.
  • Processes + threads: a process pool where each worker uses a few threads for its own I/O.

The concurrent.futures module gives threads and processes the same interface (submit(), map(), futures), which makes it easy to switch between them by changing one line. Using concurrent.futures for Simple Parallelism in Python covers it in detail.

What About Free-Threaded Python?

Python 3.13 introduced an experimental free-threaded build (often installed as python3.13t) with the GIL disabled, from PEP 703. In Python 3.14 the free-threaded build is officially supported, though it's still an optional, separate build rather than the default.

On a free-threaded build, threads can run Python code in parallel, so the CPU-bound benchmark above could speed up with threads instead of processes, without the pickling and startup costs. There are trade-offs: single-threaded code runs somewhat slower on the free-threaded build, and C extensions must be updated to support it.

For now, the advice in this post holds for the standard build that almost everyone runs. Free-threading matters most if you have CPU-bound work with lots of shared data that's painful to pickle. Understanding the GIL and Free-Threaded Python goes deeper into how the GIL works and what changes without it.

How to Choose

Work through these questions in order:

  1. Is the program actually slow because of waiting or computing? Measure first. If it's neither, for example it makes three API calls and finishes, you don't need concurrency at all.
  2. CPU-bound (number crunching, parsing, compression, image processing in pure Python)?
    • First check whether a library already does it in optimized native code. NumPy, pandas, Polars, and similar libraries release the GIL and are often fast enough on their own.
    • Otherwise, use multiprocessing, typically via ProcessPoolExecutor.
  3. I/O-bound (network, disk, databases, subprocesses)?
    • Building a new service, need many concurrent connections, and async libraries exist for your stack? Use asyncio.
    • Working with existing synchronous code or libraries, or need modest concurrency (tens of tasks)? Use threads, typically via ThreadPoolExecutor.
  4. Both? Use asyncio or threads for the I/O, and hand the CPU-heavy pieces to a process pool.
Your situationUse
Download 50 files with requests in a scriptThreadPoolExecutor
Web API handling thousands of concurrent connectionsasyncio (FastAPI, Starlette, aiohttp)
Resize 10,000 imagesProcessPoolExecutor (or a native library)
Websocket server or chat botasyncio
Run a slow blocking SDK call from async codeasyncio.to_thread()
Data analysis on large arraysNumPy, pandas, or Polars first; processes if still needed
Keep a desktop GUI responsive during a downloadA background thread

Conclusion

The three models answer different questions. Threads overlap waiting with minimal changes to synchronous code, but need locks for shared data and don't speed up computation on the standard build. Processes give real parallelism for CPU-bound work, at the cost of startup time, memory, and pickling. Asyncio overlaps huge numbers of waits in a single thread with very little overhead, provided your whole I/O path is async and nothing blocks the loop.

Figure out whether you're waiting or computing, measure before and after, and pick the simplest tool that fits. In most cases that's a ThreadPoolExecutor for I/O in a script, asyncio for a high-concurrency network service, and a ProcessPoolExecutor for heavy computation.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading