Type something to search...
Understanding the GIL and Free-Threaded Python

Understanding the GIL and Free-Threaded Python

Sooner or later every Python developer writes a multithreaded loop, runs it on an eight-core machine, and watches it finish no faster than the single-threaded version. The usual explanation is "because of the GIL", which is true but not very helpful. What is the GIL protecting, when does it actually get in your way, and what does it mean that Python now ships builds without it?

This post explains the Global Interpreter Lock from the ground up: what it is, why CPython has one, how it affects I/O-bound and CPU-bound code differently, and what thread safety looks like with and without it. Then I'll cover the free-threaded builds introduced experimentally in Python 3.13 and officially supported in Python 3.14, how to install one, how to check which build you're running, and what to watch out for before you switch.

What the GIL Is

The Global Interpreter Lock is a mutex inside CPython, the reference implementation of Python. A thread must hold it to execute Python bytecode. Only one thread can hold it at a time, so in a standard CPython process only one thread is running Python code at any given moment, no matter how many cores you have.

Threads still exist and are real OS threads. The interpreter just makes them take turns. A running thread gives up the GIL when:

  • it starts a blocking operation that releases the lock, such as reading a socket, writing a file, or calling time.sleep()
  • it has held the lock for the switch interval and another thread is waiting for it

You can inspect that interval:

import sys

print(sys.getswitchinterval())
0.005

Five milliseconds by default. After that, a waiting thread requests the GIL and the running thread drops it at the next safe point.

Why CPython Has One

The GIL is not an accident or laziness. It exists because CPython's memory management relies on reference counting. Every object carries a counter of how many references point to it, and the interpreter increments and decrements that counter constantly, on nearly every operation. If two threads updated the same counter at the same time without coordination, the count could go wrong and an object could be freed while still in use, or never freed at all.

Putting a fine-grained lock on every object would make single-threaded code much slower. A single global lock is cheap, simple, and makes the interpreter's internals (dicts, lists, the garbage collector, the import system) safe without further work. It also made writing C extensions easy, which is a large part of why Python's ecosystem of C-backed libraries grew so big.

The cost is the obvious one: pure Python code can't run on more than one core at a time within one process.

How the GIL Affects Real Code

The GIL's impact depends almost entirely on what your threads spend their time doing.

I/O-Bound Work: Threads Help

When a thread waits on the network, the disk, or a sleep, it releases the GIL. Other threads run during the wait. This is why threads work well for downloading files, calling APIs, or querying databases.

# io_threads.py
import time
from concurrent.futures import ThreadPoolExecutor


def fake_request(i: int) -> int:
    time.sleep(0.5)  # stands in for a network call; releases the GIL
    return i


start = time.perf_counter()
with ThreadPoolExecutor(max_workers=10) as pool:
    results = list(pool.map(fake_request, range(10)))
print(f"10 requests in {time.perf_counter() - start:.2f}s")
10 requests in 0.51s

Ten half-second "requests" finish in about half a second total, because all ten threads wait at the same time. The GIL is irrelevant here.

CPU-Bound Work: Threads Don't Help

Now a pure-Python, CPU-heavy function: counting primes by trial division.

# cpu_threads.py
import time
from concurrent.futures import ThreadPoolExecutor


def count_primes(limit: int) -> int:
    count = 0
    for n in range(2, limit):
        if all(n % d for d in range(2, int(n**0.5) + 1)):
            count += 1
    return count


def run(workers: int, jobs: int = 4, limit: int = 200_000) -> float:
    start = time.perf_counter()
    with ThreadPoolExecutor(max_workers=workers) as pool:
        results = list(pool.map(count_primes, [limit] * jobs))
    elapsed = time.perf_counter() - start
    print(f"{workers} thread(s): {elapsed:.2f}s  results={results[0]}")
    return elapsed


if __name__ == "__main__":
    run(workers=1)
    run(workers=4)

On a standard Python 3.13 build on an 8-core laptop:

1 thread(s): 1.11s  results=17984
4 thread(s): 1.08s  results=17984

Four threads do the same four jobs in the same time as one thread. They're all queuing for the same lock. Your exact numbers will differ, but the shape won't: on a GIL build, CPU-bound pure-Python threads don't scale.

Code That Releases the GIL

The lock only applies to running Python bytecode. C code is free to release it while it works on data that doesn't touch Python objects. Many libraries do exactly that:

  • hashlib releases the GIL while hashing large buffers
  • zlib and other compression modules release it during compression
  • NumPy releases it for many array operations
  • file and socket I/O release it while blocked

So "threads don't help CPU-bound work" really means "threads don't help CPU-bound pure Python work". If your hot loop lives inside NumPy, threads can already give you parallelism on a standard build.

The GIL Is Not a Thread-Safety Guarantee

A common misconception is that the GIL makes your code thread-safe. It doesn't. It protects the interpreter's internal state, not your program's logic. An operation like counter += 1 is several bytecode steps (load, add, store), and a thread switch can happen between them.

# race.py
import threading

counter = 0
lock = threading.Lock()


def unsafe_increment(n: int) -> None:
    global counter
    for _ in range(n):
        counter += 1


def safe_increment(n: int) -> None:
    global counter
    for _ in range(n):
        with lock:
            counter += 1


def run(target) -> int:
    global counter
    counter = 0
    threads = [threading.Thread(target=target, args=(200_000,)) for _ in range(4)]
    for t in threads:
        t.start()
    for t in threads:
        t.join()
    return counter


print("unsafe:", run(unsafe_increment))
print("safe:  ", run(safe_increment))

On a standard 3.13 build, this often prints the correct total of 800,000 for both versions. That's luck and timing, not correctness: recent CPython versions only switch threads at certain points, so the race is rare. Run it on a free-threaded build and the bug shows up immediately:

unsafe: 231716
safe:   800000

The unsafe version loses most of its increments, because four threads are now truly running at once. The locked version is correct on every build. The lesson holds regardless of which Python you use: if threads share mutable state, protect it with a threading.Lock or design it away with queues.

Free-Threaded Python

PEP 703 proposed making the GIL optional in CPython. It was accepted, and the work has landed in stages:

VersionStatus of free-threading
Python 3.13Experimental. Separate build, binary usually named python3.13t. Noticeable single-threaded slowdown.
Python 3.14Officially supported (no longer experimental, per PEP 779). Still a separate build, python3.14t. Much smaller single-threaded overhead.
FutureThe default build may eventually become free-threaded, but no version has been committed to.

The default python3 you download is still the GIL build in both 3.13 and 3.14. Free-threading is something you opt into by installing a different build.

How CPython Works Without the GIL

Removing the GIL safely required several changes to the interpreter:

  • Biased reference counting. Each object tracks references from its owning thread cheaply and uses slower atomic operations only for references from other threads.
  • Immortal objects. Objects like None, True, small integers, and interned strings never have their reference counts change, so threads don't contend on them.
  • Per-object locks on built-in containers like list and dict, so concurrent operations on the same container don't corrupt it.
  • A thread-safe memory allocator (based on mimalloc) and changes to the garbage collector.

The upshot for you: built-in types stay internally consistent under concurrent access, but compound operations in your code (check-then-set, read-modify-write) still need your own locks, exactly as before.

Installing a Free-Threaded Build

Pick whichever matches how you already install Python:

  • python.org installers for macOS and Windows have an option to install the free-threaded binaries alongside the regular ones.
  • uv can install it directly: uv python install 3.14t.
  • Linux distributions and conda-forge offer packages such as python3.14-nogil or python-freethreading, depending on the distro.
  • From source, configure CPython with --disable-gil.

With uv, you can run a script under the free-threaded interpreter without changing anything else:

uv python install 3.14t
uv run --python 3.14t cpu_threads.py

If you want to read more about uv, see Managing Python Projects with uv.

Checking Which Build You're Running

Two separate questions matter: was this interpreter built without the GIL, and is the GIL currently disabled?

# gil_check.py
import sys
import sysconfig

build_flag = sysconfig.get_config_var("Py_GIL_DISABLED")
print(f"Python {sys.version.split()[0]}")
print(f"Free-threaded build: {bool(build_flag)}")
print(f"GIL enabled right now: {sys._is_gil_enabled()}")

On a regular 3.13 install:

Python 3.13.2
Free-threaded build: False
GIL enabled right now: True

On a free-threaded 3.14 build:

Python 3.14.8
Free-threaded build: True
GIL enabled right now: False

sys._is_gil_enabled() was added in 3.13. The leading underscore marks it as an implementation detail, but it's the documented way to check at runtime. python -VV also prints "free-threading build" in the version string on these builds.

Running the Benchmark Again

Here's the same cpu_threads.py from earlier, run with python3.14t on the same machine:

1 thread(s): 1.04s  results=17984
4 thread(s): 0.40s  results=17984

Four threads now finish roughly 2.5 times faster than one. That's real parallelism from plain threading, with no multiprocessing and no pickling of arguments.

Turning the GIL Back On

A free-threaded build can re-enable the GIL at startup, which is handy for comparing behavior or working around an incompatible library:

PYTHON_GIL=1 python3.14t gil_check.py
python3.14t -X gil=1 gil_check.py

Both print GIL enabled right now: True. Setting PYTHON_GIL=0 (or -X gil=0) forces it off.

C Extensions and Compatibility

The biggest practical issue with free-threading is extension modules. A C extension written for the GIL build may assume the lock protects its global state. To be loaded without the GIL, an extension has to declare that it supports free-threading (via the Py_mod_gil slot), and it needs to be compiled specifically for the free-threaded ABI. Wheels for free-threaded Python carry a t in their ABI tag, for example cp314t.

If you import an extension that hasn't declared support, CPython doesn't crash. It re-enables the GIL for the whole process and prints a RuntimeWarning telling you which module caused it. Your code keeps working; it just loses the parallelism.

Before moving a project to a free-threaded build:

  • Check that your key dependencies publish cp313t/cp314t wheels. NumPy, Cython-based projects, and many popular packages do; smaller or older extensions may not.
  • Watch for the RuntimeWarning about the GIL being re-enabled.
  • Run your test suite under the free-threaded build. Latent race conditions in your own code, like the counter above, tend to surface fast.
  • Measure single-threaded performance. Free-threaded builds pay a small overhead for the extra safety machinery. In 3.13 it was significant; in 3.14 it's typically in the single-digit to low double-digit percent range, depending on the workload.

When to Care About Any of This

For most applications, the GIL is less of a problem than its reputation suggests:

  • Web apps and API clients are I/O-bound. Threads or asyncio already work well. See asyncio in Python: A Beginner's Guide for the async approach.
  • Numeric work usually lives in NumPy, pandas, or Polars, which release the GIL or parallelize internally.
  • CPU-bound pure Python is where the GIL hurts. On a standard build, use multiprocessing or ProcessPoolExecutor to spread the work across processes. On a free-threaded build, threads become a real option.

If you're deciding between models, the trade-offs between threads, processes, and async are covered in Threading vs Multiprocessing vs asyncio, and the simplest API for either pool is in Using concurrent.futures for Simple Parallelism.

FAQ

Does Python 3.14 remove the GIL? No. The default build still has it. 3.14 makes the separate free-threaded build officially supported.

Will free-threading make my single-threaded script faster? No. It can only help code that runs work in multiple threads at once. A single-threaded script may run slightly slower.

Do other Python implementations have a GIL? PyPy has one. Jython and IronPython never did. The GIL is a CPython design decision, not part of the language.

Do I still need locks on a GIL build? Yes. The GIL keeps the interpreter consistent, not your data. Any read-modify-write on shared state needs a lock on every build.

Conclusion

The GIL exists to keep CPython's reference counting and internals safe, and the price is that only one thread runs Python bytecode at a time. That barely matters for I/O-bound code and code that spends its time in C libraries, but it blocks parallelism for CPU-bound pure Python.

Free-threaded Python changes that. It was experimental in 3.13 and is officially supported in 3.14, installed as a separate python3.14t build. Check your build with sysconfig.get_config_var("Py_GIL_DISABLED") and sys._is_gil_enabled(), make sure your extensions ship free-threaded wheels, and use locks for shared state no matter which build you run.

Tags :
Share :

Related Posts

Abstract Base Classes in Python with the abc Module

Abstract Base Classes in Python with the abc Module

Python leans on duck typing: if an object has the method you need, you call it and move on. That works well until you have a family of classes that a

Continue Reading
*args and **kwargs in Python: Flexible Function Signatures

*args and **kwargs in Python: Flexible Function Signatures

You've seen def wrapper(*args, **kwargs): in decorators, and probably super().__init__(**kwargs) in class hierarchies. These two parameters let a

Continue Reading
Asyncio in Python: A Beginner's Guide to Asynchronous Programming

Asyncio in Python: A Beginner's Guide to Asynchronous Programming

A lot of programs spend most of their time waiting. A web scraper waits for pages to download, an API server waits for the database, a chat bot waits

Continue Reading