
Understanding the GIL and Free-Threaded Python
Sooner or later every Python developer writes a multithreaded loop, runs it on an eight-core machine, and watches it finish no faster than the single-threaded version. The usual explanation is "because of the GIL", which is true but not very helpful. What is the GIL protecting, when does it actually get in your way, and what does it mean that Python now ships builds without it?
This post explains the Global Interpreter Lock from the ground up: what it is, why CPython has one, how it affects I/O-bound and CPU-bound code differently, and what thread safety looks like with and without it. Then I'll cover the free-threaded builds introduced experimentally in Python 3.13 and officially supported in Python 3.14, how to install one, how to check which build you're running, and what to watch out for before you switch.
What the GIL Is
The Global Interpreter Lock is a mutex inside CPython, the reference implementation of Python. A thread must hold it to execute Python bytecode. Only one thread can hold it at a time, so in a standard CPython process only one thread is running Python code at any given moment, no matter how many cores you have.
Threads still exist and are real OS threads. The interpreter just makes them take turns. A running thread gives up the GIL when:
- it starts a blocking operation that releases the lock, such as reading a socket, writing a file, or calling
time.sleep() - it has held the lock for the switch interval and another thread is waiting for it
You can inspect that interval:
import sys
print(sys.getswitchinterval())
0.005
Five milliseconds by default. After that, a waiting thread requests the GIL and the running thread drops it at the next safe point.
Why CPython Has One
The GIL is not an accident or laziness. It exists because CPython's memory management relies on reference counting. Every object carries a counter of how many references point to it, and the interpreter increments and decrements that counter constantly, on nearly every operation. If two threads updated the same counter at the same time without coordination, the count could go wrong and an object could be freed while still in use, or never freed at all.
Putting a fine-grained lock on every object would make single-threaded code much slower. A single global lock is cheap, simple, and makes the interpreter's internals (dicts, lists, the garbage collector, the import system) safe without further work. It also made writing C extensions easy, which is a large part of why Python's ecosystem of C-backed libraries grew so big.
The cost is the obvious one: pure Python code can't run on more than one core at a time within one process.
How the GIL Affects Real Code
The GIL's impact depends almost entirely on what your threads spend their time doing.
I/O-Bound Work: Threads Help
When a thread waits on the network, the disk, or a sleep, it releases the GIL. Other threads run during the wait. This is why threads work well for downloading files, calling APIs, or querying databases.
# io_threads.py
import time
from concurrent.futures import ThreadPoolExecutor
def fake_request(i: int) -> int:
time.sleep(0.5) # stands in for a network call; releases the GIL
return i
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=10) as pool:
results = list(pool.map(fake_request, range(10)))
print(f"10 requests in {time.perf_counter() - start:.2f}s")
10 requests in 0.51s
Ten half-second "requests" finish in about half a second total, because all ten threads wait at the same time. The GIL is irrelevant here.
CPU-Bound Work: Threads Don't Help
Now a pure-Python, CPU-heavy function: counting primes by trial division.
# cpu_threads.py
import time
from concurrent.futures import ThreadPoolExecutor
def count_primes(limit: int) -> int:
count = 0
for n in range(2, limit):
if all(n % d for d in range(2, int(n**0.5) + 1)):
count += 1
return count
def run(workers: int, jobs: int = 4, limit: int = 200_000) -> float:
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=workers) as pool:
results = list(pool.map(count_primes, [limit] * jobs))
elapsed = time.perf_counter() - start
print(f"{workers} thread(s): {elapsed:.2f}s results={results[0]}")
return elapsed
if __name__ == "__main__":
run(workers=1)
run(workers=4)
On a standard Python 3.13 build on an 8-core laptop:
1 thread(s): 1.11s results=17984
4 thread(s): 1.08s results=17984
Four threads do the same four jobs in the same time as one thread. They're all queuing for the same lock. Your exact numbers will differ, but the shape won't: on a GIL build, CPU-bound pure-Python threads don't scale.
Code That Releases the GIL
The lock only applies to running Python bytecode. C code is free to release it while it works on data that doesn't touch Python objects. Many libraries do exactly that:
hashlibreleases the GIL while hashing large bufferszliband other compression modules release it during compression- NumPy releases it for many array operations
- file and socket I/O release it while blocked
So "threads don't help CPU-bound work" really means "threads don't help CPU-bound pure Python work". If your hot loop lives inside NumPy, threads can already give you parallelism on a standard build.
The GIL Is Not a Thread-Safety Guarantee
A common misconception is that the GIL makes your code thread-safe. It doesn't. It protects the interpreter's internal state, not your program's logic. An operation like counter += 1 is several bytecode steps (load, add, store), and a thread switch can happen between them.
# race.py
import threading
counter = 0
lock = threading.Lock()
def unsafe_increment(n: int) -> None:
global counter
for _ in range(n):
counter += 1
def safe_increment(n: int) -> None:
global counter
for _ in range(n):
with lock:
counter += 1
def run(target) -> int:
global counter
counter = 0
threads = [threading.Thread(target=target, args=(200_000,)) for _ in range(4)]
for t in threads:
t.start()
for t in threads:
t.join()
return counter
print("unsafe:", run(unsafe_increment))
print("safe: ", run(safe_increment))
On a standard 3.13 build, this often prints the correct total of 800,000 for both versions. That's luck and timing, not correctness: recent CPython versions only switch threads at certain points, so the race is rare. Run it on a free-threaded build and the bug shows up immediately:
unsafe: 231716
safe: 800000
The unsafe version loses most of its increments, because four threads are now truly running at once. The locked version is correct on every build. The lesson holds regardless of which Python you use: if threads share mutable state, protect it with a threading.Lock or design it away with queues.
Free-Threaded Python
PEP 703 proposed making the GIL optional in CPython. It was accepted, and the work has landed in stages:
| Version | Status of free-threading |
|---|---|
| Python 3.13 | Experimental. Separate build, binary usually named python3.13t. Noticeable single-threaded slowdown. |
| Python 3.14 | Officially supported (no longer experimental, per PEP 779). Still a separate build, python3.14t. Much smaller single-threaded overhead. |
| Future | The default build may eventually become free-threaded, but no version has been committed to. |
The default python3 you download is still the GIL build in both 3.13 and 3.14. Free-threading is something you opt into by installing a different build.
How CPython Works Without the GIL
Removing the GIL safely required several changes to the interpreter:
- Biased reference counting. Each object tracks references from its owning thread cheaply and uses slower atomic operations only for references from other threads.
- Immortal objects. Objects like
None,True, small integers, and interned strings never have their reference counts change, so threads don't contend on them. - Per-object locks on built-in containers like
listanddict, so concurrent operations on the same container don't corrupt it. - A thread-safe memory allocator (based on mimalloc) and changes to the garbage collector.
The upshot for you: built-in types stay internally consistent under concurrent access, but compound operations in your code (check-then-set, read-modify-write) still need your own locks, exactly as before.
Installing a Free-Threaded Build
Pick whichever matches how you already install Python:
- python.org installers for macOS and Windows have an option to install the free-threaded binaries alongside the regular ones.
- uv can install it directly:
uv python install 3.14t. - Linux distributions and conda-forge offer packages such as
python3.14-nogilorpython-freethreading, depending on the distro. - From source, configure CPython with
--disable-gil.
With uv, you can run a script under the free-threaded interpreter without changing anything else:
uv python install 3.14t
uv run --python 3.14t cpu_threads.py
If you want to read more about uv, see Managing Python Projects with uv.
Checking Which Build You're Running
Two separate questions matter: was this interpreter built without the GIL, and is the GIL currently disabled?
# gil_check.py
import sys
import sysconfig
build_flag = sysconfig.get_config_var("Py_GIL_DISABLED")
print(f"Python {sys.version.split()[0]}")
print(f"Free-threaded build: {bool(build_flag)}")
print(f"GIL enabled right now: {sys._is_gil_enabled()}")
On a regular 3.13 install:
Python 3.13.2
Free-threaded build: False
GIL enabled right now: True
On a free-threaded 3.14 build:
Python 3.14.8
Free-threaded build: True
GIL enabled right now: False
sys._is_gil_enabled() was added in 3.13. The leading underscore marks it as an implementation detail, but it's the documented way to check at runtime. python -VV also prints "free-threading build" in the version string on these builds.
Running the Benchmark Again
Here's the same cpu_threads.py from earlier, run with python3.14t on the same machine:
1 thread(s): 1.04s results=17984
4 thread(s): 0.40s results=17984
Four threads now finish roughly 2.5 times faster than one. That's real parallelism from plain threading, with no multiprocessing and no pickling of arguments.
Turning the GIL Back On
A free-threaded build can re-enable the GIL at startup, which is handy for comparing behavior or working around an incompatible library:
PYTHON_GIL=1 python3.14t gil_check.py
python3.14t -X gil=1 gil_check.py
Both print GIL enabled right now: True. Setting PYTHON_GIL=0 (or -X gil=0) forces it off.
C Extensions and Compatibility
The biggest practical issue with free-threading is extension modules. A C extension written for the GIL build may assume the lock protects its global state. To be loaded without the GIL, an extension has to declare that it supports free-threading (via the Py_mod_gil slot), and it needs to be compiled specifically for the free-threaded ABI. Wheels for free-threaded Python carry a t in their ABI tag, for example cp314t.
If you import an extension that hasn't declared support, CPython doesn't crash. It re-enables the GIL for the whole process and prints a RuntimeWarning telling you which module caused it. Your code keeps working; it just loses the parallelism.
Before moving a project to a free-threaded build:
- Check that your key dependencies publish
cp313t/cp314twheels. NumPy, Cython-based projects, and many popular packages do; smaller or older extensions may not. - Watch for the
RuntimeWarningabout the GIL being re-enabled. - Run your test suite under the free-threaded build. Latent race conditions in your own code, like the counter above, tend to surface fast.
- Measure single-threaded performance. Free-threaded builds pay a small overhead for the extra safety machinery. In 3.13 it was significant; in 3.14 it's typically in the single-digit to low double-digit percent range, depending on the workload.
When to Care About Any of This
For most applications, the GIL is less of a problem than its reputation suggests:
- Web apps and API clients are I/O-bound. Threads or
asyncioalready work well. See asyncio in Python: A Beginner's Guide for the async approach. - Numeric work usually lives in NumPy, pandas, or Polars, which release the GIL or parallelize internally.
- CPU-bound pure Python is where the GIL hurts. On a standard build, use
multiprocessingorProcessPoolExecutorto spread the work across processes. On a free-threaded build, threads become a real option.
If you're deciding between models, the trade-offs between threads, processes, and async are covered in Threading vs Multiprocessing vs asyncio, and the simplest API for either pool is in Using concurrent.futures for Simple Parallelism.
FAQ
Does Python 3.14 remove the GIL? No. The default build still has it. 3.14 makes the separate free-threaded build officially supported.
Will free-threading make my single-threaded script faster? No. It can only help code that runs work in multiple threads at once. A single-threaded script may run slightly slower.
Do other Python implementations have a GIL? PyPy has one. Jython and IronPython never did. The GIL is a CPython design decision, not part of the language.
Do I still need locks on a GIL build? Yes. The GIL keeps the interpreter consistent, not your data. Any read-modify-write on shared state needs a lock on every build.
Conclusion
The GIL exists to keep CPython's reference counting and internals safe, and the price is that only one thread runs Python bytecode at a time. That barely matters for I/O-bound code and code that spends its time in C libraries, but it blocks parallelism for CPU-bound pure Python.
Free-threaded Python changes that. It was experimental in 3.13 and is officially supported in 3.14, installed as a separate python3.14t build. Check your build with sysconfig.get_config_var("Py_GIL_DISABLED") and sys._is_gil_enabled(), make sure your extensions ship free-threaded wheels, and use locks for shared state no matter which build you run.


