- Python 3.13t free-threading achieves 5.61x speedup on 8-core CPU-bound benchmark, eliminating GIL parallelism bottleneck for pure Python code.
- Single-threaded performance degrades 17% due to atomic reference counting overhead—free-threading only wins when you can parallelize.
- C extensions like NumPy are not yet thread-safe in free-threaded builds; ecosystem compatibility remains a major blocker for production use.
The GIL Finally Dies (Sort Of)
Python 3.13t ships with experimental free-threading builds that disable the Global Interpreter Lock. After three decades of “just use multiprocessing,” you can now spawn actual parallel threads for CPU-bound work.
But does it actually work?
I ran the same CPU-intensive benchmark on Python 3.12 (GIL-locked) and 3.13t (free-threaded) to see if real parallelism delivers the speedups we’ve been promised. Spoiler: it does, but not everywhere, and the overhead is real.

What the GIL Actually Does
The Global Interpreter Lock is a mutex that protects access to Python objects. Only one thread can execute Python bytecode at a time, even on a 64-core machine. This makes CPython’s memory management simple (no concurrent reference counting nightmares), but it also means threading is useless for CPU-bound tasks.
For I/O-bound work — waiting on network, disk, database — threads release the GIL during blocking calls, so you get concurrency. For CPU-bound work — math, parsing, matrix ops — threads fight over the GIL and you end up slower than single-threaded code.
Python 3.13 introduces experimental free-threading builds (the t suffix, like python3.13t). These builds replace GIL-based synchronization with per-object locks and some atomic reference counting magic. The interpreter itself is thread-safe, so multiple threads can execute Python bytecode in parallel.
But thread safety isn’t free. Every reference count operation now needs atomic instructions or fine-grained locks. Small objects get pooled differently. The question is: does true parallelism outweigh the overhead?
The Benchmark: Prime Factorization
I picked a simple CPU-bound task: factorizing large integers. No NumPy, no C extensions, pure Python bytecode churn. Here’s the core function:
def factorize(n):
"""Return prime factors of n."""
factors = []
d = 2
while d * d <= n:
while n % d == 0:
factors.append(d)
n //= d
d += 1
if n > 1:
factors.append(n)
return factors
def benchmark_factorize(numbers):
"""Factorize a list of numbers, return total time."""
results = []
for num in numbers:
results.append(factorize(num))
return results
Nothing fancy. We’re testing raw interpreter throughput, not library performance.
I generated 1000 random 12-digit integers and timed how long it takes to factorize all of them under three strategies:
- Single-threaded: One thread, one core, no parallelism
- Threading (GIL): 8 threads on Python 3.12 (should be slower due to GIL contention)
- Free-threading: 8 threads on Python 3.13t (should scale linearly if overhead is low)
Test machine: 8-core Intel i7 (hyperthreading disabled for clean results), Ubuntu 22.04, Python 3.12.1 and 3.13.0b2 (free-threaded build).
Running the Test: 3.12 vs 3.13t
Here’s the full benchmark script:
import time
import threading
import random
def factorize(n):
factors = []
d = 2
while d * d <= n:
while n % d == 0:
factors.append(d)
n //= d
d += 1
if n > 1:
factors.append(n)
return factors
def worker(numbers, results, idx):
"""Thread worker: factorize assigned numbers."""
local_results = []
for num in numbers:
local_results.append(factorize(num))
results[idx] = local_results
def run_threaded(numbers, num_threads):
"""Split work across threads."""
chunk_size = len(numbers) // num_threads
chunks = [numbers[i*chunk_size:(i+1)*chunk_size] for i in range(num_threads)]
# Handle remainder
if len(numbers) % num_threads != 0:
chunks[-1].extend(numbers[num_threads*chunk_size:])
results = [None] * num_threads
threads = []
start = time.perf_counter()
for i, chunk in enumerate(chunks):
t = threading.Thread(target=worker, args=(chunk, results, i))
t.start()
threads.append(t)
for t in threads:
t.join()
elapsed = time.perf_counter() - start
return elapsed, results
def run_single(numbers):
"""Single-threaded baseline."""
start = time.perf_counter()
results = [factorize(n) for n in numbers]
elapsed = time.perf_counter() - start
return elapsed, results
if __name__ == "__main__":
random.seed(42)
test_numbers = [random.randint(10**11, 10**12) for _ in range(1000)]
# Single-threaded
single_time, _ = run_single(test_numbers)
print(f"Single-threaded: {single_time:.2f}s")
# 8 threads
threaded_time, _ = run_threaded(test_numbers, 8)
print(f"8 threads: {threaded_time:.2f}s")
print(f"Speedup: {single_time / threaded_time:.2f}x")
I ran this on both interpreters. Results:
Python 3.12 (GIL):
Single-threaded: 18.34s
8 threads: 19.87s
Speedup: 0.92x
Threads made it slower. Classic GIL behavior — context switching overhead with zero parallelism.
Python 3.13t (free-threaded):
Single-threaded: 21.45s
8 threads: 3.82s
Speedup: 5.61x
Now we’re talking. 5.61x speedup on 8 cores. Not perfect scaling (that would be 8x), but real parallelism.
Notice the single-threaded baseline is slower on 3.13t: 21.45s vs 18.34s. That’s the overhead of atomic reference counting and fine-grained locks. Even with one thread, you pay the cost.
Where Free-Threading Wins
The 5.61x speedup isn’t linear, but it’s massive compared to GIL Python. When does free-threading actually help?
CPU-bound loops with minimal shared state. My factorization benchmark is embarrassingly parallel — each thread operates on its own list of numbers, writes to its own slice of the results array, zero contention. Perfect case.
Long-running tasks that saturate cores. If each thread runs for seconds without touching shared objects, the per-object lock overhead gets amortized. You pay the cost upfront, then actually use all your cores.
Workloads where multiprocessing hurts. Multiprocessing copies data between processes (via pickle or shared memory). If your data is large but your tasks are small, the IPC overhead kills you. Free-threading avoids that — shared memory is just… memory.
For quick clarification: in the threaded case, I’m using a list results = [None] * num_threads where each thread writes to its own index. No lock contention because no two threads touch the same index. If I’d used a shared results.append(), I’d see serious slowdown.
Where It Still Loses
The 17% single-threaded slowdown (21.45s vs 18.34s) is brutal for workloads that can’t parallelize. If your script is inherently sequential — parsing a single JSON file, rendering a single template — you just got slower for no benefit.
And the scaling isn’t perfect. 5.61x on 8 cores means 30% of potential parallelism is lost to overhead. Some of that is atomic reference counting, some is cache contention, some is the OS scheduler not being perfect.
I also tested a pathological case: threads hammering a shared counter with counter += 1. Python 3.13t slowed down 3x compared to single-threaded, while GIL Python stayed roughly constant (because the GIL serializes everything anyway). Fine-grained locks help when contention is low, but they’re not magic.

The Math Behind Atomic Reference Counting
Python objects have a reference count field. In GIL Python, incrementing a refcount is just:
One instruction, no synchronization needed (the GIL already guarantees mutual exclusion).
In free-threaded Python, that becomes an atomic operation:
On x86, that’s a lock add instruction, which forces a memory barrier and cache coherency protocol roundtrip. Slower than a plain add, but necessary for correctness.
For small objects that get created and destroyed frequently (integers, short strings), this overhead adds up. The 3.13t interpreter mitigates this with immortal objects (common small ints never get freed) and some biased reference counting tricks, but you still pay more than GIL Python.
Compatibility Gotchas
Free-threading is experimental. The python3.13t build is not the default — you have to compile it yourself with --disable-gil or download a pre-built binary from python.org.
Most C extensions are not thread-safe. NumPy, for example, assumes the GIL protects its internal state. Running NumPy code in free-threaded mode can segfault or produce wrong results. The core team is working on a Py_GIL_DISABLED macro so extensions can detect free-threaded builds and add their own locks, but adoption will be slow.
If you import numpy in a 3.13t script, it’ll probably work fine for single-threaded code. The moment you spawn threads and call NumPy from multiple threads simultaneously, undefined behavior.
Some pure Python code also breaks. If a library uses a global dict without locks (assuming the GIL protects it), you’ll get race conditions. I haven’t personally hit this yet, but it’s a real risk.
What About asyncio?
Free-threading doesn’t replace asyncio. Async is still the right tool for I/O-bound concurrency — you get cooperative multitasking, predictable context switches, and no thread overhead.
Free-threading is for CPU-bound tasks where you want to saturate all cores. If your bottleneck is waiting on HTTP responses, asyncio is still faster and lighter than threads.
That said, you can mix them. Run an asyncio event loop in one thread, spawn worker threads for CPU-heavy tasks, and use loop.run_in_executor() to bridge them. I haven’t benchmarked this yet, but in theory it should work.
Real-World Use Case: Batch Image Processing
I tried this on a batch thumbnail generator — resize 10,000 JPEG images using Pillow. Pillow’s core operations release the GIL (the actual pixel manipulation is in C), so even GIL Python can parallelize.
Results on Python 3.12 with ThreadPoolExecutor(max_workers=8):
Processed 10000 images in 42.1s
Same code on Python 3.13t:
Processed 10000 images in 41.8s
No speedup. Why? Pillow already releases the GIL during C calls, so 3.12 was already running in parallel. Free-threading only helps when the bottleneck is Python bytecode, not C extensions.
But when I tried a pure-Python image filter (a naive convolution kernel in Python lists, no NumPy), free-threading won decisively:
Python 3.12: 128.3s
Python 3.13t: 24.7s (5.2x speedup)
So the rule is: if your workload already releases the GIL, free-threading changes nothing. If your workload is pure Python CPU churn, free-threading is a game-changer.
Does this mean you should rewrite performance-critical code in pure Python to leverage free-threading? No. NumPy is still faster. But for cases where you can’t use C extensions (sandboxed environments, deployment constraints, or just prototyping), free-threading opens new doors.
Memory Usage Comparison
I also tracked RSS memory during the benchmark runs using psutil.
Python 3.12:
– Single-threaded: 48 MB
– 8 threads: 52 MB
Python 3.13t:
– Single-threaded: 61 MB
– 8 threads: 68 MB
Free-threaded builds use ~27% more memory even for single-threaded code. Some of that is larger object headers (per-object lock pointers), some is the new memory allocator that avoids lock contention.
For most applications, an extra 20 MB is noise. But if you’re running thousands of Python processes in containers (like AWS Lambda cold starts), this adds up.
When to Actually Use 3.13t
Here’s my decision tree:
Use free-threading if:
– Your bottleneck is pure Python CPU work (no NumPy/Pandas/scikit-learn)
– You need to saturate multiple cores in a single process
– Multiprocessing is too slow (large shared data, high IPC overhead)
– You’re okay with experimental stability and C extension breakage
Stick with GIL Python (3.12) if:
– You rely on C extensions (NumPy, PyTorch, etc.) — they’re not ready yet
– Your workload is I/O-bound (use asyncio instead)
– You care about single-threaded performance (3.13t is 15-20% slower)
– You need rock-solid production stability
Use multiprocessing if:
– You need parallelism right now on any Python version
– Your tasks are coarse-grained (seconds or minutes each) so IPC overhead is negligible
– You’re okay with the memory overhead of duplicated interpreters
I’d say free-threading is production-ready for narrow use cases — batch processing pipelines, embarrassingly parallel ETL jobs, CPU-heavy web workers — but not for general application code yet.
FAQ
Q: Will the GIL be removed from default Python builds?
Not yet. Python 3.13’s default build still has the GIL. Free-threading is opt-in (the t build). The core team wants to see real-world adoption and C extension migration before making it the default. Earliest timeline: Python 3.14 or 3.15.
Q: Can I install both 3.12 and 3.13t and switch between them?
Yes. Use pyenv or compile both from source. Just make sure your virtual environments explicitly pin the interpreter version. Running pip install in a 3.13t venv won’t magically make packages thread-safe — check library compatibility first.
Q: Does free-threading help with multiprocessing.Pool?
No. multiprocessing spawns separate OS processes, each with its own interpreter and GIL (or lack thereof). Free-threading only affects threading.Thread within a single process. If you’re already using multiprocessing.Pool and it works, free-threading won’t speed it up.
My Take: Wait One More Year
Free-threading delivers on the promise — real CPU parallelism in Python. The 5.6x speedup on my benchmark is transformative for pure Python workloads.
But the ecosystem isn’t ready. NumPy, SciPy, Pandas, PyTorch — none of them officially support free-threaded builds yet. The 17% single-threaded slowdown is a steep price if you can’t parallelize everything.
I’d wait until Python 3.14 or 3.15, when C extensions have caught up and the default build ships without the GIL. For now, free-threading is a glimpse of the future, not a production tool.
That said, if you maintain a pure Python library (no C dependencies) and your users need CPU parallelism, adding thread safety and advertising 3.13t compatibility is a legitimate move. You’ll be ahead of the curve when the ecosystem flips.
One thing I’m still unsure about: how well does free-threading scale to 32 or 64 cores? My 8-core results are promising, but I suspect atomic reference counting overhead grows non-linearly. If anyone’s tested this on a threadripper or server CPU, I’d love to see the numbers. Debugging concurrency issues at that scale when they do appear? That’s where The Art of Multiprocessor Programming becomes required reading.
For now, I’m sticking with multiprocessing for production CPU work and keeping a python3.13t install around for experiments. The GIL is dying, but it’s not dead yet.
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,861 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (962 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (816 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (806 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (602 views)