21 — Concurrency: Threading & Multiprocessing
The GIL in Practice — When Threading Helps and When It Doesn't
# ── Production: benchmarking I/O-bound vs CPU-bound across threading/multiprocessing ──
# Demonstrates the GIL's impact with real timing, not theoretical claims.
import threading
import multiprocessing as mp
import time
import requests
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor, as_completed
# ── I/O-bound: threading WINS (GIL released during network waits) ──
def fetch_url(url: str) -> tuple[str, int]:
"""I/O-bound task — releases GIL during socket recv, allowing true concurrency."""
try:
resp = requests.get(url, timeout=5)
return url, resp.status_code
except Exception as e:
return url, -1
def io_bound_benchmark(urls: list[str]):
# Sequential — one at a time, total = sum of all latencies
t0 = time.perf_counter()
for url in urls:
fetch_url(url)
sequential = time.perf_counter() - t0
# Threaded — overlapping I/O waits, total ≈ slowest single request
t0 = time.perf_counter()
with ThreadPoolExecutor(max_workers=10) as pool:
list(pool.map(fetch_url, urls))
threaded = time.perf_counter() - t0
print(f"I/O-bound: sequential={sequential:.2f}s, threaded={threaded:.2f}s")
print(f" speedup: {sequential/threaded:.1f}x — threading helps (GIL released during I/O)")
# ── CPU-bound: multiprocessing WINS (GIL blocks true parallelism in threads) ──
def cpu_intensive(n: int) -> int:
"""Pure Python computation — GIL prevents parallel bytecode execution across threads."""
total = 0
for i in range(n):
total += i * i
return total
def cpu_bound_benchmark(n: int, workers: int):
tasks = [n] * workers
# Sequential
t0 = time.perf_counter()
for task in tasks:
cpu_intensive(task)
sequential = time.perf_counter() - t0
# Threaded — NO speedup (GIL serializes bytecode execution)
t0 = time.perf_counter()
with ThreadPoolExecutor(max_workers=workers) as pool:
list(pool.map(cpu_intensive, tasks))
threaded = time.perf_counter() - t0
# Multiprocessing — TRUE parallelism (separate interpreters, separate GILs)
t0 = time.perf_counter()
with ProcessPoolExecutor(max_workers=workers) as pool:
list(pool.map(cpu_intensive, tasks))
multiprocess = time.perf_counter() - t0
print(f"CPU-bound: sequential={sequential:.2f}s, threaded={threaded:.2f}s, multiprocess={multiprocess:.2f}s")
print(f" threading speedup: {sequential/threaded:.1f}x (≈1.0 — GIL blocks parallelism)")
print(f" multiprocessing speedup: {sequential/multiprocess:.1f}x (≈workers — true parallelism)")
if __name__ == "__main__": # REQUIRED for ProcessPoolExecutor (spawn re-imports module)
cpu_bound_benchmark(n=5_000_000, workers=4)
# ── Race condition: demonstrating the lost-update problem ──
# counter += 1 is NOT atomic — it's LOAD + ADD + STORE, interruptible between any step.
import threading
counter = 0
lock = threading.Lock()
def increment_unsafe(n: int):
"""RACE CONDITION: counter += 1 compiles to multiple bytecodes, GIL can switch mid-way."""
global counter
for _ in range(n):
counter += 1 # LOAD_FAST counter; LOAD_CONST 1; BINARY_OP +=; STORE_FAST counter
# GIL can release between LOAD and STORE → two threads read same old value
def increment_safe(n: int, lock: threading.Lock):
"""CORRECT: lock serializes the read-modify-write, no interleaving possible."""
global counter
for _ in range(n):
with lock: # acquire on entry, release on exit (even on exception)
counter += 1 # only one thread can be inside this block at a time
# Demonstrate the race
counter = 0
threads = [threading.Thread(target=increment_unsafe, args=(100_000,)) for _ in range(4)]
for t in threads: t.start()
for t in threads: t.join()
print(f"Unsafe (expected 400000): {counter}") # e.g. 387,341 — lost updates!
# Demonstrate the fix
counter = 0
lock = threading.Lock()
threads = [threading.Thread(target=increment_safe, args=(100_000, lock)) for _ in range(4)]
for t in threads: t.start()
for t in threads: t.join()
print(f"Safe (expected 400000): {counter}") # 400000 — always correct
threading — Shared Memory, Explicit Locks
import threading
counter = 0
def increment_unsafe():
global counter
for _ in range(100_000):
counter += 1 # NOT atomic! read-modify-write across three bytecode ops, interruptible mid-way
threads = [threading.Thread(target=increment_unsafe) for _ in range(4)]
for t in threads:
t.start()
for t in threads:
t.join()
print(counter) # expected 400,000 — often prints something LESS, e.g. 391,842 — a lost-update race condition
This is a race condition: counter += 1 is not a single atomic operation — it compiles to a load, an add, and a store, and the GIL can switch threads between any of those steps. Two threads can both read the same old value before either writes back the incremented result, silently losing an update. The fix is a Lock:
import threading
counter = 0
lock = threading.Lock()
def increment_safe():
global counter
for _ in range(100_000):
with lock: # only ONE thread can hold the lock at a time
counter += 1
threads = [threading.Thread(target=increment_safe) for _ in range(4)]
for t in threads:
t.start()
for t in threads:
t.join()
print(counter) # always exactly 400,000
with lock: acquires on entry and releases on exit (even if an exception is raised inside), exactly like the context managers from chapters 18-19 — never call lock.acquire()/lock.release() manually, since a missed release() on an exception path causes every other thread waiting on that lock to hang forever (a deadlock).
RLock, Condition, and Event
import threading
# RLock — reentrant: the SAME thread can acquire it multiple times without deadlocking itself
rlock = threading.RLock()
def outer():
with rlock:
inner() # inner() also acquires rlock — fine, same thread, RLock tracks a counter
def inner():
with rlock:
print("inner")
# A plain Lock here would deadlock: the second `with lock` in the same thread would block forever
# waiting for a lock the SAME thread already holds and will never release until this call returns.
import threading
import time
ready = threading.Event()
def waiter():
print("waiting for signal...")
ready.wait() # blocks until .set() is called from any thread
print("signal received, proceeding")
def setter():
time.sleep(1)
ready.set()
threading.Thread(target=waiter).start()
threading.Thread(target=setter).start()
Event is the simplest cross-thread signaling primitive — one or more threads wait(), another thread calls set(), and all waiters wake up. Condition extends this with the ability to wait for an arbitrary predicate and notify one (notify()) or all (notify_all()) waiters, the standard tool for producer/consumer patterns.
queue.Queue — Thread-Safe Producer/Consumer
import threading
import queue
task_queue = queue.Queue()
def producer():
for i in range(5):
task_queue.put(i)
task_queue.put(None) # sentinel value signaling "no more work"
def consumer():
while (item := task_queue.get()) is not None:
print(f"processing {item}")
task_queue.task_done()
task_queue.task_done()
threading.Thread(target=producer).start()
threading.Thread(target=consumer).start()
queue.Queue handles all locking internally — put()/get() are safe to call concurrently from any number of threads without a manual Lock, making it the preferred way to hand data between threads instead of shared mutable variables protected by hand-managed locks.
concurrent.futures — The High-Level Pool API
from concurrent.futures import ThreadPoolExecutor, as_completed
import requests
urls = ["https://example.com"] * 5
with ThreadPoolExecutor(max_workers=5) as executor:
futures = {executor.submit(requests.get, url): url for url in urls}
for future in as_completed(futures):
url = futures[future]
try:
response = future.result()
print(f"{url}: {response.status_code}")
except Exception as exc:
print(f"{url} raised {exc!r}")
ThreadPoolExecutor manages a fixed pool of worker threads and a task queue for you — submit() returns a Future immediately (non-blocking), and .result() blocks until that specific task completes, re-raising any exception the task raised. This is almost always preferable to manually creating and joining Thread objects for anything beyond a couple of one-off threads.
multiprocessing — True Parallelism via Separate Processes
from multiprocessing import Process, Pool
import os
def cpu_bound_work(n):
total = 0
for i in range(n):
total += i
return total
if __name__ == "__main__": # REQUIRED on Windows/macOS spawn start method — see gotcha below
with Pool(processes=4) as pool:
results = pool.map(cpu_bound_work, [50_000_000] * 4)
print(results) # genuinely runs on 4 cores in parallel — no GIL contention, separate interpreters
Each process spawned by multiprocessing gets its own Python interpreter and its own GIL — there is no shared memory and no lock contention between them, which is why CPU-bound work genuinely parallelizes across cores this way. The cost: data passed between processes must be pickled (chapter 18) to cross the process boundary, which has real overhead for large objects, and processes don't share global variables or objects by default.
from multiprocessing import Process, Value, Array, Manager
def worker(shared_counter, lock):
with lock:
shared_counter.value += 1
if __name__ == "__main__":
from multiprocessing import Lock
counter = Value("i", 0) # 'i' = C int, allocated in SHARED memory across processes
lock = Lock()
processes = [Process(target=worker, args=(counter, lock)) for _ in range(10)]
for p in processes:
p.start()
for p in processes:
p.join()
print(counter.value) # 10 — Value/Array are the explicit mechanisms for sharing primitive state
Value and Array are multiprocessing's narrow escape hatches for sharing simple, fixed-layout data (a single number, a fixed-size array) across process boundaries via actual shared memory — for anything more complex (dicts, lists, custom objects), use a Manager(), which runs a separate server process brokering access, at higher overhead than Value/Array but much more flexible.
if __name__ == "__main__": — Not Optional With multiprocessing
# WRONG on Windows and macOS (default 'spawn' start method) — causes infinite recursive process spawning
from multiprocessing import Process
def worker():
print("working")
p = Process(target=worker) # module-level code, NOT guarded
p.start()
# When spawn re-imports this module in the child process, THIS LINE RUNS AGAIN,
# spawning another child, which re-imports and spawns another... (RuntimeError, usually caught by a guard)
On the spawn start method (the default on Windows and macOS since Python 3.8), each new process is a fresh interpreter that re-imports the launching module from scratch — any code at module level (not inside if __name__ == "__main__":) re-executes in the child, including the Process(...).start() call itself if it isn't guarded, causing runaway recursive process creation. On Linux, the default fork start method copies the parent's already-initialized memory instead, which is why this bug is often invisible in Linux-only development and then explodes in CI or on a teammate's Mac.
from multiprocessing import Process
def worker():
print("working")
if __name__ == "__main__": # guards module-level process-spawning code — REQUIRED, not just tidy style
p = Process(target=worker)
p.start()
p.join()
ProcessPoolExecutor — The High-Level Multiprocessing API
from concurrent.futures import ProcessPoolExecutor
def is_prime(n):
if n < 2:
return False
return all(n % i for i in range(2, int(n ** 0.5) + 1))
if __name__ == "__main__":
numbers = list(range(100_000, 100_020))
with ProcessPoolExecutor(max_workers=4) as executor:
results = list(executor.map(is_prime, numbers))
print(list(zip(numbers, results)))
ProcessPoolExecutor shares the exact same .submit()/.map()/as_completed() API as ThreadPoolExecutor — swapping between thread-based and process-based parallelism for a given workload is often a one-line change, which makes it easy to benchmark both and pick whichever actually performs better for the workload at hand.
Choosing Threading vs Multiprocessing vs Async
| Workload | Best tool | Why |
|---|---|---|
| I/O-bound (network, disk, waiting) | threading or asyncio | GIL is released during I/O waits; no need for separate processes |
| CPU-bound (heavy computation) | multiprocessing | Bypasses the GIL entirely via separate interpreters/processes |
| Thousands of concurrent I/O tasks | asyncio | Threads have real memory/OS overhead per thread; coroutines are far cheaper (chapter 22) |
| Mixed / simplicity over raw throughput | concurrent.futures (either pool) | Uniform, simple API; easy to swap thread pool for process pool |
💡 Tips & Tricks
- Performance: profile before reaching for
multiprocessing— process creation and pickling data across the boundary has real overhead, and for small workloads a naive multiprocessing version can be slower than a single-threaded one; it pays off on genuinely CPU-heavy, easily-partitioned work. - Debug: a hanging program that never exits is often a
Lockacquired but never released on an exception path, or a non-daemon thread that was never.join()-ed — usethreading.enumerate()to list all currently alive threads when diagnosing a hang. - Idiom: prefer
concurrent.futures(ThreadPoolExecutor/ProcessPoolExecutor) over rawthreading.Thread/multiprocessing.Processfor anything beyond a single one-off background task — the pool API handles queuing, result collection, and exception propagation for you. - Safety: never share a mutable object across threads without a
Lock(or usequeue.Queue, which is internally safe) — "it worked in my testing" is not evidence of thread safety, since race conditions are often timing-dependent and can pass thousands of test runs before failing in production under different load. - Idiom: set
daemon=Trueon background threads that should not prevent the program from exiting (e.g., a periodic heartbeat) — a non-daemon thread that's still running keeps the entire Python process alive even aftermain()returns.
⚠️ Edge Cases & Gotchas
- The GIL means
threadingprovides no real parallelism for CPU-bound Python code — adding more threads to a compute-heavy loop can even be slower than sequential execution, due to the overhead of the GIL being repeatedly released and reacquired between threads; usemultiprocessingfor CPU-bound parallelism instead. counter += 1is not atomic, even though it looks like a single operation — it's a read, an add, and a write, and the GIL can switch threads between any of those steps, producing lost updates under concurrent access without aLock.- On Windows and macOS,
multiprocessingcode not guarded byif __name__ == "__main__":can spawn processes recursively without bound, because thespawnstart method re-imports the launching module fresh in every child process, re-running any unguarded module-level code including the process-creation call itself. - Data passed to a
multiprocessing.PoolorProcessmust be picklable — lambdas, open file handles, database connections, and many other objects can't cross the process boundary and raisePicklingErrorat the point of submission, not always where the developer expects. - A
Lockacquired in atryblock withoutfinally(or, better, without usingwith lock:) that is never released on an exception path causes every other thread waiting on that same lock to block forever — a silent deadlock with no traceback, since the waiting threads are simply parked, not crashed.
🧠 Spot the Bug
A script fans out ten "download and process" tasks across threads and accumulates results into a shared list. It occasionally produces fewer than ten results, with no exception raised. Find the bug.
import threading
results = []
def fetch_and_store(item_id):
data = f"result-{item_id}"
if len(results) < 10:
results.append(data)
threads = [threading.Thread(target=fetch_and_store, args=(i,)) for i in range(10)]
for t in threads:
t.start()
for t in threads:
t.join()
print(len(results)) # sometimes 10, sometimes less
Answer
list.append() itself is thread-safe in CPython (it's a single bytecode-level operation protected by the GIL), but if len(results) < 10: results.append(data) is two separate operations — a length check, then an append — with no atomicity guarantee across the two. Between one thread checking len(results) < 10 and it actually calling .append(), other threads can interleave, but that's not even the real bug here (all 10 threads should still append since none of them exceed 10 in this scenario) — the actual issue is subtler: len(results) < 10 is a check-then-act race in general, and in a variation of this pattern with more threads than the limit, some threads pass the check with a stale length, all append, and the list ends up larger than intended, or (as observed here) if any exception occurs inside a thread's target function it is silently swallowed by threading.Thread — printed to stderr but never re-raised to the main thread — masking a fetch_and_store failure entirely and undercounting results with no visible traceback in the main thread's flow.
The fix is a lock around the read-modify-write of the shared list, and explicit propagation of any exception raised inside a thread:
import threading
results = []
lock = threading.Lock()
errors = []
def fetch_and_store(item_id):
try:
data = f"result-{item_id}"
with lock:
results.append(data)
except Exception as exc:
errors.append(exc)
threads = [threading.Thread(target=fetch_and_store, args=(i,)) for i in range(10)]
for t in threads:
t.start()
for t in threads:
t.join()
if errors:
raise errors[0]
print(len(results)) # reliably 10
The lesson: threading.Thread swallows exceptions raised inside the target function (printing a traceback to stderr but not stopping the main thread or .join()), and any check-then-act sequence on shared mutable state needs a Lock around the whole sequence, not just the final mutation — always assume interleaving can happen between any two lines touching shared state.
Key Takeaways
- The GIL allows only one thread to execute Python bytecode at a time in CPython, which means
threadingdoes not provide real parallelism for CPU-bound work — it's only helpful for I/O-bound work, where the GIL is released during waits. multiprocessingsidesteps the GIL by using separate OS processes with their own interpreters, at the cost of pickling overhead for data crossing the process boundary and no default shared memory.- Shared mutable state accessed from multiple threads needs an explicit
Lockaround every read-modify-write sequence — operations that look atomic (counter += 1, "check then append") frequently are not. if __name__ == "__main__":is mandatory, not stylistic, aroundmultiprocessingcode on Windows/macOS — thespawnstart method re-imports the module fresh in each child process.concurrent.futures.ThreadPoolExecutor/ProcessPoolExecutorprovide a uniform, higher-level pool API that's usually preferable to hand-managing rawThread/Processobjects, and make it easy to switch between the two.- Exceptions raised inside a
threading.Threadtarget are not automatically propagated to the main thread — they must be explicitly caught and communicated back (e.g., via a shared list,queue.Queue, orconcurrent.futures.Future.result(), which does re-raise).