
The folklore about __slots__ has a mechanism attached to it, which is what makes it
persuasive: a slot is a fixed offset into the object, an instance dictionary is a hash
table, and an offset beats a hash. The first published entry in this series measured the
two on CPython 3.13.2 and found slot reads 12% slower. It printed the number and did not
explain it.
The explanation that suggests itself — the specialising interpreter picks a cheaper instruction for the plain class — is wrong in an interesting way. The instruction it picks for the slotted class is the cheaper of the two.
What the interpreter actually runs
Since 3.11, LOAD_ATTR does not stay as written. It carries a counter, and when the
counter triggers the interpreter overwrites the instruction with one specialised to the shape
of object it has been seeing. dis will show you the rewritten form if you ask for it with
adaptive=True.
$ python3 experiments/evalloop/specialise.py
python 3.13.2 on darwin (universal2)
2,000 reads per call, 200 cold and 400 warm calls per class
what LOAD_ATTR becomes
Plain LOAD_ATTR -> LOAD_ATTR_INSTANCE_VALUE on execution 2
Slotted LOAD_ATTR -> LOAD_ATTR_SLOT on execution 2
read cost best median
Plain unspecialised 5.98 6.77
Plain specialised 3.27 3.71
Slotted unspecialised 7.02 7.76
Slotted specialised 3.96 4.50
same read in a loop best
Plain 6.12
Slotted 6.61
Two executions. The first runs the generic instruction and advances its counter; on the
second the counter triggers, the specialising code overwrites the opcode and dispatches
straight back into it. From that second execution onward the loop in production is running a
different program from the one dis prints by default.
Neither of the two replacements hashes anything. Both are a type-version guard followed by
a pointer load at a constant offset, and the offset was computed once, when the instruction
was written. Here they are, from Python/bytecodes.c at the v3.13.2 tag:
macro(LOAD_ATTR_INSTANCE_VALUE) =
unused/1 + // Skip over the counter
_GUARD_TYPE_VERSION +
_CHECK_MANAGED_OBJECT_HAS_VALUES +
_LOAD_ATTR_INSTANCE_VALUE +
unused/5; // Skip over rest of cache
macro(LOAD_ATTR_SLOT) =
unused/1 +
_GUARD_TYPE_VERSION +
_LOAD_ATTR_SLOT + // NOTE: This action may also deopt
unused/5;
The slot form runs one guard where the instance-value form runs two. __slots__ does not
send the read through the slot descriptor at run time — the descriptor was consulted once,
by the specialising code, to find the offset that got baked into the instruction. So the
sentence that would explain a slower slot read, that the slot path does more work, is not
available. It does less.
What the measurement says
Specialisation is worth roughly half the read. An unspecialised read costs 6.77 ns against 3.71 for the same instruction once the interpreter has rewritten it. That is the number to keep in mind whenever a microbenchmark reports something surprising: a benchmark that never warms up is measuring a different interpreter.
There is a trap underneath that. Specialisation state lives on the code object, not on the
function. Building a fresh function out of one compiled module — exec over the same code
object, a closure, a decorator — hands back a reader whose instructions are already
specialised. The harness calls compile on every trial for exactly this reason, and it is
the only way the cold column above means anything.
The bottom block reproduces the shape the published entry used, a global read inside a
timeit loop: 6.12 ns against 6.61. The original run gave 5.95 and 6.66 — same
ordering, same size, seven weeks apart on the same interpreter.
The same two classes on five interpreters
One process is not enough to establish a 0.8 ns effect, because the gap moves with whatever
the loader did that time. across.py runs the measurement in nine separate processes per
interpreter and prints the spread beside the median:
$ python3 experiments/evalloop/across.py /opt/homebrew/bin/python3.10 /opt/homebrew/bin/python3.11 /opt/homebrew/bin/python3.13 /Library/Frameworks/Python.framework/Versions/3.13/bin/python3 /opt/homebrew/bin/python3.14
median ns per specialised read, 9 processes per interpreter
python build dict slot slot is spread over 9
3.10.21 arm64 7.25 6.38 12% faster 13% faster .. 10% faster
3.11.16 arm64 2.10 2.38 15% slower 7% slower .. 24% slower
3.13.15 arm64 3.69 3.60 2% faster 4% faster .. 2% slower
3.13.2 universal2 3.73 4.60 24% slower 20% slower .. 25% slower
3.14.7 arm64 2.42 2.46 2% slower 1% slower .. 3% slower
Read the spread column first. On 3.10.21, which has no specialisation, slot reads are faster in every one of the nine runs, and by a similar amount each time. That interpreter is where the folklore came from and on it the folklore is correct.
On 3.13.2 the gap is the other way and it is just as consistent: nine runs, all between 20 and 25% slower. And on 3.13.15 — the same minor version, the same machine, built by Homebrew for arm64 instead of by python.org as a universal2 binary — the same measurement straddles zero. 3.14.7 does not, but the whole of its spread sits between 1 and 3%. The 3.11.16 row does not hold a sign between invocations of the table at all: the run above says 15% slower, and in an earlier run of the same table the spread ran from 13% faster to 14% slower.
Two of the five builds put the difference inside 3%, and both of them are Homebrew arm64. One, the interpreter the folklore came from, is 12% in favour of slots. One cannot hold a sign. And one is 24% against slots in nine runs out of nine — the interpreter every published number on this site was taken on.
Where this leaves the decision
Use __slots__ for the memory. Forty bytes an instance is a layout fact: it travels
between builds and compilers, and at a million rows it is 40 MB.
Do not use it for read speed, and do not avoid it for read speed either. The published 12%
is a true statement about one binary on one machine and a false statement about __slots__,
and the entry it appeared in has been corrected
accordingly. If you need to know what your
own interpreter does, the answer costs one line: call the function a few times and hand it
to dis.dis with adaptive=True. The variant it prints is the one you are paying for.
Where to stop
Five builds, one machine, one architecture. The 3.13.2 row differs from the 3.13.15 row in three ways at once — patch level, vendor and universal2 against arm64 — so nothing here identifies which of the three produced the gap. It shows that the gap is not a property of the language, which is a weaker claim and the only one the measurement supports.
Every figure is a single-attribute class read through a local variable in a straight line of unrolled reads. Writes, method calls, polymorphic call sites that never keep a specialisation, the free-threaded build and the JIT are all untouched, and each of them has its own answer. And the whole effect is under one nanosecond: no production decision should turn on it in either direction.
Frequently asked
Should I stop using __slots__?
No. Use it for the 40 bytes an instance it saves, which is a layout fact and travels between builds. Do not use it for read speed in either direction — that number did not survive a change of build on the same machine.
Is the published 12% figure wrong?
It is right about the interpreter it was taken on and wrong as a statement about __slots__. The archive measured CPython 3.13.2 from python.org, and on that binary the gap reproduces at 20 to 25% across nine processes. It does not reproduce on the other four builds measured here.
How do I check this on my own interpreter?
Call a function a few times, then pass it to dis.dis with adaptive=True. The disassembler prints the instruction the interpreter rewrote rather than the one the compiler emitted, so you see which LOAD_ATTR variant your class actually got.
Does the specialised instruction ever change back?
Yes. A specialised instruction that stops seeing the shape it was written for is de-optimised back to the generic LOAD_ATTR, which then tries again later. That is why a benchmark over one object shape measures something a polymorphic call site never gets.
Where this came from
Run it yourself — every figure above came out of these:
- evalloop/specialise.py — Specialised attribute reads, and the opcodes dis reports for each
- evalloop/across.py — The same reads across five interpreters and two builds
Read, rather than assumed:


