Inside the interpreter · Part 4 of 4

Slots Did Not Get Slower, One Build Did

A read through __slots__ costs 24% more than one through an instance dictionary on CPython 3.13.2. Rebuild the same version and the gap goes away.

Elliot Sayer5 min read

The folklore about __slots__ has a mechanism attached to it, which is what makes it persuasive: a slot is a fixed offset into the object, an instance dictionary is a hash table, and an offset beats a hash. The first published entry in this series measured the two on CPython 3.13.2 and found slot reads 12% slower. It printed the number and did not explain it.

The explanation that suggests itself — the specialising interpreter picks a cheaper instruction for the plain class — is wrong in an interesting way. The instruction it picks for the slotted class is the cheaper of the two.

What the interpreter actually runs

Since 3.11, LOAD_ATTR does not stay as written. It carries a counter, and when the counter triggers the interpreter overwrites the instruction with one specialised to the shape of object it has been seeing. dis will show you the rewritten form if you ask for it with adaptive=True.

$ python3 experiments/evalloop/specialise.py
python 3.13.2 on darwin (universal2)
2,000 reads per call, 200 cold and 400 warm calls per class

what LOAD_ATTR becomes
  Plain     LOAD_ATTR -> LOAD_ATTR_INSTANCE_VALUE  on execution 2
  Slotted   LOAD_ATTR -> LOAD_ATTR_SLOT            on execution 2

read cost                   best   median
  Plain unspecialised       5.98     6.77
  Plain specialised         3.27     3.71
  Slotted unspecialised     7.02     7.76
  Slotted specialised       3.96     4.50

same read in a loop         best
  Plain                     6.12
  Slotted                   6.61

Two executions. The first runs the generic instruction and advances its counter; on the second the counter triggers, the specialising code overwrites the opcode and dispatches straight back into it. From that second execution onward the loop in production is running a different program from the one dis prints by default.

Neither of the two replacements hashes anything. Both are a type-version guard followed by a pointer load at a constant offset, and the offset was computed once, when the instruction was written. Here they are, from Python/bytecodes.c at the v3.13.2 tag:

        macro(LOAD_ATTR_INSTANCE_VALUE) =
            unused/1 + // Skip over the counter
            _GUARD_TYPE_VERSION +
            _CHECK_MANAGED_OBJECT_HAS_VALUES +
            _LOAD_ATTR_INSTANCE_VALUE +
            unused/5;  // Skip over rest of cache

        macro(LOAD_ATTR_SLOT) =
            unused/1 +
            _GUARD_TYPE_VERSION +
            _LOAD_ATTR_SLOT +  // NOTE: This action may also deopt
            unused/5;

The slot form runs one guard where the instance-value form runs two. __slots__ does not send the read through the slot descriptor at run time — the descriptor was consulted once, by the specialising code, to find the offset that got baked into the instruction. So the sentence that would explain a slower slot read, that the slot path does more work, is not available. It does less.

What the measurement says

Specialisation is worth roughly half the read. An unspecialised read costs 6.77 ns against 3.71 for the same instruction once the interpreter has rewritten it. That is the number to keep in mind whenever a microbenchmark reports something surprising: a benchmark that never warms up is measuring a different interpreter.

There is a trap underneath that. Specialisation state lives on the code object, not on the function. Building a fresh function out of one compiled module — exec over the same code object, a closure, a decorator — hands back a reader whose instructions are already specialised. The harness calls compile on every trial for exactly this reason, and it is the only way the cold column above means anything.

The bottom block reproduces the shape the published entry used, a global read inside a timeit loop: 6.12 ns against 6.61. The original run gave 5.95 and 6.66 — same ordering, same size, seven weeks apart on the same interpreter.

The same two classes on five interpreters

One process is not enough to establish a 0.8 ns effect, because the gap moves with whatever the loader did that time. across.py runs the measurement in nine separate processes per interpreter and prints the spread beside the median:

$ python3 experiments/evalloop/across.py /opt/homebrew/bin/python3.10 /opt/homebrew/bin/python3.11 /opt/homebrew/bin/python3.13 /Library/Frameworks/Python.framework/Versions/3.13/bin/python3 /opt/homebrew/bin/python3.14
median ns per specialised read, 9 processes per interpreter

python    build         dict  slot   slot is     spread over 9
3.10.21   arm64         7.25  6.38   12% faster  13% faster .. 10% faster
3.11.16   arm64         2.10  2.38   15% slower  7% slower .. 24% slower
3.13.15   arm64         3.69  3.60   2% faster   4% faster .. 2% slower
3.13.2    universal2    3.73  4.60   24% slower  20% slower .. 25% slower
3.14.7    arm64         2.42  2.46   2% slower   1% slower .. 3% slower

Read the spread column first. On 3.10.21, which has no specialisation, slot reads are faster in every one of the nine runs, and by a similar amount each time. That interpreter is where the folklore came from and on it the folklore is correct.

On 3.13.2 the gap is the other way and it is just as consistent: nine runs, all between 20 and 25% slower. And on 3.13.15 — the same minor version, the same machine, built by Homebrew for arm64 instead of by python.org as a universal2 binary — the same measurement straddles zero. 3.14.7 does not, but the whole of its spread sits between 1 and 3%. The 3.11.16 row does not hold a sign between invocations of the table at all: the run above says 15% slower, and in an earlier run of the same table the spread ran from 13% faster to 14% slower.

Two of the five builds put the difference inside 3%, and both of them are Homebrew arm64. One, the interpreter the folklore came from, is 12% in favour of slots. One cannot hold a sign. And one is 24% against slots in nine runs out of nine — the interpreter every published number on this site was taken on.

Where this leaves the decision

Use __slots__ for the memory. Forty bytes an instance is a layout fact: it travels between builds and compilers, and at a million rows it is 40 MB.

Do not use it for read speed, and do not avoid it for read speed either. The published 12% is a true statement about one binary on one machine and a false statement about __slots__, and the entry it appeared in has been corrected accordingly. If you need to know what your own interpreter does, the answer costs one line: call the function a few times and hand it to dis.dis with adaptive=True. The variant it prints is the one you are paying for.

Where to stop

Five builds, one machine, one architecture. The 3.13.2 row differs from the 3.13.15 row in three ways at once — patch level, vendor and universal2 against arm64 — so nothing here identifies which of the three produced the gap. It shows that the gap is not a property of the language, which is a weaker claim and the only one the measurement supports.

Every figure is a single-attribute class read through a local variable in a straight line of unrolled reads. Writes, method calls, polymorphic call sites that never keep a specialisation, the free-threaded build and the JIT are all untouched, and each of them has its own answer. And the whole effect is under one nanosecond: no production decision should turn on it in either direction.

Frequently asked

Should I stop using __slots__?

No. Use it for the 40 bytes an instance it saves, which is a layout fact and travels between builds. Do not use it for read speed in either direction — that number did not survive a change of build on the same machine.

Is the published 12% figure wrong?

It is right about the interpreter it was taken on and wrong as a statement about __slots__. The archive measured CPython 3.13.2 from python.org, and on that binary the gap reproduces at 20 to 25% across nine processes. It does not reproduce on the other four builds measured here.

How do I check this on my own interpreter?

Call a function a few times, then pass it to dis.dis with adaptive=True. The disassembler prints the instruction the interpreter rewrote rather than the one the compiler emitted, so you see which LOAD_ATTR variant your class actually got.

Does the specialised instruction ever change back?

Yes. A specialised instruction that stops seeing the shape it was written for is de-optimised back to the generic LOAD_ATTR, which then tries again later. That is why a benchmark over one object shape measures something a polymorphic call site never gets.

Where this came from

Run it yourself — every figure above came out of these:

  1. evalloop/specialise.py — Specialised attribute reads, and the opcodes dis reports for each
  2. evalloop/across.py — The same reads across five interpreters and two builds

Read, rather than assumed:

  1. PEP 659 — Specializing Adaptive Interpreter
  2. Python/bytecodes.c at CPython v3.13.2
  3. __slots__, Python 3.13 language reference
Share

Arrow keys to move, Enter to open.