# Slots Did Not Get Slower, One Build Did

> A read through __slots__ costs 24% more than one through an instance dictionary on CPython 3.13.2. Rebuild the same version and the gap goes away.

- Published: 2026-08-30
- Category: runtime
- Tags: object-layout, benchmarking
- Sources: https://github.com/CognatePress/erperiments.pydepth.com/blob/main/evalloop/specialise.py, https://github.com/CognatePress/erperiments.pydepth.com/blob/main/evalloop/across.py, https://peps.python.org/pep-0659/, https://github.com/python/cpython/blob/v3.13.2/Python/bytecodes.c, https://docs.python.org/3.13/reference/datamodel.html#object.__slots__
- Source: https://pydepth.com/blog/slots-did-not-get-slower/
- Language: en-US
- Author: Elliot Sayer

---
The folklore about `__slots__` has a mechanism attached to it, which is what makes it
persuasive: a slot is a fixed offset into the object, an instance dictionary is a hash
table, and an offset beats a hash. The first published entry in this series measured the
two on CPython 3.13.2 and found slot reads 12% slower. It printed the number and did not
explain it.

The explanation that suggests itself — the specialising interpreter picks a cheaper
instruction for the plain class — is wrong in an interesting way. The instruction it picks
for the slotted class is the cheaper of the two.

## What the interpreter actually runs

Since 3.11, `LOAD_ATTR` does not stay as written. It carries a counter, and when the
counter triggers the interpreter overwrites the instruction with one specialised to the shape
of object it has been seeing. `dis` will show you the rewritten form if you ask for it with
`adaptive=True`.

```
$ python3 experiments/evalloop/specialise.py
python 3.13.2 on darwin (universal2)
2,000 reads per call, 200 cold and 400 warm calls per class

what LOAD_ATTR becomes
  Plain     LOAD_ATTR -> LOAD_ATTR_INSTANCE_VALUE  on execution 2
  Slotted   LOAD_ATTR -> LOAD_ATTR_SLOT            on execution 2

read cost                   best   median
  Plain unspecialised       5.98     6.77
  Plain specialised         3.27     3.71
  Slotted unspecialised     7.02     7.76
  Slotted specialised       3.96     4.50

same read in a loop         best
  Plain                     6.12
  Slotted                   6.61
```

Two executions. The first runs the generic instruction and advances its counter; on the
second the counter triggers, the specialising code overwrites the opcode and dispatches
straight back into it. From that second execution onward the loop in production is running a
different program from the one `dis` prints by default.

Neither of the two replacements hashes anything. Both are a type-version guard followed by
a pointer load at a constant offset, and the offset was computed once, when the instruction
was written. Here they are, from `Python/bytecodes.c` at the v3.13.2 tag:

```c
        macro(LOAD_ATTR_INSTANCE_VALUE) =
            unused/1 + // Skip over the counter
            _GUARD_TYPE_VERSION +
            _CHECK_MANAGED_OBJECT_HAS_VALUES +
            _LOAD_ATTR_INSTANCE_VALUE +
            unused/5;  // Skip over rest of cache

        macro(LOAD_ATTR_SLOT) =
            unused/1 +
            _GUARD_TYPE_VERSION +
            _LOAD_ATTR_SLOT +  // NOTE: This action may also deopt
            unused/5;
```

The slot form runs one guard where the instance-value form runs two. `__slots__` does not
send the read through the slot descriptor at run time — the descriptor was consulted once,
by the specialising code, to find the offset that got baked into the instruction. So the
sentence that would explain a slower slot read, that the slot path does more work, is not
available. It does less.

## What the measurement says

Specialisation is worth roughly half the read. An unspecialised read costs 6.77 ns against
3.71 for the same instruction once the interpreter has rewritten it. That is the number to
keep in mind whenever a microbenchmark reports something surprising: a benchmark that never
warms up is measuring a different interpreter.

There is a trap underneath that. Specialisation state lives on the code object, not on the
function. Building a fresh function out of one compiled module — `exec` over the same code
object, a closure, a decorator — hands back a reader whose instructions are already
specialised. The harness calls `compile` on every trial for exactly this reason, and it is
the only way the cold column above means anything.

The bottom block reproduces the shape the published entry used, a global read inside a
`timeit` loop: 6.12 ns against 6.61. The original run gave 5.95 and 6.66 — same
ordering, same size, seven weeks apart on the same interpreter.

## The same two classes on five interpreters

One process is not enough to establish a 0.8 ns effect, because the gap moves with whatever
the loader did that time. `across.py` runs the measurement in nine separate processes per
interpreter and prints the spread beside the median:

```
$ python3 experiments/evalloop/across.py /opt/homebrew/bin/python3.10 /opt/homebrew/bin/python3.11 /opt/homebrew/bin/python3.13 /Library/Frameworks/Python.framework/Versions/3.13/bin/python3 /opt/homebrew/bin/python3.14
median ns per specialised read, 9 processes per interpreter

python    build         dict  slot   slot is     spread over 9
3.10.21   arm64         7.25  6.38   12% faster  13% faster .. 10% faster
3.11.16   arm64         2.10  2.38   15% slower  7% slower .. 24% slower
3.13.15   arm64         3.69  3.60   2% faster   4% faster .. 2% slower
3.13.2    universal2    3.73  4.60   24% slower  20% slower .. 25% slower
3.14.7    arm64         2.42  2.46   2% slower   1% slower .. 3% slower
```

Read the spread column first. On 3.10.21, which has no specialisation, slot reads are
faster in every one of the nine runs, and by a similar amount each time. That interpreter is
where the folklore came from and on it the folklore is correct.

On 3.13.2 the gap is the other way and it is just as consistent: nine runs, all between 20
and 25% slower. And on 3.13.15 — the same minor version, the same machine, built by
Homebrew for arm64 instead of by python.org as a universal2 binary — the same measurement
straddles zero. 3.14.7 does not, but the whole of its spread sits between 1 and 3%. The 3.11.16 row does not
hold a sign between invocations of the table at all: the run above says 15% slower, and in an
earlier run of the same table the spread ran from 13% faster to 14% slower.

Two of the five builds put the difference inside 3%, and both of them are Homebrew arm64.
One, the interpreter the folklore came from, is 12% in favour of slots. One cannot hold a
sign. And one is 24% against slots in nine runs out of nine — the interpreter
every published number on this site was taken on.

## Where this leaves the decision

Use `__slots__` for the memory. Forty bytes an instance is a layout fact: it travels
between builds and compilers, and at a million rows it is 40 MB.

Do not use it for read speed, and do not avoid it for read speed either. The published 12%
is a true statement about one binary on one machine and a false statement about `__slots__`,
and [the entry it appeared in](/blog/the-shape-of-a-python-object/) has been corrected
accordingly. If you need to know what your
own interpreter does, the answer costs one line: call the function a few times and hand it
to `dis.dis` with `adaptive=True`. The variant it prints is the one you are paying for.

## Where to stop

Five builds, one machine, one architecture. The 3.13.2 row differs from the 3.13.15 row in
three ways at once — patch level, vendor and universal2 against arm64 — so nothing here
identifies which of the three produced the gap. It shows that the gap is not a property of
the language, which is a weaker claim and the only one the measurement supports.

Every figure is a single-attribute class read through a local variable in a straight line of
unrolled reads. Writes, method calls, polymorphic call sites that never keep a
specialisation, the free-threaded build and the JIT are all untouched, and each of them has
its own answer. And the whole effect is under one nanosecond: no production decision should
turn on it in either direction.
