# A NamedTuple Read Cost 0.81x a Plain One, Then 1.9x

> On CPython 3.10 a NamedTuple attribute read cost 0.81x a read through an instance dictionary. On 3.11 it costs 1.90x. Nothing about NamedTuple changed.

- Published: 2026-09-13
- Category: runtime
- Tags: object-layout, memory, versions
- Sources: https://github.com/CognatePress/erperiments.pydepth.com/blob/main/objshape/containers.py, https://peps.python.org/pep-412/, https://github.com/python/cpython/blob/v3.13.2/Objects/dictobject.c, https://github.com/python/cpython/blob/v3.13.2/Include/internal/pycore_dict.h, https://github.com/python/cpython/blob/v3.13.2/Python/specialize.c
- Source: https://pydepth.com/blog/the-cost-of-a-field/
- Language: en-US
- Author: Elliot Sayer

---
Two tables decide most of these choices. A widely cited one puts a NamedTuple's attribute
access at 43.2 ns against a dataclass's 49.7 and its instance at 56 bytes against 152 —
smaller and faster on both counts, measured on CPython 3.8. A second reports that
`slots=True` cuts memory by a factor of eight and makes attribute access a quarter faster,
measured on 3.11.

Measure the same six containers on 3.10 and the NamedTuple read is the cheapest of them:
8.06 ns against 9.95 for a read through an instance dictionary, a ratio of 0.81. Measure it on
3.11 and the read costs 6.74 ns against 3.54 — 1.90x, and by 3.13 it is the dearest read in
the table. Nothing about NamedTuple changed between those two numbers.

## What the six containers cost

A plain class, a `dataclass`, a `dataclass(slots=True)`, a `NamedTuple`, a `TypedDict` and a
bare tuple, all holding the same three fields. The bytes come from a `tracemalloc` delta over
200,000 instances rather than from `getsizeof`, which
[reports a figure that moves with how many instances the class has made](/blog/what-getsizeof-does-not-count/); the reads are the median of nine processes per interpreter, because the
gap between two of these containers is smaller than the gap between two runs of the same one:

```
$ python3 experiments/objshape/containers.py /opt/homebrew/bin/python3.10 \
    /opt/homebrew/bin/python3.11 /opt/homebrew/bin/python3.13 /opt/homebrew/bin/python3.14
python 3.13.2 on darwin (universal2)
200,000 instances of three fields each

container             bytes/inst   read is
plain class                 96.0   LOAD_ATTR_INSTANCE_VALUE
dataclass                   96.0   LOAD_ATTR_INSTANCE_VALUE
dataclass(slots)            56.0   LOAD_ATTR_SLOT
NamedTuple                  72.0   LOAD_ATTR
TypedDict                  184.0   obj['y']
tuple                       64.0   obj[1]

ns per read, 9 processes per interpreter, best of 7 x 1,000,000 in each

python 3.13.2 (universal2)
  container             median  slowest  fastest  vs plain class
  plain class             5.84     6.06     5.59         1.00x
  dataclass               5.76     6.02     5.64         0.99x
  dataclass(slots)        6.44     6.59     6.24         1.10x
  NamedTuple              9.52     9.63     9.42         1.63x
  TypedDict               6.48     6.87     6.25         1.11x
  tuple                   6.15     6.36     5.97         1.05x

python 3.10.21 (arm64)
  container             median  slowest  fastest  vs plain class
  plain class             9.95    10.10     9.74         1.00x
  dataclass               9.79    10.05     9.67         0.98x
  dataclass(slots)        8.68     8.84     8.58         0.87x
  NamedTuple              8.06     8.24     7.78         0.81x
  TypedDict               8.85     9.10     8.74         0.89x
  tuple                   8.42     8.70     8.29         0.85x

python 3.11.16 (arm64)
  container             median  slowest  fastest  vs plain class
  plain class             3.54     3.59     3.41         1.00x
  dataclass               3.57     3.59     3.53         1.01x
  dataclass(slots)        3.47     3.51     3.32         0.98x
  NamedTuple              6.74     6.91     6.66         1.90x
  TypedDict               7.43     7.89     7.15         2.10x
  tuple                   4.75     4.95     4.56         1.34x

python 3.13.15 (arm64)
  container             median  slowest  fastest  vs plain class
  plain class             3.80     4.14     3.63         1.00x
  dataclass               3.89     3.98     3.73         1.02x
  dataclass(slots)        3.92     4.10     3.65         1.03x
  NamedTuple              7.80     7.99     7.57         2.05x
  TypedDict               7.74     8.42     7.42         2.04x
  tuple                   4.54     4.75     4.46         1.19x

python 3.14.7 (arm64)
  container             median  slowest  fastest  vs plain class
  plain class             4.63     5.16     4.35         1.00x
  dataclass               4.83     5.11     3.96         1.04x
  dataclass(slots)        6.62     6.93     5.32         1.43x
  NamedTuple              8.76     8.85     8.59         1.89x
  TypedDict               7.47     7.85     7.30         1.61x
  tuple                   4.85     4.89     4.66         1.05x

the field-count cliff, bytes per instance, 200,000 instances each

fields                          3       10       20       28       29       30       31       40
dict-based, 3.13.2             96      160      248      320      320     1904     1904     1904
dict-based, 3.10.21           152      192      280      448      448      448      448      448
dict-based, 3.11.16            96      160      248      320      320     1640     1640     1640
dict-based, 3.13.15            96      160      248      320      320     1904     1904     1904
dict-based, 3.14.7             96      160      248      320      328      328     1160     1160
__slots__, 3.13.2              56      112      192      256      264      272      280      352
```

The byte ranking is stable across every interpreter that has inline values, and it spans a
factor of 3.3: 56 bytes for a slotted dataclass, 184 for a `TypedDict`, which is a plain
dictionary at run time and pays a dictionary's header. The read ranking is not stable, and
the container with the fewest bytes is never the one with the cheapest read.

## The instruction that did not get faster

The `read is` column says what happened. Since 3.11 the interpreter rewrites `LOAD_ATTR` in
place once it has seen the shape of the object a few times: a plain attribute becomes
`LOAD_ATTR_INSTANCE_VALUE`, a slot becomes `LOAD_ATTR_SLOT`. A `NamedTuple` field is neither.
It is a `_tuplegetter`, a C descriptor defined in `collections`, and it carries both `__get__`
and `__set__`. That makes it an overriding descriptor. `analyze_descriptor` in
`Python/specialize.c` classifies it as `OVERRIDING`, and the `LOAD_ATTR` specialiser rejects
that case outright with `SPEC_FAIL_ATTR_OVERRIDING_DESCRIPTOR`. The instruction stays generic:

```python
# namedread.py
import dis
from typing import NamedTuple

class Point(NamedTuple):
    x: int
    y: int

def read(p):
    p.y

point = Point(1, 2)
for _ in range(16):          # well past the specialisation threshold
    read(point)

print(type(Point.y).__name__)
print([i.opname for i in dis.get_instructions(read, adaptive=True)
       if i.opname.startswith("LOAD_ATTR")])
```

```
$ python3 namedread.py
_tuplegetter
['LOAD_ATTR']
```

So the read that got cheaper is the one it is measured against. Between 3.10 and 3.11 the
plain instance read fell from 9.95 ns to 3.54, a factor of 2.8. The NamedTuple read went from
8.06 to 6.74 over the same step, which is 16%. The advice was correct when it was written and
the instruction underneath the comparison changed.

## Where the byte column stops being a curve

The byte column looks like eight bytes a field, and for twenty-nine fields it is. The
thirtieth costs 1,584.

That number is the instance dictionary the class stopped sharing. `_PyDict_NewKeysForClass`
in `Objects/dictobject.c` creates one shared keys table per class and sets `dk_usable` to
`SHARED_KEYS_MAX_SIZE`, which `Include/internal/pycore_dict.h` defines as 30. Every new
instance spends one of those before `__init__` runs — `_PyObject_InitInlineValues` decrements
`dk_usable` and then sizes the object's inline values array from what is left — so the first
instance of a class has 29 slots to fill, not 30. `insert_split_key` refuses a name when
`dk_usable` has reached 0, and a refused name sends the instance down
`_PyObject_StoreInstanceAttribute`, which builds a real dictionary for it. That happens to
every instance of the class, not only the first.

The arithmetic checks out: 1,904 bytes is the 320 the object already cost, inline values array
included, plus 1,584 for the dictionary it now also carries. Nothing is freed. The instance
pays for both.

On 3.14 the same table steps at the thirty-first field instead, and lands on 1,160 rather than
1,904. `_PyDict_NewKeysForClass` there takes the class and pre-loads the shared keys from its
`__static_attributes__` — a tuple, present since 3.13, of the names assigned through `self.X`
from any function in the class body. Those inserts happen at class creation, before any
instance has spent anything, so all thirty fit. On 3.10, which has no managed dict and so no
inline values, there is no step at all: 28 fields and 40 fields both cost 448 bytes an
instance.

## Choosing on the axis that holds

Choose on bytes. That column is a layout fact, it reproduces on every interpreter measured
here, and it is the one that multiplies: 40 bytes an instance between a dataclass and a
slotted one is 40 MB at a million rows and 400 KB at ten thousand. Below ten thousand
instances the whole table is worth less than half a megabyte, and the decision is not a
memory decision.

Use `NamedTuple` for immutability and unpacking, not for reads. If a hot path reads fields by
name a million times, indexing the same tuple costs 6.15 ns against 9.52 on the build measured
here. And if a class is dict-based and has grown past twenty-five fields, count them: the
next few are 8 bytes each until one of them is 1,584.

## Where to stop

One machine, macOS on Apple silicon, five interpreters. 3.12 was not measured and this entry
claims nothing about it. Nothing here is a Linux or x86_64 number.

The harness measures two things: bytes attributable to 200,000 instances, and one attribute
read in a `timeit` loop. Construction, hashing, equality and copying are not measured.
Construction matters most of those: the 3.8 table quoted at the top puts a NamedTuple's at
465.3 ns against a dataclass's 846.9, and this entry did not re-run it on anything.

The slot-against-dict read gap is the one row in the read table not to act on. It is 1.10x on
the python.org 3.13.2 universal2 build, 1.03x on the Homebrew 3.13.15 build and 1.43x on
3.14.7, with the 3.14 spread running from 5.32 to 6.93 ns. An earlier entry
[took that measurement apart across nine processes per build](/blog/slots-did-not-get-slower/)
and found the sign is a property of the binary rather than of `__slots__`.
