A NamedTuple Read Cost 0.81x a Plain One, Then 1.9x

On CPython 3.10 a NamedTuple attribute read cost 0.81x a read through an instance dictionary. On 3.11 it costs 1.90x. Nothing about NamedTuple changed.

Elliot Sayer5 min read

Two tables decide most of these choices. A widely cited one puts a NamedTuple’s attribute access at 43.2 ns against a dataclass’s 49.7 and its instance at 56 bytes against 152 — smaller and faster on both counts, measured on CPython 3.8. A second reports that slots=True cuts memory by a factor of eight and makes attribute access a quarter faster, measured on 3.11.

Measure the same six containers on 3.10 and the NamedTuple read is the cheapest of them: 8.06 ns against 9.95 for a read through an instance dictionary, a ratio of 0.81. Measure it on 3.11 and the read costs 6.74 ns against 3.54 — 1.90x, and by 3.13 it is the dearest read in the table. Nothing about NamedTuple changed between those two numbers.

What the six containers cost

A plain class, a dataclass, a dataclass(slots=True), a NamedTuple, a TypedDict and a bare tuple, all holding the same three fields. The bytes come from a tracemalloc delta over 200,000 instances rather than from getsizeof, which reports a figure that moves with how many instances the class has made; the reads are the median of nine processes per interpreter, because the gap between two of these containers is smaller than the gap between two runs of the same one:

$ python3 experiments/objshape/containers.py /opt/homebrew/bin/python3.10 \
    /opt/homebrew/bin/python3.11 /opt/homebrew/bin/python3.13 /opt/homebrew/bin/python3.14
python 3.13.2 on darwin (universal2)
200,000 instances of three fields each

container             bytes/inst   read is
plain class                 96.0   LOAD_ATTR_INSTANCE_VALUE
dataclass                   96.0   LOAD_ATTR_INSTANCE_VALUE
dataclass(slots)            56.0   LOAD_ATTR_SLOT
NamedTuple                  72.0   LOAD_ATTR
TypedDict                  184.0   obj['y']
tuple                       64.0   obj[1]

ns per read, 9 processes per interpreter, best of 7 x 1,000,000 in each

python 3.13.2 (universal2)
  container             median  slowest  fastest  vs plain class
  plain class             5.84     6.06     5.59         1.00x
  dataclass               5.76     6.02     5.64         0.99x
  dataclass(slots)        6.44     6.59     6.24         1.10x
  NamedTuple              9.52     9.63     9.42         1.63x
  TypedDict               6.48     6.87     6.25         1.11x
  tuple                   6.15     6.36     5.97         1.05x

python 3.10.21 (arm64)
  container             median  slowest  fastest  vs plain class
  plain class             9.95    10.10     9.74         1.00x
  dataclass               9.79    10.05     9.67         0.98x
  dataclass(slots)        8.68     8.84     8.58         0.87x
  NamedTuple              8.06     8.24     7.78         0.81x
  TypedDict               8.85     9.10     8.74         0.89x
  tuple                   8.42     8.70     8.29         0.85x

python 3.11.16 (arm64)
  container             median  slowest  fastest  vs plain class
  plain class             3.54     3.59     3.41         1.00x
  dataclass               3.57     3.59     3.53         1.01x
  dataclass(slots)        3.47     3.51     3.32         0.98x
  NamedTuple              6.74     6.91     6.66         1.90x
  TypedDict               7.43     7.89     7.15         2.10x
  tuple                   4.75     4.95     4.56         1.34x

python 3.13.15 (arm64)
  container             median  slowest  fastest  vs plain class
  plain class             3.80     4.14     3.63         1.00x
  dataclass               3.89     3.98     3.73         1.02x
  dataclass(slots)        3.92     4.10     3.65         1.03x
  NamedTuple              7.80     7.99     7.57         2.05x
  TypedDict               7.74     8.42     7.42         2.04x
  tuple                   4.54     4.75     4.46         1.19x

python 3.14.7 (arm64)
  container             median  slowest  fastest  vs plain class
  plain class             4.63     5.16     4.35         1.00x
  dataclass               4.83     5.11     3.96         1.04x
  dataclass(slots)        6.62     6.93     5.32         1.43x
  NamedTuple              8.76     8.85     8.59         1.89x
  TypedDict               7.47     7.85     7.30         1.61x
  tuple                   4.85     4.89     4.66         1.05x

the field-count cliff, bytes per instance, 200,000 instances each

fields                          3       10       20       28       29       30       31       40
dict-based, 3.13.2             96      160      248      320      320     1904     1904     1904
dict-based, 3.10.21           152      192      280      448      448      448      448      448
dict-based, 3.11.16            96      160      248      320      320     1640     1640     1640
dict-based, 3.13.15            96      160      248      320      320     1904     1904     1904
dict-based, 3.14.7             96      160      248      320      328      328     1160     1160
__slots__, 3.13.2              56      112      192      256      264      272      280      352

The byte ranking is stable across every interpreter that has inline values, and it spans a factor of 3.3: 56 bytes for a slotted dataclass, 184 for a TypedDict, which is a plain dictionary at run time and pays a dictionary’s header. The read ranking is not stable, and the container with the fewest bytes is never the one with the cheapest read.

The instruction that did not get faster

The read is column says what happened. Since 3.11 the interpreter rewrites LOAD_ATTR in place once it has seen the shape of the object a few times: a plain attribute becomes LOAD_ATTR_INSTANCE_VALUE, a slot becomes LOAD_ATTR_SLOT. A NamedTuple field is neither. It is a _tuplegetter, a C descriptor defined in collections, and it carries both __get__ and __set__. That makes it an overriding descriptor. analyze_descriptor in Python/specialize.c classifies it as OVERRIDING, and the LOAD_ATTR specialiser rejects that case outright with SPEC_FAIL_ATTR_OVERRIDING_DESCRIPTOR. The instruction stays generic:

# namedread.py
import dis
from typing import NamedTuple

class Point(NamedTuple):
    x: int
    y: int

def read(p):
    p.y

point = Point(1, 2)
for _ in range(16):          # well past the specialisation threshold
    read(point)

print(type(Point.y).__name__)
print([i.opname for i in dis.get_instructions(read, adaptive=True)
       if i.opname.startswith("LOAD_ATTR")])
$ python3 namedread.py
_tuplegetter
['LOAD_ATTR']

So the read that got cheaper is the one it is measured against. Between 3.10 and 3.11 the plain instance read fell from 9.95 ns to 3.54, a factor of 2.8. The NamedTuple read went from 8.06 to 6.74 over the same step, which is 16%. The advice was correct when it was written and the instruction underneath the comparison changed.

Where the byte column stops being a curve

The byte column looks like eight bytes a field, and for twenty-nine fields it is. The thirtieth costs 1,584.

That number is the instance dictionary the class stopped sharing. _PyDict_NewKeysForClass in Objects/dictobject.c creates one shared keys table per class and sets dk_usable to SHARED_KEYS_MAX_SIZE, which Include/internal/pycore_dict.h defines as 30. Every new instance spends one of those before __init__ runs — _PyObject_InitInlineValues decrements dk_usable and then sizes the object’s inline values array from what is left — so the first instance of a class has 29 slots to fill, not 30. insert_split_key refuses a name when dk_usable has reached 0, and a refused name sends the instance down _PyObject_StoreInstanceAttribute, which builds a real dictionary for it. That happens to every instance of the class, not only the first.

The arithmetic checks out: 1,904 bytes is the 320 the object already cost, inline values array included, plus 1,584 for the dictionary it now also carries. Nothing is freed. The instance pays for both.

On 3.14 the same table steps at the thirty-first field instead, and lands on 1,160 rather than 1,904. _PyDict_NewKeysForClass there takes the class and pre-loads the shared keys from its __static_attributes__ — a tuple, present since 3.13, of the names assigned through self.X from any function in the class body. Those inserts happen at class creation, before any instance has spent anything, so all thirty fit. On 3.10, which has no managed dict and so no inline values, there is no step at all: 28 fields and 40 fields both cost 448 bytes an instance.

Choosing on the axis that holds

Choose on bytes. That column is a layout fact, it reproduces on every interpreter measured here, and it is the one that multiplies: 40 bytes an instance between a dataclass and a slotted one is 40 MB at a million rows and 400 KB at ten thousand. Below ten thousand instances the whole table is worth less than half a megabyte, and the decision is not a memory decision.

Use NamedTuple for immutability and unpacking, not for reads. If a hot path reads fields by name a million times, indexing the same tuple costs 6.15 ns against 9.52 on the build measured here. And if a class is dict-based and has grown past twenty-five fields, count them: the next few are 8 bytes each until one of them is 1,584.

Where to stop

One machine, macOS on Apple silicon, five interpreters. 3.12 was not measured and this entry claims nothing about it. Nothing here is a Linux or x86_64 number.

The harness measures two things: bytes attributable to 200,000 instances, and one attribute read in a timeit loop. Construction, hashing, equality and copying are not measured. Construction matters most of those: the 3.8 table quoted at the top puts a NamedTuple’s at 465.3 ns against a dataclass’s 846.9, and this entry did not re-run it on anything.

The slot-against-dict read gap is the one row in the read table not to act on. It is 1.10x on the python.org 3.13.2 universal2 build, 1.03x on the Homebrew 3.13.15 build and 1.43x on 3.14.7, with the 3.14 spread running from 5.32 to 6.93 ns. An earlier entry took that measurement apart across nine processes per build and found the sign is a property of the binary rather than of __slots__.

Frequently asked

Should I stop using NamedTuple?

No — stop choosing it for read speed. It is 72 bytes against 96 for a dataclass, it is immutable, and it unpacks. Those are the reasons that survived the specialising interpreter. The attribute-read advantage did not: it was 0.81x on 3.10 and is 1.90x on 3.11.

Why does the NamedTuple read not specialise?

A NamedTuple field is a _tuplegetter, a C descriptor from the collections module. The specialiser has variants for instance values, slots and properties; a generic descriptor is not one of them, so the instruction stays LOAD_ATTR and pays for the descriptor protocol on every read. Pass a warmed function to dis.dis with adaptive=True and the disassembler shows which one you got.

Is the thirty-field limit documented anywhere?

Not as a documented limit. It falls out of SHARED_KEYS_MAX_SIZE in Include/internal/pycore_dict.h, which is 30, and out of the order in which the value is spent: the first instance of a class takes one before __init__ runs. PEP 412 describes key sharing and does not name a ceiling.

Does indexing a NamedTuple avoid the cost?

Yes, and that is the point of the comparison with a bare tuple in the same table. Reading by index costs 6.15 ns on the build measured here against 9.52 ns by name. If a hot loop reads the same field a million times, the name is what you are paying for.

Which number should decide the container?

The byte column. It is a layout fact, it reproduces on every interpreter measured here, and it is what multiplies by a million rows. The read column moved by a factor of three between 3.10 and 3.11 and, for the slot-against-dict pair, changes sign between two builds of the same version.

Where this came from

Run it yourself — every figure above came out of these:

  1. objshape/containers.py — Six container shapes on two axes, across five interpreters

Read, rather than assumed:

  1. PEP 412 — Key-Sharing Dictionary
  2. Objects/dictobject.c at CPython v3.13.2
  3. Include/internal/pycore_dict.h at CPython v3.13.2
  4. Python/specialize.c at CPython v3.13.2
Share

Arrow keys to move, Enter to open.