# getsizeof Reports 296 Bytes, Then 96, for the Same Class

> sys.getsizeof reports 296 bytes for a class's first instance dict and 96 once twenty-six exist. The instance itself costs 96 bytes and getsizeof says 48.

- Published: 2026-09-06
- Category: runtime
- Tags: memory, object-layout
- Sources: https://github.com/CognatePress/erperiments.pydepth.com/blob/main/objshape/deepsize.py, https://peps.python.org/pep-412/, https://github.com/python/cpython/blob/v3.13.2/Objects/dictobject.c, https://github.com/python/cpython/blob/v3.13.2/Include/internal/pycore_dict.h
- Source: https://pydepth.com/blog/what-getsizeof-does-not-count/
- Language: en-US
- Author: Elliot Sayer

---
The usual answer to "how much memory is this object" is `sys.getsizeof`. The usual correction
to that answer is that `getsizeof` is shallow — it does not follow references — so the fix is
to add the instance dictionary to it. Run both readings against the same class and they are
wrong in opposite directions: the shallow one misses half the object, and the corrected one
reports a number that is not a property of the class at all.

## The number that moves

Build one instance of a three-field class and ask what its `__dict__` costs. Then build
twenty-six and ask again:

```
getsizeof(row.__dict__), by how many instances of the class exist
        1 instances     296
        2 instances     288
        3 instances     280
       10 instances     224
       26 instances      96
       50 instances      96
  200,000 instances      96
```

Same class, same three attributes, same interpreter, one process. The reported size falls by
exactly 8 bytes for every instance built, from 296 down to a floor of 96, and stays there.

The mechanism is in `Objects/dictobject.c`. For a dictionary with a split table — which is what
an instance dictionary is, since PEP 412 — `sizeof_lock_held` adds
`shared_keys_usable_size(mp->ma_keys) * sizeof(PyObject*)` to the dict header, and that helper,
in `Include/internal/pycore_dict.h`, returns `dk_nentries + dk_usable`. `dk_nentries` is the
number of attribute names in the shared keys table and does not move once `__init__` has run.
`dk_usable` is how many more entries the table could take, and it is decremented as instances
are created. So the term shrinks with every instance, and `getsizeof` reports the size of a
values array sized for a class that has just been defined rather than one that has been used.

Twenty-five instances is not a lot. Any class in a service has passed that count before the
first request is served, which means the 296 is not what a real program pays and the 96 is —
and that the figure a REPL session hands you is the one taken at the least representative
moment in the class's life. That 296 is the number
[this archive published for a three-field instance](/blog/the-shape-of-a-python-object/), and it
was taken on the first instance the class ever built.

## The other direction

The shallow reading has the opposite problem. `sys.getsizeof` on the instance itself reports
48 bytes, and that is not the object either:

```python
>>> import sys
>>> row = Row(1, 2, "row")
>>> row.__sizeof__(), sys.getsizeof(row)
(16, 48)
```

`object.__sizeof__` returns `tp_basicsize`, which for this class is 16 bytes: a refcount and a
type pointer. `sys.getsizeof` adds `_PyType_PreHeaderSize` on top — the GC header and the two
pointers CPython keeps in front of an object with a managed dict — for 48. Neither term
includes the values array where `x`, `y` and `label` actually live. That array is allocated
with the object, in `_PyType_AllocNoTrack`, by adding `_PyInlineValuesSize(type)` to the
allocation request. It is part of the same block; `getsizeof` has no term for it.

`tracemalloc`, which counts the block rather than the layout, puts a row at 96 bytes. The
shallow number understates by half; the corrected number, taken on the first instance,
overstates by 3.6x.

## Four questions, one set of objects

The harness allocates 200,000 rows twice — once where every row points at the same label
string, once where each row owns a string nothing else references — and asks four different
questions about them, each in its own process:

```
$ python3 experiments/objshape/deepsize.py
python 3.13.2 (v3.13.2:4f8bb3947cf, Feb  4 2025, 11:51:10) on darwin
200,000 rows of three fields, one measurement per process

bytes per row                shared label   distinct labels
sys.getsizeof(row)                   48.0              48.0
  + getsizeof(__dict__)             344.0             344.0
recursive walk                      374.0             387.0
recursive walk, deduped             144.0             201.0
tracemalloc                          96.0             153.0
RSS growth                           96.3             160.6
```

Read the two columns against each other before reading either one down. The payload changed
between them — 200,000 shared references became 200,000 strings of 57 bytes each — and the
first two rows do not move at all. Whatever `getsizeof` is measuring, it is not what the
process is holding.

The recursive walk is the fix people reach for next, and it fails twice. Walked naively, it
counts every shared object once per row, which puts the shared-label case at 374 bytes against
a real 96. Walked with an identity set, it stops double-counting, and still reports 144 for a
96-byte row, because it includes the instance dict — a dict that did not exist until the walk
asked for it.

`tracemalloc` and RSS are the two that track the payload, and they agree with each other:
96.0 against 96.3 bytes a row, and 153.0 against 160.6. They are answering different questions
that happen to have close answers here — allocated Python blocks against pages the operating
system is holding — and the gap between them is where the allocator's own overhead would show
up at larger scales.

## The measurement that allocates

An instance with inline values has no dictionary until something asks for one. Asking is what a
measurement does:

```
reading row.__dict__ on every row: 96.0 -> 160.0 bytes/row (12.8 MB to ask the question)
```

Sixty-four bytes a row, for the dict header each materialised `__dict__` needs. A script that
walks a million objects to find out where the memory went adds 64 MB while looking, and the
report it prints includes the bytes it spent writing the report.

## Which number answers which question

Use `sys.getsizeof` for one allocation whose layout you already understand: a bytes object, a
list's own storage, a single dict you built yourself. It is exact for those and it is cheap.

Use `tracemalloc` when the question is where a block of code's memory went. Start it, run the
work, take the delta, and attribute it to lines. That is the tool that answers "the worker grew
by 400 MB" — it counts what was allocated rather than what a layout would suggest.

Use RSS when the question is whether the process fits, because RSS is the number the operating
system enforces and no Python-level figure is.

And do not add `getsizeof(obj)` to `getsizeof(obj.__dict__)`. The first term misses the values
array, the second term depends on how many instances the class has made, and the act of reading
the second one allocates a dictionary that did not exist.

## Where to stop

Every figure above is CPython 3.13.2 on macOS on Apple silicon, and the whole table reproduces
byte for byte on Homebrew 3.14.7. On 3.11.16 the shape holds and the constants do not: the
object header reads 56 rather than 48, and on that build RSS growth is 112.4 and 192.7 bytes a row
against tracemalloc's 96 and 161.

The RSS row is `ru_maxrss`, which is a peak rather than a current figure, and this harness only
allocates — it never frees, so peak and current coincide. It has not been run on Linux, where
`ru_maxrss` is reported in kibibytes and the allocator's arena behaviour differs.

None of it applies to a slotted class, which has no instance dictionary and no values array to
misread. What that class buys is the 40 bytes, not the read speed: the read gap usually quoted
alongside it [did not survive a change of build](/blog/slots-did-not-get-slower/).

`tracemalloc` counts allocations that go through CPython's allocator. A process whose memory is
in NumPy arrays or an extension module's own `malloc` will show a gap to RSS that none of this
explains, and the tool to reach for there is not in the standard library.
