getsizeof Reports 296 Bytes, Then 96, for the Same Class

sys.getsizeof reports 296 bytes for a class's first instance dict and 96 once twenty-six exist. The instance itself costs 96 bytes and getsizeof says 48.

Elliot Sayer6 min read

The usual answer to “how much memory is this object” is sys.getsizeof. The usual correction to that answer is that getsizeof is shallow — it does not follow references — so the fix is to add the instance dictionary to it. Run both readings against the same class and they are wrong in opposite directions: the shallow one misses half the object, and the corrected one reports a number that is not a property of the class at all.

The number that moves

Build one instance of a three-field class and ask what its __dict__ costs. Then build twenty-six and ask again:

getsizeof(row.__dict__), by how many instances of the class exist
        1 instances     296
        2 instances     288
        3 instances     280
       10 instances     224
       26 instances      96
       50 instances      96
  200,000 instances      96

Same class, same three attributes, same interpreter, one process. The reported size falls by exactly 8 bytes for every instance built, from 296 down to a floor of 96, and stays there.

The mechanism is in Objects/dictobject.c. For a dictionary with a split table — which is what an instance dictionary is, since PEP 412 — sizeof_lock_held adds shared_keys_usable_size(mp->ma_keys) * sizeof(PyObject*) to the dict header, and that helper, in Include/internal/pycore_dict.h, returns dk_nentries + dk_usable. dk_nentries is the number of attribute names in the shared keys table and does not move once __init__ has run. dk_usable is how many more entries the table could take, and it is decremented as instances are created. So the term shrinks with every instance, and getsizeof reports the size of a values array sized for a class that has just been defined rather than one that has been used.

Twenty-five instances is not a lot. Any class in a service has passed that count before the first request is served, which means the 296 is not what a real program pays and the 96 is — and that the figure a REPL session hands you is the one taken at the least representative moment in the class’s life. That 296 is the number this archive published for a three-field instance, and it was taken on the first instance the class ever built.

The other direction

The shallow reading has the opposite problem. sys.getsizeof on the instance itself reports 48 bytes, and that is not the object either:

>>> import sys
>>> row = Row(1, 2, "row")
>>> row.__sizeof__(), sys.getsizeof(row)
(16, 48)

object.__sizeof__ returns tp_basicsize, which for this class is 16 bytes: a refcount and a type pointer. sys.getsizeof adds _PyType_PreHeaderSize on top — the GC header and the two pointers CPython keeps in front of an object with a managed dict — for 48. Neither term includes the values array where x, y and label actually live. That array is allocated with the object, in _PyType_AllocNoTrack, by adding _PyInlineValuesSize(type) to the allocation request. It is part of the same block; getsizeof has no term for it.

tracemalloc, which counts the block rather than the layout, puts a row at 96 bytes. The shallow number understates by half; the corrected number, taken on the first instance, overstates by 3.6x.

Four questions, one set of objects

The harness allocates 200,000 rows twice — once where every row points at the same label string, once where each row owns a string nothing else references — and asks four different questions about them, each in its own process:

$ python3 experiments/objshape/deepsize.py
python 3.13.2 (v3.13.2:4f8bb3947cf, Feb  4 2025, 11:51:10) on darwin
200,000 rows of three fields, one measurement per process

bytes per row                shared label   distinct labels
sys.getsizeof(row)                   48.0              48.0
  + getsizeof(__dict__)             344.0             344.0
recursive walk                      374.0             387.0
recursive walk, deduped             144.0             201.0
tracemalloc                          96.0             153.0
RSS growth                           96.3             160.6

Read the two columns against each other before reading either one down. The payload changed between them — 200,000 shared references became 200,000 strings of 57 bytes each — and the first two rows do not move at all. Whatever getsizeof is measuring, it is not what the process is holding.

The recursive walk is the fix people reach for next, and it fails twice. Walked naively, it counts every shared object once per row, which puts the shared-label case at 374 bytes against a real 96. Walked with an identity set, it stops double-counting, and still reports 144 for a 96-byte row, because it includes the instance dict — a dict that did not exist until the walk asked for it.

tracemalloc and RSS are the two that track the payload, and they agree with each other: 96.0 against 96.3 bytes a row, and 153.0 against 160.6. They are answering different questions that happen to have close answers here — allocated Python blocks against pages the operating system is holding — and the gap between them is where the allocator’s own overhead would show up at larger scales.

The measurement that allocates

An instance with inline values has no dictionary until something asks for one. Asking is what a measurement does:

reading row.__dict__ on every row: 96.0 -> 160.0 bytes/row (12.8 MB to ask the question)

Sixty-four bytes a row, for the dict header each materialised __dict__ needs. A script that walks a million objects to find out where the memory went adds 64 MB while looking, and the report it prints includes the bytes it spent writing the report.

Which number answers which question

Use sys.getsizeof for one allocation whose layout you already understand: a bytes object, a list’s own storage, a single dict you built yourself. It is exact for those and it is cheap.

Use tracemalloc when the question is where a block of code’s memory went. Start it, run the work, take the delta, and attribute it to lines. That is the tool that answers “the worker grew by 400 MB” — it counts what was allocated rather than what a layout would suggest.

Use RSS when the question is whether the process fits, because RSS is the number the operating system enforces and no Python-level figure is.

And do not add getsizeof(obj) to getsizeof(obj.__dict__). The first term misses the values array, the second term depends on how many instances the class has made, and the act of reading the second one allocates a dictionary that did not exist.

Where to stop

Every figure above is CPython 3.13.2 on macOS on Apple silicon, and the whole table reproduces byte for byte on Homebrew 3.14.7. On 3.11.16 the shape holds and the constants do not: the object header reads 56 rather than 48, and on that build RSS growth is 112.4 and 192.7 bytes a row against tracemalloc’s 96 and 161.

The RSS row is ru_maxrss, which is a peak rather than a current figure, and this harness only allocates — it never frees, so peak and current coincide. It has not been run on Linux, where ru_maxrss is reported in kibibytes and the allocator’s arena behaviour differs.

None of it applies to a slotted class, which has no instance dictionary and no values array to misread. What that class buys is the 40 bytes, not the read speed: the read gap usually quoted alongside it did not survive a change of build.

tracemalloc counts allocations that go through CPython’s allocator. A process whose memory is in NumPy arrays or an extension module’s own malloc will show a gap to RSS that none of this explains, and the tool to reach for there is not in the standard library.

Frequently asked

Is sys.getsizeof broken?

No. It answers the question it documents — the memory attributed directly to one object — and that question is a poor fit for an instance, whose fields live in an array allocated beside it and whose dict, if it has one, shares a keys table with its siblings. The number is right; the question is the wrong one.

A worker grew by 400 MB and nobody knows where it went. Which tool?

tracemalloc, started before the growth and snapshotted after it. It attributes bytes to the line that allocated them, which is the question being asked. sys.getsizeof answers a question about a single allocation you already have in your hand, and RSS tells you the size of the problem without telling you where it is.

Why does the reported instance-dict size stop falling at 96 bytes?

Because the term that shrinks is (dk_nentries + dk_usable) * 8, and dk_usable stops at 1 on the class measured here rather than reaching 0. Three entries plus one usable slot is 32 bytes on top of the 64-byte dict header, and that is the floor every class of three attributes lands on.

Does any of this change with __slots__?

It removes the moving part. A slotted class has no instance dict to materialise and no keys table to share, so sys.getsizeof on the instance is the whole story and reports 56 bytes, which matches what tracemalloc measures. The trade is the one the earlier entry measured: 40 bytes an instance, and no dict to inspect at run time.

Where this came from

Run it yourself — every figure above came out of these:

  1. objshape/deepsize.py — Recursive object size against what getsizeof reports

Read, rather than assumed:

  1. PEP 412 — Key-Sharing Dictionary
  2. Objects/dictobject.c at CPython v3.13.2
  3. Include/internal/pycore_dict.h at CPython v3.13.2
Share

Arrow keys to move, Enter to open.