
The usual answer to “how much memory is this object” is sys.getsizeof. The usual correction
to that answer is that getsizeof is shallow — it does not follow references — so the fix is
to add the instance dictionary to it. Run both readings against the same class and they are
wrong in opposite directions: the shallow one misses half the object, and the corrected one
reports a number that is not a property of the class at all.
The number that moves
Build one instance of a three-field class and ask what its __dict__ costs. Then build
twenty-six and ask again:
getsizeof(row.__dict__), by how many instances of the class exist
1 instances 296
2 instances 288
3 instances 280
10 instances 224
26 instances 96
50 instances 96
200,000 instances 96
Same class, same three attributes, same interpreter, one process. The reported size falls by exactly 8 bytes for every instance built, from 296 down to a floor of 96, and stays there.
The mechanism is in Objects/dictobject.c. For a dictionary with a split table — which is what
an instance dictionary is, since PEP 412 — sizeof_lock_held adds
shared_keys_usable_size(mp->ma_keys) * sizeof(PyObject*) to the dict header, and that helper,
in Include/internal/pycore_dict.h, returns dk_nentries + dk_usable. dk_nentries is the
number of attribute names in the shared keys table and does not move once __init__ has run.
dk_usable is how many more entries the table could take, and it is decremented as instances
are created. So the term shrinks with every instance, and getsizeof reports the size of a
values array sized for a class that has just been defined rather than one that has been used.
Twenty-five instances is not a lot. Any class in a service has passed that count before the first request is served, which means the 296 is not what a real program pays and the 96 is — and that the figure a REPL session hands you is the one taken at the least representative moment in the class’s life. That 296 is the number this archive published for a three-field instance, and it was taken on the first instance the class ever built.
The other direction
The shallow reading has the opposite problem. sys.getsizeof on the instance itself reports
48 bytes, and that is not the object either:
>>> import sys
>>> row = Row(1, 2, "row")
>>> row.__sizeof__(), sys.getsizeof(row)
(16, 48)
object.__sizeof__ returns tp_basicsize, which for this class is 16 bytes: a refcount and a
type pointer. sys.getsizeof adds _PyType_PreHeaderSize on top — the GC header and the two
pointers CPython keeps in front of an object with a managed dict — for 48. Neither term
includes the values array where x, y and label actually live. That array is allocated
with the object, in _PyType_AllocNoTrack, by adding _PyInlineValuesSize(type) to the
allocation request. It is part of the same block; getsizeof has no term for it.
tracemalloc, which counts the block rather than the layout, puts a row at 96 bytes. The
shallow number understates by half; the corrected number, taken on the first instance,
overstates by 3.6x.
Four questions, one set of objects
The harness allocates 200,000 rows twice — once where every row points at the same label string, once where each row owns a string nothing else references — and asks four different questions about them, each in its own process:
$ python3 experiments/objshape/deepsize.py
python 3.13.2 (v3.13.2:4f8bb3947cf, Feb 4 2025, 11:51:10) on darwin
200,000 rows of three fields, one measurement per process
bytes per row shared label distinct labels
sys.getsizeof(row) 48.0 48.0
+ getsizeof(__dict__) 344.0 344.0
recursive walk 374.0 387.0
recursive walk, deduped 144.0 201.0
tracemalloc 96.0 153.0
RSS growth 96.3 160.6
Read the two columns against each other before reading either one down. The payload changed
between them — 200,000 shared references became 200,000 strings of 57 bytes each — and the
first two rows do not move at all. Whatever getsizeof is measuring, it is not what the
process is holding.
The recursive walk is the fix people reach for next, and it fails twice. Walked naively, it counts every shared object once per row, which puts the shared-label case at 374 bytes against a real 96. Walked with an identity set, it stops double-counting, and still reports 144 for a 96-byte row, because it includes the instance dict — a dict that did not exist until the walk asked for it.
tracemalloc and RSS are the two that track the payload, and they agree with each other:
96.0 against 96.3 bytes a row, and 153.0 against 160.6. They are answering different questions
that happen to have close answers here — allocated Python blocks against pages the operating
system is holding — and the gap between them is where the allocator’s own overhead would show
up at larger scales.
The measurement that allocates
An instance with inline values has no dictionary until something asks for one. Asking is what a measurement does:
reading row.__dict__ on every row: 96.0 -> 160.0 bytes/row (12.8 MB to ask the question)
Sixty-four bytes a row, for the dict header each materialised __dict__ needs. A script that
walks a million objects to find out where the memory went adds 64 MB while looking, and the
report it prints includes the bytes it spent writing the report.
Which number answers which question
Use sys.getsizeof for one allocation whose layout you already understand: a bytes object, a
list’s own storage, a single dict you built yourself. It is exact for those and it is cheap.
Use tracemalloc when the question is where a block of code’s memory went. Start it, run the
work, take the delta, and attribute it to lines. That is the tool that answers “the worker grew
by 400 MB” — it counts what was allocated rather than what a layout would suggest.
Use RSS when the question is whether the process fits, because RSS is the number the operating system enforces and no Python-level figure is.
And do not add getsizeof(obj) to getsizeof(obj.__dict__). The first term misses the values
array, the second term depends on how many instances the class has made, and the act of reading
the second one allocates a dictionary that did not exist.
Where to stop
Every figure above is CPython 3.13.2 on macOS on Apple silicon, and the whole table reproduces byte for byte on Homebrew 3.14.7. On 3.11.16 the shape holds and the constants do not: the object header reads 56 rather than 48, and on that build RSS growth is 112.4 and 192.7 bytes a row against tracemalloc’s 96 and 161.
The RSS row is ru_maxrss, which is a peak rather than a current figure, and this harness only
allocates — it never frees, so peak and current coincide. It has not been run on Linux, where
ru_maxrss is reported in kibibytes and the allocator’s arena behaviour differs.
None of it applies to a slotted class, which has no instance dictionary and no values array to misread. What that class buys is the 40 bytes, not the read speed: the read gap usually quoted alongside it did not survive a change of build.
tracemalloc counts allocations that go through CPython’s allocator. A process whose memory is
in NumPy arrays or an extension module’s own malloc will show a gap to RSS that none of this
explains, and the tool to reach for there is not in the standard library.
Frequently asked
Is sys.getsizeof broken?
No. It answers the question it documents — the memory attributed directly to one object — and that question is a poor fit for an instance, whose fields live in an array allocated beside it and whose dict, if it has one, shares a keys table with its siblings. The number is right; the question is the wrong one.
A worker grew by 400 MB and nobody knows where it went. Which tool?
tracemalloc, started before the growth and snapshotted after it. It attributes bytes to the line that allocated them, which is the question being asked. sys.getsizeof answers a question about a single allocation you already have in your hand, and RSS tells you the size of the problem without telling you where it is.
Why does the reported instance-dict size stop falling at 96 bytes?
Because the term that shrinks is (dk_nentries + dk_usable) * 8, and dk_usable stops at 1 on the class measured here rather than reaching 0. Three entries plus one usable slot is 32 bytes on top of the 64-byte dict header, and that is the floor every class of three attributes lands on.
Does any of this change with __slots__?
It removes the moving part. A slotted class has no instance dict to materialise and no keys table to share, so sys.getsizeof on the instance is the whole story and reports 56 bytes, which matches what tracemalloc measures. The trade is the one the earlier entry measured: 40 bytes an instance, and no dict to inspect at run time.
Where this came from
Run it yourself — every figure above came out of these:
- objshape/deepsize.py — Recursive object size against what getsizeof reports
Read, rather than assumed:


