gigl.scripts.verify_glt_patches#

Assert that the patches in gigl/scripts/patches/ are live in the INSTALLED graphlearn_torch.

Run by install_glt.sh immediately after uv pip install dist/*.whl, and exits non-zero if any patched behaviour is missing.

WHY THIS EXISTS

install_glt.sh already fails loudly if a patch file does not APPLY. That is a weaker guarantee than it looks: applying is a property of the source tree, while what ships is a compiled .so. A patch can apply to a file the build then excludes, a stale build directory can be reused, or a wheel can be installed from somewhere other than the tree that was patched – and every one of those failures is silent, because the unpatched code paths work. They just work at 3x the memory (the bitmap count) or reject the int32 topology the trainer is about to build (int32 support), several hours into a 16-GPU job.

The unit tests in tests/unit/utils/glt_int32_indices_test.py cover the same ground in far more detail, but they SKIP on an unpatched wheel by design – GiGL’s CI runs against the released graphlearn_torch, where failing would be wrong. So the test suite cannot be the gate for the image, and this script cannot be the detailed test. Both exist, deliberately.

Deliberately dependency-free (torch and graphlearn_torch only): it runs inside the image build before the rest of the repo’s test dependencies are necessarily importable.

TODO (dsaini2-sc): once a Snapchat-maintained GLT fork carries these changes as commits, this

script outlives the patch files – it verifies the installed wheel, wherever it was built from – but its docstrings and failure messages should be reworded away from “patches”.

Usage:

python gigl/scripts/verify_glt_patches.py

Functions#

build_cpu_csr_graph(indptr, indices)

Build a CPU Graph over a ready-made CSR, bypassing Topology.__init__.

main()

verify_bitmap_col_count()

The hardened distinct-count: exact on valid input, loud on invalid input.

verify_int32_indices()

int32 columns must be accepted AND sample identically to int64.

verify_queue_teardown()

The teardown-unpin patch must be COMPILED IN, and teardown must stay leak-free.

Module Contents#

gigl.scripts.verify_glt_patches.build_cpu_csr_graph(indptr, indices)[source]#

Build a CPU Graph over a ready-made CSR, bypassing Topology.__init__.

Populating the attributes directly keeps the fixture free of the arange(num_edges) edge ids Topology.__init__ would attach, so the CSR reaches the compiled extension exactly as given – including a dtype upstream’s constructor would upcast.

Lives here rather than in the test suite because this module must stay importable during an image build, before the test dependencies are installed.

Parameters:
  • indptr (torch.Tensor)

  • indices (torch.Tensor)

Return type:

graphlearn_torch.data.Graph

gigl.scripts.verify_glt_patches.main()[source]#
Return type:

int

gigl.scripts.verify_glt_patches.verify_bitmap_col_count()[source]#

The hardened distinct-count: exact on valid input, loud on invalid input.

Return type:

bool

gigl.scripts.verify_glt_patches.verify_int32_indices()[source]#

int32 columns must be accepted AND sample identically to int64.

Return type:

bool

gigl.scripts.verify_glt_patches.verify_queue_teardown()[source]#

The teardown-unpin patch must be COMPILED IN, and teardown must stay leak-free.

The runtime payload – the deleter calling cudaHostUnregister for a mapping this process pinned – cannot be exercised without a CUDA device, and image builds have none. Worse, pin_memory() on a driverless host does not raise: GLT’s CUDACheckError calls exit(EXIT_FAILURE), which would kill this verifier and the build with it. So this check NEVER pins unless a device is actually available. Presence of the patched BINARY is proven by a compiled capability marker instead (supports_unpin_on_teardown, added by the same patch to the same .so as the deleter); the unregister BEHAVIOUR is proven by the GPU canary probe (pin -> destroy -> repin cycles), which fails on an unpatched wheel.

What runs everywhere: the deleter still detaches and removes the SysV segment across repeated cycles, and the zero-copy ShmData path still holds the mapping open while a dequeued message is in flight.

Return type:

bool