Skip to content
Preprint

TorchMorph: CUDA-accelerated Morphological Transforms

Aug 2026 · 0 citations · 24 references
Computer Science

Abstract

Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e. scipy.ndimage, is CPU-only, single-array, and therefore unusable inside a GPU training loop without an expensive device-to-host round trip. GPU vision libraries built on PyTorch cover a narrow subset of these operators, typically restricted to two spatial dimensions and flat structuring elements. We present TorchMorph, a lightweight PyTorch extension that closes this gap. TorchMorph exposes 22 public operators covering binary morphology, greyscale morphology, exact and approximate distance transforms, and entropy-regularised optimal transport, all implemented as fused CUDA kernels that operate directly on (B, C, Spatial...) CUDA tensors with up to eight spatial dimensions. The API deliberately mirrors scipy.ndimage argument-for-argument, including border modes, structuring-element origins and pre-allocated outputs, so that existing pipelines port with a change of import. We describe the layered architecture and the kernel designs behind each operator family. Against single-threaded CPU references, batched execution reaches up to 1.1e3 times the throughput of scipy.ndimage on greyscale morphology and up to 350x on exact Euclidean distance transforms, while the Sinkhorn solver runs up to 42x faster than POT. Binary and chamfer operators reproduce their SciPy counterparts exactly, and every float-valued operator agrees with the CPU reference to within 1.8e-6 absolute error. TorchMorph is released under the MIT licence at https://intcomp.github.io/tm.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer

Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conventional accelerator...

Hao-Ran Jin, Kang-Qi Zhang, Ji-Rong Yang et al. · 0 citations
Preprint Sep 2026

CuACD: A Fully GPU-Resident Approximate Convex Decomposition

Approximate convex decomposition (ACD) converts triangle meshes into small sets of convex parts and is a standard preprocessing step for physics simulation, collision detection, and large-scale robot learning. The majority of modern ACD methods produce high-quality decompositions through an expensive search over candid...

Ruo-Xi Shi, Xin-Yue Wei, Fan-Bo Xiang et al. · 0 citations
Preprint Sep 2026

Taming Bitwise Behavior in GPU Kernels with Tensor Core: Black-Box Reconstruction, Compiler Enforcement, and Static Verification

Determinism and numerical reproducibility are increasingly required of GPU kernels in machine learning systems, yet deterministic implementations of the same kernel can still differ bit for bit. Floating-point reduction order is the primary cause, alongside partial-sum precision, fused multiply-add operations, and roun...

Zi-Teng Yang, Nicholas J. Riasanovsky, Warren Deng et al. · 1 citation
Preprint Sep 2026

RGB Input Pipelines: Throughput, GPU Memory, and Transformation Coverage

An image-augmentation pipeline must deliver a complete batch before a model can use it. We compare seven input paths from five libraries, starting with RGB JPEG files and ending with a synchronized CUDA float16 batch. We manually matched transformation recipes and parameters across libraries to make the workloads as co...

V. Iglovikov · 0 citations
Preprint Aug 2026

FastKron: Efficient Quantization with Kronecker-Factored Hessians

We accelerate a family of algorithms for neural network quantization which utilize a two-sided version of the GPTQ/LDLQ algorithm. Standard GPTQ-style adaptive rounding uses one-sided correlation information derived from input activations. A natural two-sided extension can additionally capture correlations across outpu...

Johann Birnick, R. Saab · 1 citation
Preprint Aug 2026

BaKron: Efficient Quantization with Kronecker-Factored Hessians

BaKron is an efficient solver that combines anti-diagonal parallelism with a recursive divide-and-conquer construction that matches the cubic scaling of GPTQ while exploiting richer curvature information.

Johann Birnick, R. Saab · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.