Dynamic Aware Biopsy Needle Identification in Ultrasound Images Using Temporal Prior Guided U-Net Cross Transformer With Limited Training Data.
Authors
Affiliations (4)
Affiliations (4)
- Department of Intelligent Electronics and Computer Engineering, Chonnam National University, Gwangju, Republic of Korea.
- Department of Internal Medicine, Chonnam National University Hospital, Gwangju, Republic of Korea; Department of Internal Medicine, Chonnam National University Medical School, Gwangju, Republic of Korea.
- Department of Anesthesiology and Pain Medicine, Chosun University Hospital, Gwangju, Republic of Korea; Department of Anesthesiology and Pain Medicine, Chosun University Medical School, Gwangju, Republic of Korea.
- Department of Intelligent Electronics and Computer Engineering, Chonnam National University, Gwangju, Republic of Korea. Electronic address: [email protected].
Abstract
Ultrasound-guided needle placement has been commonly used for minimally invasive clinical procedures, including biopsy, regional anesthesia and localized drug administration. This study aimed to enhance existing deep learning frameworks by incorporating a classical background subtraction, which enriches the inductive bias and thereby enables more reliable needle detection even when the available training dataset is small. Although deep learning methods such as U-Net and its derivatives have substantially advanced needle localization performance, they remain constrained by a limited receptive field that prevents effective modeling of long-range spatial dependencies. Although Vision Transformer overcomes this limitation through self-attention mechanisms, it demands large-scale training data and tends to sacrifice fine-grained spatial locality. To overcome these limitations, we propose U-Net Cross Transformer (UXFormer), a dynamic-aware hybrid architecture that combines classical background subtraction with an integrated U-Net × Vision Transformer fusion framework. The model comprises three key components: (i) null subspace-based extraction of temporal prior information, (ii) a temporal-to-spatial cross-attention module during encoding and (iii) a global-to-local cross-convolutional block attention module during decoding, enabling continuous bidirectional communication between localized temporal dynamics and globally contextualized spatial representations. Experimental results demonstrate that the proposed method outperforms multiple competing approaches, achieving significant improvements: a 14.2% increase in Jaccard index, a 9.0% increase in Dice score, 8.5% increase in recall, 5.5% increase in precision, and 63.5% increase in the 95th percentile Hausdorff distance, thus leading to a 44.0% improvement in tip position error and a 17.7% improvement in trajectory angle error, even under varying needle visibility conditions. By explicitly bridging classical signal processing and deep learning within a unified framework, this work strengthens inductive bias while substantially reducing the massive data requirements inherent in transformer-based models, yielding superior detection performance across diverse needle-tissue interaction conditions, even when training data are limited in volume.