Conceptual

nnY-Net: Swin-NeXt with Metadata Cross-Attention for 3D Medical Image Segmentation

A 3D medical image segmentation architecture pairing a Swin Transformer encoder with a ConvNeXt decoder (Swin-NeXt) and adding a bottleneck cross-attention module (the Y) that uses the encoder's lowest-level feature map as Key/Value and patient metadata (pathology, treatment) as Query, wrapped in the self-configuring nnU-Net framework and trained with a combined DiceFocalCELoss for imbalanced voxel classification.