{
  "id": 653876,
  "title": "CSA-Net [LB-0.472] - Cross attention across layers",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/653876",
  "author_name": "",
  "post_date": "2025-12-07T06:00:58.616099900Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>So this discussion is about the model I am working with. I was trying to find architectures that have some form of attention along an axis, something similar to an LSTM where information flows linearly. That’s when I found this architecture called CSA-Net ( <a href=\"https://arxiv.org/abs/2405.00130\" target=\"_blank\">https://arxiv.org/abs/2405.00130</a> ).</p>\n<p>In short, it takes three contiguous slices along some axis (let’s say the z-axis) as input: a-1, a, a+1. It extracts features and applies cross-attention of slice a with a-1 and a+1, and applies self-attention on aitself. After this encoding, everything is passed through a decoder similar to a U-Net.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2Fe6ebe47cfef45e8fcd9e60fb858ad6b5%2Fcsanet.png?generation=1765178543956546&amp;alt=media\" alt=\"architecture\">\nMy intuition (though I could be completely wrong) is that such a model could perform well in this case, because we want continuity along one axis, while also having continuity in the perpendicular plane.</p>\n<p>I trained the model for 5 epochs with very standard methods. I used two metrics: Weighted BCE (with 5× weighting for voxels labeled 1) and Tversky loss to penalize false positive label-1 outputs, I didn't try to incorporate something related to fix topological errors, because I couldn't find anything which did the thing in reasonable time, So i planned to do deal with that in post processing.</p>\n<p>I’m sharing this architecture and making my pretrained models and notebooks public because I only train on Kaggle, and the model is huge (or maybe I have a trash pipeline), and each epoch takes me 4 hours. So I can’t iterate as quickly as I’d like. If someone else with better resources is interested, they can experiment further.</p>\n<p>A few ideas I have are:</p>\n<p>Train the encoder using the unlabeled data and then fine-tune the decoder (for example, predicting slice a given a-1 and a+1 as input).</p>\n<p>Taking a larger context window, for example a-k to a-1 and symmetrically on the other side, instead of just a-1 and a+1.</p>\n<p>Add some form of information flow along the z-axis like in LSTMs (which was my original intuition).</p>\n<p>So yeah, if anyone has suggestions (especially regarding my pipeline, if my pipeline is the issue and the bottleneck isn’t the model 😭😭), feel free to discuss. I would also love to hear about similar approaches, better metrics and post processing related stuff. I'm just a beginner, so please pardon any naive mistakes</p>",
  "messages": [
    {
      "id": "3365371",
      "postDate": "12/07/2025 06:00:58",
      "content": "<p>So this discussion is about the model I am working with. I was trying to find architectures that have some form of attention along an axis, something similar to an LSTM where information flows linearly. That’s when I found this architecture called CSA-Net ( <a href=\"https://arxiv.org/abs/2405.00130\" target=\"_blank\">https://arxiv.org/abs/2405.00130</a> ).</p>\n<p>In short, it takes three contiguous slices along some axis (let’s say the z-axis) as input: a-1, a, a+1. It extracts features and applies cross-attention of slice a with a-1 and a+1, and applies self-attention on aitself. After this encoding, everything is passed through a decoder similar to a U-Net.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2Fe6ebe47cfef45e8fcd9e60fb858ad6b5%2Fcsanet.png?generation=1765178543956546&amp;alt=media\" alt=\"architecture\">\nMy intuition (though I could be completely wrong) is that such a model could perform well in this case, because we want continuity along one axis, while also having continuity in the perpendicular plane.</p>\n<p>I trained the model for 5 epochs with very standard methods. I used two metrics: Weighted BCE (with 5× weighting for voxels labeled 1) and Tversky loss to penalize false positive label-1 outputs, I didn't try to incorporate something related to fix topological errors, because I couldn't find anything which did the thing in reasonable time, So i planned to do deal with that in post processing.</p>\n<p>I’m sharing this architecture and making my pretrained models and notebooks public because I only train on Kaggle, and the model is huge (or maybe I have a trash pipeline), and each epoch takes me 4 hours. So I can’t iterate as quickly as I’d like. If someone else with better resources is interested, they can experiment further.</p>\n<p>A few ideas I have are:</p>\n<p>Train the encoder using the unlabeled data and then fine-tune the decoder (for example, predicting slice a given a-1 and a+1 as input).</p>\n<p>Taking a larger context window, for example a-k to a-1 and symmetrically on the other side, instead of just a-1 and a+1.</p>\n<p>Add some form of information flow along the z-axis like in LSTMs (which was my original intuition).</p>\n<p>So yeah, if anyone has suggestions (especially regarding my pipeline, if my pipeline is the issue and the bottleneck isn’t the model 😭😭), feel free to discuss. I would also love to hear about similar approaches, better metrics and post processing related stuff. I'm just a beginner, so please pardon any naive mistakes</p>",
      "rawMarkdown": "So this discussion is about the model I am working with. I was trying to find architectures that have some form of attention along an axis, something similar to an LSTM where information flows linearly. That’s when I found this architecture called CSA-Net ( https://arxiv.org/abs/2405.00130 ).\n\nIn short, it takes three contiguous slices along some axis (let’s say the z-axis) as input: a-1, a, a+1. It extracts features and applies cross-attention of slice a with a-1 and a+1, and applies self-attention on aitself. After this encoding, everything is passed through a decoder similar to a U-Net.\n![architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2Fe6ebe47cfef45e8fcd9e60fb858ad6b5%2Fcsanet.png?generation=1765178543956546&alt=media)\nMy intuition (though I could be completely wrong) is that such a model could perform well in this case, because we want continuity along one axis, while also having continuity in the perpendicular plane.\n\nI trained the model for 5 epochs with very standard methods. I used two metrics: Weighted BCE (with 5× weighting for voxels labeled 1) and Tversky loss to penalize false positive label-1 outputs, I didn't try to incorporate something related to fix topological errors, because I couldn't find anything which did the thing in reasonable time, So i planned to do deal with that in post processing.\n\nI’m sharing this architecture and making my pretrained models and notebooks public because I only train on Kaggle, and the model is huge (or maybe I have a trash pipeline), and each epoch takes me 4 hours. So I can’t iterate as quickly as I’d like. If someone else with better resources is interested, they can experiment further.\n\nA few ideas I have are:\n\nTrain the encoder using the unlabeled data and then fine-tune the decoder (for example, predicting slice a given a-1 and a+1 as input).\n\nTaking a larger context window, for example a-k to a-1 and symmetrically on the other side, instead of just a-1 and a+1.\n\nAdd some form of information flow along the z-axis like in LSTMs (which was my original intuition).\n\nSo yeah, if anyone has suggestions (especially regarding my pipeline, if my pipeline is the issue and the bottleneck isn’t the model 😭😭), feel free to discuss. I would also love to hear about similar approaches, better metrics and post processing related stuff. I'm just a beginner, so please pardon any naive mistakes",
      "votes": null
    },
    {
      "id": "3365722",
      "postDate": "12/07/2025 10:21:51",
      "content": "<p>Have a look at this: <a href=\"https://arxiv.org/pdf/2102.05095\" target=\"_blank\">https://arxiv.org/pdf/2102.05095</a></p>",
      "rawMarkdown": "Have a look at this: https://arxiv.org/pdf/2102.05095",
      "votes": null
    },
    {
      "id": "3378156",
      "postDate": "12/17/2025 16:30:36",
      "content": "<p>Hello, nice going through your thoughts.\nWas wondering if there is a reason why you chose this model over a swinUnetR? \nWouldn't attention applied across a window of pixels be better than one applied across different axis?</p>",
      "rawMarkdown": "Hello, nice going through your thoughts.\nWas wondering if there is a reason why you chose this model over a swinUnetR? \nWouldn't attention applied across a window of pixels be better than one applied across different axis?",
      "votes": null
    },
    {
      "id": "3378206",
      "postDate": "12/17/2025 18:23:10",
      "content": "<p>In CSA, attention is applied in two ways, Cross attention (between different Layers) and In-Slice attention which is in the same layer itself. Also we have to choose models that understand 3d features of scroll, because It is important for models to learn the \"flow\" of sheets in all directions. The structure in x-y is very different from that along z axis, which is why i was trying to find such architectures.</p>",
      "rawMarkdown": "In CSA, attention is applied in two ways, Cross attention (between different Layers) and In-Slice attention which is in the same layer itself. Also we have to choose models that understand 3d features of scroll, because It is important for models to learn the \"flow\" of sheets in all directions. The structure in x-y is very different from that along z axis, which is why i was trying to find such architectures.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3365722,
      "author_name": "giorgioangelotti",
      "author_url": "",
      "post_date": "12/07/2025 10:21:51",
      "content": "<p>Have a look at this: <a href=\"https://arxiv.org/pdf/2102.05095\" target=\"_blank\">https://arxiv.org/pdf/2102.05095</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3378156,
      "author_name": "arjunashokbhandary",
      "author_url": "",
      "post_date": "12/17/2025 16:30:36",
      "content": "<p>Hello, nice going through your thoughts.\nWas wondering if there is a reason why you chose this model over a swinUnetR? \nWouldn't attention applied across a window of pixels be better than one applied across different axis?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3378206,
          "author_name": "choudharymanas",
          "author_url": "",
          "post_date": "12/17/2025 18:23:10",
          "content": "<p>In CSA, attention is applied in two ways, Cross attention (between different Layers) and In-Slice attention which is in the same layer itself. Also we have to choose models that understand 3d features of scroll, because It is important for models to learn the \"flow\" of sheets in all directions. The structure in x-y is very different from that along z axis, which is why i was trying to find such architectures.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3365371": "So this discussion is about the model I am working with. I was trying to find architectures that have some form of attention along an axis, something similar to an LSTM where information flows linearly. That’s when I found this architecture called CSA-Net ( https://arxiv.org/abs/2405.00130 ).\n\nIn short, it takes three contiguous slices along some axis (let’s say the z-axis) as input: a-1, a, a+1. It extracts features and applies cross-attention of slice a with a-1 and a+1, and applies self-attention on aitself. After this encoding, everything is passed through a decoder similar to a U-Net.\n![architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2Fe6ebe47cfef45e8fcd9e60fb858ad6b5%2Fcsanet.png?generation=1765178543956546&alt=media)\nMy intuition (though I could be completely wrong) is that such a model could perform well in this case, because we want continuity along one axis, while also having continuity in the perpendicular plane.\n\nI trained the model for 5 epochs with very standard methods. I used two metrics: Weighted BCE (with 5× weighting for voxels labeled 1) and Tversky loss to penalize false positive label-1 outputs, I didn't try to incorporate something related to fix topological errors, because I couldn't find anything which did the thing in reasonable time, So i planned to do deal with that in post processing.\n\nI’m sharing this architecture and making my pretrained models and notebooks public because I only train on Kaggle, and the model is huge (or maybe I have a trash pipeline), and each epoch takes me 4 hours. So I can’t iterate as quickly as I’d like. If someone else with better resources is interested, they can experiment further.\n\nA few ideas I have are:\n\nTrain the encoder using the unlabeled data and then fine-tune the decoder (for example, predicting slice a given a-1 and a+1 as input).\n\nTaking a larger context window, for example a-k to a-1 and symmetrically on the other side, instead of just a-1 and a+1.\n\nAdd some form of information flow along the z-axis like in LSTMs (which was my original intuition).\n\nSo yeah, if anyone has suggestions (especially regarding my pipeline, if my pipeline is the issue and the bottleneck isn’t the model 😭😭), feel free to discuss. I would also love to hear about similar approaches, better metrics and post processing related stuff. I'm just a beginner, so please pardon any naive mistakes",
    "3365722": "Have a look at this: https://arxiv.org/pdf/2102.05095",
    "3378156": "Hello, nice going through your thoughts.\nWas wondering if there is a reason why you chose this model over a swinUnetR? \nWouldn't attention applied across a window of pixels be better than one applied across different axis?",
    "3378206": "In CSA, attention is applied in two ways, Cross attention (between different Layers) and In-Slice attention which is in the same layer itself. Also we have to choose models that understand 3d features of scroll, because It is important for models to learn the \"flow\" of sheets in all directions. The structure in x-y is very different from that along z axis, which is why i was trying to find such architectures."
  },
  "source": "meta"
}