{
  "id": 539686,
  "title": "My Siamese Network Solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/rasoul-my-siamese-network-solution",
  "author_name": "",
  "post_date": "2024-11-04T13:48:18.087Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p><strong>My Core Solution</strong></p>\n<p>I am presenting one of my solutions that I believe it has been less explored within this competition by other participants, Siamese Network. The analysis of the labels distributions revealed that we are dealing with a severe unbalanced dataset with very few samples per some of the classes. Below shows a sample of a distribution for Axial T2 series.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4392121%2F117500e6c90bf80a09cdbfa26221a6dd%2Fdistribution_bar_graph_for_l1_l2.png?generation=1728551246683795&amp;alt=media\" alt=\"\"><br>\nI designed a two-stage solution, the first stage employs YOLOv9 to extract patches centered at the key points,<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4392121%2F29604f1cfd50f0117e8c43167dfa09b8%2Fall_series_yolo_detection.png?generation=1728551739645456&amp;alt=media\" alt=\"\"><br>\nThe patches are then fed to the trained Siamese Network for being classified as <code>Normal/Mild</code>, <code>Moderate</code> or <code>Severe</code>.</p>\n<p><strong>Siamese Network</strong></p>\n<p>I used a standard Siamese Network architecture for image similarity task, with a few different backbones, namely, <code>ResNet18</code>, <code>DenseNet121</code>, <code>EfficientNet</code> and <code>VGG11</code> and the patches were resized to 224x224 size.<br>\nAt the training phase, two patches are selected randomly and then depending on their true classes, a meta-label of similarity created on the fly (which is 1.0 if two samples are from the same class and 0.0 otherwise). For loss, I implemented a custom cosine similarity loss function,</p>\n<pre><code> (nn.Module):\n     ():\n        (ContrastiveLossCosine, ).__init__()\n        .margin = margin\n\n     ():\n        pos_loss = ( - similarity_score) * label\n        neg_loss = (similarity_score - .margin) * ( - label)\n        loss = pos_loss + torch.relu(neg_loss)\n         loss.mean()\n</code></pre>\n<p>For inference, first I cluster the images of each class for the maximum <strong>dissimilarity</strong> scores, then I select randomly few images per cluster such that there will be N images per class representing the diversity and variation in the patches. I call these selected images as <strong>reference images</strong>. I have added a pre-computing embedding vectors function to the model so that it uses the reference images per class to pre-compute the embedding vectors for each class. At inference time, extracting a patch from YOLOv9 pipeline and feeding it to the Siamese Network, it will output the similarity scores of the patch with respect to all reference images per class. An averaging of the scores give us the final score per class.<br>\nNow we have similarity scores for all classes, the question is how to generate a probability distribution out of these similarity scores? I examined different methods (such as softmax) and found out a simple vector normalization works the best, <code>probabilities = scores / scores.sum()</code>.</p>\n<p><strong>Other Solutions</strong><br>\nI used an ensemble of different models: Single-stage multi-task models using different backbones and slice sampling, a 3D volume approach for Sagittal T2 series inspired by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> work.</p>\n<p><strong>Training Code</strong><br>\n<a href=\"https://github.com/Rasoul77/SiameseSimNet\" target=\"_blank\">https://github.com/Rasoul77/SiameseSimNet</a></p>\n<p><strong>Axial T2 Inference Code</strong><br>\n<a href=\"https://www.kaggle.com/code/rasoulmojtahedzadeh/siamese-inference-axial-t2\" target=\"_blank\">https://www.kaggle.com/code/rasoulmojtahedzadeh/siamese-inference-axial-t2</a></p>",
  "messages": [
    {
      "id": "3013585",
      "postDate": "10/10/2024 09:52:45",
      "content": "<p><strong>My Core Solution</strong></p>\n<p>I am presenting one of my solutions that I believe it has been less explored within this competition by other participants, Siamese Network. The analysis of the labels distributions revealed that we are dealing with a severe unbalanced dataset with very few samples per some of the classes. Below shows a sample of a distribution for Axial T2 series.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4392121%2F117500e6c90bf80a09cdbfa26221a6dd%2Fdistribution_bar_graph_for_l1_l2.png?generation=1728551246683795&amp;alt=media\" alt=\"\"><br>\nI designed a two-stage solution, the first stage employs YOLOv9 to extract patches centered at the key points,<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4392121%2F29604f1cfd50f0117e8c43167dfa09b8%2Fall_series_yolo_detection.png?generation=1728551739645456&amp;alt=media\" alt=\"\"><br>\nThe patches are then fed to the trained Siamese Network for being classified as <code>Normal/Mild</code>, <code>Moderate</code> or <code>Severe</code>.</p>\n<p><strong>Siamese Network</strong></p>\n<p>I used a standard Siamese Network architecture for image similarity task, with a few different backbones, namely, <code>ResNet18</code>, <code>DenseNet121</code>, <code>EfficientNet</code> and <code>VGG11</code> and the patches were resized to 224x224 size.<br>\nAt the training phase, two patches are selected randomly and then depending on their true classes, a meta-label of similarity created on the fly (which is 1.0 if two samples are from the same class and 0.0 otherwise). For loss, I implemented a custom cosine similarity loss function,</p>\n<pre><code> (nn.Module):\n     ():\n        (ContrastiveLossCosine, ).__init__()\n        .margin = margin\n\n     ():\n        pos_loss = ( - similarity_score) * label\n        neg_loss = (similarity_score - .margin) * ( - label)\n        loss = pos_loss + torch.relu(neg_loss)\n         loss.mean()\n</code></pre>\n<p>For inference, first I cluster the images of each class for the maximum <strong>dissimilarity</strong> scores, then I select randomly few images per cluster such that there will be N images per class representing the diversity and variation in the patches. I call these selected images as <strong>reference images</strong>. I have added a pre-computing embedding vectors function to the model so that it uses the reference images per class to pre-compute the embedding vectors for each class. At inference time, extracting a patch from YOLOv9 pipeline and feeding it to the Siamese Network, it will output the similarity scores of the patch with respect to all reference images per class. An averaging of the scores give us the final score per class.<br>\nNow we have similarity scores for all classes, the question is how to generate a probability distribution out of these similarity scores? I examined different methods (such as softmax) and found out a simple vector normalization works the best, <code>probabilities = scores / scores.sum()</code>.</p>\n<p><strong>Other Solutions</strong><br>\nI used an ensemble of different models: Single-stage multi-task models using different backbones and slice sampling, a 3D volume approach for Sagittal T2 series inspired by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> work.</p>\n<p><strong>Training Code</strong><br>\n<a href=\"https://github.com/Rasoul77/SiameseSimNet\" target=\"_blank\">https://github.com/Rasoul77/SiameseSimNet</a></p>\n<p><strong>Axial T2 Inference Code</strong><br>\n<a href=\"https://www.kaggle.com/code/rasoulmojtahedzadeh/siamese-inference-axial-t2\" target=\"_blank\">https://www.kaggle.com/code/rasoulmojtahedzadeh/siamese-inference-axial-t2</a></p>",
      "rawMarkdown": "**My Core Solution**\n\nI am presenting one of my solutions that I believe it has been less explored within this competition by other participants, Siamese Network. The analysis of the labels distributions revealed that we are dealing with a severe unbalanced dataset with very few samples per some of the classes. Below shows a sample of a distribution for Axial T2 series.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4392121%2F117500e6c90bf80a09cdbfa26221a6dd%2Fdistribution_bar_graph_for_l1_l2.png?generation=1728551246683795&alt=media)\nI designed a two-stage solution, the first stage employs YOLOv9 to extract patches centered at the key points,\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4392121%2F29604f1cfd50f0117e8c43167dfa09b8%2Fall_series_yolo_detection.png?generation=1728551739645456&alt=media)\nThe patches are then fed to the trained Siamese Network for being classified as `Normal/Mild`, `Moderate` or `Severe`.\n\n**Siamese Network**\n\nI used a standard Siamese Network architecture for image similarity task, with a few different backbones, namely, `ResNet18`, `DenseNet121`, `EfficientNet` and `VGG11` and the patches were resized to 224x224 size.\nAt the training phase, two patches are selected randomly and then depending on their true classes, a meta-label of similarity created on the fly (which is 1.0 if two samples are from the same class and 0.0 otherwise). For loss, I implemented a custom cosine similarity loss function,\n```python\nclass ContrastiveLossCosine(nn.Module):\n    def __init__(self, margin=0.5):\n        super(ContrastiveLossCosine, self).__init__()\n        self.margin = margin\n   \n    def forward(self, similarity_score, label):\n        pos_loss = (1 - similarity_score) * label\n        neg_loss = (similarity_score - self.margin) * (1 - label)\n        loss = pos_loss + torch.relu(neg_loss)\n        return loss.mean()\n\n```\nFor inference, first I cluster the images of each class for the maximum **dissimilarity** scores, then I select randomly few images per cluster such that there will be N images per class representing the diversity and variation in the patches. I call these selected images as **reference images**. I have added a pre-computing embedding vectors function to the model so that it uses the reference images per class to pre-compute the embedding vectors for each class. At inference time, extracting a patch from YOLOv9 pipeline and feeding it to the Siamese Network, it will output the similarity scores of the patch with respect to all reference images per class. An averaging of the scores give us the final score per class.\nNow we have similarity scores for all classes, the question is how to generate a probability distribution out of these similarity scores? I examined different methods (such as softmax) and found out a simple vector normalization works the best, ` probabilities = scores / scores.sum()`.\n\n**Other Solutions**\nI used an ensemble of different models: Single-stage multi-task models using different backbones and slice sampling, a 3D volume approach for Sagittal T2 series inspired by @hengck23 work.\n\n**Training Code**\nhttps://github.com/Rasoul77/SiameseSimNet\n\n**Axial T2 Inference Code**\nhttps://www.kaggle.com/code/rasoulmojtahedzadeh/siamese-inference-axial-t2",
      "votes": null
    },
    {
      "id": "3054097",
      "postDate": "11/24/2024 10:22:40",
      "content": "<p>hey <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a> can we get the private dataset of reference image please?</p>",
      "rawMarkdown": "hey @rasoulmojtahedzadeh can we get the private dataset of reference image please?",
      "votes": null
    },
    {
      "id": "3054125",
      "postDate": "11/24/2024 10:59:42",
      "content": "<p>Hello,<br>\nThe dataset of reference images is already public. You can access it to download or use in your own notebook from below link,<br>\n<a href=\"https://www.kaggle.com/datasets/rasoulmojtahedzadeh/rsna24-refimages-axial-t2-saimese/versions/1\" target=\"_blank\">https://www.kaggle.com/datasets/rasoulmojtahedzadeh/rsna24-refimages-axial-t2-saimese/versions/1</a></p>",
      "rawMarkdown": "Hello,\nThe dataset of reference images is already public. You can access it to download or use in your own notebook from below link,\nhttps://www.kaggle.com/datasets/rasoulmojtahedzadeh/rsna24-refimages-axial-t2-saimese/versions/1",
      "votes": null
    },
    {
      "id": "3056346",
      "postDate": "11/26/2024 21:20:37",
      "content": "<p>Thank you sir.</p>",
      "rawMarkdown": "Thank you sir.",
      "votes": null
    },
    {
      "id": "3056347",
      "postDate": "11/26/2024 21:21:57",
      "content": "<p>Did you suffer from the problem of no any detection during patch extraction in some of the test images?</p>",
      "rawMarkdown": "Did you suffer from the problem of no any detection during patch extraction in some of the test images?",
      "votes": null
    },
    {
      "id": "3056351",
      "postDate": "11/26/2024 21:29:38",
      "content": "<p>Yes, the YOLOv9 model that I trained is not perfect and for some test images fails to extract any patch for a level/side.</p>",
      "rawMarkdown": "Yes, the YOLOv9 model that I trained is not perfect and for some test images fails to extract any patch for a level/side.",
      "votes": null
    },
    {
      "id": "3056915",
      "postDate": "11/27/2024 14:00:57",
      "content": "<p>ok thank you <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a> I was worried maybe I did something wrong or what…Thank you for this clarification, wish we had some way to make it extract patches for all the test images but ok what can we do….</p>",
      "rawMarkdown": "ok thank you @rasoulmojtahedzadeh I was worried maybe I did something wrong or what...Thank you for this clarification, wish we had some way to make it extract patches for all the test images but ok what can we do....",
      "votes": null
    },
    {
      "id": "3057233",
      "postDate": "11/27/2024 22:17:39",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a>, what are the confidence and detection threshold you used?</p>",
      "rawMarkdown": "Hello @rasoulmojtahedzadeh, what are the confidence and detection threshold you used?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3054097,
      "author_name": "aliqyanabid21",
      "author_url": "",
      "post_date": "11/24/2024 10:22:40",
      "content": "<p>hey <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a> can we get the private dataset of reference image please?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3054125,
          "author_name": "rasoulmojtahedzadeh",
          "author_url": "",
          "post_date": "11/24/2024 10:59:42",
          "content": "<p>Hello,<br>\nThe dataset of reference images is already public. You can access it to download or use in your own notebook from below link,<br>\n<a href=\"https://www.kaggle.com/datasets/rasoulmojtahedzadeh/rsna24-refimages-axial-t2-saimese/versions/1\" target=\"_blank\">https://www.kaggle.com/datasets/rasoulmojtahedzadeh/rsna24-refimages-axial-t2-saimese/versions/1</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3056346,
              "author_name": "aliqyanabid21",
              "author_url": "",
              "post_date": "11/26/2024 21:20:37",
              "content": "<p>Thank you sir.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3056347,
              "author_name": "aliqyanabid21",
              "author_url": "",
              "post_date": "11/26/2024 21:21:57",
              "content": "<p>Did you suffer from the problem of no any detection during patch extraction in some of the test images?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3056351,
                  "author_name": "rasoulmojtahedzadeh",
                  "author_url": "",
                  "post_date": "11/26/2024 21:29:38",
                  "content": "<p>Yes, the YOLOv9 model that I trained is not perfect and for some test images fails to extract any patch for a level/side.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3056915,
                      "author_name": "aliqyanabid21",
                      "author_url": "",
                      "post_date": "11/27/2024 14:00:57",
                      "content": "<p>ok thank you <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a> I was worried maybe I did something wrong or what…Thank you for this clarification, wish we had some way to make it extract patches for all the test images but ok what can we do….</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3057233,
      "author_name": "aliqyanabid21",
      "author_url": "",
      "post_date": "11/27/2024 22:17:39",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a>, what are the confidence and detection threshold you used?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3013585": "**My Core Solution**\n\nI am presenting one of my solutions that I believe it has been less explored within this competition by other participants, Siamese Network. The analysis of the labels distributions revealed that we are dealing with a severe unbalanced dataset with very few samples per some of the classes. Below shows a sample of a distribution for Axial T2 series.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4392121%2F117500e6c90bf80a09cdbfa26221a6dd%2Fdistribution_bar_graph_for_l1_l2.png?generation=1728551246683795&alt=media)\nI designed a two-stage solution, the first stage employs YOLOv9 to extract patches centered at the key points,\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4392121%2F29604f1cfd50f0117e8c43167dfa09b8%2Fall_series_yolo_detection.png?generation=1728551739645456&alt=media)\nThe patches are then fed to the trained Siamese Network for being classified as `Normal/Mild`, `Moderate` or `Severe`.\n\n**Siamese Network**\n\nI used a standard Siamese Network architecture for image similarity task, with a few different backbones, namely, `ResNet18`, `DenseNet121`, `EfficientNet` and `VGG11` and the patches were resized to 224x224 size.\nAt the training phase, two patches are selected randomly and then depending on their true classes, a meta-label of similarity created on the fly (which is 1.0 if two samples are from the same class and 0.0 otherwise). For loss, I implemented a custom cosine similarity loss function,\n```python\nclass ContrastiveLossCosine(nn.Module):\n    def __init__(self, margin=0.5):\n        super(ContrastiveLossCosine, self).__init__()\n        self.margin = margin\n   \n    def forward(self, similarity_score, label):\n        pos_loss = (1 - similarity_score) * label\n        neg_loss = (similarity_score - self.margin) * (1 - label)\n        loss = pos_loss + torch.relu(neg_loss)\n        return loss.mean()\n\n```\nFor inference, first I cluster the images of each class for the maximum **dissimilarity** scores, then I select randomly few images per cluster such that there will be N images per class representing the diversity and variation in the patches. I call these selected images as **reference images**. I have added a pre-computing embedding vectors function to the model so that it uses the reference images per class to pre-compute the embedding vectors for each class. At inference time, extracting a patch from YOLOv9 pipeline and feeding it to the Siamese Network, it will output the similarity scores of the patch with respect to all reference images per class. An averaging of the scores give us the final score per class.\nNow we have similarity scores for all classes, the question is how to generate a probability distribution out of these similarity scores? I examined different methods (such as softmax) and found out a simple vector normalization works the best, ` probabilities = scores / scores.sum()`.\n\n**Other Solutions**\nI used an ensemble of different models: Single-stage multi-task models using different backbones and slice sampling, a 3D volume approach for Sagittal T2 series inspired by @hengck23 work.\n\n**Training Code**\nhttps://github.com/Rasoul77/SiameseSimNet\n\n**Axial T2 Inference Code**\nhttps://www.kaggle.com/code/rasoulmojtahedzadeh/siamese-inference-axial-t2",
    "3054097": "hey @rasoulmojtahedzadeh can we get the private dataset of reference image please?",
    "3054125": "Hello,\nThe dataset of reference images is already public. You can access it to download or use in your own notebook from below link,\nhttps://www.kaggle.com/datasets/rasoulmojtahedzadeh/rsna24-refimages-axial-t2-saimese/versions/1",
    "3056346": "Thank you sir.",
    "3056347": "Did you suffer from the problem of no any detection during patch extraction in some of the test images?",
    "3056351": "Yes, the YOLOv9 model that I trained is not perfect and for some test images fails to extract any patch for a level/side.",
    "3056915": "ok thank you @rasoulmojtahedzadeh I was worried maybe I did something wrong or what...Thank you for this clarification, wish we had some way to make it extract patches for all the test images but ok what can we do....",
    "3057233": "Hello @rasoulmojtahedzadeh, what are the confidence and detection threshold you used?"
  },
  "source": "meta"
}