{
  "id": 270832,
  "title": " Transfer Learning  , Does it work for you ?",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/270832",
  "author_name": "",
  "post_date": "2021-09-07T07:04:17.619945Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi , I have been trying to use transfer learning for this competition as most of us here but strangely the model doesn't learn much since the data is so small , I just wanted to know if anyone here is able to get to score of 0.7+ on public lb using 3d pretrained models ??</p>\n<p>Do you think it's better to create our own model and train them again ? </p>",
  "messages": [
    {
      "id": "1505335",
      "postDate": "09/07/2021 07:04:17",
      "content": "<p>Hi , I have been trying to use transfer learning for this competition as most of us here but strangely the model doesn't learn much since the data is so small , I just wanted to know if anyone here is able to get to score of 0.7+ on public lb using 3d pretrained models ??</p>\n<p>Do you think it's better to create our own model and train them again ? </p>",
      "rawMarkdown": "Hi , I have been trying to use transfer learning for this competition as most of us here but strangely the model doesn't learn much since the data is so small , I just wanted to know if anyone here is able to get to score of 0.7+ on public lb using 3d pretrained models ??\n\nDo you think it's better to create our own model and train them again ?",
      "votes": null
    },
    {
      "id": "1508801",
      "postDate": "09/10/2021 15:06:42",
      "content": "<p>Hi. I would suggest checking your pipeline well. Are you using the public one that has the highest score on the public LB? If yes, keep in mind that it has several flaws and the results might be very misleading. Personally I made it over 0.7 (both CV and LB) without using any pretrained models</p>",
      "rawMarkdown": "Hi. I would suggest checking your pipeline well. Are you using the public one that has the highest score on the public LB? If yes, keep in mind that it has several flaws and the results might be very misleading. Personally I made it over 0.7 (both CV and LB) without using any pretrained models",
      "votes": null
    },
    {
      "id": "1511174",
      "postDate": "09/13/2021 07:18:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mikecho\" target=\"_blank\">@mikecho</a> , Thank you for putting up your suggestion , May I ask a few things like - </p>\n<p>When you say flaw , do you mean that the model loss is virtually not going down ? or something else ?<br>\nCongratulations on getting such high AUC on the CV &amp; Lb , may you share some guidance about how can i approach this problem , using a 2d CNN or 3d CNN or is there something I am missing out.<br>\nI really appreciate your help .  </p>",
      "rawMarkdown": "Hi @mikecho , Thank you for putting up your suggestion , May I ask a few things like - \n\nWhen you say flaw , do you mean that the model loss is virtually not going down ? or something else ?\nCongratulations on getting such high AUC on the CV & Lb , may you share some guidance about how can i approach this problem , using a 2d CNN or 3d CNN or is there something I am missing out.\nI really appreciate your help .",
      "votes": null
    },
    {
      "id": "1511641",
      "postDate": "09/13/2021 15:24:26",
      "content": "<p>By flaws I mean that:</p>\n<ol>\n<li>for some reason Batch Normalization layers in the used model don't work well and the network outputs the same probabilities for each sample within the batch in the eval mode during training. This is why you can rely neither on the computed loss nor auc,</li>\n<li>The notebook does not use sigmoid layers - it relies purely on the output of the linear layers, so that the returned value is not a probability - which is expected while calculating AUC</li>\n</ol>\n<p>Regarding 3d CNN, you can have a look at my <a href=\"https://www.kaggle.com/mikecho/rsna-miccai-monai-ensemble\" target=\"_blank\">notebook</a></p>\n<p>I am updating it to include more techniques that should prevent overfitting (like data augmentation). You can check it out soon to see them in action. I guess that the LB score achieved by this notebook will be improving after applying them.</p>",
      "rawMarkdown": "By flaws I mean that:\n1. for some reason Batch Normalization layers in the used model don't work well and the network outputs the same probabilities for each sample within the batch in the eval mode during training. This is why you can rely neither on the computed loss nor auc,\n2. The notebook does not use sigmoid layers - it relies purely on the output of the linear layers, so that the returned value is not a probability - which is expected while calculating AUC\n\nRegarding 3d CNN, you can have a look at my [notebook](https://www.kaggle.com/mikecho/rsna-miccai-monai-ensemble)\n\nI am updating it to include more techniques that should prevent overfitting (like data augmentation). You can check it out soon to see them in action. I guess that the LB score achieved by this notebook will be improving after applying them.",
      "votes": null
    },
    {
      "id": "1514151",
      "postDate": "09/15/2021 18:18:16",
      "content": "<p>Hi, I can't understand what you said about sigmoid function and returned value is not probability.<br>\nwhat do you mean?<br>\ndata augmentation is useful or not and we should save augmented images as new image or not use it when read new batch? <a href=\"https://www.kaggle.com/mikecho\" target=\"_blank\">@mikecho</a> </p>",
      "rawMarkdown": "Hi, I can't understand what you said about sigmoid function and returned value is not probability.\nwhat do you mean?\ndata augmentation is useful or not and we should save augmented images as new image or not use it when read new batch? @mikecho",
      "votes": null
    },
    {
      "id": "1514213",
      "postDate": "09/15/2021 19:26:07",
      "content": "<p>The predictions used in the notebook are not converted into probabilities that needs to be in the range between 0 and 1. Usually it is performed using softmax or sigmoid activation functions, but as this modeled problem is based on binary classification, sigmoid is the proper choice here. This is not done in the best performing public notebook.</p>\n<p>Regarding data augmentation, it is definitely useful and without it it would be difficult ti reach a high score in this competition as the provided dataset is very small.</p>\n<p>I think that if you don't want to apply the same augmentations again, you will not have to store the augmented images in the disk. Generating random augmentations in every batch would also increase variance in the input data, so it might lower overfitting a little bit more compared to the limited, stored amount of images, and with a low learning rate your model should end up trained on plenty of various (although similar) data samples. However, perhaps some optimization could be to augment the images and store them on the disk first. I think that this is not a good approach in this competition as the dataset, even after preprocessing requires a lot of storage memory. And resources are very limited if, both here (datasets), and on Google Colab (notebook disk space). </p>",
      "rawMarkdown": "The predictions used in the notebook are not converted into probabilities that needs to be in the range between 0 and 1. Usually it is performed using softmax or sigmoid activation functions, but as this modeled problem is based on binary classification, sigmoid is the proper choice here. This is not done in the best performing public notebook.\n\nRegarding data augmentation, it is definitely useful and without it it would be difficult ti reach a high score in this competition as the provided dataset is very small.\n\nI think that if you don't want to apply the same augmentations again, you will not have to store the augmented images in the disk. Generating random augmentations in every batch would also increase variance in the input data, so it might lower overfitting a little bit more compared to the limited, stored amount of images, and with a low learning rate your model should end up trained on plenty of various (although similar) data samples. However, perhaps some optimization could be to augment the images and store them on the disk first. I think that this is not a good approach in this competition as the dataset, even after preprocessing requires a lot of storage memory. And resources are very limited if, both here (datasets), and on Google Colab (notebook disk space).",
      "votes": null
    },
    {
      "id": "1514571",
      "postDate": "09/16/2021 07:47:02",
      "content": "<p>Thanks a lot. In your notebook i can't understand why you used LSTM model and train with this.<br>\nwhat is the advantage of use LSTM and  why you use this?</p>",
      "rawMarkdown": "Thanks a lot. In your notebook i can't understand why you used LSTM model and train with this.\nwhat is the advantage of use LSTM and  why you use this?",
      "votes": null
    },
    {
      "id": "1514582",
      "postDate": "09/16/2021 07:59:00",
      "content": "<p><a href=\"https://www.kaggle.com/mohammadhosein1998\" target=\"_blank\">@mohammadhosein1998</a> I didn't use LSTM in my notebook, but it's an interesting idea to use Recurrent CNNs instead of 3d CNNs. In <a href=\"https://www.kaggle.com/mikecho/rsna-miccai-monai-ensemble\" target=\"_blank\">my notebook</a> I am using an implementation Densenet121 3d provided by <a href=\"https://www.kaggle.com/mikecho/monai-v060-deep-learning-in-healthcare-imaging\" target=\"_blank\">framework Monai</a></p>",
      "rawMarkdown": "mohammadhosein1998 I didn't use LSTM in my notebook, but it's an interesting idea to use Recurrent CNNs instead of 3d CNNs. In [my notebook](https://www.kaggle.com/mikecho/rsna-miccai-monai-ensemble) I am using an implementation Densenet121 3d provided by [framework Monai](https://www.kaggle.com/mikecho/monai-v060-deep-learning-in-healthcare-imaging)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1508801,
      "author_name": "mikecho",
      "author_url": "",
      "post_date": "09/10/2021 15:06:42",
      "content": "<p>Hi. I would suggest checking your pipeline well. Are you using the public one that has the highest score on the public LB? If yes, keep in mind that it has several flaws and the results might be very misleading. Personally I made it over 0.7 (both CV and LB) without using any pretrained models</p>",
      "votes": null,
      "replies": [
        {
          "id": 1511174,
          "author_name": "avikrams",
          "author_url": "",
          "post_date": "09/13/2021 07:18:59",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mikecho\" target=\"_blank\">@mikecho</a> , Thank you for putting up your suggestion , May I ask a few things like - </p>\n<p>When you say flaw , do you mean that the model loss is virtually not going down ? or something else ?<br>\nCongratulations on getting such high AUC on the CV &amp; Lb , may you share some guidance about how can i approach this problem , using a 2d CNN or 3d CNN or is there something I am missing out.<br>\nI really appreciate your help .  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1511641,
          "author_name": "mikecho",
          "author_url": "",
          "post_date": "09/13/2021 15:24:26",
          "content": "<p>By flaws I mean that:</p>\n<ol>\n<li>for some reason Batch Normalization layers in the used model don't work well and the network outputs the same probabilities for each sample within the batch in the eval mode during training. This is why you can rely neither on the computed loss nor auc,</li>\n<li>The notebook does not use sigmoid layers - it relies purely on the output of the linear layers, so that the returned value is not a probability - which is expected while calculating AUC</li>\n</ol>\n<p>Regarding 3d CNN, you can have a look at my <a href=\"https://www.kaggle.com/mikecho/rsna-miccai-monai-ensemble\" target=\"_blank\">notebook</a></p>\n<p>I am updating it to include more techniques that should prevent overfitting (like data augmentation). You can check it out soon to see them in action. I guess that the LB score achieved by this notebook will be improving after applying them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1514151,
          "author_name": "mohammadhosein1998",
          "author_url": "",
          "post_date": "09/15/2021 18:18:16",
          "content": "<p>Hi, I can't understand what you said about sigmoid function and returned value is not probability.<br>\nwhat do you mean?<br>\ndata augmentation is useful or not and we should save augmented images as new image or not use it when read new batch? <a href=\"https://www.kaggle.com/mikecho\" target=\"_blank\">@mikecho</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1514213,
          "author_name": "mikecho",
          "author_url": "",
          "post_date": "09/15/2021 19:26:07",
          "content": "<p>The predictions used in the notebook are not converted into probabilities that needs to be in the range between 0 and 1. Usually it is performed using softmax or sigmoid activation functions, but as this modeled problem is based on binary classification, sigmoid is the proper choice here. This is not done in the best performing public notebook.</p>\n<p>Regarding data augmentation, it is definitely useful and without it it would be difficult ti reach a high score in this competition as the provided dataset is very small.</p>\n<p>I think that if you don't want to apply the same augmentations again, you will not have to store the augmented images in the disk. Generating random augmentations in every batch would also increase variance in the input data, so it might lower overfitting a little bit more compared to the limited, stored amount of images, and with a low learning rate your model should end up trained on plenty of various (although similar) data samples. However, perhaps some optimization could be to augment the images and store them on the disk first. I think that this is not a good approach in this competition as the dataset, even after preprocessing requires a lot of storage memory. And resources are very limited if, both here (datasets), and on Google Colab (notebook disk space). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1514571,
          "author_name": "mohammadhosein1998",
          "author_url": "",
          "post_date": "09/16/2021 07:47:02",
          "content": "<p>Thanks a lot. In your notebook i can't understand why you used LSTM model and train with this.<br>\nwhat is the advantage of use LSTM and  why you use this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1514582,
          "author_name": "mikecho",
          "author_url": "",
          "post_date": "09/16/2021 07:59:00",
          "content": "<p><a href=\"https://www.kaggle.com/mohammadhosein1998\" target=\"_blank\">@mohammadhosein1998</a> I didn't use LSTM in my notebook, but it's an interesting idea to use Recurrent CNNs instead of 3d CNNs. In <a href=\"https://www.kaggle.com/mikecho/rsna-miccai-monai-ensemble\" target=\"_blank\">my notebook</a> I am using an implementation Densenet121 3d provided by <a href=\"https://www.kaggle.com/mikecho/monai-v060-deep-learning-in-healthcare-imaging\" target=\"_blank\">framework Monai</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1505335": "Hi , I have been trying to use transfer learning for this competition as most of us here but strangely the model doesn't learn much since the data is so small , I just wanted to know if anyone here is able to get to score of 0.7+ on public lb using 3d pretrained models ??\n\nDo you think it's better to create our own model and train them again ?",
    "1508801": "Hi. I would suggest checking your pipeline well. Are you using the public one that has the highest score on the public LB? If yes, keep in mind that it has several flaws and the results might be very misleading. Personally I made it over 0.7 (both CV and LB) without using any pretrained models",
    "1511174": "Hi @mikecho , Thank you for putting up your suggestion , May I ask a few things like - \n\nWhen you say flaw , do you mean that the model loss is virtually not going down ? or something else ?\nCongratulations on getting such high AUC on the CV & Lb , may you share some guidance about how can i approach this problem , using a 2d CNN or 3d CNN or is there something I am missing out.\nI really appreciate your help .",
    "1511641": "By flaws I mean that:\n1. for some reason Batch Normalization layers in the used model don't work well and the network outputs the same probabilities for each sample within the batch in the eval mode during training. This is why you can rely neither on the computed loss nor auc,\n2. The notebook does not use sigmoid layers - it relies purely on the output of the linear layers, so that the returned value is not a probability - which is expected while calculating AUC\n\nRegarding 3d CNN, you can have a look at my [notebook](https://www.kaggle.com/mikecho/rsna-miccai-monai-ensemble)\n\nI am updating it to include more techniques that should prevent overfitting (like data augmentation). You can check it out soon to see them in action. I guess that the LB score achieved by this notebook will be improving after applying them.",
    "1514151": "Hi, I can't understand what you said about sigmoid function and returned value is not probability.\nwhat do you mean?\ndata augmentation is useful or not and we should save augmented images as new image or not use it when read new batch? @mikecho",
    "1514213": "The predictions used in the notebook are not converted into probabilities that needs to be in the range between 0 and 1. Usually it is performed using softmax or sigmoid activation functions, but as this modeled problem is based on binary classification, sigmoid is the proper choice here. This is not done in the best performing public notebook.\n\nRegarding data augmentation, it is definitely useful and without it it would be difficult ti reach a high score in this competition as the provided dataset is very small.\n\nI think that if you don't want to apply the same augmentations again, you will not have to store the augmented images in the disk. Generating random augmentations in every batch would also increase variance in the input data, so it might lower overfitting a little bit more compared to the limited, stored amount of images, and with a low learning rate your model should end up trained on plenty of various (although similar) data samples. However, perhaps some optimization could be to augment the images and store them on the disk first. I think that this is not a good approach in this competition as the dataset, even after preprocessing requires a lot of storage memory. And resources are very limited if, both here (datasets), and on Google Colab (notebook disk space).",
    "1514571": "Thanks a lot. In your notebook i can't understand why you used LSTM model and train with this.\nwhat is the advantage of use LSTM and  why you use this?",
    "1514582": "mohammadhosein1998 I didn't use LSTM in my notebook, but it's an interesting idea to use Recurrent CNNs instead of 3d CNNs. In [my notebook](https://www.kaggle.com/mikecho/rsna-miccai-monai-ensemble) I am using an implementation Densenet121 3d provided by [framework Monai](https://www.kaggle.com/mikecho/monai-v060-deep-learning-in-healthcare-imaging)"
  },
  "source": "meta"
}