{
  "id": 165352,
  "title": "4 Things that did not work!",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165352",
  "author_name": "Amin",
  "post_date": "2020-07-09T10:57:07.127000",
  "votes": 62,
  "comment_count": 103,
  "views": 0,
  "content": "<p>I would like to share and discuss the experiments that I tried and didn't improve the LB score:</p>\n\n<p><strong>1. Focal loss:</strong> Many people reported that focal loss is giving them better results, in my case BCE significantly increases the score. <a href=\"https://github.com/artemmavrin/focal-loss\"><em>(The focal loss I am using)</em></a></p>\n\n<p><strong>2. Heavier head of the net (more layers on the top of the backbone):</strong> Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score. <strong>[JUST FOR IMAGE DATA]</strong></p>\n\n<p><strong>3. Augmentation with CutMix:</strong> (without CutMix: 0.947, with CutMix: 0.941). This discussion explains the reason behind this: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160784\">CutMix is tricky in this competition</a>. </p>\n\n<p><strong>4. Using XGBoost with tabular data:</strong> I reach LB=0.75 but it does worse than <a href=\"https://www.kaggle.com/titericz/simple-baseline\">giba's 0.70 notebook</a> when I ensemble with the CNN (Ensemble += giba's kernel=0.951 / Ensemble += XGBoost=0.941).</p>\n\n<p>I would like to have your feedback about the above, if those things worked for you or not. Or if there are any other things that did not work for you!</p>\n\n<p>&gt;### Setup:\n* Models: seresnext50, seresnext101 and all EfficientNets\n* Image sizes: 256, 512, 768 and 1024\n* Hyperparameters: tweaking dropout, batchNorm, focal loss gamma and alpha parameters…\n* I am using the custom scheduler everybody is using from the flowers comp. </p>",
  "messages": [
    {
      "id": 921522,
      "postDate": "2020-07-09T10:57:07.127Z",
      "content": "<p>I would like to share and discuss the experiments that I tried and didn't improve the LB score:</p>\n\n<p><strong>1. Focal loss:</strong> Many people reported that focal loss is giving them better results, in my case BCE significantly increases the score. <a href=\"https://github.com/artemmavrin/focal-loss\"><em>(The focal loss I am using)</em></a></p>\n\n<p><strong>2. Heavier head of the net (more layers on the top of the backbone):</strong> Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score. <strong>[JUST FOR IMAGE DATA]</strong></p>\n\n<p><strong>3. Augmentation with CutMix:</strong> (without CutMix: 0.947, with CutMix: 0.941). This discussion explains the reason behind this: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160784\">CutMix is tricky in this competition</a>. </p>\n\n<p><strong>4. Using XGBoost with tabular data:</strong> I reach LB=0.75 but it does worse than <a href=\"https://www.kaggle.com/titericz/simple-baseline\">giba's 0.70 notebook</a> when I ensemble with the CNN (Ensemble += giba's kernel=0.951 / Ensemble += XGBoost=0.941).</p>\n\n<p>I would like to have your feedback about the above, if those things worked for you or not. Or if there are any other things that did not work for you!</p>\n\n<p>&gt;### Setup:\n* Models: seresnext50, seresnext101 and all EfficientNets\n* Image sizes: 256, 512, 768 and 1024\n* Hyperparameters: tweaking dropout, batchNorm, focal loss gamma and alpha parameters…\n* I am using the custom scheduler everybody is using from the flowers comp. </p>",
      "rawMarkdown": "I would like to share and discuss the experiments that I tried and didn't improve the LB score:\n\n\n**1. Focal loss:** Many people reported that focal loss is giving them better results, in my case BCE significantly increases the score. [*(The focal loss I am using)*](https://github.com/artemmavrin/focal-loss)\n\n**2. Heavier head of the net (more layers on the top of the backbone):** Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score. **[JUST FOR IMAGE DATA]**\n\n\n**3. Augmentation with CutMix:** (without CutMix: 0.947, with CutMix: 0.941). This discussion explains the reason behind this: [CutMix is tricky in this competition](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160784). \n\n\n**4. Using XGBoost with tabular data:** I reach LB=0.75 but it does worse than [giba's 0.70 notebook](https://www.kaggle.com/titericz/simple-baseline) when I ensemble with the CNN (Ensemble += giba's kernel=0.951 / Ensemble += XGBoost=0.941).\n\nI would like to have your feedback about the above, if those things worked for you or not. Or if there are any other things that did not work for you!\n\n&gt;### Setup:\n* Models: seresnext50, seresnext101 and all EfficientNets\n* Image sizes: 256, 512, 768 and 1024\n* Hyperparameters: tweaking dropout, batchNorm, focal loss gamma and alpha parameters…\n* I am using the custom scheduler everybody is using from the flowers comp. ",
      "votes": 61
    },
    {
      "id": 923601,
      "postDate": "2020-07-11T02:38:23.133Z",
      "content": "<p><strong>2. Dense head: Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score.</strong>\nB6 with Dense (256,256,1) : LB 0.939\nB7 with Dense (256,256,1) : LB 0.941\nI think it also depends on how you setup your CV and lr-scheduler. \nThose results are using full data. </p>",
      "rawMarkdown": "**2. Dense head: Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score.**\nB6 with Dense (256,256,1) : LB 0.939\nB7 with Dense (256,256,1) : LB 0.941\nI think it also depends on how you setup your CV and lr-scheduler. \nThose results are using full data. \n",
      "votes": 3,
      "replies": [
        {
          "id": 923948,
          "postDate": "2020-07-11T07:38:39.060Z",
          "content": "<p><a href=\"/tretrausaigon\">@tretrausaigon</a> What I meant is adding layers on the top of <strong>the same model.</strong> In my case: </p>\n\n<p>B7 with 1 Dense(1 output, sigmoid) LB= 0.937\nB7 with 2 Dense (128, 1) LB= 0.935\nB7 with 3 Dense (256, 128, 1) LB=0.934</p>\n\n<p>Did you experiment with adding/removing layers to the same model and compare the LB score?</p>",
          "rawMarkdown": "@tretrausaigon What I meant is adding layers on the top of **the same model.** In my case: \n\nB7 with 1 Dense(1 output, sigmoid) LB= 0.937\nB7 with 2 Dense (128, 1) LB= 0.935\nB7 with 3 Dense (256, 128, 1) LB=0.934\n\nDid you experiment with adding/removing layers to the same model and compare the LB score?"
        },
        {
          "id": 925350,
          "postDate": "2020-07-12T03:07:53.320Z",
          "content": "<p><a href=\"/amiiiney\">@amiiiney</a> I can confirm what you're saying. Only one thing is different in my case: I get better results using focal_loss. I have tested switching to BCE and I get, with exactly the same net, 0.91 (with FL is 0.935). DNNs are complicated beasts when you don't have millions of images.</p>",
          "rawMarkdown": "@amiiiney I can confirm what you're saying. Only one thing is different in my case: I get better results using focal_loss. I have tested switching to BCE and I get, with exactly the same net, 0.91 (with FL is 0.935). DNNs are complicated beasts when you don't have millions of images.",
          "votes": 1
        },
        {
          "id": 929622,
          "postDate": "2020-07-14T19:52:33.167Z",
          "content": "<p>Thanks <a href=\"/luigisaetta\">@luigisaetta</a>! I experimented with dense (1024 nodes) + 0.4 dropout and it did better than the other smaller layers.</p>",
          "rawMarkdown": "Thanks @luigisaetta! I experimented with dense (1024 nodes) + 0.4 dropout and it did better than the other smaller layers.",
          "votes": 2
        },
        {
          "id": 929742,
          "postDate": "2020-07-14T22:58:48.260Z",
          "content": "<p>When you scale your network to B7 are you also scaling image size to the native image size of B7 (600x600) or are you using a smaller size like 256?  I find B7 with native size of 600 is just so slow, and obviously I have to compromise on batch_size.</p>",
          "rawMarkdown": "When you scale your network to B7 are you also scaling image size to the native image size of B7 (600x600) or are you using a smaller size like 256?  I find B7 with native size of 600 is just so slow, and obviously I have to compromise on batch_size.",
          "votes": 1
        },
        {
          "id": 938180,
          "postDate": "2020-07-21T11:47:21.340Z",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> I tried both and B7 with image size 600x600 is way slower (900s per epoch on TPU). For the moment I am just working with smaller image sizes to experiment more things. I couldn't compare their performance.</p>",
          "rawMarkdown": "@brianfeeny I tried both and B7 with image size 600x600 is way slower (900s per epoch on TPU). For the moment I am just working with smaller image sizes to experiment more things. I couldn't compare their performance."
        }
      ]
    },
    {
      "id": 921582,
      "postDate": "2020-07-09T11:59:02.077Z",
      "content": "<p>Hey <a href=\"/amiiiney\">@amiiiney</a> \nThanks for the report!\nHow many experiments have you done? \nForgive me otherwise, but it sounds like you have trained 1 model on each variant, but DL has intrinsic random variance, so you need to run several runs with different random states and/or hyper parameters to take any conclusion, IMHO\nE.g. a different loss may very well require a different learning rate, and it may also react differently to different regularisation techniques (more hyper params).</p>",
      "rawMarkdown": "Hey @amiiiney \nThanks for the report!\nHow many experiments have you done? \nForgive me otherwise, but it sounds like you have trained 1 model on each variant, but DL has intrinsic random variance, so you need to run several runs with different random states and/or hyper parameters to take any conclusion, IMHO\nE.g. a different loss may very well require a different learning rate, and it may also react differently to different regularisation techniques (more hyper params).",
      "votes": 4,
      "replies": [
        {
          "id": 921613,
          "postDate": "2020-07-09T12:25:21.510Z",
          "content": "<p>Thanks <a href=\"/hmendonca\">@hmendonca</a> for your feedback!\nYou are 100% right! Actually that's exactly why I started this topic because I know that for example focal loss should work better and would like to have the feedback of people working with focal loss and know the combination of models/parameters that made it work.\nRegarding the experiments, I am experimenting with:\n* Models:  seresnext50, seresnext101 and all EfficientNets\n* Image sizes: 256, 512, 768 and 1024\n* Hyperparameters:  tweaking dropout, batchNorm, focal loss gamma and alpha parameters...\n* I am using the custom scheduler everybody is using from the flowers comp.</p>\n\n<p>Obviously, there are more experiments to run that's why I am relying on fellow kagglers to share their experiences but in my case the conclusions I have reached with the combination of models, image sizes and hyperparameters tweaking showed a similar pattern.</p>",
          "rawMarkdown": "Thanks @hmendonca for your feedback!\nYou are 100% right! Actually that's exactly why I started this topic because I know that for example focal loss should work better and would like to have the feedback of people working with focal loss and know the combination of models/parameters that made it work.\nRegarding the experiments, I am experimenting with:\n* Models:  seresnext50, seresnext101 and all EfficientNets\n* Image sizes: 256, 512, 768 and 1024\n* Hyperparameters:  tweaking dropout, batchNorm, focal loss gamma and alpha parameters...\n* I am using the custom scheduler everybody is using from the flowers comp.\n\nObviously, there are more experiments to run that's why I am relying on fellow kagglers to share their experiences but in my case the conclusions I have reached with the combination of models, image sizes and hyperparameters tweaking showed a similar pattern.",
          "votes": 1
        },
        {
          "id": 921639,
          "postDate": "2020-07-09T12:49:06.087Z",
          "content": "<p>Nice <a href=\"/amiiiney\">@amiiiney</a> !\nI've done some +- extensive hyper search for this comp: <a href=\"https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\">https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model</a>\nIt was also binary and did get some good models with both bce and focal, they seem to be more or less equivalent for binary AUC, in that problem at least\nHowever, the output scales vary a lot between the 2, so you need to do some kind of scaling e.g. ranking for ensembling</p>",
          "rawMarkdown": "Nice @amiiiney !\nI've done some +- extensive hyper search for this comp: https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\nIt was also binary and did get some good models with both bce and focal, they seem to be more or less equivalent for binary AUC, in that problem at least\nHowever, the output scales vary a lot between the 2, so you need to do some kind of scaling e.g. ranking for ensembling",
          "votes": 1
        }
      ]
    },
    {
      "id": 942457,
      "postDate": "2020-07-23T18:47:48.850Z",
      "content": "<p>I agree with your 4th point. I reached an lb of 0.72 with xgb but that pushed my lb to 0.94 against my current 0.958 against the same notebook. I am yet to understand the reason behind that.\nI also experimented with adding cnn+tabular features and the results were pretty bad on cv and lb so i dropped off that idea.\nCurrently i am relying of EffB6-7 and an ensemble model of these two and experimenting with some hyper parameters\nDid anyone try extracting the weights and boosting it with tabular features?</p>",
      "rawMarkdown": "I agree with your 4th point. I reached an lb of 0.72 with xgb but that pushed my lb to 0.94 against my current 0.958 against the same notebook. I am yet to understand the reason behind that.\nI also experimented with adding cnn+tabular features and the results were pretty bad on cv and lb so i dropped off that idea.\nCurrently i am relying of EffB6-7 and an ensemble model of these two and experimenting with some hyper parameters\nDid anyone try extracting the weights and boosting it with tabular features?",
      "votes": 1,
      "replies": [
        {
          "id": 952286,
          "postDate": "2020-07-30T19:26:05.293Z",
          "content": "<p><a href=\"/akshatshreemali91\">@akshatshreemali91</a> what image size or sizes (in case of multiple image size) are you using with B6 and B7 ?</p>",
          "rawMarkdown": "@akshatshreemali91 what image size or sizes (in case of multiple image size) are you using with B6 and B7 ?"
        }
      ]
    },
    {
      "id": 937145,
      "postDate": "2020-07-20T18:56:45.763Z",
      "content": "<p><a href=\"/amiiiney\">@amiiiney</a> Thansk for sharing the points. Can you kindly elaborate how did you use <code>CutMix</code>in this competition as when apply cut mix, we get <code>multi-label</code> for an image. For example, if you combine <code>cut(10%) of benign</code> and mix it into malignant image then label will be <code>(0.1, 0.9) -&gt; (benign, malignant)</code>. As you mentioned cutmix is tricky, so did you apply cutmix on images randomly without considering the <code>target(benign, malignant)</code> or you apply cutmix on similar label images (cutmix of malignant with malignant and benign with bengin)?</p>",
      "rawMarkdown": "@amiiiney Thansk for sharing the points. Can you kindly elaborate how did you use `CutMix `in this competition as when apply cut mix, we get `multi-label` for an image. For example, if you combine `cut(10%) of benign` and mix it into malignant image then label will be `(0.1, 0.9) -&gt; (benign, malignant)`. As you mentioned cutmix is tricky, so did you apply cutmix on images randomly without considering the `target(benign, malignant)` or you apply cutmix on similar label images (cutmix of malignant with malignant and benign with bengin)?",
      "votes": 1,
      "replies": [
        {
          "id": 938169,
          "postDate": "2020-07-21T11:36:42.850Z",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> I apply CutMix considering the benign and malignant images so that the output image is a mixture of both. I tried a 10%, 15% and 20% cuts but they all did bad. I am trying now to experiment with a combination of Cutmix/GridMask <a href=\"https://www.kaggle.com/aziz69/efficientnets-meta-data-augs\">from this notebook.</a></p>",
          "rawMarkdown": "@abdurrehman245 I apply CutMix considering the benign and malignant images so that the output image is a mixture of both. I tried a 10%, 15% and 20% cuts but they all did bad. I am trying now to experiment with a combination of Cutmix/GridMask [from this notebook.](https://www.kaggle.com/aziz69/efficientnets-meta-data-augs)",
          "votes": 1
        }
      ]
    },
    {
      "id": 932956,
      "postDate": "2020-07-17T11:57:20.803Z",
      "content": "<p>Thank you for sharing, valuable for considering new approaches with Melanoma Classification.</p>",
      "rawMarkdown": "Thank you for sharing, valuable for considering new approaches with Melanoma Classification.",
      "votes": 1
    },
    {
      "id": 930402,
      "postDate": "2020-07-15T12:41:59.943Z",
      "content": "<p>I am using EfficientNet B0.</p>\n\n<p>I am a bit mitigate with the dropout because the  results can change +/- 0.01 depending of the dropout/external dataset:\n- with External dataset  : (2018) dropout of 0.2 seems better\n- without External dataset : no dropout is better </p>\n\n<p>But it does not seem reproductible, I mean each time I launch a training I can have a different results .</p>",
      "rawMarkdown": "I am using EfficientNet B0.\n\nI am a bit mitigate with the dropout because the  results can change +/- 0.01 depending of the dropout/external dataset:\n- with External dataset  : (2018) dropout of 0.2 seems better\n- without External dataset : no dropout is better \n\nBut it does not seem reproductible, I mean each time I launch a training I can have a different results .",
      "votes": 1
    },
    {
      "id": 929218,
      "postDate": "2020-07-14T14:37:45.457Z",
      "content": "<p>Thank you! I am thinking of going for either resnet or efficient net for this! Any suggestions?</p>",
      "rawMarkdown": "Thank you! I am thinking of going for either resnet or efficient net for this! Any suggestions?",
      "votes": 1,
      "replies": [
        {
          "id": 929571,
          "postDate": "2020-07-14T18:50:16.767Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 929767,
          "postDate": "2020-07-14T23:39:07.867Z",
          "content": "<p><a href=\"/epocxy\">@epocxy</a>    You do:</p>\n\n<pre><code>def forward(self,x):\n    x = self.block(x)\n    #x = x.view(*(x.shape[:-2]),-1).mean(-1) ## For GAP\n    x = torch.cat([self.mp(x), self.ap(x)], 1)\n    x = x.view(x.size(0), -1)\n    x = self.fc(x)\n    return x\n</code></pre>\n\n<p>What do you think \"x = x.view(x.size(0), -1)\" is doing there for you?  I don't think its doing anything.  After your cat(), your data is BatchSize, (Size of mp + size of ap).  After your view() call, it should be the same.</p>",
          "rawMarkdown": "@epocxy    You do:\n\n    def forward(self,x):\n        x = self.block(x)\n        #x = x.view(*(x.shape[:-2]),-1).mean(-1) ## For GAP\n        x = torch.cat([self.mp(x), self.ap(x)], 1)\n        x = x.view(x.size(0), -1)\n        x = self.fc(x)\n        return x\n\nWhat do you think \"x = x.view(x.size(0), -1)\" is doing there for you?  I don't think its doing anything.  After your cat(), your data is BatchSize, (Size of mp + size of ap).  After your view() call, it should be the same."
        },
        {
          "id": 930073,
          "postDate": "2020-07-15T07:10:10.953Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 930096,
          "postDate": "2020-07-15T07:35:41.473Z",
          "content": "<p>@<em>dash</em> Oh I see, I thought you just clipped the fc layer from ResNet (Which would have already left you a flattened output), but I see you took out the part where it flattens as well.  </p>",
          "rawMarkdown": "@_dash_ Oh I see, I thought you just clipped the fc layer from ResNet (Which would have already left you a flattened output), but I see you took out the part where it flattens as well.  "
        }
      ]
    },
    {
      "id": 926078,
      "postDate": "2020-07-12T13:30:40.333Z",
      "content": "<p>Loved it!\nEspecially since Focal loss is highly spoken yet no one mentions that it didn't work for them.</p>",
      "rawMarkdown": "Loved it!\nEspecially since Focal loss is highly spoken yet no one mentions that it didn't work for them.",
      "votes": 1
    },
    {
      "id": 923550,
      "postDate": "2020-07-11T01:01:06.067Z",
      "content": "<p>Regarding FocalLoss, what implementation did you use? Did you set gamma? If so to what?</p>\n\n<p>I have had some trouble just \"dropping in\" Focal Loss.  What other changes did you have to make when you swapped in Focal Loss for BCE?  For example, I had to make sure my ground truths (y) was int64, because it didn't like them as float32.  </p>",
      "rawMarkdown": "Regarding FocalLoss, what implementation did you use? Did you set gamma? If so to what?\n\nI have had some trouble just \"dropping in\" Focal Loss.  What other changes did you have to make when you swapped in Focal Loss for BCE?  For example, I had to make sure my ground truths (y) was int64, because it didn't like them as float32.  ",
      "votes": 1,
      "replies": [
        {
          "id": 924174,
          "postDate": "2020-07-11T09:51:35.267Z",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> I am using <a href=\"https://github.com/artemmavrin/focal-loss\">this focal loss</a> with gamma=2. The other changes that I tried to make are: \n* Tweak the learning rate and try different optimizers: Adam and RMSProp. But I am definitely doing something wrong.\nDid focal loss give you better results than BCE? </p>",
          "rawMarkdown": "@brianfeeny I am using [this focal loss](https://github.com/artemmavrin/focal-loss) with gamma=2. The other changes that I tried to make are: \n* Tweak the learning rate and try different optimizers: Adam and RMSProp. But I am definitely doing something wrong.\nDid focal loss give you better results than BCE? "
        },
        {
          "id": 924251,
          "postDate": "2020-07-11T10:50:55.530Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks\n"
        },
        {
          "id": 926471,
          "postDate": "2020-07-12T18:39:24.280Z",
          "content": "<p>@amin slightly better with Focal Loss.  I am also trying Focal Loss with class weights.  If you are going to try Adam why not try AdamW, and did you turn on Amsgrad (for Adam)?</p>",
          "rawMarkdown": "@amin slightly better with Focal Loss.  I am also trying Focal Loss with class weights.  If you are going to try Adam why not try AdamW, and did you turn on Amsgrad (for Adam)?",
          "votes": 1
        },
        {
          "id": 929624,
          "postDate": "2020-07-14T19:54:35.613Z",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> no I didn't turn on Amsgrad for Adam, does it help with focal loss? </p>\n\n<p>And what learning rate are you using? </p>",
          "rawMarkdown": "@brianfeeny no I didn't turn on Amsgrad for Adam, does it help with focal loss? \n\nAnd what learning rate are you using? "
        },
        {
          "id": 929711,
          "postDate": "2020-07-14T22:00:46.933Z",
          "content": "<p>@Amin I have done numerous LR experiments.  The thing is, LR is very sensitive to other parameters such as batch size, and probably which model and image size, etc.  For me .0001 is best........I am actually using a scheduler (OneCycleLR) and just set .0003 as the max.  I have tried other schedulers too and they really help but require some tuning to find the best range.</p>\n\n<p>LR is pretty easy to nail down though.  Just run a few epochs and look at loss and roc/auc.  If its unstable lower LR.  If its not going anywhere quickly, increase LR.  I do like .01, .005, .001, .0005, .0001 etc to start with my \"lr search\".  How about you, what have you found is good for LR?</p>",
          "rawMarkdown": "@Amin I have done numerous LR experiments.  The thing is, LR is very sensitive to other parameters such as batch size, and probably which model and image size, etc.  For me .0001 is best........I am actually using a scheduler (OneCycleLR) and just set .0003 as the max.  I have tried other schedulers too and they really help but require some tuning to find the best range.\n\nLR is pretty easy to nail down though.  Just run a few epochs and look at loss and roc/auc.  If its unstable lower LR.  If its not going anywhere quickly, increase LR.  I do like .01, .005, .001, .0005, .0001 etc to start with my \"lr search\".  How about you, what have you found is good for LR?"
        }
      ]
    },
    {
      "id": 922184,
      "postDate": "2020-07-09T21:36:53.813Z",
      "content": "<p>Could you explain 2:</p>\n\n<p>\"2. Dense head: Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score.\"</p>\n\n<p>I assume you take some outputs from CNN then take features from csv and then use Dense layer, and your experiences are that this should be single linear layer without Dropout or BN? Can you clarify?</p>",
      "rawMarkdown": "Could you explain 2:\n\n\"2. Dense head: Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score.\"\n\nI assume you take some outputs from CNN then take features from csv and then use Dense layer, and your experiences are that this should be single linear layer without Dropout or BN? Can you clarify?\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 922542,
          "postDate": "2020-07-10T07:27:10.040Z",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> By dense head I mean a heavier head of the net, more and different layers on the top of the backbone. <strong>Just for image data!</strong> I haven't tried to concatenate the metadata+image yet.</p>\n\n<p><code>\nmodel = tf.keras.Sequential([\n        efn.EfficientNetB7(\n            input_shape=(*IMAGE_SIZE, 3),\n            weights='imagenet',\n            include_top=False\n        ),\n        tf.keras.layers.GlobalAveragePooling2D()\n        tf.keras.layers.Dense(256, activation='relu'), \n        tf.keras.layers.Dropout(0.2), \n        tf.keras.layers.Dense(128, activation='relu'), \n        tf.keras.layers.Dropout(0.2), \n        tf.keras.layers.Dense(1, activation='sigmoid')\n    ])\n</code>\nWhat I meant is: Adding more dense layers lowered CV/LB and introducing Dropout/BatchNorm in between dense layers didn't improve the score.</p>",
          "rawMarkdown": "@jacekpoplawski By dense head I mean a heavier head of the net, more and different layers on the top of the backbone. **Just for image data!** I haven't tried to concatenate the metadata+image yet.\n\n\n```\nmodel = tf.keras.Sequential([\n        efn.EfficientNetB7(\n            input_shape=(*IMAGE_SIZE, 3),\n            weights='imagenet',\n            include_top=False\n        ),\n        tf.keras.layers.GlobalAveragePooling2D()\n        tf.keras.layers.Dense(256, activation='relu'), \n        tf.keras.layers.Dropout(0.2), \n        tf.keras.layers.Dense(128, activation='relu'), \n        tf.keras.layers.Dropout(0.2), \n        tf.keras.layers.Dense(1, activation='sigmoid')\n    ])\n```\nWhat I meant is: Adding more dense layers lowered CV/LB and introducing Dropout/BatchNorm in between dense layers didn't improve the score."
        },
        {
          "id": 922559,
          "postDate": "2020-07-10T07:46:12.323Z",
          "content": "<p><a href=\"/amiiiney\">@amiiiney</a> yes I understand, however my question was - what is the \"good\" version of code, and which one is \"too much\", in your code there are THREE Dense layers plus two dropouts, so it is version which is best or is it version which is too much?</p>",
          "rawMarkdown": "@amiiiney yes I understand, however my question was - what is the \"good\" version of code, and which one is \"too much\", in your code there are THREE Dense layers plus two dropouts, so it is version which is best or is it version which is too much?"
        },
        {
          "id": 922688,
          "postDate": "2020-07-10T09:18:00.813Z",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> Those results are for an efficientNet B7 trained over 23 epochs</p>\n\n<p>1 Dense(1 output, sigmoid) LB= 0.937\n2 Dense (128, 1) LB= 0.935\n3 Dense (256, 128, 1) LB=0.934</p>\n\n<p>All of them are trained over 23 epochs. I have a feeling that maybe models with heavier heads need more training epochs to give better results, this is something I haven't tried yet because I am limited to the kaggle 3H TPU sessions. I would like to know the experience of the others with [heavier head + more epochs]</p>\n\n<p>Adding dropouts (p=0.1, 0.2, 0.3) and Batch normalization gave exactly the same LB scores!</p>\n\n<p>I would like to know your opinion about this! :)</p>",
          "rawMarkdown": "@jacekpoplawski Those results are for an efficientNet B7 trained over 23 epochs\n\n1 Dense(1 output, sigmoid) LB= 0.937\n2 Dense (128, 1) LB= 0.935\n3 Dense (256, 128, 1) LB=0.934\n\nAll of them are trained over 23 epochs. I have a feeling that maybe models with heavier heads need more training epochs to give better results, this is something I haven't tried yet because I am limited to the kaggle 3H TPU sessions. I would like to know the experience of the others with [heavier head + more epochs]\n\nAdding dropouts (p=0.1, 0.2, 0.3) and Batch normalization gave exactly the same LB scores!\n\nI would like to know your opinion about this! :)\n",
          "votes": 1
        },
        {
          "id": 922697,
          "postDate": "2020-07-10T09:24:22.023Z",
          "content": "<p>Thanks for the numbers, could you show the code for first case? Because I want to be sure we are talking about same thing. Is it just GlobalAveragePooling2D and Dense with 1 output?</p>\n\n<p>One more thing - in your code there is just Dense operating on efficientnet output, what about metadata (sex, age, area)? I would like to see whole dense code of your solution otherwise I think some information may be incomplete, like maybe you area talking about one Dense here then another Dense later.</p>",
          "rawMarkdown": "Thanks for the numbers, could you show the code for first case? Because I want to be sure we are talking about same thing. Is it just GlobalAveragePooling2D and Dense with 1 output?\n\nOne more thing - in your code there is just Dense operating on efficientnet output, what about metadata (sex, age, area)? I would like to see whole dense code of your solution otherwise I think some information may be incomplete, like maybe you area talking about one Dense here then another Dense later."
        },
        {
          "id": 922717,
          "postDate": "2020-07-10T09:41:33.103Z",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> Exactly!! just GlobalAveragePooling2D + Dense with 1 output\n<code>\n model = tf.keras.Sequential([\n        efn.EfficientNetB0(\n            input_shape=(*IMAGE_SIZE, 3),\n            weights='imagenet',\n            include_top=False\n        ),\n        L.GlobalAveragePooling2D(),\n        L.Dense(1, activation='sigmoid')\n    ])\n</code></p>\n\n<p>I didn't include metadata, the dense layers I am using are <strong>JUST FOR IMAGE DATA</strong>. </p>\n\n<p>I run the same experiment with B0 over 45 epochs and again the the models with heavier heads are always behind the model with just GlobalAveragePooling2D + Dense with 1 output.</p>",
          "rawMarkdown": "@jacekpoplawski Exactly!! just GlobalAveragePooling2D + Dense with 1 output\n```\n model = tf.keras.Sequential([\n        efn.EfficientNetB0(\n            input_shape=(*IMAGE_SIZE, 3),\n            weights='imagenet',\n            include_top=False\n        ),\n        L.GlobalAveragePooling2D(),\n        L.Dense(1, activation='sigmoid')\n    ])\n```\n\nI didn't include metadata, the dense layers I am using are **JUST FOR IMAGE DATA**. \n\nI run the same experiment with B0 over 45 epochs and again the the models with heavier heads are always behind the model with just GlobalAveragePooling2D + Dense with 1 output."
        },
        {
          "id": 922719,
          "postDate": "2020-07-10T09:42:05.743Z",
          "content": "<p>\"I didn't include metadata, the dense layers I am using are JUST FOR IMAGE DATA. \"</p>\n\n<p>so how do you use meta data?</p>",
          "rawMarkdown": "\"I didn't include metadata, the dense layers I am using are JUST FOR IMAGE DATA. \"\n\nso how do you use meta data?"
        },
        {
          "id": 922748,
          "postDate": "2020-07-10T10:05:17.263Z",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  I am training a separate XGB model and ensemble with the CNN trained on image data. </p>",
          "rawMarkdown": "@jacekpoplawski  I am training a separate XGB model and ensemble with the CNN trained on image data. "
        },
        {
          "id": 922753,
          "postDate": "2020-07-10T10:13:37.157Z",
          "content": "<p>I see, I was thinking you are talking about Dense layers used to merge CNN data and meta data, in case of CNN data only I see no point in adding more Dense layers.</p>",
          "rawMarkdown": "I see, I was thinking you are talking about Dense layers used to merge CNN data and meta data, in case of CNN data only I see no point in adding more Dense layers."
        },
        {
          "id": 922757,
          "postDate": "2020-07-10T10:18:25.703Z",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>! Do you think merging CNN data and meta data is better than training 2 separate models as I am doing?</p>",
          "rawMarkdown": "@jacekpoplawski! Do you think merging CNN data and meta data is better than training 2 separate models as I am doing?"
        },
        {
          "id": 922786,
          "postDate": "2020-07-10T10:36:21.367Z",
          "content": "<p>Theoretically it makes more sense to train with the metadata <a href=\"/amiiiney\">@amiiiney</a> (edit: so no separate models), as your model could condition (e.g. use different parts of the network) based on the metadata. But then again, this is deep learning, it's all alchemy and there's no such thing as theory.</p>",
          "rawMarkdown": "Theoretically it makes more sense to train with the metadata @amiiiney (edit: so no separate models), as your model could condition (e.g. use different parts of the network) based on the metadata. But then again, this is deep learning, it's all alchemy and there's no such thing as theory.",
          "votes": 2
        },
        {
          "id": 922794,
          "postDate": "2020-07-10T10:42:00.833Z",
          "content": "<p>in my opinion training separate model for meta data is pointness, but this is just my opinion, I never tried it</p>",
          "rawMarkdown": "in my opinion training separate model for meta data is pointness, but this is just my opinion, I never tried it",
          "votes": 1
        },
        {
          "id": 922863,
          "postDate": "2020-07-10T11:21:21.633Z",
          "content": "<p>Thanks <a href=\"/group16\">@group16</a> that explains a lot of things!\nDo you mind if I ask what is your best single model (CNN with metadata) LB score?</p>",
          "rawMarkdown": "Thanks @group16 that explains a lot of things!\nDo you mind if I ask what is your best single model (CNN with metadata) LB score?",
          "votes": 1
        },
        {
          "id": 922899,
          "postDate": "2020-07-10T12:02:11.073Z",
          "content": "<p>You should generate multiple outputs from CNN then merge with multiple features then Dense layers can think how to combine them - that was my point all the time, do you mean only one Dense layer is enough here, in my opinion no.</p>",
          "rawMarkdown": "You should generate multiple outputs from CNN then merge with multiple features then Dense layers can think how to combine them - that was my point all the time, do you mean only one Dense layer is enough here, in my opinion no.",
          "votes": 1
        },
        {
          "id": 922902,
          "postDate": "2020-07-10T12:04:39.183Z",
          "content": "<p>Of course you may ask <a href=\"/amiiiney\">@amiiiney</a>  :). Currently I'm only at 0.937</p>",
          "rawMarkdown": "Of course you may ask @amiiiney  :). Currently I'm only at 0.937",
          "votes": 1
        },
        {
          "id": 923551,
          "postDate": "2020-07-11T01:04:31.993Z",
          "content": "<p>Yes, that would make it more clear.  I assume perhaps he is saying that he took the date right out of the  main model and directly into 1 feature? </p>",
          "rawMarkdown": "Yes, that would make it more clear.  I assume perhaps he is saying that he took the date right out of the  main model and directly into 1 feature? ",
          "votes": 1
        },
        {
          "id": 923553,
          "postDate": "2020-07-11T01:07:15.510Z",
          "content": "<p>so when you say:</p>\n\n<p>1 Dense(1 output, sigmoid) LB= 0.937</p>\n\n<p>you mean you took the output of B7 which is 2560 and put it right into a single feature and that did better than adding additional layers?</p>",
          "rawMarkdown": "so when you say:\n\n1 Dense(1 output, sigmoid) LB= 0.937\n\n\nyou mean you took the output of B7 which is 2560 and put it right into a single feature and that did better than adding additional layers?",
          "votes": 1
        },
        {
          "id": 924068,
          "postDate": "2020-07-11T08:37:05.637Z",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> <a href=\"/jacekpoplawski\">@jacekpoplawski</a> Sorry if my explanation wasn't clear enough. Here what's I am doing:</p>\n\n<ul>\n<li><p>1 Dense(1, sigmoid) just like this public notebook: <a href=\"https://www.kaggle.com/manojprabhaakr/melanoma-tpu-starter-efficientnet-b0\">Melanoma TPU Starter EfficientNet B0: Version2</a> gives LB=0.937</p></li>\n<li><p>A heavier head of the net more or less like this public notebook: <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">Melanoma TPU EfficientNet B5_dense_head</a> gives LB=0.935 and LB=0.934 (more layers= lower score).</p></li>\n</ul>\n\n<p>I am training all of the them over the same number of epochs so probably the models with heavier heads need more epochs!?! \nI tried layers of (512, 256 and 128). Now I am trying with more neurons (1024) as discussed above and <a href=\"https://www.kaggle.com/aziz69/efficientnets-meta-data-augs\">in this notebook</a> and apparently adding one layer of 1024 neurons + 0.4 dropout does better.</p>",
          "rawMarkdown": "@brianfeeny @jacekpoplawski Sorry if my explanation wasn't clear enough. Here what's I am doing:\n\n* 1 Dense(1, sigmoid) just like this public notebook: [Melanoma TPU Starter EfficientNet B0: Version2](https://www.kaggle.com/manojprabhaakr/melanoma-tpu-starter-efficientnet-b0) gives LB=0.937\n\n* A heavier head of the net more or less like this public notebook: [Melanoma TPU EfficientNet B5_dense_head](https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head) gives LB=0.935 and LB=0.934 (more layers= lower score).\n\nI am training all of the them over the same number of epochs so probably the models with heavier heads need more epochs!?! \nI tried layers of (512, 256 and 128). Now I am trying with more neurons (1024) as discussed above and [in this notebook](https://www.kaggle.com/aziz69/efficientnets-meta-data-augs) and apparently adding one layer of 1024 neurons + 0.4 dropout does better."
        },
        {
          "id": 929747,
          "postDate": "2020-07-14T23:16:26.623Z",
          "content": "<p>@Amin you mention training for 23 epochs (on more involved Dense heads).  Do you find you are seeing improvements even at 23 epochs? What are you measuring to decide how many epochs: loss or roc/auc?  I assume you are not using any sort of early stopping/patience.  I rarely see improvements past 10 epochs, I typically deploy some sort of early stopping.  I AM using a scheduler which modifies/drops learning rates if no improvement.  How about you?</p>",
          "rawMarkdown": "@Amin you mention training for 23 epochs (on more involved Dense heads).  Do you find you are seeing improvements even at 23 epochs? What are you measuring to decide how many epochs: loss or roc/auc?  I assume you are not using any sort of early stopping/patience.  I rarely see improvements past 10 epochs, I typically deploy some sort of early stopping.  I AM using a scheduler which modifies/drops learning rates if no improvement.  How about you?"
        }
      ]
    },
    {
      "id": 921981,
      "postDate": "2020-07-09T17:41:58.933Z",
      "content": "<p>Regarding #4, is your XGB using just meta features from the provided CSV file (like Giba's model), or are you making more meta features like image width and height?</p>",
      "rawMarkdown": "Regarding #4, is your XGB using just meta features from the provided CSV file (like Giba's model), or are you making more meta features like image width and height?",
      "votes": 1,
      "replies": [
        {
          "id": 921993,
          "postDate": "2020-07-09T18:00:34.707Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> the XGB using just the meta features provided, no feature engineering. The thing is not just my XGB but also the other public kernels dealing with metadata, they all have higher LB scores +70 but they don't do well in the ensemble. </p>",
          "rawMarkdown": "@cdeotte the XGB using just the meta features provided, no feature engineering. The thing is not just my XGB but also the other public kernels dealing with metadata, they all have higher LB scores +70 but they don't do well in the ensemble. "
        },
        {
          "id": 922001,
          "postDate": "2020-07-09T18:09:29.923Z",
          "content": "<p>It sounds like your XGB may be overfitting the train data. And many of these public notebooks probably overfit too.</p>\n\n<p>Also regarding other public notebooks, I have noticed that there are strong correlations between image width height and target in the train data. But these same correlations are not present in the test data.</p>\n\n<p>For example most images in train data with original height width <code>4000x6000</code> have no malignant (and there are 14703 of these images!) but this is not true for test data (which has 4162 of these images).</p>",
          "rawMarkdown": "It sounds like your XGB may be overfitting the train data. And many of these public notebooks probably overfit too.\n\nAlso regarding other public notebooks, I have noticed that there are strong correlations between image width height and target in the train data. But these same correlations are not present in the test data.\n\nFor example most images in train data with original height width `4000x6000` have no malignant (and there are 14703 of these images!) but this is not true for test data (which has 4162 of these images).",
          "votes": 2
        },
        {
          "id": 922005,
          "postDate": "2020-07-09T18:14:06.680Z",
          "content": "<p>Giba's <a href=\"https://www.kaggle.com/titericz/simple-baseline\">notebook</a> uses Bayesian smoothing to prevent overfitting. His models ignore feature values which occur rarely. And only consider the most prominent patterns.</p>\n\n<pre><code>te['ll'] = ((te['mean']*te['count'])+(M*L))/(te['count']+L)\n</code></pre>",
          "rawMarkdown": "Giba's [notebook][1] uses Bayesian smoothing to prevent overfitting. His models ignore feature values which occur rarely. And only consider the most prominent patterns.\n\n    te['ll'] = ((te['mean']*te['count'])+(M*L))/(te['count']+L)\n\n[1]: https://www.kaggle.com/titericz/simple-baseline",
          "votes": 3
        },
        {
          "id": 922044,
          "postDate": "2020-07-09T18:54:21.810Z",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> that explains why the public notebooks are not doing well. I tried some other features, they all increase CV but decrease LB, it's frustrating but I still believe that there should be some features out there that could make the metadata more useful!</p>",
          "rawMarkdown": "Thanks @cdeotte that explains why the public notebooks are not doing well. I tried some other features, they all increase CV but decrease LB, it's frustrating but I still believe that there should be some features out there that could make the metadata more useful!"
        }
      ]
    },
    {
      "id": 921904,
      "postDate": "2020-07-09T16:34:59.270Z",
      "content": "<p>I'm confused. Before your 4 numbered points, you have 4 bullet points. The bullet points include using EfficientNet and sizes 256, 512, 768, 1024. Are you saying that EfficientNet and sizes 256, 512, 768, 1024 don't help?</p>",
      "rawMarkdown": "I'm confused. Before your 4 numbered points, you have 4 bullet points. The bullet points include using EfficientNet and sizes 256, 512, 768, 1024. Are you saying that EfficientNet and sizes 256, 512, 768, 1024 don't help?",
      "votes": 1,
      "replies": [
        {
          "id": 921910,
          "postDate": "2020-07-09T16:42:20.607Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 921924,
          "postDate": "2020-07-09T16:52:32.490Z",
          "content": "<p>I had to restructure my thread because I was asked what kind of models and parameters led the the things that didn't work. \nSo the bullet points are the \"materials and methods\".\nAnd the numbered points are the things that didn't work.\nSorry <a href=\"/cdeotte\">@cdeotte</a> for confusing you :) I will reorganize the thread again :)</p>\n\n<p>I would love to know your opinion about this topic, especially point number 4 :)</p>",
          "rawMarkdown": "I had to restructure my thread because I was asked what kind of models and parameters led the the things that didn't work. \nSo the bullet points are the \"materials and methods\".\nAnd the numbered points are the things that didn't work.\nSorry @cdeotte for confusing you :) I will reorganize the thread again :)\n\nI would love to know your opinion about this topic, especially point number 4 :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 921730,
      "postDate": "2020-07-09T14:01:27.650Z",
      "content": "<p>oh , well done bro</p>",
      "rawMarkdown": "oh , well done bro",
      "votes": 1
    },
    {
      "id": 929806,
      "postDate": "2020-07-15T01:18:03.957Z",
      "content": "<p><a href=\"/amiiiney\">@amiiiney</a> Thanks for presenting results and observations of so many simulations. </p>",
      "rawMarkdown": "@amiiiney Thanks for presenting results and observations of so many simulations. ",
      "votes": 2
    },
    {
      "id": 925268,
      "postDate": "2020-07-12T00:47:19.283Z",
      "content": "<p>I have to tell that this post has in some way ruined my already sleepless nights (I'm a newcomer to these competitions,). In the sense that it has challenged some (wrong ?)  results I had arrived to (but probably due to timing not fully checked).\nI can confirm that with a very simple network (EfficientNet B7 and only a Dense(1, sigmoid), without metadata (only image) you can get LB = 0.935. That is more all less the best single model (no ensemble) result I have achieved so far (well I have 0.936 with metadata and a more complicated network, but the difference could be not significant).</p>\n\n<p>But in my case, with smaller (B4) I have checked a 10% improvement with FocalLoss, and now I'm running with FocalLoss. \nThe biggest problem I see is that all the networks from B4 are overfitting the training set. And I don't see improvements with more than 10 epochs!!\nIf you add a longer Dense head it is worse, unless you go with dropout=0.5. With a so simple net (well, simple since you can't modify B7 and add dropout inside) the only thing is to add images to the training set (and I'm using the full Deotte's set).\nI'm going to try switching the loss function. If I get an improvement... what should I say?: 'Less is more' (and I have been always a fan of the old Apple motto).\nI'll soon post the entire Notebook.</p>",
      "rawMarkdown": "I have to tell that this post has in some way ruined my already sleepless nights (I'm a newcomer to these competitions,). In the sense that it has challenged some (wrong ?)  results I had arrived to (but probably due to timing not fully checked).\nI can confirm that with a very simple network (EfficientNet B7 and only a Dense(1, sigmoid), without metadata (only image) you can get LB = 0.935. That is more all less the best single model (no ensemble) result I have achieved so far (well I have 0.936 with metadata and a more complicated network, but the difference could be not significant).\n\nBut in my case, with smaller (B4) I have checked a 10% improvement with FocalLoss, and now I'm running with FocalLoss. \nThe biggest problem I see is that all the networks from B4 are overfitting the training set. And I don't see improvements with more than 10 epochs!!\nIf you add a longer Dense head it is worse, unless you go with dropout=0.5. With a so simple net (well, simple since you can't modify B7 and add dropout inside) the only thing is to add images to the training set (and I'm using the full Deotte's set).\nI'm going to try switching the loss function. If I get an improvement... what should I say?: 'Less is more' (and I have been always a fan of the old Apple motto).\nI'll soon post the entire Notebook.",
      "votes": 2,
      "replies": [
        {
          "id": 925289,
          "postDate": "2020-07-12T01:47:12.033Z",
          "content": "<p>hello, could you clarify:</p>\n\n<ul>\n<li>\"a very simple network (EfficientNet B7\"</li>\n<li>\"all the networks from B4 are overfitting\"</li>\n</ul>\n\n<p>why do you call B7 simple? do you mean B7 or B0? and if you are saying that all from B4 are overfitting I assume you mean B7 too? do you mean B4-B7 or B4-B0?</p>",
          "rawMarkdown": "hello, could you clarify:\n\n- \"a very simple network (EfficientNet B7\"\n- \"all the networks from B4 are overfitting\"\n\nwhy do you call B7 simple? do you mean B7 or B0? and if you are saying that all from B4 are overfitting I assume you mean B7 too? do you mean B4-B7 or B4-B0?\n\n\n",
          "votes": 2
        },
        {
          "id": 925343,
          "postDate": "2020-07-12T03:04:19.607Z",
          "content": "<p>HI Jacek. I see your point: by itself B7 is NOT simple. What is simple is my network for this post (since the magic of transfer learning obviously): it is simply B7 followed by a Dense(1)</p>\n\n<pre><code>with strategy.scope():\n    inp = tf.keras.layers.Input(shape=(dim, dim,3))\n\n    base = efns7(input_shape=(dim, dim, 3),weights='imagenet',include_top=False)\n\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(1,activation='sigmoid')(x)\n\n    model = tf.keras.Model(inputs=inp,outputs=x)\n</code></pre>\n\n<p>from my tests all the networks from B4 to B7 are overfitting: at the end of the training, the training accuracy is 0.99, and validation accuracy is 0.93. From this point of view adding capacity to the network, with a more complex Dense head normally doesn't help. I have seen very small improvements using dropout(5) but results vary and I don't think are useful. </p>\n\n<p>So, for now I agree that simpler is better.</p>",
          "rawMarkdown": "HI Jacek. I see your point: by itself B7 is NOT simple. What is simple is my network for this post (since the magic of transfer learning obviously): it is simply B7 followed by a Dense(1)\n\n    with strategy.scope():\n        inp = tf.keras.layers.Input(shape=(dim, dim,3))\n        \n        base = efns7(input_shape=(dim, dim, 3),weights='imagenet',include_top=False)\n    \n        x = base(inp)\n        x = tf.keras.layers.GlobalAveragePooling2D()(x)\n        x = tf.keras.layers.Dense(1,activation='sigmoid')(x)\n    \n        model = tf.keras.Model(inputs=inp,outputs=x)\n    \n from my tests all the networks from B4 to B7 are overfitting: at the end of the training, the training accuracy is 0.99, and validation accuracy is 0.93. From this point of view adding capacity to the network, with a more complex Dense head normally doesn't help. I have seen very small improvements using dropout(5) but results vary and I don't think are useful. \n\nSo, for now I agree that simpler is better.\n",
          "votes": 1
        },
        {
          "id": 939695,
          "postDate": "2020-07-22T11:39:44.090Z",
          "content": "<p><a href=\"/luigisaetta\">@luigisaetta</a>  Thanks for clearing up things..! Can you please tell that is this ( 0.93 Acc. ) model is your single best model and you are now using ensembles to score higher. If yes then how many models you are using ?\nIn my case I achieved 93.70 using B6 with 2018 external data and focal loss + little aug.</p>",
          "rawMarkdown": "@luigisaetta  Thanks for clearing up things..! Can you please tell that is this ( 0.93 Acc. ) model is your single best model and you are now using ensembles to score higher. If yes then how many models you are using ?\nIn my case I achieved 93.70 using B6 with 2018 external data and focal loss + little aug."
        }
      ]
    },
    {
      "id": 922578,
      "postDate": "2020-07-10T07:58:48.557Z",
      "content": "<p>Thanks for sharing. I'm quite the opposite of you 0.4 ratio of CutMix with gridMask works well for me, focal loss improves the score too, adding dense head improves but it took me a lot of time to tweak it. My strategy is the get the best score I can get with a single model (0.933 so far ) then ensemble but I'm kinda stuck I don't know what else to try now (I've tried hair augmentations =&gt; improves a little bit, I also tried microscope aug =&gt; doesn't improve) Any help on how to improve my single model to get over 0.933 would be appreciated </p>",
      "rawMarkdown": "Thanks for sharing. I'm quite the opposite of you 0.4 ratio of CutMix with gridMask works well for me, focal loss improves the score too, adding dense head improves but it took me a lot of time to tweak it. My strategy is the get the best score I can get with a single model (0.933 so far ) then ensemble but I'm kinda stuck I don't know what else to try now (I've tried hair augmentations =&gt; improves a little bit, I also tried microscope aug =&gt; doesn't improve) Any help on how to improve my single model to get over 0.933 would be appreciated ",
      "votes": 2,
      "replies": [
        {
          "id": 922740,
          "postDate": "2020-07-10T10:00:52.470Z",
          "content": "<p>Thanks <a href=\"/aziz69\">@aziz69</a> for you feedback. I am glad to see that those things worked for you!\n* Do you mind if I ask which focal loss are you using? I am using <a href=\"https://github.com/artemmavrin/focal-loss\">this focal loss</a> and which learning rate, gamma and alpha parameters gave you the best result?</p>\n\n<ul>\n<li><p>How many layers and which kind of layers do you have in your head? In my case more layers lowered the score and including dropouts and batch normalization didn't help much (you can see my dense head in the conversation below).</p></li>\n<li><p>I was planning to experiment with gridMask next week. Did the combo CutMix+GridMask give a better score than GridMask alone?</p></li>\n</ul>\n\n<p>Which model and for how many epochs are you training your 0.933 model? Ideas that you didn't mention and could increase your score:\n* Did you try <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">CAM CutMix explained in this post</a> \n* including the metadata to your CNN by concatenating the image+metadata \n* Training for longer hours </p>",
          "rawMarkdown": "Thanks @aziz69 for you feedback. I am glad to see that those things worked for you!\n* Do you mind if I ask which focal loss are you using? I am using [this focal loss](https://github.com/artemmavrin/focal-loss) and which learning rate, gamma and alpha parameters gave you the best result?\n\n* How many layers and which kind of layers do you have in your head? In my case more layers lowered the score and including dropouts and batch normalization didn't help much (you can see my dense head in the conversation below).\n\n* I was planning to experiment with gridMask next week. Did the combo CutMix+GridMask give a better score than GridMask alone?\n\nWhich model and for how many epochs are you training your 0.933 model? Ideas that you didn't mention and could increase your score:\n* Did you try [CAM CutMix explained in this post](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021) \n* including the metadata to your CNN by concatenating the image+metadata \n* Training for longer hours \n",
          "votes": 1
        },
        {
          "id": 922792,
          "postDate": "2020-07-10T10:41:03.587Z",
          "content": "<p>hello <a href=\"/aziz69\">@aziz69</a> \ncan you share the info about CutMix? how do you use it?</p>",
          "rawMarkdown": "hello @aziz69 \ncan you share the info about CutMix? how do you use it?"
        },
        {
          "id": 923005,
          "postDate": "2020-07-10T13:29:44.657Z",
          "content": "<p>Hey guys so this is the focal loss I used : \n* loss = tfa.losses.SigmoidFocalCrossEntropy(reduction=tf.keras.losses.Reduction.AUTO) with lr = 1e-5  (tfa =&gt; tensrflow addons) \n* using only one dense layer after you model ( number of neurons has to be tuned for 1024 + 0.4 dropout worked best ) \n* Yes gridMask + Cutmix combo is best then working with each one solo.\n* I'm using efficientNets for 20 epochs.</p>\n\n<p>I have tried including meta data model (xgb 0.73) with my best cnn submission (0.94) somehow it didn't add anything the score was lower even when i tried different coeffients. </p>\n\n<p>I haven't tried CAM cutmix so thanks for that I'll try it. \nThis is my notebook (I haven't commited for a while but most of my work is based on this) <a href=\"https://www.kaggle.com/aziz69/efficientnets-meta-data-augs\">https://www.kaggle.com/aziz69/efficientnets-meta-data-augs</a></p>",
          "rawMarkdown": "Hey guys so this is the focal loss I used : \n* loss = tfa.losses.SigmoidFocalCrossEntropy(reduction=tf.keras.losses.Reduction.AUTO) with lr = 1e-5  (tfa =&gt; tensrflow addons) \n* using only one dense layer after you model ( number of neurons has to be tuned for 1024 + 0.4 dropout worked best ) \n* Yes gridMask + Cutmix combo is best then working with each one solo.\n* I'm using efficientNets for 20 epochs.\n\n I have tried including meta data model (xgb 0.73) with my best cnn submission (0.94) somehow it didn't add anything the score was lower even when i tried different coeffients. \n\nI haven't tried CAM cutmix so thanks for that I'll try it. \nThis is my notebook (I haven't commited for a while but most of my work is based on this) https://www.kaggle.com/aziz69/efficientnets-meta-data-augs",
          "votes": 1
        }
      ]
    },
    {
      "id": 921686,
      "postDate": "2020-07-09T13:31:51.663Z",
      "content": "<p>Your 4th point makes me fear that we are only overfitting the Public LB.</p>\n\n<p>None of my experiments was a real success, CV vs LB is not consistent for me and similar CV give different LB...</p>\n\n<p>Would you mind explaining what's the custom scheduler everybody is using from the flowers competition?</p>",
      "rawMarkdown": "Your 4th point makes me fear that we are only overfitting the Public LB.\n\nNone of my experiments was a real success, CV vs LB is not consistent for me and similar CV give different LB...\n\nWould you mind explaining what's the custom scheduler everybody is using from the flowers competition?",
      "votes": 2,
      "replies": [
        {
          "id": 921715,
          "postDate": "2020-07-09T13:53:54.273Z",
          "content": "<p>Hey <a href=\"/optimo\">@optimo</a> \n* The <strong>4th point</strong> also got me thinking about what I might be doing wrong! I would like to know the experience of others with tabular data, because most of the discussion is about architectures and augmentations.\n* The custom scheduler from this baseline kernel in the flower comp.: <a href=\"https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule\">https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule</a> . Most public kernels are using it:\n```\nLR_START = 0.00001\nLR_MAX = 0.00005 * strategy.num_replicas_in_sync\nLR_MIN = 0.00001\nLR_RAMPUP_EPOCHS = 5\nLR_SUSTAIN_EPOCHS = 0\nLR_EXP_DECAY = .8</p>\n\n<p>def lrfn(epoch):\n    if epoch &lt; LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch &lt; LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN\n    return lr</p>\n\n<p>lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=True)\n```</p>",
          "rawMarkdown": "Hey @optimo \n* The **4th point** also got me thinking about what I might be doing wrong! I would like to know the experience of others with tabular data, because most of the discussion is about architectures and augmentations.\n* The custom scheduler from this baseline kernel in the flower comp.: https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule . Most public kernels are using it:\n```\nLR_START = 0.00001\nLR_MAX = 0.00005 * strategy.num_replicas_in_sync\nLR_MIN = 0.00001\nLR_RAMPUP_EPOCHS = 5\nLR_SUSTAIN_EPOCHS = 0\nLR_EXP_DECAY = .8\n\ndef lrfn(epoch):\n    if epoch &lt; LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch &lt; LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN\n    return lr\n    \nlr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=True)\n```"
        },
        {
          "id": 921719,
          "postDate": "2020-07-09T13:57:51.623Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 921721,
          "postDate": "2020-07-09T13:59:16.733Z",
          "content": "<p>thanks I'll have a look!</p>",
          "rawMarkdown": "thanks I'll have a look!"
        },
        {
          "id": 921733,
          "postDate": "2020-07-09T14:04:18.793Z",
          "content": "<p>hey <a href=\"/synked\">@synked</a> what do you mean by \"the other(image only) experiments\"?</p>\n\n<p>The unstability I was referring to is not specific to metadata usage. The CV vs LB seems unstable for image and image+metadata experiments I did.</p>",
          "rawMarkdown": "hey @synked what do you mean by \"the other(image only) experiments\"?\n\nThe unstability I was referring to is not specific to metadata usage. The CV vs LB seems unstable for image and image+metadata experiments I did."
        },
        {
          "id": 921747,
          "postDate": "2020-07-09T14:12:38.050Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 958092,
      "postDate": "2020-08-04T19:00:44.833Z",
      "content": "<p>great job !!</p>",
      "rawMarkdown": "great job !!"
    },
    {
      "id": 952285,
      "postDate": "2020-07-30T19:24:27.230Z",
      "content": "<p><a href=\"/amiiiney\">@amiiiney</a> are you using <code>kaggle TPU</code> for <code>1024</code> images? If yes, what <code>batch size</code> you are using and does your model train completely in one <code>TPU session(3 hrs)</code> or do you resume training ?</p>\n\n<p>Another thing is you mentioned that you are experimenting with 4 different image sizes so is there any specific reason of choosing these image sizes as you can choose other sizes as well like <code>128, 192, 384</code>?</p>",
      "rawMarkdown": "@amiiiney are you using `kaggle TPU` for `1024` images? If yes, what `batch size` you are using and does your model train completely in one `TPU session(3 hrs)` or do you resume training ?\n\nAnother thing is you mentioned that you are experimenting with 4 different image sizes so is there any specific reason of choosing these image sizes as you can choose other sizes as well like `128, 192, 384 `?"
    },
    {
      "id": 924215,
      "postDate": "2020-07-11T10:32:13.080Z",
      "content": "<p>ok</p>",
      "rawMarkdown": "ok"
    },
    {
      "id": 924154,
      "postDate": "2020-07-11T09:29:02.883Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice"
    },
    {
      "id": 923741,
      "postDate": "2020-07-11T05:56:17.177Z",
      "content": "<p>Nice work :)</p>",
      "rawMarkdown": "Nice work :)"
    },
    {
      "id": 929556,
      "postDate": "2020-07-14T18:42:41.203Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 929563,
          "postDate": "2020-07-14T18:45:49.100Z",
          "content": "<p>The 2019 data contains the 2018 and 2017 data. So half of 2019 data is new, and half of 2019 is those old comp data.</p>\n\n<p>Do you use all 2019? Some people have observed a difference between using the new half and old half. In my TFRecords, the even numbered records are 2018 2017 and the odd numbered records are the new 2019 half.</p>",
          "rawMarkdown": "The 2019 data contains the 2018 and 2017 data. So half of 2019 data is new, and half of 2019 is those old comp data.\n\nDo you use all 2019? Some people have observed a difference between using the new half and old half. In my TFRecords, the even numbered records are 2018 2017 and the odd numbered records are the new 2019 half.",
          "votes": 2
        },
        {
          "id": 929576,
          "postDate": "2020-07-14T18:57:52.073Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 929583,
          "postDate": "2020-07-14T19:04:32.193Z",
          "content": "<p>If you have the downloaded external 2019 JPEGS dataset, then you can separate the images using the <code>train.csv</code> contained within. All images with original size <code>1024x1024</code> as designated by the <code>width</code> and <code>height</code> columns are the new 2019 data and the other images are the old 2019 which is say 2018 2017 data</p>",
          "rawMarkdown": "If you have the downloaded external 2019 JPEGS dataset, then you can separate the images using the `train.csv` contained within. All images with original size `1024x1024` as designated by the `width` and `height` columns are the new 2019 data and the other images are the old 2019 which is say 2018 2017 data"
        },
        {
          "id": 929594,
          "postDate": "2020-07-14T19:16:39.800Z",
          "content": "<p>Hi Chris </p>\n\n<p>Thanks again for all amazing contributions. </p>\n\n<p>It seems the train csv file I found <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\">here</a> contains only 2020 data.  </p>\n\n<p>Could you please provided the link of CSV file containing 2020 + external data ?   Thanks again.   </p>",
          "rawMarkdown": "Hi Chris \n\nThanks again for all amazing contributions. \n\nIt seems the train csv file I found [here](https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv) contains only 2020 data.  \n\nCould you please provided the link of CSV file containing 2020 + external data ?   Thanks again.   \n\n",
          "votes": 1
        },
        {
          "id": 929595,
          "postDate": "2020-07-14T19:17:11.040Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 929603,
          "postDate": "2020-07-14T19:22:28.473Z",
          "content": "<p>&gt; Looklike there are 12414 records which have (h,w)--&gt;(1024,1024)so they are all related to 2019 only ?</p>\n\n<p>Yes. Those are the images that are in 2019 but not in 2018 2017. If you display them, you will see that they look different than the others. They are more zoomed in and have large dark spots.</p>",
          "rawMarkdown": "&gt; Looklike there are 12414 records which have (h,w)--&gt;(1024,1024)so they are all related to 2019 only ?\n\nYes. Those are the images that are in 2019 but not in 2018 2017. If you display them, you will see that they look different than the others. They are more zoomed in and have large dark spots."
        },
        {
          "id": 929605,
          "postDate": "2020-07-14T19:25:04.590Z",
          "content": "<p><a href=\"/serigne\">@serigne</a> My csv file naming is confusing because i call all my cvs <code>train.csv</code>. You need to use a <code>train.csv</code> from inside my external dataset (either JPEG or TFRecord). A direct link is <a href=\"https://www.kaggle.com/cdeotte/isic2019-256x256?select=train.csv\">here</a></p>",
          "rawMarkdown": "@serigne My csv file naming is confusing because i call all my cvs `train.csv`. You need to use a `train.csv` from inside my external dataset (either JPEG or TFRecord). A direct link is [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/isic2019-256x256?select=train.csv",
          "votes": 1
        },
        {
          "id": 929611,
          "postDate": "2020-07-14T19:41:01.027Z",
          "content": "<p>Thank <a href=\"/cdeotte\">@cdeotte</a>  !</p>\n\n<p>I forgot you made separated datasets. I got confused by another discussion Topic</p>",
          "rawMarkdown": "Thank @cdeotte  !\n\nI forgot you made separated datasets. I got confused by another discussion Topic",
          "votes": 1
        },
        {
          "id": 929628,
          "postDate": "2020-07-14T20:00:46.870Z",
          "content": "<p><a href=\"/epocxy\">@epocxy</a> yes I am using external data but in my case they increase the LB score, however the CV/LB with external data is not stable.</p>",
          "rawMarkdown": "@epocxy yes I am using external data but in my case they increase the LB score, however the CV/LB with external data is not stable."
        },
        {
          "id": 939683,
          "postDate": "2020-07-22T11:32:13.640Z",
          "content": "<p><a href=\"/amiiiney\">@amiiiney</a> my cv is 90.10 and lb is 93.70 when i use 2018 data + 2020 data. Is that CV that we can trust and except not much shakeup in the private lB?\nAlso, I want to know why there is so much gap? \n- There could not be any leak because I am using <a href=\"/cdeotte\">@cdeotte</a> 's triple stratified dataset.\nI think that after training using an external dataset the test data is simple for our model to classify. what do you say ?\n<a href=\"/cdeotte\">@cdeotte</a> <a href=\"/amiiiney\">@amiiiney</a> </p>",
          "rawMarkdown": "@amiiiney my cv is 90.10 and lb is 93.70 when i use 2018 data + 2020 data. Is that CV that we can trust and except not much shakeup in the private lB?\nAlso, I want to know why there is so much gap? \n- There could not be any leak because I am using @cdeotte 's triple stratified dataset.\nI think that after training using an external dataset the test data is simple for our model to classify. what do you say ?\n@cdeotte @amiiiney "
        }
      ]
    },
    {
      "id": 927496,
      "postDate": "2020-07-13T12:25:55.623Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 922372,
      "postDate": "2020-07-10T04:07:27.273Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    },
    {
      "id": 921726,
      "postDate": "2020-07-09T13:59:57.830Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 921739,
          "postDate": "2020-07-09T14:07:04.407Z",
          "content": "<p>Same as <a href=\"/optimo\">@optimo</a>, CV and LB are not stable, especially when I use external data, the gap gets bigger.</p>",
          "rawMarkdown": "Same as @optimo, CV and LB are not stable, especially when I use external data, the gap gets bigger.",
          "votes": 1
        },
        {
          "id": 921740,
          "postDate": "2020-07-09T14:07:52.517Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 921745,
          "postDate": "2020-07-09T14:12:19.163Z",
          "content": "<p>Yes both of them! I am using 5-folds. I should probably try 10-folds.</p>",
          "rawMarkdown": "Yes both of them! I am using 5-folds. I should probably try 10-folds."
        },
        {
          "id": 921749,
          "postDate": "2020-07-09T14:13:43.320Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 921758,
          "postDate": "2020-07-09T14:25:12.443Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 921765,
          "postDate": "2020-07-09T14:32:17.183Z",
          "content": "<p>I didn't because in this discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">How To Use Last Years 2019 Comp Data (and 2018, 2017)</a> Chris detected the duplicates and said it's safe to leave them in because they are just 59 out of 30,000 images. I think increasing the folds number might help to stabilize CV/LB.</p>",
          "rawMarkdown": "I didn't because in this discussion [How To Use Last Years 2019 Comp Data (and 2018, 2017)]( https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910) Chris detected the duplicates and said it's safe to leave them in because they are just 59 out of 30,000 images. I think increasing the folds number might help to stabilize CV/LB.\n"
        },
        {
          "id": 921887,
          "postDate": "2020-07-09T16:19:57.613Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 922058,
          "postDate": "2020-07-09T19:10:50.103Z",
          "content": "<p><a href=\"/synked\">@synked</a> Oh you are right! I will try to remove the duplicates from my train data!</p>",
          "rawMarkdown": "@synked Oh you are right! I will try to remove the duplicates from my train data!"
        },
        {
          "id": 922185,
          "postDate": "2020-07-09T21:38:37.223Z",
          "content": "<p>you have not mentioned external data on the list, I assume it increased your LB score but makes validation less stable, am I correct?</p>",
          "rawMarkdown": "you have not mentioned external data on the list, I assume it increased your LB score but makes validation less stable, am I correct?",
          "votes": 1
        },
        {
          "id": 922798,
          "postDate": "2020-07-10T10:47:36.920Z",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  Exactly!</p>",
          "rawMarkdown": "@jacekpoplawski  Exactly!",
          "votes": 1
        },
        {
          "id": 929749,
          "postDate": "2020-07-14T23:18:57.067Z",
          "content": "<p>Let us know if you try 10-folds and it helps.  I am using StratifiedGroupKFold @ 5 Folds.  I find using 5-fold CV over just a basic train/test split definitely makes my CV tighter.</p>",
          "rawMarkdown": "Let us know if you try 10-folds and it helps.  I am using StratifiedGroupKFold @ 5 Folds.  I find using 5-fold CV over just a basic train/test split definitely makes my CV tighter."
        }
      ]
    },
    {
      "id": 921528,
      "postDate": "2020-07-09T11:02:03.480Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 937433,
      "postDate": "2020-07-21T02:47:54.847Z",
      "content": "<p>Thank you for sharing,</p>",
      "rawMarkdown": "Thank you for sharing,",
      "votes": 1
    },
    {
      "id": 924451,
      "postDate": "2020-07-11T12:33:13.920Z",
      "content": "<p>Great, thanks for sharing!</p>",
      "rawMarkdown": "Great, thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 923601,
      "author_name": "tretrausaigon",
      "author_url": "",
      "post_date": "2020-07-11T02:38:23.133000",
      "content": "<p><strong>2. Dense head: Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score.</strong>\nB6 with Dense (256,256,1) : LB 0.939\nB7 with Dense (256,256,1) : LB 0.941\nI think it also depends on how you setup your CV and lr-scheduler. \nThose results are using full data. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 923948,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-11T07:38:39.060000",
          "content": "<p><a href=\"/tretrausaigon\">@tretrausaigon</a> What I meant is adding layers on the top of <strong>the same model.</strong> In my case: </p>\n\n<p>B7 with 1 Dense(1 output, sigmoid) LB= 0.937\nB7 with 2 Dense (128, 1) LB= 0.935\nB7 with 3 Dense (256, 128, 1) LB=0.934</p>\n\n<p>Did you experiment with adding/removing layers to the same model and compare the LB score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 925350,
          "author_name": "Luigi Saetta",
          "author_url": "",
          "post_date": "2020-07-12T03:07:53.320000",
          "content": "<p><a href=\"/amiiiney\">@amiiiney</a> I can confirm what you're saying. Only one thing is different in my case: I get better results using focal_loss. I have tested switching to BCE and I get, with exactly the same net, 0.91 (with FL is 0.935). DNNs are complicated beasts when you don't have millions of images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929622,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-14T19:52:33.167000",
          "content": "<p>Thanks <a href=\"/luigisaetta\">@luigisaetta</a>! I experimented with dense (1024 nodes) + 0.4 dropout and it did better than the other smaller layers.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 929742,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-14T22:58:48.260000",
          "content": "<p>When you scale your network to B7 are you also scaling image size to the native image size of B7 (600x600) or are you using a smaller size like 256?  I find B7 with native size of 600 is just so slow, and obviously I have to compromise on batch_size.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 938180,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-21T11:47:21.340000",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> I tried both and B7 with image size 600x600 is way slower (900s per epoch on TPU). For the moment I am just working with smaller image sizes to experiment more things. I couldn't compare their performance.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 921582,
      "author_name": "Henrique Mendonça",
      "author_url": "",
      "post_date": "2020-07-09T11:59:02.077000",
      "content": "<p>Hey <a href=\"/amiiiney\">@amiiiney</a> \nThanks for the report!\nHow many experiments have you done? \nForgive me otherwise, but it sounds like you have trained 1 model on each variant, but DL has intrinsic random variance, so you need to run several runs with different random states and/or hyper parameters to take any conclusion, IMHO\nE.g. a different loss may very well require a different learning rate, and it may also react differently to different regularisation techniques (more hyper params).</p>",
      "votes": 4,
      "replies": [
        {
          "id": 921613,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T12:25:21.510000",
          "content": "<p>Thanks <a href=\"/hmendonca\">@hmendonca</a> for your feedback!\nYou are 100% right! Actually that's exactly why I started this topic because I know that for example focal loss should work better and would like to have the feedback of people working with focal loss and know the combination of models/parameters that made it work.\nRegarding the experiments, I am experimenting with:\n* Models:  seresnext50, seresnext101 and all EfficientNets\n* Image sizes: 256, 512, 768 and 1024\n* Hyperparameters:  tweaking dropout, batchNorm, focal loss gamma and alpha parameters...\n* I am using the custom scheduler everybody is using from the flowers comp.</p>\n\n<p>Obviously, there are more experiments to run that's why I am relying on fellow kagglers to share their experiences but in my case the conclusions I have reached with the combination of models, image sizes and hyperparameters tweaking showed a similar pattern.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 921639,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-07-09T12:49:06.087000",
          "content": "<p>Nice <a href=\"/amiiiney\">@amiiiney</a> !\nI've done some +- extensive hyper search for this comp: <a href=\"https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\">https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model</a>\nIt was also binary and did get some good models with both bce and focal, they seem to be more or less equivalent for binary AUC, in that problem at least\nHowever, the output scales vary a lot between the 2, so you need to do some kind of scaling e.g. ranking for ensembling</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 942457,
      "author_name": "AkshatShreemali",
      "author_url": "",
      "post_date": "2020-07-23T18:47:48.850000",
      "content": "<p>I agree with your 4th point. I reached an lb of 0.72 with xgb but that pushed my lb to 0.94 against my current 0.958 against the same notebook. I am yet to understand the reason behind that.\nI also experimented with adding cnn+tabular features and the results were pretty bad on cv and lb so i dropped off that idea.\nCurrently i am relying of EffB6-7 and an ensemble model of these two and experimenting with some hyper parameters\nDid anyone try extracting the weights and boosting it with tabular features?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 952286,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2020-07-30T19:26:05.293000",
          "content": "<p><a href=\"/akshatshreemali91\">@akshatshreemali91</a> what image size or sizes (in case of multiple image size) are you using with B6 and B7 ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 937145,
      "author_name": "Abdur Rehman",
      "author_url": "",
      "post_date": "2020-07-20T18:56:45.763000",
      "content": "<p><a href=\"/amiiiney\">@amiiiney</a> Thansk for sharing the points. Can you kindly elaborate how did you use <code>CutMix</code>in this competition as when apply cut mix, we get <code>multi-label</code> for an image. For example, if you combine <code>cut(10%) of benign</code> and mix it into malignant image then label will be <code>(0.1, 0.9) -&gt; (benign, malignant)</code>. As you mentioned cutmix is tricky, so did you apply cutmix on images randomly without considering the <code>target(benign, malignant)</code> or you apply cutmix on similar label images (cutmix of malignant with malignant and benign with bengin)?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 938169,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-21T11:36:42.850000",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> I apply CutMix considering the benign and malignant images so that the output image is a mixture of both. I tried a 10%, 15% and 20% cuts but they all did bad. I am trying now to experiment with a combination of Cutmix/GridMask <a href=\"https://www.kaggle.com/aziz69/efficientnets-meta-data-augs\">from this notebook.</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 932956,
      "author_name": "Uljan",
      "author_url": "",
      "post_date": "2020-07-17T11:57:20.803000",
      "content": "<p>Thank you for sharing, valuable for considering new approaches with Melanoma Classification.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 930402,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2020-07-15T12:41:59.943000",
      "content": "<p>I am using EfficientNet B0.</p>\n\n<p>I am a bit mitigate with the dropout because the  results can change +/- 0.01 depending of the dropout/external dataset:\n- with External dataset  : (2018) dropout of 0.2 seems better\n- without External dataset : no dropout is better </p>\n\n<p>But it does not seem reproductible, I mean each time I launch a training I can have a different results .</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 929218,
      "author_name": "Aditya Baurai",
      "author_url": "",
      "post_date": "2020-07-14T14:37:45.457000",
      "content": "<p>Thank you! I am thinking of going for either resnet or efficient net for this! Any suggestions?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 929571,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-14T18:50:16.767000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 929767,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-14T23:39:07.867000",
          "content": "<p><a href=\"/epocxy\">@epocxy</a>    You do:</p>\n\n<pre><code>def forward(self,x):\n    x = self.block(x)\n    #x = x.view(*(x.shape[:-2]),-1).mean(-1) ## For GAP\n    x = torch.cat([self.mp(x), self.ap(x)], 1)\n    x = x.view(x.size(0), -1)\n    x = self.fc(x)\n    return x\n</code></pre>\n\n<p>What do you think \"x = x.view(x.size(0), -1)\" is doing there for you?  I don't think its doing anything.  After your cat(), your data is BatchSize, (Size of mp + size of ap).  After your view() call, it should be the same.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930073,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-15T07:10:10.953000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 930096,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-15T07:35:41.473000",
          "content": "<p>@<em>dash</em> Oh I see, I thought you just clipped the fc layer from ResNet (Which would have already left you a flattened output), but I see you took out the part where it flattens as well.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 926078,
      "author_name": "chenshov",
      "author_url": "",
      "post_date": "2020-07-12T13:30:40.333000",
      "content": "<p>Loved it!\nEspecially since Focal loss is highly spoken yet no one mentions that it didn't work for them.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 923550,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-07-11T01:01:06.067000",
      "content": "<p>Regarding FocalLoss, what implementation did you use? Did you set gamma? If so to what?</p>\n\n<p>I have had some trouble just \"dropping in\" Focal Loss.  What other changes did you have to make when you swapped in Focal Loss for BCE?  For example, I had to make sure my ground truths (y) was int64, because it didn't like them as float32.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 924174,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-11T09:51:35.267000",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> I am using <a href=\"https://github.com/artemmavrin/focal-loss\">this focal loss</a> with gamma=2. The other changes that I tried to make are: \n* Tweak the learning rate and try different optimizers: Adam and RMSProp. But I am definitely doing something wrong.\nDid focal loss give you better results than BCE? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 924251,
          "author_name": "Sanchit Gupta",
          "author_url": "",
          "post_date": "2020-07-11T10:50:55.530000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926471,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-12T18:39:24.280000",
          "content": "<p>@amin slightly better with Focal Loss.  I am also trying Focal Loss with class weights.  If you are going to try Adam why not try AdamW, and did you turn on Amsgrad (for Adam)?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929624,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-14T19:54:35.613000",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> no I didn't turn on Amsgrad for Adam, does it help with focal loss? </p>\n\n<p>And what learning rate are you using? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929711,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-14T22:00:46.933000",
          "content": "<p>@Amin I have done numerous LR experiments.  The thing is, LR is very sensitive to other parameters such as batch size, and probably which model and image size, etc.  For me .0001 is best........I am actually using a scheduler (OneCycleLR) and just set .0003 as the max.  I have tried other schedulers too and they really help but require some tuning to find the best range.</p>\n\n<p>LR is pretty easy to nail down though.  Just run a few epochs and look at loss and roc/auc.  If its unstable lower LR.  If its not going anywhere quickly, increase LR.  I do like .01, .005, .001, .0005, .0001 etc to start with my \"lr search\".  How about you, what have you found is good for LR?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 922184,
      "author_name": "Jacek Poplawski",
      "author_url": "",
      "post_date": "2020-07-09T21:36:53.813000",
      "content": "<p>Could you explain 2:</p>\n\n<p>\"2. Dense head: Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score.\"</p>\n\n<p>I assume you take some outputs from CNN then take features from csv and then use Dense layer, and your experiences are that this should be single linear layer without Dropout or BN? Can you clarify?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 922542,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-10T07:27:10.040000",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> By dense head I mean a heavier head of the net, more and different layers on the top of the backbone. <strong>Just for image data!</strong> I haven't tried to concatenate the metadata+image yet.</p>\n\n<p><code>\nmodel = tf.keras.Sequential([\n        efn.EfficientNetB7(\n            input_shape=(*IMAGE_SIZE, 3),\n            weights='imagenet',\n            include_top=False\n        ),\n        tf.keras.layers.GlobalAveragePooling2D()\n        tf.keras.layers.Dense(256, activation='relu'), \n        tf.keras.layers.Dropout(0.2), \n        tf.keras.layers.Dense(128, activation='relu'), \n        tf.keras.layers.Dropout(0.2), \n        tf.keras.layers.Dense(1, activation='sigmoid')\n    ])\n</code>\nWhat I meant is: Adding more dense layers lowered CV/LB and introducing Dropout/BatchNorm in between dense layers didn't improve the score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922559,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-10T07:46:12.323000",
          "content": "<p><a href=\"/amiiiney\">@amiiiney</a> yes I understand, however my question was - what is the \"good\" version of code, and which one is \"too much\", in your code there are THREE Dense layers plus two dropouts, so it is version which is best or is it version which is too much?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922688,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-10T09:18:00.813000",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> Those results are for an efficientNet B7 trained over 23 epochs</p>\n\n<p>1 Dense(1 output, sigmoid) LB= 0.937\n2 Dense (128, 1) LB= 0.935\n3 Dense (256, 128, 1) LB=0.934</p>\n\n<p>All of them are trained over 23 epochs. I have a feeling that maybe models with heavier heads need more training epochs to give better results, this is something I haven't tried yet because I am limited to the kaggle 3H TPU sessions. I would like to know the experience of the others with [heavier head + more epochs]</p>\n\n<p>Adding dropouts (p=0.1, 0.2, 0.3) and Batch normalization gave exactly the same LB scores!</p>\n\n<p>I would like to know your opinion about this! :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922697,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-10T09:24:22.023000",
          "content": "<p>Thanks for the numbers, could you show the code for first case? Because I want to be sure we are talking about same thing. Is it just GlobalAveragePooling2D and Dense with 1 output?</p>\n\n<p>One more thing - in your code there is just Dense operating on efficientnet output, what about metadata (sex, age, area)? I would like to see whole dense code of your solution otherwise I think some information may be incomplete, like maybe you area talking about one Dense here then another Dense later.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922717,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-10T09:41:33.103000",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> Exactly!! just GlobalAveragePooling2D + Dense with 1 output\n<code>\n model = tf.keras.Sequential([\n        efn.EfficientNetB0(\n            input_shape=(*IMAGE_SIZE, 3),\n            weights='imagenet',\n            include_top=False\n        ),\n        L.GlobalAveragePooling2D(),\n        L.Dense(1, activation='sigmoid')\n    ])\n</code></p>\n\n<p>I didn't include metadata, the dense layers I am using are <strong>JUST FOR IMAGE DATA</strong>. </p>\n\n<p>I run the same experiment with B0 over 45 epochs and again the the models with heavier heads are always behind the model with just GlobalAveragePooling2D + Dense with 1 output.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922719,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-10T09:42:05.743000",
          "content": "<p>\"I didn't include metadata, the dense layers I am using are JUST FOR IMAGE DATA. \"</p>\n\n<p>so how do you use meta data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922748,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-10T10:05:17.263000",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  I am training a separate XGB model and ensemble with the CNN trained on image data. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922753,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-10T10:13:37.157000",
          "content": "<p>I see, I was thinking you are talking about Dense layers used to merge CNN data and meta data, in case of CNN data only I see no point in adding more Dense layers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922757,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-10T10:18:25.703000",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>! Do you think merging CNN data and meta data is better than training 2 separate models as I am doing?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922786,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-07-10T10:36:21.367000",
          "content": "<p>Theoretically it makes more sense to train with the metadata <a href=\"/amiiiney\">@amiiiney</a> (edit: so no separate models), as your model could condition (e.g. use different parts of the network) based on the metadata. But then again, this is deep learning, it's all alchemy and there's no such thing as theory.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 922794,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-10T10:42:00.833000",
          "content": "<p>in my opinion training separate model for meta data is pointness, but this is just my opinion, I never tried it</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922863,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-10T11:21:21.633000",
          "content": "<p>Thanks <a href=\"/group16\">@group16</a> that explains a lot of things!\nDo you mind if I ask what is your best single model (CNN with metadata) LB score?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922899,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-10T12:02:11.073000",
          "content": "<p>You should generate multiple outputs from CNN then merge with multiple features then Dense layers can think how to combine them - that was my point all the time, do you mean only one Dense layer is enough here, in my opinion no.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922902,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-07-10T12:04:39.183000",
          "content": "<p>Of course you may ask <a href=\"/amiiiney\">@amiiiney</a>  :). Currently I'm only at 0.937</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 923551,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-11T01:04:31.993000",
          "content": "<p>Yes, that would make it more clear.  I assume perhaps he is saying that he took the date right out of the  main model and directly into 1 feature? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 923553,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-11T01:07:15.510000",
          "content": "<p>so when you say:</p>\n\n<p>1 Dense(1 output, sigmoid) LB= 0.937</p>\n\n<p>you mean you took the output of B7 which is 2560 and put it right into a single feature and that did better than adding additional layers?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 924068,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-11T08:37:05.637000",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> <a href=\"/jacekpoplawski\">@jacekpoplawski</a> Sorry if my explanation wasn't clear enough. Here what's I am doing:</p>\n\n<ul>\n<li><p>1 Dense(1, sigmoid) just like this public notebook: <a href=\"https://www.kaggle.com/manojprabhaakr/melanoma-tpu-starter-efficientnet-b0\">Melanoma TPU Starter EfficientNet B0: Version2</a> gives LB=0.937</p></li>\n<li><p>A heavier head of the net more or less like this public notebook: <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">Melanoma TPU EfficientNet B5_dense_head</a> gives LB=0.935 and LB=0.934 (more layers= lower score).</p></li>\n</ul>\n\n<p>I am training all of the them over the same number of epochs so probably the models with heavier heads need more epochs!?! \nI tried layers of (512, 256 and 128). Now I am trying with more neurons (1024) as discussed above and <a href=\"https://www.kaggle.com/aziz69/efficientnets-meta-data-augs\">in this notebook</a> and apparently adding one layer of 1024 neurons + 0.4 dropout does better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929747,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-14T23:16:26.623000",
          "content": "<p>@Amin you mention training for 23 epochs (on more involved Dense heads).  Do you find you are seeing improvements even at 23 epochs? What are you measuring to decide how many epochs: loss or roc/auc?  I assume you are not using any sort of early stopping/patience.  I rarely see improvements past 10 epochs, I typically deploy some sort of early stopping.  I AM using a scheduler which modifies/drops learning rates if no improvement.  How about you?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 921981,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-09T17:41:58.933000",
      "content": "<p>Regarding #4, is your XGB using just meta features from the provided CSV file (like Giba's model), or are you making more meta features like image width and height?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 921993,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T18:00:34.707000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> the XGB using just the meta features provided, no feature engineering. The thing is not just my XGB but also the other public kernels dealing with metadata, they all have higher LB scores +70 but they don't do well in the ensemble. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922001,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-09T18:09:29.923000",
          "content": "<p>It sounds like your XGB may be overfitting the train data. And many of these public notebooks probably overfit too.</p>\n\n<p>Also regarding other public notebooks, I have noticed that there are strong correlations between image width height and target in the train data. But these same correlations are not present in the test data.</p>\n\n<p>For example most images in train data with original height width <code>4000x6000</code> have no malignant (and there are 14703 of these images!) but this is not true for test data (which has 4162 of these images).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 922005,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-09T18:14:06.680000",
          "content": "<p>Giba's <a href=\"https://www.kaggle.com/titericz/simple-baseline\">notebook</a> uses Bayesian smoothing to prevent overfitting. His models ignore feature values which occur rarely. And only consider the most prominent patterns.</p>\n\n<pre><code>te['ll'] = ((te['mean']*te['count'])+(M*L))/(te['count']+L)\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 922044,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T18:54:21.810000",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> that explains why the public notebooks are not doing well. I tried some other features, they all increase CV but decrease LB, it's frustrating but I still believe that there should be some features out there that could make the metadata more useful!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 921904,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-09T16:34:59.270000",
      "content": "<p>I'm confused. Before your 4 numbered points, you have 4 bullet points. The bullet points include using EfficientNet and sizes 256, 512, 768, 1024. Are you saying that EfficientNet and sizes 256, 512, 768, 1024 don't help?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 921910,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-09T16:42:20.607000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 921924,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T16:52:32.490000",
          "content": "<p>I had to restructure my thread because I was asked what kind of models and parameters led the the things that didn't work. \nSo the bullet points are the \"materials and methods\".\nAnd the numbered points are the things that didn't work.\nSorry <a href=\"/cdeotte\">@cdeotte</a> for confusing you :) I will reorganize the thread again :)</p>\n\n<p>I would love to know your opinion about this topic, especially point number 4 :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 921730,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-09T14:01:27.650000",
      "content": "<p>oh , well done bro</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 929806,
      "author_name": "Solve for Fun",
      "author_url": "",
      "post_date": "2020-07-15T01:18:03.957000",
      "content": "<p><a href=\"/amiiiney\">@amiiiney</a> Thanks for presenting results and observations of so many simulations. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 925268,
      "author_name": "Luigi Saetta",
      "author_url": "",
      "post_date": "2020-07-12T00:47:19.283000",
      "content": "<p>I have to tell that this post has in some way ruined my already sleepless nights (I'm a newcomer to these competitions,). In the sense that it has challenged some (wrong ?)  results I had arrived to (but probably due to timing not fully checked).\nI can confirm that with a very simple network (EfficientNet B7 and only a Dense(1, sigmoid), without metadata (only image) you can get LB = 0.935. That is more all less the best single model (no ensemble) result I have achieved so far (well I have 0.936 with metadata and a more complicated network, but the difference could be not significant).</p>\n\n<p>But in my case, with smaller (B4) I have checked a 10% improvement with FocalLoss, and now I'm running with FocalLoss. \nThe biggest problem I see is that all the networks from B4 are overfitting the training set. And I don't see improvements with more than 10 epochs!!\nIf you add a longer Dense head it is worse, unless you go with dropout=0.5. With a so simple net (well, simple since you can't modify B7 and add dropout inside) the only thing is to add images to the training set (and I'm using the full Deotte's set).\nI'm going to try switching the loss function. If I get an improvement... what should I say?: 'Less is more' (and I have been always a fan of the old Apple motto).\nI'll soon post the entire Notebook.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 925289,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-12T01:47:12.033000",
          "content": "<p>hello, could you clarify:</p>\n\n<ul>\n<li>\"a very simple network (EfficientNet B7\"</li>\n<li>\"all the networks from B4 are overfitting\"</li>\n</ul>\n\n<p>why do you call B7 simple? do you mean B7 or B0? and if you are saying that all from B4 are overfitting I assume you mean B7 too? do you mean B4-B7 or B4-B0?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 925343,
          "author_name": "Luigi Saetta",
          "author_url": "",
          "post_date": "2020-07-12T03:04:19.607000",
          "content": "<p>HI Jacek. I see your point: by itself B7 is NOT simple. What is simple is my network for this post (since the magic of transfer learning obviously): it is simply B7 followed by a Dense(1)</p>\n\n<pre><code>with strategy.scope():\n    inp = tf.keras.layers.Input(shape=(dim, dim,3))\n\n    base = efns7(input_shape=(dim, dim, 3),weights='imagenet',include_top=False)\n\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(1,activation='sigmoid')(x)\n\n    model = tf.keras.Model(inputs=inp,outputs=x)\n</code></pre>\n\n<p>from my tests all the networks from B4 to B7 are overfitting: at the end of the training, the training accuracy is 0.99, and validation accuracy is 0.93. From this point of view adding capacity to the network, with a more complex Dense head normally doesn't help. I have seen very small improvements using dropout(5) but results vary and I don't think are useful. </p>\n\n<p>So, for now I agree that simpler is better.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 939695,
          "author_name": "Prateek Mishra",
          "author_url": "",
          "post_date": "2020-07-22T11:39:44.090000",
          "content": "<p><a href=\"/luigisaetta\">@luigisaetta</a>  Thanks for clearing up things..! Can you please tell that is this ( 0.93 Acc. ) model is your single best model and you are now using ensembles to score higher. If yes then how many models you are using ?\nIn my case I achieved 93.70 using B6 with 2018 external data and focal loss + little aug.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 922578,
      "author_name": "Aziz_Belaweid",
      "author_url": "",
      "post_date": "2020-07-10T07:58:48.557000",
      "content": "<p>Thanks for sharing. I'm quite the opposite of you 0.4 ratio of CutMix with gridMask works well for me, focal loss improves the score too, adding dense head improves but it took me a lot of time to tweak it. My strategy is the get the best score I can get with a single model (0.933 so far ) then ensemble but I'm kinda stuck I don't know what else to try now (I've tried hair augmentations =&gt; improves a little bit, I also tried microscope aug =&gt; doesn't improve) Any help on how to improve my single model to get over 0.933 would be appreciated </p>",
      "votes": 2,
      "replies": [
        {
          "id": 922740,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-10T10:00:52.470000",
          "content": "<p>Thanks <a href=\"/aziz69\">@aziz69</a> for you feedback. I am glad to see that those things worked for you!\n* Do you mind if I ask which focal loss are you using? I am using <a href=\"https://github.com/artemmavrin/focal-loss\">this focal loss</a> and which learning rate, gamma and alpha parameters gave you the best result?</p>\n\n<ul>\n<li><p>How many layers and which kind of layers do you have in your head? In my case more layers lowered the score and including dropouts and batch normalization didn't help much (you can see my dense head in the conversation below).</p></li>\n<li><p>I was planning to experiment with gridMask next week. Did the combo CutMix+GridMask give a better score than GridMask alone?</p></li>\n</ul>\n\n<p>Which model and for how many epochs are you training your 0.933 model? Ideas that you didn't mention and could increase your score:\n* Did you try <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">CAM CutMix explained in this post</a> \n* including the metadata to your CNN by concatenating the image+metadata \n* Training for longer hours </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922792,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-10T10:41:03.587000",
          "content": "<p>hello <a href=\"/aziz69\">@aziz69</a> \ncan you share the info about CutMix? how do you use it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 923005,
          "author_name": "Aziz_Belaweid",
          "author_url": "",
          "post_date": "2020-07-10T13:29:44.657000",
          "content": "<p>Hey guys so this is the focal loss I used : \n* loss = tfa.losses.SigmoidFocalCrossEntropy(reduction=tf.keras.losses.Reduction.AUTO) with lr = 1e-5  (tfa =&gt; tensrflow addons) \n* using only one dense layer after you model ( number of neurons has to be tuned for 1024 + 0.4 dropout worked best ) \n* Yes gridMask + Cutmix combo is best then working with each one solo.\n* I'm using efficientNets for 20 epochs.</p>\n\n<p>I have tried including meta data model (xgb 0.73) with my best cnn submission (0.94) somehow it didn't add anything the score was lower even when i tried different coeffients. </p>\n\n<p>I haven't tried CAM cutmix so thanks for that I'll try it. \nThis is my notebook (I haven't commited for a while but most of my work is based on this) <a href=\"https://www.kaggle.com/aziz69/efficientnets-meta-data-augs\">https://www.kaggle.com/aziz69/efficientnets-meta-data-augs</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 921686,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2020-07-09T13:31:51.663000",
      "content": "<p>Your 4th point makes me fear that we are only overfitting the Public LB.</p>\n\n<p>None of my experiments was a real success, CV vs LB is not consistent for me and similar CV give different LB...</p>\n\n<p>Would you mind explaining what's the custom scheduler everybody is using from the flowers competition?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 921715,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T13:53:54.273000",
          "content": "<p>Hey <a href=\"/optimo\">@optimo</a> \n* The <strong>4th point</strong> also got me thinking about what I might be doing wrong! I would like to know the experience of others with tabular data, because most of the discussion is about architectures and augmentations.\n* The custom scheduler from this baseline kernel in the flower comp.: <a href=\"https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule\">https://www.kaggle.com/mgornergoogle/gpu-test-flowers-on-tpu-ensemble-lr-schedule</a> . Most public kernels are using it:\n```\nLR_START = 0.00001\nLR_MAX = 0.00005 * strategy.num_replicas_in_sync\nLR_MIN = 0.00001\nLR_RAMPUP_EPOCHS = 5\nLR_SUSTAIN_EPOCHS = 0\nLR_EXP_DECAY = .8</p>\n\n<p>def lrfn(epoch):\n    if epoch &lt; LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch &lt; LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN\n    return lr</p>\n\n<p>lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=True)\n```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 921719,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-09T13:57:51.623000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 921721,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-09T13:59:16.733000",
          "content": "<p>thanks I'll have a look!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 921733,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-09T14:04:18.793000",
          "content": "<p>hey <a href=\"/synked\">@synked</a> what do you mean by \"the other(image only) experiments\"?</p>\n\n<p>The unstability I was referring to is not specific to metadata usage. The CV vs LB seems unstable for image and image+metadata experiments I did.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 921747,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-09T14:12:38.050000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 958092,
      "author_name": "Shaitender Singh",
      "author_url": "",
      "post_date": "2020-08-04T19:00:44.833000",
      "content": "<p>great job !!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 952285,
      "author_name": "Abdur Rehman",
      "author_url": "",
      "post_date": "2020-07-30T19:24:27.230000",
      "content": "<p><a href=\"/amiiiney\">@amiiiney</a> are you using <code>kaggle TPU</code> for <code>1024</code> images? If yes, what <code>batch size</code> you are using and does your model train completely in one <code>TPU session(3 hrs)</code> or do you resume training ?</p>\n\n<p>Another thing is you mentioned that you are experimenting with 4 different image sizes so is there any specific reason of choosing these image sizes as you can choose other sizes as well like <code>128, 192, 384</code>?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 924215,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T10:32:13.080000",
      "content": "<p>ok</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 924154,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T09:29:02.883000",
      "content": "<p>nice</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 923741,
      "author_name": "Surya Prakash",
      "author_url": "",
      "post_date": "2020-07-11T05:56:17.177000",
      "content": "<p>Nice work :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 929556,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T18:42:41.203000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 929563,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-14T18:45:49.100000",
          "content": "<p>The 2019 data contains the 2018 and 2017 data. So half of 2019 data is new, and half of 2019 is those old comp data.</p>\n\n<p>Do you use all 2019? Some people have observed a difference between using the new half and old half. In my TFRecords, the even numbered records are 2018 2017 and the odd numbered records are the new 2019 half.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 929576,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-14T18:57:52.073000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929583,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-14T19:04:32.193000",
          "content": "<p>If you have the downloaded external 2019 JPEGS dataset, then you can separate the images using the <code>train.csv</code> contained within. All images with original size <code>1024x1024</code> as designated by the <code>width</code> and <code>height</code> columns are the new 2019 data and the other images are the old 2019 which is say 2018 2017 data</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929594,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-14T19:16:39.800000",
          "content": "<p>Hi Chris </p>\n\n<p>Thanks again for all amazing contributions. </p>\n\n<p>It seems the train csv file I found <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256?select=train.csv\">here</a> contains only 2020 data.  </p>\n\n<p>Could you please provided the link of CSV file containing 2020 + external data ?   Thanks again.   </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929595,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-14T19:17:11.040000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929603,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-14T19:22:28.473000",
          "content": "<p>&gt; Looklike there are 12414 records which have (h,w)--&gt;(1024,1024)so they are all related to 2019 only ?</p>\n\n<p>Yes. Those are the images that are in 2019 but not in 2018 2017. If you display them, you will see that they look different than the others. They are more zoomed in and have large dark spots.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929605,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-14T19:25:04.590000",
          "content": "<p><a href=\"/serigne\">@serigne</a> My csv file naming is confusing because i call all my cvs <code>train.csv</code>. You need to use a <code>train.csv</code> from inside my external dataset (either JPEG or TFRecord). A direct link is <a href=\"https://www.kaggle.com/cdeotte/isic2019-256x256?select=train.csv\">here</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929611,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-14T19:41:01.027000",
          "content": "<p>Thank <a href=\"/cdeotte\">@cdeotte</a>  !</p>\n\n<p>I forgot you made separated datasets. I got confused by another discussion Topic</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929628,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-14T20:00:46.870000",
          "content": "<p><a href=\"/epocxy\">@epocxy</a> yes I am using external data but in my case they increase the LB score, however the CV/LB with external data is not stable.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 939683,
          "author_name": "Prateek Mishra",
          "author_url": "",
          "post_date": "2020-07-22T11:32:13.640000",
          "content": "<p><a href=\"/amiiiney\">@amiiiney</a> my cv is 90.10 and lb is 93.70 when i use 2018 data + 2020 data. Is that CV that we can trust and except not much shakeup in the private lB?\nAlso, I want to know why there is so much gap? \n- There could not be any leak because I am using <a href=\"/cdeotte\">@cdeotte</a> 's triple stratified dataset.\nI think that after training using an external dataset the test data is simple for our model to classify. what do you say ?\n<a href=\"/cdeotte\">@cdeotte</a> <a href=\"/amiiiney\">@amiiiney</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 927496,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-13T12:25:55.623000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 922372,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-10T04:07:27.273000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 921726,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-09T13:59:57.830000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 921739,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T14:07:04.407000",
          "content": "<p>Same as <a href=\"/optimo\">@optimo</a>, CV and LB are not stable, especially when I use external data, the gap gets bigger.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 921740,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-09T14:07:52.517000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 921745,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T14:12:19.163000",
          "content": "<p>Yes both of them! I am using 5-folds. I should probably try 10-folds.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 921749,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-09T14:13:43.320000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 921758,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-09T14:25:12.443000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 921765,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T14:32:17.183000",
          "content": "<p>I didn't because in this discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\">How To Use Last Years 2019 Comp Data (and 2018, 2017)</a> Chris detected the duplicates and said it's safe to leave them in because they are just 59 out of 30,000 images. I think increasing the folds number might help to stabilize CV/LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 921887,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-09T16:19:57.613000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922058,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-09T19:10:50.103000",
          "content": "<p><a href=\"/synked\">@synked</a> Oh you are right! I will try to remove the duplicates from my train data!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 922185,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-09T21:38:37.223000",
          "content": "<p>you have not mentioned external data on the list, I assume it increased your LB score but makes validation less stable, am I correct?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 922798,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-07-10T10:47:36.920000",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  Exactly!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929749,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-14T23:18:57.067000",
          "content": "<p>Let us know if you try 10-folds and it helps.  I am using StratifiedGroupKFold @ 5 Folds.  I find using 5-fold CV over just a basic train/test split definitely makes my CV tighter.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 921528,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-09T11:02:03.480000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937433,
      "author_name": "DynmiWg",
      "author_url": "",
      "post_date": "2020-07-21T02:47:54.847000",
      "content": "<p>Thank you for sharing,</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 924451,
      "author_name": "Researcher 002",
      "author_url": "",
      "post_date": "2020-07-11T12:33:13.920000",
      "content": "<p>Great, thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "921522": "I would like to share and discuss the experiments that I tried and didn't improve the LB score:\n\n\n**1. Focal loss:** Many people reported that focal loss is giving them better results, in my case BCE significantly increases the score. [*(The focal loss I am using)*](https://github.com/artemmavrin/focal-loss)\n\n**2. Heavier head of the net (more layers on the top of the backbone):** Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score. **[JUST FOR IMAGE DATA]**\n\n\n**3. Augmentation with CutMix:** (without CutMix: 0.947, with CutMix: 0.941). This discussion explains the reason behind this: [CutMix is tricky in this competition](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160784). \n\n\n**4. Using XGBoost with tabular data:** I reach LB=0.75 but it does worse than [giba's 0.70 notebook](https://www.kaggle.com/titericz/simple-baseline) when I ensemble with the CNN (Ensemble += giba's kernel=0.951 / Ensemble += XGBoost=0.941).\n\nI would like to have your feedback about the above, if those things worked for you or not. Or if there are any other things that did not work for you!\n\n&gt;### Setup:\n* Models: seresnext50, seresnext101 and all EfficientNets\n* Image sizes: 256, 512, 768 and 1024\n* Hyperparameters: tweaking dropout, batchNorm, focal loss gamma and alpha parameters…\n* I am using the custom scheduler everybody is using from the flowers comp. ",
    "923601": "**2. Dense head: Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score.**\nB6 with Dense (256,256,1) : LB 0.939\nB7 with Dense (256,256,1) : LB 0.941\nI think it also depends on how you setup your CV and lr-scheduler. \nThose results are using full data. \n",
    "921582": "Hey @amiiiney \nThanks for the report!\nHow many experiments have you done? \nForgive me otherwise, but it sounds like you have trained 1 model on each variant, but DL has intrinsic random variance, so you need to run several runs with different random states and/or hyper parameters to take any conclusion, IMHO\nE.g. a different loss may very well require a different learning rate, and it may also react differently to different regularisation techniques (more hyper params).",
    "942457": "I agree with your 4th point. I reached an lb of 0.72 with xgb but that pushed my lb to 0.94 against my current 0.958 against the same notebook. I am yet to understand the reason behind that.\nI also experimented with adding cnn+tabular features and the results were pretty bad on cv and lb so i dropped off that idea.\nCurrently i am relying of EffB6-7 and an ensemble model of these two and experimenting with some hyper parameters\nDid anyone try extracting the weights and boosting it with tabular features?",
    "937145": "@amiiiney Thansk for sharing the points. Can you kindly elaborate how did you use `CutMix `in this competition as when apply cut mix, we get `multi-label` for an image. For example, if you combine `cut(10%) of benign` and mix it into malignant image then label will be `(0.1, 0.9) -&gt; (benign, malignant)`. As you mentioned cutmix is tricky, so did you apply cutmix on images randomly without considering the `target(benign, malignant)` or you apply cutmix on similar label images (cutmix of malignant with malignant and benign with bengin)?",
    "932956": "Thank you for sharing, valuable for considering new approaches with Melanoma Classification.",
    "930402": "I am using EfficientNet B0.\n\nI am a bit mitigate with the dropout because the  results can change +/- 0.01 depending of the dropout/external dataset:\n- with External dataset  : (2018) dropout of 0.2 seems better\n- without External dataset : no dropout is better \n\nBut it does not seem reproductible, I mean each time I launch a training I can have a different results .",
    "929218": "Thank you! I am thinking of going for either resnet or efficient net for this! Any suggestions?",
    "926078": "Loved it!\nEspecially since Focal loss is highly spoken yet no one mentions that it didn't work for them.",
    "923550": "Regarding FocalLoss, what implementation did you use? Did you set gamma? If so to what?\n\nI have had some trouble just \"dropping in\" Focal Loss.  What other changes did you have to make when you swapped in Focal Loss for BCE?  For example, I had to make sure my ground truths (y) was int64, because it didn't like them as float32.  ",
    "922184": "Could you explain 2:\n\n\"2. Dense head: Adding one more layer doesn't increase the score, adding more than 2 layers lowers the score. Dropout and Batch normalization also lower the LB score.\"\n\nI assume you take some outputs from CNN then take features from csv and then use Dense layer, and your experiences are that this should be single linear layer without Dropout or BN? Can you clarify?\n\n",
    "921981": "Regarding #4, is your XGB using just meta features from the provided CSV file (like Giba's model), or are you making more meta features like image width and height?",
    "921904": "I'm confused. Before your 4 numbered points, you have 4 bullet points. The bullet points include using EfficientNet and sizes 256, 512, 768, 1024. Are you saying that EfficientNet and sizes 256, 512, 768, 1024 don't help?",
    "921730": "oh , well done bro",
    "929806": "@amiiiney Thanks for presenting results and observations of so many simulations. ",
    "925268": "I have to tell that this post has in some way ruined my already sleepless nights (I'm a newcomer to these competitions,). In the sense that it has challenged some (wrong ?)  results I had arrived to (but probably due to timing not fully checked).\nI can confirm that with a very simple network (EfficientNet B7 and only a Dense(1, sigmoid), without metadata (only image) you can get LB = 0.935. That is more all less the best single model (no ensemble) result I have achieved so far (well I have 0.936 with metadata and a more complicated network, but the difference could be not significant).\n\nBut in my case, with smaller (B4) I have checked a 10% improvement with FocalLoss, and now I'm running with FocalLoss. \nThe biggest problem I see is that all the networks from B4 are overfitting the training set. And I don't see improvements with more than 10 epochs!!\nIf you add a longer Dense head it is worse, unless you go with dropout=0.5. With a so simple net (well, simple since you can't modify B7 and add dropout inside) the only thing is to add images to the training set (and I'm using the full Deotte's set).\nI'm going to try switching the loss function. If I get an improvement... what should I say?: 'Less is more' (and I have been always a fan of the old Apple motto).\nI'll soon post the entire Notebook.",
    "922578": "Thanks for sharing. I'm quite the opposite of you 0.4 ratio of CutMix with gridMask works well for me, focal loss improves the score too, adding dense head improves but it took me a lot of time to tweak it. My strategy is the get the best score I can get with a single model (0.933 so far ) then ensemble but I'm kinda stuck I don't know what else to try now (I've tried hair augmentations =&gt; improves a little bit, I also tried microscope aug =&gt; doesn't improve) Any help on how to improve my single model to get over 0.933 would be appreciated ",
    "921686": "Your 4th point makes me fear that we are only overfitting the Public LB.\n\nNone of my experiments was a real success, CV vs LB is not consistent for me and similar CV give different LB...\n\nWould you mind explaining what's the custom scheduler everybody is using from the flowers competition?",
    "958092": "great job !!",
    "952285": "@amiiiney are you using `kaggle TPU` for `1024` images? If yes, what `batch size` you are using and does your model train completely in one `TPU session(3 hrs)` or do you resume training ?\n\nAnother thing is you mentioned that you are experimenting with 4 different image sizes so is there any specific reason of choosing these image sizes as you can choose other sizes as well like `128, 192, 384 `?",
    "924215": "ok",
    "924154": "nice",
    "923741": "Nice work :)",
    "929556": "",
    "927496": "",
    "922372": "",
    "921726": "",
    "921528": "",
    "937433": "Thank you for sharing,",
    "924451": "Great, thanks for sharing!"
  }
}