{
  "id": 88016,
  "title": "My 97.23 with DoG, U-net, ResNeXt",
  "url": "/competitions/histopathologic-cancer-detection/discussion/88016",
  "author_name": "",
  "post_date": "2019-04-05T07:31:04.520215Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<h2>Result = 97.23 private LB</h2>\n\n<p>with a very strange score of 96.93 on public LB...</p>\n\n<p><strong>- Augmentation :</strong> DoG (per-channel difference of gaussian) + per-channel normalization + vert/horiz flip + hue + per-channel random gain</p>\n\n<p><strong>Model 1 :</strong>\n- ResNeXt50 + ResNet50 + custom pure ConvNet</p>\n\n<p><strong>Model 2 :</strong>\n- custom \"U-Net\" + custom \"ResNet\" (sort of auto-encoder + residual convnet), slow to train from scratch, but very cool results.\npaper :\n<a href=\"https://lmb.informatik.uni-freiburg.de/people/ronneber/u-net/\">https://lmb.informatik.uni-freiburg.de/people/ronneber/u-net/</a></p>\n\n<ul>\n<li>each model 16 fold</li>\n</ul>\n\n<p><strong>Training :</strong>\n- simple brute force Adam (0.0001) with 0.95x decay after each epoch for each model\n- keeping a eye on overfitting helped deciding when to stop learning.</p>\n\n<p><strong>Ensembling :</strong>\nBest result was achieved with simple averaging (60/40%) of ResNeXt50 + U-net\nThese 2 models seem to have the most different \"undertstanding\" of the features.\nEven though U-Net alone had poor results (about 97%ROC on local data), mixing it with ResNeXt50 (about 99.2ROC) was a good idea. These 2 models don't overlap a lot, and averaging the 2 provided clearly the best results. Funny fact!</p>\n\n<p><strong>What was good :</strong>\n- Difference of Gaussian was a big improvement, yielding sometimes strange results on some sample images. Improved both learning speeds and results.</p>\n\n<p><strong>What was not good :</strong>\n- Ensembling too similar models yields no better results (ex. ResNet50 + VGG)\n- Upscaling (to 224x224) didn't improve overall performance, especially on custom designed networks.</p>\n\n<p><strong>What I should have done :</strong>\n- Try to get rid of the (HUGE!) JPEG artefacts, since it really polluted the data.\npaper to implement :\n<a href=\"https://arxiv.org/abs/1605.00366\">https://arxiv.org/abs/1605.00366</a>\nBTW :  no real-life medical documents should ever be JPG-compressed... (and converted back to TIF!)\nI'm pretty sure this would have added robustness, as some image samples were heavily distorded by JPEGing, especially visible after DoG and per-channel normalization.</p>\n\n<ul>\n<li>Maybe Pseudo-labelling on \"U-Net\" way have improved results on 'unsure' results.</li>\n</ul>\n\n<p><strong>What I didn't want to do :</strong>\n- cheating with the original Cam16 files\n- WSI things... sorry, not interesting for me.</p>\n\n<p>This was my first Kaggle competition. A lot of fun, a lot of things learned.\nThanks to all this cool community!</p>\n\n<p>Simon (France)</p>",
  "messages": [
    {
      "id": "507779",
      "postDate": "04/05/2019 07:31:04",
      "content": "<h2>Result = 97.23 private LB</h2>\n\n<p>with a very strange score of 96.93 on public LB...</p>\n\n<p><strong>- Augmentation :</strong> DoG (per-channel difference of gaussian) + per-channel normalization + vert/horiz flip + hue + per-channel random gain</p>\n\n<p><strong>Model 1 :</strong>\n- ResNeXt50 + ResNet50 + custom pure ConvNet</p>\n\n<p><strong>Model 2 :</strong>\n- custom \"U-Net\" + custom \"ResNet\" (sort of auto-encoder + residual convnet), slow to train from scratch, but very cool results.\npaper :\n<a href=\"https://lmb.informatik.uni-freiburg.de/people/ronneber/u-net/\">https://lmb.informatik.uni-freiburg.de/people/ronneber/u-net/</a></p>\n\n<ul>\n<li>each model 16 fold</li>\n</ul>\n\n<p><strong>Training :</strong>\n- simple brute force Adam (0.0001) with 0.95x decay after each epoch for each model\n- keeping a eye on overfitting helped deciding when to stop learning.</p>\n\n<p><strong>Ensembling :</strong>\nBest result was achieved with simple averaging (60/40%) of ResNeXt50 + U-net\nThese 2 models seem to have the most different \"undertstanding\" of the features.\nEven though U-Net alone had poor results (about 97%ROC on local data), mixing it with ResNeXt50 (about 99.2ROC) was a good idea. These 2 models don't overlap a lot, and averaging the 2 provided clearly the best results. Funny fact!</p>\n\n<p><strong>What was good :</strong>\n- Difference of Gaussian was a big improvement, yielding sometimes strange results on some sample images. Improved both learning speeds and results.</p>\n\n<p><strong>What was not good :</strong>\n- Ensembling too similar models yields no better results (ex. ResNet50 + VGG)\n- Upscaling (to 224x224) didn't improve overall performance, especially on custom designed networks.</p>\n\n<p><strong>What I should have done :</strong>\n- Try to get rid of the (HUGE!) JPEG artefacts, since it really polluted the data.\npaper to implement :\n<a href=\"https://arxiv.org/abs/1605.00366\">https://arxiv.org/abs/1605.00366</a>\nBTW :  no real-life medical documents should ever be JPG-compressed... (and converted back to TIF!)\nI'm pretty sure this would have added robustness, as some image samples were heavily distorded by JPEGing, especially visible after DoG and per-channel normalization.</p>\n\n<ul>\n<li>Maybe Pseudo-labelling on \"U-Net\" way have improved results on 'unsure' results.</li>\n</ul>\n\n<p><strong>What I didn't want to do :</strong>\n- cheating with the original Cam16 files\n- WSI things... sorry, not interesting for me.</p>\n\n<p>This was my first Kaggle competition. A lot of fun, a lot of things learned.\nThanks to all this cool community!</p>\n\n<p>Simon (France)</p>",
      "rawMarkdown": "## Result = 97.23 private LB\nwith a very strange score of 96.93 on public LB...\n\n**- Augmentation :** DoG (per-channel difference of gaussian) + per-channel normalization + vert/horiz flip + hue + per-channel random gain\n\n**Model 1 :**\n- ResNeXt50 + ResNet50 + custom pure ConvNet\n\n**Model 2 :**\n- custom \"U-Net\" + custom \"ResNet\" (sort of auto-encoder + residual convnet), slow to train from scratch, but very cool results.\npaper :\nhttps://lmb.informatik.uni-freiburg.de/people/ronneber/u-net/\n\n- each model 16 fold\n\n**Training :**\n- simple brute force Adam (0.0001) with 0.95x decay after each epoch for each model\n- keeping a eye on overfitting helped deciding when to stop learning.\n\n**Ensembling :**\nBest result was achieved with simple averaging (60/40%) of ResNeXt50 + U-net\nThese 2 models seem to have the most different \"undertstanding\" of the features.\nEven though U-Net alone had poor results (about 97%ROC on local data), mixing it with ResNeXt50 (about 99.2ROC) was a good idea. These 2 models don't overlap a lot, and averaging the 2 provided clearly the best results. Funny fact!\n\n**What was good :**\n- Difference of Gaussian was a big improvement, yielding sometimes strange results on some sample images. Improved both learning speeds and results.\n\n**What was not good :**\n- Ensembling too similar models yields no better results (ex. ResNet50 + VGG)\n- Upscaling (to 224x224) didn't improve overall performance, especially on custom designed networks.\n\n**What I should have done :**\n- Try to get rid of the (HUGE!) JPEG artefacts, since it really polluted the data.\npaper to implement :\nhttps://arxiv.org/abs/1605.00366\nBTW :  no real-life medical documents should ever be JPG-compressed... (and converted back to TIF!)\nI'm pretty sure this would have added robustness, as some image samples were heavily distorded by JPEGing, especially visible after DoG and per-channel normalization.\n\n- Maybe Pseudo-labelling on \"U-Net\" way have improved results on 'unsure' results.\n\n**What I didn't want to do :**\n- cheating with the original Cam16 files\n- WSI things... sorry, not interesting for me.\n\nThis was my first Kaggle competition. A lot of fun, a lot of things learned.\nThanks to all this cool community!\n\nSimon (France)",
      "votes": null
    },
    {
      "id": "507954",
      "postDate": "04/05/2019 12:39:46",
      "content": "<p>Interesting to use U-net. Thanks for sharing, Simon! </p>",
      "rawMarkdown": "Interesting to use U-net. Thanks for sharing, Simon!",
      "votes": null
    },
    {
      "id": "508274",
      "postDate": "04/05/2019 23:00:40",
      "content": "<p>Hi Simon, thanks for shaing your approach.  What was your encoder on your unet? Did you use the same ResNeXt50? I was planning to try a unet as the this is really a pixel classification problem and though many of the pixels in a patch may be non-tumor we still want to classify it as tumor. I tried this on the skin cancer data set and could not get better results than a regular resnet but I hadn't thought of ensembling the results. What masks did you use to train the unet?</p>",
      "rawMarkdown": "Hi Simon, thanks for shaing your approach.  What was your encoder on your unet? Did you use the same ResNeXt50? I was planning to try a unet as the this is really a pixel classification problem and though many of the pixels in a patch may be non-tumor we still want to classify it as tumor. I tried this on the skin cancer data set and could not get better results than a regular resnet but I hadn't thought of ensembling the results. What masks did you use to train the unet?",
      "votes": null
    },
    {
      "id": "509004",
      "postDate": "04/07/2019 07:08:32",
      "content": "<p>Hi Mark.\nAbout U-Net(s) : I love autoencoder-like designs, since they provide robustness, noise filtering and ability to retain valuable information in most cases. U-Nets are used sometimes for providing masks, and focus model's attention on specific features. But when used as mask generator, they tend to remove \"side\" or \"contextual\" informations, but provide faster and deeper understanding of the features.\nLike you, I tried generating masks, but it provided no benefit (about 96%ROC).\nWhat I like in U-Nets is their ability to combine low and high-level features. My final \"U-net/ResNet\" was a bit like:\n- U-net from 224x224x3 to 7x7x1024 in 5 steps, combining each features with \"reconstructed\" features (see <a href=\"http://deeplearning.net/tutorial/_images/unet.jpg\">http://deeplearning.net/tutorial/_images/unet.jpg</a>)\n- After reconstructing the 4th step (112x112x64), adding features to the output first Conv layer of a classical ResNet (<a href=\"https://i.stack.imgur.com/XTo6Q.png\">https://i.stack.imgur.com/XTo6Q.png</a>) with feature depth=64.\n- Total depth is then 128, inheriting 64 from U-Net and 64 from first conv layer of ResNet. From here, it is fed to the rest of a classical ResNet.\n- Total final number of layers is quite big. Model is very slow to train from scratch.\n- Main idea was to provide deep layers with different \"views\" and \"reconstructions\" of the features, hoping that mixing these would provide a deeper and wider understanding of the features.</p>\n\n<p>It didn't really succeeded, with results about 1% under the resNeXt50, BUT : averaging the 2 gave the best results !!! I think these 2 different networks had very different opinions about hard samples (only about 95% predictions overlapping)</p>",
      "rawMarkdown": "Hi Mark.\nAbout U-Net(s) : I love autoencoder-like designs, since they provide robustness, noise filtering and ability to retain valuable information in most cases. U-Nets are used sometimes for providing masks, and focus model's attention on specific features. But when used as mask generator, they tend to remove \"side\" or \"contextual\" informations, but provide faster and deeper understanding of the features.\nLike you, I tried generating masks, but it provided no benefit (about 96%ROC).\nWhat I like in U-Nets is their ability to combine low and high-level features. My final \"U-net/ResNet\" was a bit like:\n- U-net from 224x224x3 to 7x7x1024 in 5 steps, combining each features with \"reconstructed\" features (see http://deeplearning.net/tutorial/_images/unet.jpg)\n- After reconstructing the 4th step (112x112x64), adding features to the output first Conv layer of a classical ResNet (https://i.stack.imgur.com/XTo6Q.png) with feature depth=64.\n- Total depth is then 128, inheriting 64 from U-Net and 64 from first conv layer of ResNet. From here, it is fed to the rest of a classical ResNet.\n- Total final number of layers is quite big. Model is very slow to train from scratch.\n- Main idea was to provide deep layers with different \"views\" and \"reconstructions\" of the features, hoping that mixing these would provide a deeper and wider understanding of the features.\n\nIt didn't really succeeded, with results about 1% under the resNeXt50, BUT : averaging the 2 gave the best results !!! I think these 2 different networks had very different opinions about hard samples (only about 95% predictions overlapping)",
      "votes": null
    },
    {
      "id": "510138",
      "postDate": "04/08/2019 17:47:19",
      "content": "<p><a href=\"/simoninparis\">@simoninparis</a> thanks for the detail, a very interesting approach..</p>",
      "rawMarkdown": "simoninparis thanks for the detail, a very interesting approach..",
      "votes": null
    },
    {
      "id": "570984",
      "postDate": "07/09/2019 02:21:41",
      "content": "<p>great summary Simon!</p>",
      "rawMarkdown": "great summary Simon!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 507954,
      "author_name": "ivanpan",
      "author_url": "",
      "post_date": "04/05/2019 12:39:46",
      "content": "<p>Interesting to use U-net. Thanks for sharing, Simon! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 508274,
      "author_name": "markdesimone",
      "author_url": "",
      "post_date": "04/05/2019 23:00:40",
      "content": "<p>Hi Simon, thanks for shaing your approach.  What was your encoder on your unet? Did you use the same ResNeXt50? I was planning to try a unet as the this is really a pixel classification problem and though many of the pixels in a patch may be non-tumor we still want to classify it as tumor. I tried this on the skin cancer data set and could not get better results than a regular resnet but I hadn't thought of ensembling the results. What masks did you use to train the unet?</p>",
      "votes": null,
      "replies": [
        {
          "id": 509004,
          "author_name": "simoninparis",
          "author_url": "",
          "post_date": "04/07/2019 07:08:32",
          "content": "<p>Hi Mark.\nAbout U-Net(s) : I love autoencoder-like designs, since they provide robustness, noise filtering and ability to retain valuable information in most cases. U-Nets are used sometimes for providing masks, and focus model's attention on specific features. But when used as mask generator, they tend to remove \"side\" or \"contextual\" informations, but provide faster and deeper understanding of the features.\nLike you, I tried generating masks, but it provided no benefit (about 96%ROC).\nWhat I like in U-Nets is their ability to combine low and high-level features. My final \"U-net/ResNet\" was a bit like:\n- U-net from 224x224x3 to 7x7x1024 in 5 steps, combining each features with \"reconstructed\" features (see <a href=\"http://deeplearning.net/tutorial/_images/unet.jpg\">http://deeplearning.net/tutorial/_images/unet.jpg</a>)\n- After reconstructing the 4th step (112x112x64), adding features to the output first Conv layer of a classical ResNet (<a href=\"https://i.stack.imgur.com/XTo6Q.png\">https://i.stack.imgur.com/XTo6Q.png</a>) with feature depth=64.\n- Total depth is then 128, inheriting 64 from U-Net and 64 from first conv layer of ResNet. From here, it is fed to the rest of a classical ResNet.\n- Total final number of layers is quite big. Model is very slow to train from scratch.\n- Main idea was to provide deep layers with different \"views\" and \"reconstructions\" of the features, hoping that mixing these would provide a deeper and wider understanding of the features.</p>\n\n<p>It didn't really succeeded, with results about 1% under the resNeXt50, BUT : averaging the 2 gave the best results !!! I think these 2 different networks had very different opinions about hard samples (only about 95% predictions overlapping)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 510138,
          "author_name": "markdesimone",
          "author_url": "",
          "post_date": "04/08/2019 17:47:19",
          "content": "<p><a href=\"/simoninparis\">@simoninparis</a> thanks for the detail, a very interesting approach..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 570984,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "07/09/2019 02:21:41",
      "content": "<p>great summary Simon!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "507779": "## Result = 97.23 private LB\nwith a very strange score of 96.93 on public LB...\n\n**- Augmentation :** DoG (per-channel difference of gaussian) + per-channel normalization + vert/horiz flip + hue + per-channel random gain\n\n**Model 1 :**\n- ResNeXt50 + ResNet50 + custom pure ConvNet\n\n**Model 2 :**\n- custom \"U-Net\" + custom \"ResNet\" (sort of auto-encoder + residual convnet), slow to train from scratch, but very cool results.\npaper :\nhttps://lmb.informatik.uni-freiburg.de/people/ronneber/u-net/\n\n- each model 16 fold\n\n**Training :**\n- simple brute force Adam (0.0001) with 0.95x decay after each epoch for each model\n- keeping a eye on overfitting helped deciding when to stop learning.\n\n**Ensembling :**\nBest result was achieved with simple averaging (60/40%) of ResNeXt50 + U-net\nThese 2 models seem to have the most different \"undertstanding\" of the features.\nEven though U-Net alone had poor results (about 97%ROC on local data), mixing it with ResNeXt50 (about 99.2ROC) was a good idea. These 2 models don't overlap a lot, and averaging the 2 provided clearly the best results. Funny fact!\n\n**What was good :**\n- Difference of Gaussian was a big improvement, yielding sometimes strange results on some sample images. Improved both learning speeds and results.\n\n**What was not good :**\n- Ensembling too similar models yields no better results (ex. ResNet50 + VGG)\n- Upscaling (to 224x224) didn't improve overall performance, especially on custom designed networks.\n\n**What I should have done :**\n- Try to get rid of the (HUGE!) JPEG artefacts, since it really polluted the data.\npaper to implement :\nhttps://arxiv.org/abs/1605.00366\nBTW :  no real-life medical documents should ever be JPG-compressed... (and converted back to TIF!)\nI'm pretty sure this would have added robustness, as some image samples were heavily distorded by JPEGing, especially visible after DoG and per-channel normalization.\n\n- Maybe Pseudo-labelling on \"U-Net\" way have improved results on 'unsure' results.\n\n**What I didn't want to do :**\n- cheating with the original Cam16 files\n- WSI things... sorry, not interesting for me.\n\nThis was my first Kaggle competition. A lot of fun, a lot of things learned.\nThanks to all this cool community!\n\nSimon (France)",
    "507954": "Interesting to use U-net. Thanks for sharing, Simon!",
    "508274": "Hi Simon, thanks for shaing your approach.  What was your encoder on your unet? Did you use the same ResNeXt50? I was planning to try a unet as the this is really a pixel classification problem and though many of the pixels in a patch may be non-tumor we still want to classify it as tumor. I tried this on the skin cancer data set and could not get better results than a regular resnet but I hadn't thought of ensembling the results. What masks did you use to train the unet?",
    "509004": "Hi Mark.\nAbout U-Net(s) : I love autoencoder-like designs, since they provide robustness, noise filtering and ability to retain valuable information in most cases. U-Nets are used sometimes for providing masks, and focus model's attention on specific features. But when used as mask generator, they tend to remove \"side\" or \"contextual\" informations, but provide faster and deeper understanding of the features.\nLike you, I tried generating masks, but it provided no benefit (about 96%ROC).\nWhat I like in U-Nets is their ability to combine low and high-level features. My final \"U-net/ResNet\" was a bit like:\n- U-net from 224x224x3 to 7x7x1024 in 5 steps, combining each features with \"reconstructed\" features (see http://deeplearning.net/tutorial/_images/unet.jpg)\n- After reconstructing the 4th step (112x112x64), adding features to the output first Conv layer of a classical ResNet (https://i.stack.imgur.com/XTo6Q.png) with feature depth=64.\n- Total depth is then 128, inheriting 64 from U-Net and 64 from first conv layer of ResNet. From here, it is fed to the rest of a classical ResNet.\n- Total final number of layers is quite big. Model is very slow to train from scratch.\n- Main idea was to provide deep layers with different \"views\" and \"reconstructions\" of the features, hoping that mixing these would provide a deeper and wider understanding of the features.\n\nIt didn't really succeeded, with results about 1% under the resNeXt50, BUT : averaging the 2 gave the best results !!! I think these 2 different networks had very different opinions about hard samples (only about 95% predictions overlapping)",
    "510138": "simoninparis thanks for the detail, a very interesting approach..",
    "570984": "great summary Simon!"
  },
  "source": "meta"
}