{
  "id": 176382,
  "title": "Report of 408 submissions!",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/176382",
  "author_name": "Amin",
  "post_date": "2020-08-21T14:26:18.549000",
  "votes": 23,
  "comment_count": 21,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers, Kaggle and my teammates <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a>, <a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> and <a href=\"https://www.kaggle.com/nicohrubec\" target=\"_blank\">@nicohrubec</a> for this amazing experience.</p>\n<p>I am happy that I picked SIIM-Melanoma to be my first competition because it had everything: A great leader <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, active teammates, unstable public LB, a terrible shake-up and finally, the drama related to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174635\" target=\"_blank\">the participation of the brilliant students of the medal farming academy</a>.</p>\n<p>We won a bronze medal thanks to the clean up done by kaggle and the investigation started by <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a>. Big thanks!</p>\n<p>As most of the participants, we didn’t select our best submission that would have landed us a silver medal. Actually 60 submissions scored higher than our best selected sub. This is part of the game, taking good decisions is an essential skill in ML/DL. Here is a first glimpse of our submissions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F45cdd5353e253758de95d199217756ce%2FScreen%20Shot%202020-08-21%20at%2017.03.07.png?generation=1598015028261645&amp;alt=media\" alt=\"\"></p>\n<h3>First time crossing the medal zone:</h3>\n<p>The first time I crossed the medal zone was the 15th of July with private LB= 0.9375 (Slightly higher than our best selected submission).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F7ccedce9337d62222933c5ae7acf95ea%2FScreen%20Shot%202020-08-21%20at%2017.00.32.png?generation=1598014866149997&amp;alt=media\" alt=\"\"><br>\nWhy did this sub score higher than our selected submission? Because it had an ingredient that was missing in the final sub: Higher <code>image_size= 768</code>. I was experimenting with higher image sizes but keeping the CNN's <code>input_size</code> smaller for faster training, it's a clear information loss/training time trade-off. </p>\n<pre><code>Image size= 512. Input size= 226\n\nImage size= 768. Input size= 384\n\nImage size= 1024 Input size= 512\n</code></pre>\n<p>It scored higher than the models with <code>image_size</code> 512 or 384, which is good news, but I decided not to go further with this experiment because it was risky since parts of the images are lost with the center crops.</p>\n<p>I was happy with this decision until I saw <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344\" target=\"_blank\">solution</a> where he used smaller <code>input_size</code> with random crops instead of center crops at each epoch to make sure the model loops over all the pixels instead of center cropping the same pixels over the epochs. Another trick learned to be used in future competitions ;)</p>\n<h3>Teaming-up:</h3>\n<p>The day we teamed-up has seen a boost in our private LB score. It was just a rank ensemble of our best submissions. This shows the clear benefit of teaming up late because you will definitely add diversity to the final solution.</p>\n<h3>The team era:</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F0ff03771186de3fafc3b9eed04860d8c%2FScreen%20Shot%202020-08-21%20at%2017.12.11.png?generation=1598015689624500&amp;alt=media\" alt=\"\"><br>\nYou can see in the plot that the best era was after teaming up because we experimented many ideas together that helped us reach the silver zone at our peak. I am not gonna lie, at that point I thought we were overfitting to public LB because most of the silver zone submissions are nested blends of our +0.95 submissions. However, the correlation of our public LB scores in the range (0.955, 0.9642) with private LB was 0.89! The question here is: When did we start overfitting?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fa319913972edb44e93fc1c68c56dd4c2%2FScreen%20Shot%202020-08-21%20at%2017.24.13.png?generation=1598016297260469&amp;alt=media\" alt=\"\"><br>\nOverfitting (the decay) starts after including the <a href=\"https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619\" target=\"_blank\">last public notebook in our blend that was terribly overfitting (Public= 0.96, Private= 0.91)</a>. The blends scoring +0.9643 included overfitted public notebooks</p>\n<p>It is worth mentioning that the highest public/private LB correlation we had was with subs in the medal zone, so our subs definitely had something special to score high private scores without any CV strategy. What ideas or ingredients did the silver submissions have? Simple answer: Post-processing.</p>\n<h3>Post-processing:</h3>\n<p>We had 2 post-processing techniques, they both boost the public and private LB score.<br>\nThe first was a finding by <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> where he set all the palms/soles and oral/genital predictions to 0, code below:</p>\n<pre><code>ps= test.loc[test['anatom_site_general_challenge']=='palms/soles'].image_name\nog= test.loc[test['anatom_site_general_challenge']=='oral/genital'].image_name\nsubmission= submission.set_index('image_name')\nsubmission.loc[ps]=0\nsubmission.loc[og]=0\n</code></pre>\n<p>This pp gives a +0.001 boost in private LB (Try it at home! +0.001 boost guaranteed 😉).</p>\n<p>The second was image embeddings extraction using RAPIDS cuML TSNE shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks Chris for everything you shared in this competition.<br>\n<a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> tried to detect the malignant images in the test set by extracting the patients around the red area (malignants) as they are more susceptible. This pp was unstable but led to some high private LB scores.<br>\n(If you have any questions about this method, please tag <a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> in the comment)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F55bebda96a72c7bca0fc39dfa2bae19b%2FScreen%20Shot%202020-08-21%20at%2016.36.14.png?generation=1598013477503902&amp;alt=media\" alt=\"\">:</p>\n<h3>The final solution :</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F6312bdd9b6c769f81efb3dccdeefaccd%2FScreen%20Shot%202020-08-23%20at%2010.36.49.png?generation=1598164656057343&amp;alt=media\" alt=\"\"><br>\n<strong>The blue line marks the switch from trusting(climbing) LB to trusting CV.</strong></p>\n<p>The final step in this competition was to prepare a solid solution.  Even though we had some high scoring subs, interesting ideas (training with high resolution images and smaller CNN input size + post-processing), we decided to do it the safest way and not use any of them and stack models trained with different architecture (B3-B7 + seresnext50) and different image sizes (384, 512).<br>\nWith stacking, we managed to reduce the gap between CV/LB to 0 in our best selected submission, where CV=LB=0.9483!</p>\n<h3>Conclusion:</h3>\n<p>We ended up with a very correct solution that is missing the special ingredient to make it a top solution. However, I am happy to see that the ideas we have been testing actually work, even though not using them in our final solution penalized our private LB score. Those are techniques learned to be used in future competitions.</p>\n<p><strong>Best private and public subs:</strong></p>\n<table>\n  <tbody><tr>\n    <th>Submission</th>\n    <th>CV</th>\n    <th>Public LB</th>\n    <th>Private LB</th>\n\n  </tr>\n  <tr>\n    <td>Best public </td>\n    <td>N/A</td>\n    <td>0.9657</td>\n    <td>0.9374</td>\n  </tr>\n  <tr>\n    <td>Best private</td>\n    <td>N/A</td>\n    <td>0.9613</td>\n    <td>0.9410</td>\n  </tr>\n\n\n</tbody></table>\n<p><strong>3 Selected submissions</strong>    </p>\n<table>\n  <tbody><tr>\n    <th>Submission</th>\n    <th>CV</th>\n    <th>Public LB</th>\n    <th>Private LB</th>\n\n  </tr>\n\n\n\n   <tr>\n    <td>Most stable</td>\n    <td>0.9483</td>\n    <td>0.9483</td>\n    <td>0.9374</td>\n  </tr>\n\n   <tr>\n    <td>Best CV</td>\n    <td>0.9541</td>\n    <td>0.9510</td>\n    <td>0.9359</td>\n  </tr>\n  <tr>\n    <td>2 Best CV</td>\n    <td>0.9535</td>\n    <td>0.9526</td>\n    <td>0.9337</td>\n  </tr>\n</tbody></table>",
  "messages": [
    {
      "id": 980349,
      "postDate": "2020-08-21T14:26:18.550Z",
      "content": "<p>First of all, I would like to thank the organizers, Kaggle and my teammates <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a>, <a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> and <a href=\"https://www.kaggle.com/nicohrubec\" target=\"_blank\">@nicohrubec</a> for this amazing experience.</p>\n<p>I am happy that I picked SIIM-Melanoma to be my first competition because it had everything: A great leader <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, active teammates, unstable public LB, a terrible shake-up and finally, the drama related to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174635\" target=\"_blank\">the participation of the brilliant students of the medal farming academy</a>.</p>\n<p>We won a bronze medal thanks to the clean up done by kaggle and the investigation started by <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a>. Big thanks!</p>\n<p>As most of the participants, we didn’t select our best submission that would have landed us a silver medal. Actually 60 submissions scored higher than our best selected sub. This is part of the game, taking good decisions is an essential skill in ML/DL. Here is a first glimpse of our submissions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F45cdd5353e253758de95d199217756ce%2FScreen%20Shot%202020-08-21%20at%2017.03.07.png?generation=1598015028261645&amp;alt=media\" alt=\"\"></p>\n<h3>First time crossing the medal zone:</h3>\n<p>The first time I crossed the medal zone was the 15th of July with private LB= 0.9375 (Slightly higher than our best selected submission).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F7ccedce9337d62222933c5ae7acf95ea%2FScreen%20Shot%202020-08-21%20at%2017.00.32.png?generation=1598014866149997&amp;alt=media\" alt=\"\"><br>\nWhy did this sub score higher than our selected submission? Because it had an ingredient that was missing in the final sub: Higher <code>image_size= 768</code>. I was experimenting with higher image sizes but keeping the CNN's <code>input_size</code> smaller for faster training, it's a clear information loss/training time trade-off. </p>\n<pre><code>Image size= 512. Input size= 226\n\nImage size= 768. Input size= 384\n\nImage size= 1024 Input size= 512\n</code></pre>\n<p>It scored higher than the models with <code>image_size</code> 512 or 384, which is good news, but I decided not to go further with this experiment because it was risky since parts of the images are lost with the center crops.</p>\n<p>I was happy with this decision until I saw <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344\" target=\"_blank\">solution</a> where he used smaller <code>input_size</code> with random crops instead of center crops at each epoch to make sure the model loops over all the pixels instead of center cropping the same pixels over the epochs. Another trick learned to be used in future competitions ;)</p>\n<h3>Teaming-up:</h3>\n<p>The day we teamed-up has seen a boost in our private LB score. It was just a rank ensemble of our best submissions. This shows the clear benefit of teaming up late because you will definitely add diversity to the final solution.</p>\n<h3>The team era:</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F0ff03771186de3fafc3b9eed04860d8c%2FScreen%20Shot%202020-08-21%20at%2017.12.11.png?generation=1598015689624500&amp;alt=media\" alt=\"\"><br>\nYou can see in the plot that the best era was after teaming up because we experimented many ideas together that helped us reach the silver zone at our peak. I am not gonna lie, at that point I thought we were overfitting to public LB because most of the silver zone submissions are nested blends of our +0.95 submissions. However, the correlation of our public LB scores in the range (0.955, 0.9642) with private LB was 0.89! The question here is: When did we start overfitting?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fa319913972edb44e93fc1c68c56dd4c2%2FScreen%20Shot%202020-08-21%20at%2017.24.13.png?generation=1598016297260469&amp;alt=media\" alt=\"\"><br>\nOverfitting (the decay) starts after including the <a href=\"https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619\" target=\"_blank\">last public notebook in our blend that was terribly overfitting (Public= 0.96, Private= 0.91)</a>. The blends scoring +0.9643 included overfitted public notebooks</p>\n<p>It is worth mentioning that the highest public/private LB correlation we had was with subs in the medal zone, so our subs definitely had something special to score high private scores without any CV strategy. What ideas or ingredients did the silver submissions have? Simple answer: Post-processing.</p>\n<h3>Post-processing:</h3>\n<p>We had 2 post-processing techniques, they both boost the public and private LB score.<br>\nThe first was a finding by <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> where he set all the palms/soles and oral/genital predictions to 0, code below:</p>\n<pre><code>ps= test.loc[test['anatom_site_general_challenge']=='palms/soles'].image_name\nog= test.loc[test['anatom_site_general_challenge']=='oral/genital'].image_name\nsubmission= submission.set_index('image_name')\nsubmission.loc[ps]=0\nsubmission.loc[og]=0\n</code></pre>\n<p>This pp gives a +0.001 boost in private LB (Try it at home! +0.001 boost guaranteed 😉).</p>\n<p>The second was image embeddings extraction using RAPIDS cuML TSNE shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks Chris for everything you shared in this competition.<br>\n<a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> tried to detect the malignant images in the test set by extracting the patients around the red area (malignants) as they are more susceptible. This pp was unstable but led to some high private LB scores.<br>\n(If you have any questions about this method, please tag <a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> in the comment)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F55bebda96a72c7bca0fc39dfa2bae19b%2FScreen%20Shot%202020-08-21%20at%2016.36.14.png?generation=1598013477503902&amp;alt=media\" alt=\"\">:</p>\n<h3>The final solution :</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F6312bdd9b6c769f81efb3dccdeefaccd%2FScreen%20Shot%202020-08-23%20at%2010.36.49.png?generation=1598164656057343&amp;alt=media\" alt=\"\"><br>\n<strong>The blue line marks the switch from trusting(climbing) LB to trusting CV.</strong></p>\n<p>The final step in this competition was to prepare a solid solution.  Even though we had some high scoring subs, interesting ideas (training with high resolution images and smaller CNN input size + post-processing), we decided to do it the safest way and not use any of them and stack models trained with different architecture (B3-B7 + seresnext50) and different image sizes (384, 512).<br>\nWith stacking, we managed to reduce the gap between CV/LB to 0 in our best selected submission, where CV=LB=0.9483!</p>\n<h3>Conclusion:</h3>\n<p>We ended up with a very correct solution that is missing the special ingredient to make it a top solution. However, I am happy to see that the ideas we have been testing actually work, even though not using them in our final solution penalized our private LB score. Those are techniques learned to be used in future competitions.</p>\n<p><strong>Best private and public subs:</strong></p>\n<table>\n  <tbody><tr>\n    <th>Submission</th>\n    <th>CV</th>\n    <th>Public LB</th>\n    <th>Private LB</th>\n\n  </tr>\n  <tr>\n    <td>Best public </td>\n    <td>N/A</td>\n    <td>0.9657</td>\n    <td>0.9374</td>\n  </tr>\n  <tr>\n    <td>Best private</td>\n    <td>N/A</td>\n    <td>0.9613</td>\n    <td>0.9410</td>\n  </tr>\n\n\n</tbody></table>\n<p><strong>3 Selected submissions</strong>    </p>\n<table>\n  <tbody><tr>\n    <th>Submission</th>\n    <th>CV</th>\n    <th>Public LB</th>\n    <th>Private LB</th>\n\n  </tr>\n\n\n\n   <tr>\n    <td>Most stable</td>\n    <td>0.9483</td>\n    <td>0.9483</td>\n    <td>0.9374</td>\n  </tr>\n\n   <tr>\n    <td>Best CV</td>\n    <td>0.9541</td>\n    <td>0.9510</td>\n    <td>0.9359</td>\n  </tr>\n  <tr>\n    <td>2 Best CV</td>\n    <td>0.9535</td>\n    <td>0.9526</td>\n    <td>0.9337</td>\n  </tr>\n</tbody></table>",
      "rawMarkdown": "First of all, I would like to thank the organizers, Kaggle and my teammates @underwearfitting, @hawkey and @nicohrubec for this amazing experience.\n\nI am happy that I picked SIIM-Melanoma to be my first competition because it had everything: A great leader @cdeotte, active teammates, unstable public LB, a terrible shake-up and finally, the drama related to [the participation of the brilliant students of the medal farming academy](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174635).\n\nWe won a bronze medal thanks to the clean up done by kaggle and the investigation started by @group16. Big thanks!\n\nAs most of the participants, we didn’t select our best submission that would have landed us a silver medal. Actually 60 submissions scored higher than our best selected sub. This is part of the game, taking good decisions is an essential skill in ML/DL. Here is a first glimpse of our submissions.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F45cdd5353e253758de95d199217756ce%2FScreen%20Shot%202020-08-21%20at%2017.03.07.png?generation=1598015028261645&alt=media)\n### First time crossing the medal zone:\nThe first time I crossed the medal zone was the 15th of July with private LB= 0.9375 (Slightly higher than our best selected submission).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F7ccedce9337d62222933c5ae7acf95ea%2FScreen%20Shot%202020-08-21%20at%2017.00.32.png?generation=1598014866149997&alt=media)\nWhy did this sub score higher than our selected submission? Because it had an ingredient that was missing in the final sub: Higher `image_size= 768`. I was experimenting with higher image sizes but keeping the CNN's `input_size` smaller for faster training, it's a clear information loss/training time trade-off. \n```\nImage size= 512. Input size= 226\n\nImage size= 768. Input size= 384\n\nImage size= 1024 Input size= 512\n```\nIt scored higher than the models with `image_size` 512 or 384, which is good news, but I decided not to go further with this experiment because it was risky since parts of the images are lost with the center crops.\n\nI was happy with this decision until I saw @cdeotte's [solution](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344) where he used smaller `input_size` with random crops instead of center crops at each epoch to make sure the model loops over all the pixels instead of center cropping the same pixels over the epochs. Another trick learned to be used in future competitions ;)\n\n### Teaming-up:\nThe day we teamed-up has seen a boost in our private LB score. It was just a rank ensemble of our best submissions. This shows the clear benefit of teaming up late because you will definitely add diversity to the final solution.\n\n### The team era:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F0ff03771186de3fafc3b9eed04860d8c%2FScreen%20Shot%202020-08-21%20at%2017.12.11.png?generation=1598015689624500&alt=media)\nYou can see in the plot that the best era was after teaming up because we experimented many ideas together that helped us reach the silver zone at our peak. I am not gonna lie, at that point I thought we were overfitting to public LB because most of the silver zone submissions are nested blends of our +0.95 submissions. However, the correlation of our public LB scores in the range (0.955, 0.9642) with private LB was 0.89! The question here is: When did we start overfitting?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fa319913972edb44e93fc1c68c56dd4c2%2FScreen%20Shot%202020-08-21%20at%2017.24.13.png?generation=1598016297260469&alt=media)\nOverfitting (the decay) starts after including the [last public notebook in our blend that was terribly overfitting (Public= 0.96, Private= 0.91)](https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619). The blends scoring +0.9643 included overfitted public notebooks\n \nIt is worth mentioning that the highest public/private LB correlation we had was with subs in the medal zone, so our subs definitely had something special to score high private scores without any CV strategy. What ideas or ingredients did the silver submissions have? Simple answer: Post-processing.\n\n\n### Post-processing:\nWe had 2 post-processing techniques, they both boost the public and private LB score.\nThe first was a finding by @underwearfitting where he set all the palms/soles and oral/genital predictions to 0, code below:\n```\nps= test.loc[test['anatom_site_general_challenge']=='palms/soles'].image_name\nog= test.loc[test['anatom_site_general_challenge']=='oral/genital'].image_name\nsubmission= submission.set_index('image_name')\nsubmission.loc[ps]=0\nsubmission.loc[og]=0\n```\n\n\nThis pp gives a +0.001 boost in private LB (Try it at home! +0.001 boost guaranteed 😉).\n\nThe second was image embeddings extraction using RAPIDS cuML TSNE shared by @cdeotte, thanks Chris for everything you shared in this competition.\n@hawkey tried to detect the malignant images in the test set by extracting the patients around the red area (malignants) as they are more susceptible. This pp was unstable but led to some high private LB scores.\n(If you have any questions about this method, please tag @hawkey in the comment)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F55bebda96a72c7bca0fc39dfa2bae19b%2FScreen%20Shot%202020-08-21%20at%2016.36.14.png?generation=1598013477503902&alt=media):\n\n### The final solution :\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F6312bdd9b6c769f81efb3dccdeefaccd%2FScreen%20Shot%202020-08-23%20at%2010.36.49.png?generation=1598164656057343&alt=media)\n**The blue line marks the switch from trusting(climbing) LB to trusting CV.**\n\nThe final step in this competition was to prepare a solid solution.  Even though we had some high scoring subs, interesting ideas (training with high resolution images and smaller CNN input size + post-processing), we decided to do it the safest way and not use any of them and stack models trained with different architecture (B3-B7 + seresnext50) and different image sizes (384, 512).\nWith stacking, we managed to reduce the gap between CV/LB to 0 in our best selected submission, where CV=LB=0.9483!\n\n\n### Conclusion:\nWe ended up with a very correct solution that is missing the special ingredient to make it a top solution. However, I am happy to see that the ideas we have been testing actually work, even though not using them in our final solution penalized our private LB score. Those are techniques learned to be used in future competitions.\n\n**Best private and public subs:**\n<table >\n  <tr>\n    <th>Submission</th>\n    <th>CV</th>\n    <th>Public LB</th>\n    <th>Private LB</th>\n    \n  </tr>\n  <tr>\n    <td>Best public </td>\n    <td>N/A</td>\n    <td>0.9657</td>\n    <td>0.9374</td>\n  </tr>\n  <tr>\n    <td>Best private</td>\n    <td>N/A</td>\n    <td>0.9613</td>\n    <td>0.9410</td>\n  </tr>\n    \n\n</table>\n**3 Selected submissions**    \n\n<table >\n  <tr>\n    <th>Submission</th>\n    <th>CV</th>\n    <th>Public LB</th>\n    <th>Private LB</th>\n    \n  </tr>\n\n    \n\n   <tr>\n    <td>Most stable</td>\n    <td>0.9483</td>\n    <td>0.9483</td>\n    <td>0.9374</td>\n  </tr>\n    \n   <tr>\n    <td>Best CV</td>\n    <td>0.9541</td>\n    <td>0.9510</td>\n    <td>0.9359</td>\n  </tr>\n  <tr>\n    <td>2 Best CV</td>\n    <td>0.9535</td>\n    <td>0.9526</td>\n    <td>0.9337</td>\n  </tr>\n</table>\n",
      "votes": 21
    },
    {
      "id": 981964,
      "postDate": "2020-08-22T22:19:29.583Z",
      "content": "<p>It was a great pleasure teaming up with you Amin. I definitely learned a lot from you and I really enjoyed our teamwork. We relentlessly trained so many models here and there, expecting to ensemble to get good results. However, this time we lost to high quality models, in other words, we couldn't build models that score around 0.94 CV due to our lack of experience. We learned a lot in this competition from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> about the usage of TPU and TSNE projections, now the winning solutions from all the winners; I can't say I am not satisfied😄. It was great to have diverse teammates like <a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> <a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> <a href=\"https://www.kaggle.com/nicohrubec\" target=\"_blank\">@nicohrubec</a>, coming from different background and cultures. Hoping to connect with you all soon.😬</p>",
      "rawMarkdown": "It was a great pleasure teaming up with you Amin. I definitely learned a lot from you and I really enjoyed our teamwork. We relentlessly trained so many models here and there, expecting to ensemble to get good results. However, this time we lost to high quality models, in other words, we couldn't build models that score around 0.94 CV due to our lack of experience. We learned a lot in this competition from @cdeotte about the usage of TPU and TSNE projections, now the winning solutions from all the winners; I can't say I am not satisfied😄. It was great to have diverse teammates like @hawkey @amiiiney @nicohrubec, coming from different background and cultures. Hoping to connect with you all soon.😬",
      "votes": 3,
      "replies": [
        {
          "id": 981969,
          "postDate": "2020-08-22T22:31:23.040Z",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> Thanks Eric for your kind words! Teaming up was the best thing I did in this competition,  I definitely learned a lot from you, you would make a great professor after your PhD 😊<br>\nI also wouldn't say I am not satisfied with the results because it was a solid solution that was missing some ingredients that we could find and learn from the top solutions, now our baseline/pipeline is richer and ready for new challenges :)</p>",
          "rawMarkdown": "@underwearfitting Thanks Eric for your kind words! Teaming up was the best thing I did in this competition,  I definitely learned a lot from you, you would make a great professor after your PhD 😊\nI also wouldn't say I am not satisfied with the results because it was a solid solution that was missing some ingredients that we could find and learn from the top solutions, now our baseline/pipeline is richer and ready for new challenges :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 980877,
      "postDate": "2020-08-22T00:30:11.910Z",
      "content": "<p>Wow, that's magic. My best (unselected submission) was Private LB 0.9450. When I use your trick it increases to Private LB 0.9459 Gold Medal !!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe6cbf0c0f839fe713317ddbf17a1d161%2Fbest.png?generation=1598055407574078&amp;alt=media\" alt=\"\"></p>\n<p>I did some quick probing where I set all targets to zero and then all of a specific body part to 1. From the submissions we can see that it is rare for there to be positive cases on soles, palms, oral, and genital. Great detective work!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fdaea882ba0b991086abdcc7142decef2%2Fprobe.png?generation=1598056151310064&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Wow, that's magic. My best (unselected submission) was Private LB 0.9450. When I use your trick it increases to Private LB 0.9459 Gold Medal !!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe6cbf0c0f839fe713317ddbf17a1d161%2Fbest.png?generation=1598055407574078&alt=media)\n\nI did some quick probing where I set all targets to zero and then all of a specific body part to 1. From the submissions we can see that it is rare for there to be positive cases on soles, palms, oral, and genital. Great detective work!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fdaea882ba0b991086abdcc7142decef2%2Fprobe.png?generation=1598056151310064&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 980878,
          "postDate": "2020-08-22T00:38:14.563Z",
          "content": "<p>Actually the above probes prove there are no malignant palm, sole, oral, nor genital in test.</p>\n<p>There are approximately 115 images out of 10,000 test images with palm/sole (based on fact that train has 1.1% palm/sole). And there are approximately 38 test images with oral/genital. So these probe submissions, imply there are no malignant palm, sole, oral, nor genital in the test images. Since a benign will lower LB score by approx 0.00015 and a malignant would increase 0.006</p>",
          "rawMarkdown": "Actually the above probes prove there are no malignant palm, sole, oral, nor genital in test.\n\nThere are approximately 115 images out of 10,000 test images with palm/sole (based on fact that train has 1.1% palm/sole). And there are approximately 38 test images with oral/genital. So these probe submissions, imply there are no malignant palm, sole, oral, nor genital in the test images. Since a benign will lower LB score by approx 0.00015 and a malignant would increase 0.006",
          "votes": 2
        },
        {
          "id": 981913,
          "postDate": "2020-08-22T20:46:33.790Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for trying our method (by detective <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> :) )! We haven't used it either in our final solution, unfortunately, it turned out that there was really no palm, sole, oral nor genital in test, but in case there were, as you said, this could lower significantly our private score, maybe we should have risked with this pp in one of the 3 subs.</p>",
          "rawMarkdown": "Thanks @cdeotte for trying our method (by detective @underwearfitting :) )! We haven't used it either in our final solution, unfortunately, it turned out that there was really no palm, sole, oral nor genital in test, but in case there were, as you said, this could lower significantly our private score, maybe we should have risked with this pp in one of the 3 subs.",
          "votes": 1
        },
        {
          "id": 981955,
          "postDate": "2020-08-22T22:08:46.170Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for trying this. We also learned a lesson about trying out something risky!</p>",
          "rawMarkdown": "@cdeotte thanks for trying this. We also learned a lesson about trying out something risky!",
          "votes": 3
        }
      ]
    },
    {
      "id": 980868,
      "postDate": "2020-08-22T00:09:31.703Z",
      "content": "<p>Great analysis. I love all the plots. It tells a great story. Ensembling and teaming help increase score. What's the blue vertical line in the bottom plot? Your private LB scores were not after the blue line.</p>\n<p>When you say <code>Image size= 512. Input size= 226</code>, how did you reduce 512 to 226? Did you resize or crop? was it random or centered?</p>",
      "rawMarkdown": "Great analysis. I love all the plots. It tells a great story. Ensembling and teaming help increase score. What's the blue vertical line in the bottom plot? Your private LB scores were not after the blue line.\n\nWhen you say `Image size= 512. Input size= 226`, how did you reduce 512 to 226? Did you resize or crop? was it random or centered?",
      "votes": 1,
      "replies": [
        {
          "id": 981908,
          "postDate": "2020-08-22T20:30:57.803Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your compliment, it means a lot coming from you! And thank you again for all the knowledge you shared in this competition.</p>\n<ul>\n<li><p>Regarding the blue vertical line: Before the blue line our strategy was training models and adding them to the blend, this way we were climbing LB. After the blue line we stopped blending, thus stopped climbing LB and started saving the oofs for stacking for a final stable solution. <strong>The blue line marks the switch from trusting LB to trusting CV.</strong></p></li>\n<li><p>I reduced 512 images to 226 but setting the CNN input_size to 226. I used <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\" target=\"_blank\">AgentAuers's notebook</a>, he used 3 parameters,</p></li>\n</ul>\n<pre><code>  read_size         = 256, \n  crop_size         = 250, \n  net_size          = 224,\n</code></pre>\n<p>This is the set of parameters he used. I tweaked the crop_size (in his <code>prepare_image</code>function he used random_crop and centered_crop) but what made my training faster was not tweaking the crop_size but the net_size, I set the CNN <code>net_size =224</code> with <code>read_size=512</code> and I got better results than <code>image_size=256</code> with <code>net_size=224</code>, probably I was doing something wrong but it had a good private score.</p>\n<p>I would like to learn more about your approach because this is a technique I would like to use in future competitions, how did you make you model random_crop different crops at each epoch? <br>\nIs this augmentation line of code enough to make the training go correctly?<br>\n<code>img = tf.image.random_crop(img, crop_size, crop_size, 3])</code></p>",
          "rawMarkdown": "Thank you @cdeotte for your compliment, it means a lot coming from you! And thank you again for all the knowledge you shared in this competition.\n\n* Regarding the blue vertical line: Before the blue line our strategy was training models and adding them to the blend, this way we were climbing LB. After the blue line we stopped blending, thus stopped climbing LB and started saving the oofs for stacking for a final stable solution. **The blue line marks the switch from trusting LB to trusting CV.**\n\n* I reduced 512 images to 226 but setting the CNN input_size to 226. I used [AgentAuers's notebook](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once), he used 3 parameters,\n```  \n  read_size         = 256, \n  crop_size         = 250, \n  net_size          = 224,\n```\nThis is the set of parameters he used. I tweaked the crop_size (in his `prepare_image`function he used random_crop and centered_crop) but what made my training faster was not tweaking the crop_size but the net_size, I set the CNN `net_size =224` with `read_size=512` and I got better results than `image_size=256` with `net_size=224`, probably I was doing something wrong but it had a good private score.\n\nI would like to learn more about your approach because this is a technique I would like to use in future competitions, how did you make you model random_crop different crops at each epoch? \nIs this augmentation line of code enough to make the training go correctly?\n`img = tf.image.random_crop(img, crop_size, crop_size, 3])`\n"
        },
        {
          "id": 981941,
          "postDate": "2020-08-22T21:41:59.443Z",
          "content": "<p>Using AgentAuers's terminology, your read size was 512 and your net size was 224. What what your crop size?</p>\n<p>The single line <code>img = tf.image.random_crop(img, crop_size, crop_size, 3])</code> does random cropping each epoch if you put it inside the block <code>if augment:</code> inside <code>prepare_image()</code> function.</p>\n<p>Looking closely at AgentAuers's notebook now, I see he uses random cropping. His \"read_size\" are the TFRecords that are used. Then his \"crop_size\" is a random crop done each epoch. Lastly, his \"net_size\" is where he resizes (scales) the random crops. This serves as a blurring effect. The high resolution is reduced to lower resolution.</p>\n<p>Afterward, your network is effectively using <code>X = read_size * net_size / crop_size</code> as it's \"image input size\". So once you set those 3 variables, you can say that your network is learning from images sized <code>X</code>. And to make a diverse set of models you want to build multiple models using <code>X = 256, 384, 512, 768, 1024</code>.</p>",
          "rawMarkdown": "Using AgentAuers's terminology, your read size was 512 and your net size was 224. What what your crop size?\n\nThe single line `img = tf.image.random_crop(img, crop_size, crop_size, 3])` does random cropping each epoch if you put it inside the block `if augment:` inside ` prepare_image()` function.\n\nLooking closely at AgentAuers's notebook now, I see he uses random cropping. His \"read_size\" are the TFRecords that are used. Then his \"crop_size\" is a random crop done each epoch. Lastly, his \"net_size\" is where he resizes (scales) the random crops. This serves as a blurring effect. The high resolution is reduced to lower resolution.\n\nAfterward, your network is effectively using `X = read_size * net_size / crop_size` as it's \"image input size\". So once you set those 3 variables, you can say that your network is learning from images sized `X`. And to make a diverse set of models you want to build multiple models using `X = 256, 384, 512, 768, 1024`.",
          "votes": 1
        },
        {
          "id": 981948,
          "postDate": "2020-08-22T21:57:18.953Z",
          "content": "<p>Thanks again <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.  My crop_size was 256 when image_size=512.<br>\nLets consider just the read_size and crop-size, I understand that if you set the crop size smaller than the read_size and inside the block <code>if augment</code> you use random_crop then <strong>ALL</strong> the images are random cropped differently at every epoch?</p>\n<p>In other words, was your approach similar to AgentAuer's notebook if you set read_size=512 and crop_size=256, with random_crop inside the <code>if augment</code> block?</p>",
          "rawMarkdown": "Thanks again @cdeotte.  My crop_size was 256 when image_size=512.\nLets consider just the read_size and crop-size, I understand that if you set the crop size smaller than the read_size and inside the block `if augment` you use random_crop then **ALL** the images are random cropped differently at every epoch?\n\nIn other words, was your approach similar to AgentAuer's notebook if you set read_size=512 and crop_size=256, with random_crop inside the `if augment` block?\n\n",
          "votes": 1
        },
        {
          "id": 981952,
          "postDate": "2020-08-22T22:06:00.017Z",
          "content": "<p>Yes, my approach was exactly AgentAuer's notebook. (However i didn't realize that AgentAuer's was doing this when i added random crop to my private notebook forked from my public notebook).</p>\n<p>In my final writeup <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344\" target=\"_blank\">here</a>, when i say <code>read_size = 512</code> and <code>crop_size = 256</code>, then that is AgentAuer's <code>read_size = 512</code>, <code>crop_size = 256</code>, <code>net_size = crop_size</code>.</p>\n<p>Each epoch, our models will see <strong>different crops</strong> from the train images. (And after many epochs, it will have learned about the entire images). It looks like you were using random cropping too. And then when you used <code>net_size = 224</code> after <code>crop_size = 256</code>, then you were slightly blurring the images and effectively reducing the image size from 512 to 448.</p>\n<p>(So when you used 512, 256, 224, then you were effectively reading 448 and cropping 224).</p>",
          "rawMarkdown": "Yes, my approach was exactly AgentAuer's notebook. (However i didn't realize that AgentAuer's was doing this when i added random crop to my private notebook forked from my public notebook).\n\nIn my final writeup [here][1], when i say `read_size = 512` and `crop_size = 256`, then that is AgentAuer's `read_size = 512`, `crop_size = 256`, `net_size = crop_size`.\n\nEach epoch, our models will see **different crops** from the train images. (And after many epochs, it will have learned about the entire images). It looks like you were using random cropping too. And then when you used `net_size = 224` after `crop_size = 256`, then you were slightly blurring the images and effectively reducing the image size from 512 to 448.\n\n(So when you used 512, 256, 224, then you were effectively reading 448 and cropping 224).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344",
          "votes": 1
        },
        {
          "id": 981966,
          "postDate": "2020-08-22T22:20:15.633Z",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your explanation. That's exactly what I needed to understand, I actually gave up on this method because I didn't fully understand it, now that I understand, training stronger models will be more affordable in the future.</p>\n<p>One last question: in this competition, the target was small (the Mole), so random cropping <strong>mostly</strong> will crop and include the whole Mole. Do you think this method can be effective in competitions where the target covers a larger area in the image? </p>",
          "rawMarkdown": "Thanks a lot @cdeotte for your explanation. That's exactly what I needed to understand, I actually gave up on this method because I didn't fully understand it, now that I understand, training stronger models will be more affordable in the future.\n\nOne last question: in this competition, the target was small (the Mole), so random cropping **mostly** will crop and include the whole Mole. Do you think this method can be effective in competitions where the target covers a larger area in the image? ",
          "votes": 1
        },
        {
          "id": 981992,
          "postDate": "2020-08-22T23:39:44.700Z",
          "content": "<p>Random cropping is one type of augmentation. In every image comp, you need to discover what are the best types of augmentation. The goal of augmentation is to make more training data. When experimenting with augmentation, just look at the result and ask \"are these new images helpful if i were training a human?\"</p>\n<p>Random cropping is nearly always effective because you can apply it gently. For example, when you have 512x512, you can start by using 480x480 crops. (Note that this is similar to shift by 32 pixels augmentation). There's no harm there. Using 256x256 crops within 512 is aggressive and it worked in this comp because there was so much unnecessary blank space in each image and the mole was centered (so every random crop captured a piece of mole). Such an aggressive crop doesn't always work.</p>",
          "rawMarkdown": "Random cropping is one type of augmentation. In every image comp, you need to discover what are the best types of augmentation. The goal of augmentation is to make more training data. When experimenting with augmentation, just look at the result and ask \"are these new images helpful if i were training a human?\"\n\nRandom cropping is nearly always effective because you can apply it gently. For example, when you have 512x512, you can start by using 480x480 crops. (Note that this is similar to shift by 32 pixels augmentation). There's no harm there. Using 256x256 crops within 512 is aggressive and it worked in this comp because there was so much unnecessary blank space in each image and the mole was centered (so every random crop captured a piece of mole). Such an aggressive crop doesn't always work.",
          "votes": 1
        }
      ]
    },
    {
      "id": 980545,
      "postDate": "2020-08-21T17:11:54.897Z",
      "content": "<p>Very nice meta analysis, thanks for sharing. Interesting to see the result of stacking.</p>",
      "rawMarkdown": "Very nice meta analysis, thanks for sharing. Interesting to see the result of stacking.",
      "votes": 1,
      "replies": [
        {
          "id": 980575,
          "postDate": "2020-08-21T17:39:11.600Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/fchmiel\" target=\"_blank\">@fchmiel</a>! I think stacking would have scored higher with deeper architectures and higher image resolutions.<br>\nAnother interesting thing is that in our case, we could have trust public LB because it had high correlation with private LB (if we don't consider the subs including overfitted public notebooks).</p>",
          "rawMarkdown": "Thank you @fchmiel! I think stacking would have scored higher with deeper architectures and higher image resolutions.\nAnother interesting thing is that in our case, we could have trust public LB because it had high correlation with private LB (if we don't consider the subs including overfitted public notebooks)."
        },
        {
          "id": 981052,
          "postDate": "2020-08-22T06:15:24.270Z",
          "content": "<p>Have did you do the cross-validation for stacking? </p>\n<p>I thought about trying some stacking, but settled for a blend,  because I couldn't properly validate it without using a hold-out test set and I didn't want to waste the data. I guess one could use leaderboard feeback as your  hold out test set but this prooved risky in this competition!</p>",
          "rawMarkdown": "Have did you do the cross-validation for stacking? \n\nI thought about trying some stacking, but settled for a blend,  because I couldn't properly validate it without using a hold-out test set and I didn't want to waste the data. I guess one could use leaderboard feeback as your  hold out test set but this prooved risky in this competition!"
        },
        {
          "id": 981910,
          "postDate": "2020-08-22T20:42:03.463Z",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/fchmiel\" target=\"_blank\">@fchmiel</a> We used 5 fold strategy and did rank ensemble the 5 folds to get the final CV score, which was very stable compared to LB. Stacking can give you a stable CV that you can trust. In blending you just trust the LB feedback, which in this worked out well for you, congrats for the medal :) ( blending worked well for us as well, but we didn't trust LB).</p>",
          "rawMarkdown": "Yes @fchmiel We used 5 fold strategy and did rank ensemble the 5 folds to get the final CV score, which was very stable compared to LB. Stacking can give you a stable CV that you can trust. In blending you just trust the LB feedback, which in this worked out well for you, congrats for the medal :) ( blending worked well for us as well, but we didn't trust LB)."
        }
      ]
    },
    {
      "id": 980879,
      "postDate": "2020-08-22T00:44:49.343Z",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> tried to detect the malignant images in the test set by extracting the patients around the red area (malignants) as they are more susceptible. This pp was unstable but led to some high private LB scores.</p>\n</blockquote>\n<p>Great use of t-SNE. I really like this idea.</p>",
      "rawMarkdown": "> @hawkey tried to detect the malignant images in the test set by extracting the patients around the red area (malignants) as they are more susceptible. This pp was unstable but led to some high private LB scores.\n\nGreat use of t-SNE. I really like this idea.",
      "votes": 2,
      "replies": [
        {
          "id": 981953,
          "postDate": "2020-08-22T22:06:07.407Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Chris, we abandoned this because we extracted embedding from different architectures and they seem to produce different embeddings, therefore unstable.</p>",
          "rawMarkdown": "@cdeotte Chris, we abandoned this because we extracted embedding from different architectures and they seem to produce different embeddings, therefore unstable.",
          "votes": 3
        }
      ]
    },
    {
      "id": 981306,
      "postDate": "2020-08-22T11:05:16.697Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 981917,
          "postDate": "2020-08-22T20:50:58.207Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/synked\" target=\"_blank\">@synked</a> for your feedback! Regarding stacking. We used:</p>\n<ul>\n<li>5 folds CV strategy.</li>\n<li>Optuna for hyperparameters tuning</li>\n<li>Metamodels: LGB, XGB and catboost to stack our best oofs.</li>\n</ul>",
          "rawMarkdown": "Thanks @synked for your feedback! Regarding stacking. We used:\n*  5 folds CV strategy.\n*  Optuna for hyperparameters tuning\n* Metamodels: LGB, XGB and catboost to stack our best oofs.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 981964,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2020-08-22T22:19:29.583000",
      "content": "<p>It was a great pleasure teaming up with you Amin. I definitely learned a lot from you and I really enjoyed our teamwork. We relentlessly trained so many models here and there, expecting to ensemble to get good results. However, this time we lost to high quality models, in other words, we couldn't build models that score around 0.94 CV due to our lack of experience. We learned a lot in this competition from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> about the usage of TPU and TSNE projections, now the winning solutions from all the winners; I can't say I am not satisfied😄. It was great to have diverse teammates like <a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> <a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> <a href=\"https://www.kaggle.com/nicohrubec\" target=\"_blank\">@nicohrubec</a>, coming from different background and cultures. Hoping to connect with you all soon.😬</p>",
      "votes": 3,
      "replies": [
        {
          "id": 981969,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-08-22T22:31:23.040000",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> Thanks Eric for your kind words! Teaming up was the best thing I did in this competition,  I definitely learned a lot from you, you would make a great professor after your PhD 😊<br>\nI also wouldn't say I am not satisfied with the results because it was a solid solution that was missing some ingredients that we could find and learn from the top solutions, now our baseline/pipeline is richer and ready for new challenges :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 980877,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-22T00:30:11.910000",
      "content": "<p>Wow, that's magic. My best (unselected submission) was Private LB 0.9450. When I use your trick it increases to Private LB 0.9459 Gold Medal !!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe6cbf0c0f839fe713317ddbf17a1d161%2Fbest.png?generation=1598055407574078&amp;alt=media\" alt=\"\"></p>\n<p>I did some quick probing where I set all targets to zero and then all of a specific body part to 1. From the submissions we can see that it is rare for there to be positive cases on soles, palms, oral, and genital. Great detective work!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fdaea882ba0b991086abdcc7142decef2%2Fprobe.png?generation=1598056151310064&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 980878,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-22T00:38:14.563000",
          "content": "<p>Actually the above probes prove there are no malignant palm, sole, oral, nor genital in test.</p>\n<p>There are approximately 115 images out of 10,000 test images with palm/sole (based on fact that train has 1.1% palm/sole). And there are approximately 38 test images with oral/genital. So these probe submissions, imply there are no malignant palm, sole, oral, nor genital in the test images. Since a benign will lower LB score by approx 0.00015 and a malignant would increase 0.006</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 981913,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-08-22T20:46:33.790000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for trying our method (by detective <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> :) )! We haven't used it either in our final solution, unfortunately, it turned out that there was really no palm, sole, oral nor genital in test, but in case there were, as you said, this could lower significantly our private score, maybe we should have risked with this pp in one of the 3 subs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 981955,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-08-22T22:08:46.170000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for trying this. We also learned a lesson about trying out something risky!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 980868,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-22T00:09:31.703000",
      "content": "<p>Great analysis. I love all the plots. It tells a great story. Ensembling and teaming help increase score. What's the blue vertical line in the bottom plot? Your private LB scores were not after the blue line.</p>\n<p>When you say <code>Image size= 512. Input size= 226</code>, how did you reduce 512 to 226? Did you resize or crop? was it random or centered?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 981908,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-08-22T20:30:57.803000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your compliment, it means a lot coming from you! And thank you again for all the knowledge you shared in this competition.</p>\n<ul>\n<li><p>Regarding the blue vertical line: Before the blue line our strategy was training models and adding them to the blend, this way we were climbing LB. After the blue line we stopped blending, thus stopped climbing LB and started saving the oofs for stacking for a final stable solution. <strong>The blue line marks the switch from trusting LB to trusting CV.</strong></p></li>\n<li><p>I reduced 512 images to 226 but setting the CNN input_size to 226. I used <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\" target=\"_blank\">AgentAuers's notebook</a>, he used 3 parameters,</p></li>\n</ul>\n<pre><code>  read_size         = 256, \n  crop_size         = 250, \n  net_size          = 224,\n</code></pre>\n<p>This is the set of parameters he used. I tweaked the crop_size (in his <code>prepare_image</code>function he used random_crop and centered_crop) but what made my training faster was not tweaking the crop_size but the net_size, I set the CNN <code>net_size =224</code> with <code>read_size=512</code> and I got better results than <code>image_size=256</code> with <code>net_size=224</code>, probably I was doing something wrong but it had a good private score.</p>\n<p>I would like to learn more about your approach because this is a technique I would like to use in future competitions, how did you make you model random_crop different crops at each epoch? <br>\nIs this augmentation line of code enough to make the training go correctly?<br>\n<code>img = tf.image.random_crop(img, crop_size, crop_size, 3])</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981941,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-22T21:41:59.443000",
          "content": "<p>Using AgentAuers's terminology, your read size was 512 and your net size was 224. What what your crop size?</p>\n<p>The single line <code>img = tf.image.random_crop(img, crop_size, crop_size, 3])</code> does random cropping each epoch if you put it inside the block <code>if augment:</code> inside <code>prepare_image()</code> function.</p>\n<p>Looking closely at AgentAuers's notebook now, I see he uses random cropping. His \"read_size\" are the TFRecords that are used. Then his \"crop_size\" is a random crop done each epoch. Lastly, his \"net_size\" is where he resizes (scales) the random crops. This serves as a blurring effect. The high resolution is reduced to lower resolution.</p>\n<p>Afterward, your network is effectively using <code>X = read_size * net_size / crop_size</code> as it's \"image input size\". So once you set those 3 variables, you can say that your network is learning from images sized <code>X</code>. And to make a diverse set of models you want to build multiple models using <code>X = 256, 384, 512, 768, 1024</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 981948,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-08-22T21:57:18.953000",
          "content": "<p>Thanks again <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.  My crop_size was 256 when image_size=512.<br>\nLets consider just the read_size and crop-size, I understand that if you set the crop size smaller than the read_size and inside the block <code>if augment</code> you use random_crop then <strong>ALL</strong> the images are random cropped differently at every epoch?</p>\n<p>In other words, was your approach similar to AgentAuer's notebook if you set read_size=512 and crop_size=256, with random_crop inside the <code>if augment</code> block?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 981952,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-22T22:06:00.017000",
          "content": "<p>Yes, my approach was exactly AgentAuer's notebook. (However i didn't realize that AgentAuer's was doing this when i added random crop to my private notebook forked from my public notebook).</p>\n<p>In my final writeup <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344\" target=\"_blank\">here</a>, when i say <code>read_size = 512</code> and <code>crop_size = 256</code>, then that is AgentAuer's <code>read_size = 512</code>, <code>crop_size = 256</code>, <code>net_size = crop_size</code>.</p>\n<p>Each epoch, our models will see <strong>different crops</strong> from the train images. (And after many epochs, it will have learned about the entire images). It looks like you were using random cropping too. And then when you used <code>net_size = 224</code> after <code>crop_size = 256</code>, then you were slightly blurring the images and effectively reducing the image size from 512 to 448.</p>\n<p>(So when you used 512, 256, 224, then you were effectively reading 448 and cropping 224).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 981966,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-08-22T22:20:15.633000",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your explanation. That's exactly what I needed to understand, I actually gave up on this method because I didn't fully understand it, now that I understand, training stronger models will be more affordable in the future.</p>\n<p>One last question: in this competition, the target was small (the Mole), so random cropping <strong>mostly</strong> will crop and include the whole Mole. Do you think this method can be effective in competitions where the target covers a larger area in the image? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 981992,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-22T23:39:44.700000",
          "content": "<p>Random cropping is one type of augmentation. In every image comp, you need to discover what are the best types of augmentation. The goal of augmentation is to make more training data. When experimenting with augmentation, just look at the result and ask \"are these new images helpful if i were training a human?\"</p>\n<p>Random cropping is nearly always effective because you can apply it gently. For example, when you have 512x512, you can start by using 480x480 crops. (Note that this is similar to shift by 32 pixels augmentation). There's no harm there. Using 256x256 crops within 512 is aggressive and it worked in this comp because there was so much unnecessary blank space in each image and the mole was centered (so every random crop captured a piece of mole). Such an aggressive crop doesn't always work.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 980545,
      "author_name": "FChmiel",
      "author_url": "",
      "post_date": "2020-08-21T17:11:54.897000",
      "content": "<p>Very nice meta analysis, thanks for sharing. Interesting to see the result of stacking.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 980575,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-08-21T17:39:11.600000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/fchmiel\" target=\"_blank\">@fchmiel</a>! I think stacking would have scored higher with deeper architectures and higher image resolutions.<br>\nAnother interesting thing is that in our case, we could have trust public LB because it had high correlation with private LB (if we don't consider the subs including overfitted public notebooks).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981052,
          "author_name": "FChmiel",
          "author_url": "",
          "post_date": "2020-08-22T06:15:24.270000",
          "content": "<p>Have did you do the cross-validation for stacking? </p>\n<p>I thought about trying some stacking, but settled for a blend,  because I couldn't properly validate it without using a hold-out test set and I didn't want to waste the data. I guess one could use leaderboard feeback as your  hold out test set but this prooved risky in this competition!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 981910,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-08-22T20:42:03.463000",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/fchmiel\" target=\"_blank\">@fchmiel</a> We used 5 fold strategy and did rank ensemble the 5 folds to get the final CV score, which was very stable compared to LB. Stacking can give you a stable CV that you can trust. In blending you just trust the LB feedback, which in this worked out well for you, congrats for the medal :) ( blending worked well for us as well, but we didn't trust LB).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 980879,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-22T00:44:49.343000",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/hawkey\" target=\"_blank\">@hawkey</a> tried to detect the malignant images in the test set by extracting the patients around the red area (malignants) as they are more susceptible. This pp was unstable but led to some high private LB scores.</p>\n</blockquote>\n<p>Great use of t-SNE. I really like this idea.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 981953,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-08-22T22:06:07.407000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Chris, we abandoned this because we extracted embedding from different architectures and they seem to produce different embeddings, therefore unstable.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 981306,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-22T11:05:16.697000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 981917,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2020-08-22T20:50:58.207000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/synked\" target=\"_blank\">@synked</a> for your feedback! Regarding stacking. We used:</p>\n<ul>\n<li>5 folds CV strategy.</li>\n<li>Optuna for hyperparameters tuning</li>\n<li>Metamodels: LGB, XGB and catboost to stack our best oofs.</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "980349": "First of all, I would like to thank the organizers, Kaggle and my teammates @underwearfitting, @hawkey and @nicohrubec for this amazing experience.\n\nI am happy that I picked SIIM-Melanoma to be my first competition because it had everything: A great leader @cdeotte, active teammates, unstable public LB, a terrible shake-up and finally, the drama related to [the participation of the brilliant students of the medal farming academy](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174635).\n\nWe won a bronze medal thanks to the clean up done by kaggle and the investigation started by @group16. Big thanks!\n\nAs most of the participants, we didn’t select our best submission that would have landed us a silver medal. Actually 60 submissions scored higher than our best selected sub. This is part of the game, taking good decisions is an essential skill in ML/DL. Here is a first glimpse of our submissions.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F45cdd5353e253758de95d199217756ce%2FScreen%20Shot%202020-08-21%20at%2017.03.07.png?generation=1598015028261645&alt=media)\n### First time crossing the medal zone:\nThe first time I crossed the medal zone was the 15th of July with private LB= 0.9375 (Slightly higher than our best selected submission).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F7ccedce9337d62222933c5ae7acf95ea%2FScreen%20Shot%202020-08-21%20at%2017.00.32.png?generation=1598014866149997&alt=media)\nWhy did this sub score higher than our selected submission? Because it had an ingredient that was missing in the final sub: Higher `image_size= 768`. I was experimenting with higher image sizes but keeping the CNN's `input_size` smaller for faster training, it's a clear information loss/training time trade-off. \n```\nImage size= 512. Input size= 226\n\nImage size= 768. Input size= 384\n\nImage size= 1024 Input size= 512\n```\nIt scored higher than the models with `image_size` 512 or 384, which is good news, but I decided not to go further with this experiment because it was risky since parts of the images are lost with the center crops.\n\nI was happy with this decision until I saw @cdeotte's [solution](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175344) where he used smaller `input_size` with random crops instead of center crops at each epoch to make sure the model loops over all the pixels instead of center cropping the same pixels over the epochs. Another trick learned to be used in future competitions ;)\n\n### Teaming-up:\nThe day we teamed-up has seen a boost in our private LB score. It was just a rank ensemble of our best submissions. This shows the clear benefit of teaming up late because you will definitely add diversity to the final solution.\n\n### The team era:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F0ff03771186de3fafc3b9eed04860d8c%2FScreen%20Shot%202020-08-21%20at%2017.12.11.png?generation=1598015689624500&alt=media)\nYou can see in the plot that the best era was after teaming up because we experimented many ideas together that helped us reach the silver zone at our peak. I am not gonna lie, at that point I thought we were overfitting to public LB because most of the silver zone submissions are nested blends of our +0.95 submissions. However, the correlation of our public LB scores in the range (0.955, 0.9642) with private LB was 0.89! The question here is: When did we start overfitting?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fa319913972edb44e93fc1c68c56dd4c2%2FScreen%20Shot%202020-08-21%20at%2017.24.13.png?generation=1598016297260469&alt=media)\nOverfitting (the decay) starts after including the [last public notebook in our blend that was terribly overfitting (Public= 0.96, Private= 0.91)](https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619). The blends scoring +0.9643 included overfitted public notebooks\n \nIt is worth mentioning that the highest public/private LB correlation we had was with subs in the medal zone, so our subs definitely had something special to score high private scores without any CV strategy. What ideas or ingredients did the silver submissions have? Simple answer: Post-processing.\n\n\n### Post-processing:\nWe had 2 post-processing techniques, they both boost the public and private LB score.\nThe first was a finding by @underwearfitting where he set all the palms/soles and oral/genital predictions to 0, code below:\n```\nps= test.loc[test['anatom_site_general_challenge']=='palms/soles'].image_name\nog= test.loc[test['anatom_site_general_challenge']=='oral/genital'].image_name\nsubmission= submission.set_index('image_name')\nsubmission.loc[ps]=0\nsubmission.loc[og]=0\n```\n\n\nThis pp gives a +0.001 boost in private LB (Try it at home! +0.001 boost guaranteed 😉).\n\nThe second was image embeddings extraction using RAPIDS cuML TSNE shared by @cdeotte, thanks Chris for everything you shared in this competition.\n@hawkey tried to detect the malignant images in the test set by extracting the patients around the red area (malignants) as they are more susceptible. This pp was unstable but led to some high private LB scores.\n(If you have any questions about this method, please tag @hawkey in the comment)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F55bebda96a72c7bca0fc39dfa2bae19b%2FScreen%20Shot%202020-08-21%20at%2016.36.14.png?generation=1598013477503902&alt=media):\n\n### The final solution :\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F6312bdd9b6c769f81efb3dccdeefaccd%2FScreen%20Shot%202020-08-23%20at%2010.36.49.png?generation=1598164656057343&alt=media)\n**The blue line marks the switch from trusting(climbing) LB to trusting CV.**\n\nThe final step in this competition was to prepare a solid solution.  Even though we had some high scoring subs, interesting ideas (training with high resolution images and smaller CNN input size + post-processing), we decided to do it the safest way and not use any of them and stack models trained with different architecture (B3-B7 + seresnext50) and different image sizes (384, 512).\nWith stacking, we managed to reduce the gap between CV/LB to 0 in our best selected submission, where CV=LB=0.9483!\n\n\n### Conclusion:\nWe ended up with a very correct solution that is missing the special ingredient to make it a top solution. However, I am happy to see that the ideas we have been testing actually work, even though not using them in our final solution penalized our private LB score. Those are techniques learned to be used in future competitions.\n\n**Best private and public subs:**\n<table >\n  <tr>\n    <th>Submission</th>\n    <th>CV</th>\n    <th>Public LB</th>\n    <th>Private LB</th>\n    \n  </tr>\n  <tr>\n    <td>Best public </td>\n    <td>N/A</td>\n    <td>0.9657</td>\n    <td>0.9374</td>\n  </tr>\n  <tr>\n    <td>Best private</td>\n    <td>N/A</td>\n    <td>0.9613</td>\n    <td>0.9410</td>\n  </tr>\n    \n\n</table>\n**3 Selected submissions**    \n\n<table >\n  <tr>\n    <th>Submission</th>\n    <th>CV</th>\n    <th>Public LB</th>\n    <th>Private LB</th>\n    \n  </tr>\n\n    \n\n   <tr>\n    <td>Most stable</td>\n    <td>0.9483</td>\n    <td>0.9483</td>\n    <td>0.9374</td>\n  </tr>\n    \n   <tr>\n    <td>Best CV</td>\n    <td>0.9541</td>\n    <td>0.9510</td>\n    <td>0.9359</td>\n  </tr>\n  <tr>\n    <td>2 Best CV</td>\n    <td>0.9535</td>\n    <td>0.9526</td>\n    <td>0.9337</td>\n  </tr>\n</table>\n",
    "981964": "It was a great pleasure teaming up with you Amin. I definitely learned a lot from you and I really enjoyed our teamwork. We relentlessly trained so many models here and there, expecting to ensemble to get good results. However, this time we lost to high quality models, in other words, we couldn't build models that score around 0.94 CV due to our lack of experience. We learned a lot in this competition from @cdeotte about the usage of TPU and TSNE projections, now the winning solutions from all the winners; I can't say I am not satisfied😄. It was great to have diverse teammates like @hawkey @amiiiney @nicohrubec, coming from different background and cultures. Hoping to connect with you all soon.😬",
    "980877": "Wow, that's magic. My best (unselected submission) was Private LB 0.9450. When I use your trick it increases to Private LB 0.9459 Gold Medal !!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fe6cbf0c0f839fe713317ddbf17a1d161%2Fbest.png?generation=1598055407574078&alt=media)\n\nI did some quick probing where I set all targets to zero and then all of a specific body part to 1. From the submissions we can see that it is rare for there to be positive cases on soles, palms, oral, and genital. Great detective work!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fdaea882ba0b991086abdcc7142decef2%2Fprobe.png?generation=1598056151310064&alt=media)",
    "980868": "Great analysis. I love all the plots. It tells a great story. Ensembling and teaming help increase score. What's the blue vertical line in the bottom plot? Your private LB scores were not after the blue line.\n\nWhen you say `Image size= 512. Input size= 226`, how did you reduce 512 to 226? Did you resize or crop? was it random or centered?",
    "980545": "Very nice meta analysis, thanks for sharing. Interesting to see the result of stacking.",
    "980879": "> @hawkey tried to detect the malignant images in the test set by extracting the patients around the red area (malignants) as they are more susceptible. This pp was unstable but led to some high private LB scores.\n\nGreat use of t-SNE. I really like this idea.",
    "981306": ""
  }
}