{
  "id": 175352,
  "title": "27th overview - I Survived the massive shakedown and be in the top 1%",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/27th-overview-i-survived-the-massive-shakedown-and",
  "author_name": "",
  "post_date": "2020-08-20T10:21:02.990Z",
  "votes": 29,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Since I had been so much afraid of shakedown, I tried to make an <strong>extremely low variance ensemble and it turned out to be a not so bad approach.</strong></p>\n<p><strong>With efficient net B0 ~ B7 imagenet  + B0 ~ B7 noisy-student</strong><br>\n<strong>Size 256 , 384, 512, 768, 1024, it amounts to total 95 MODEL!</strong></p>\n<p>I ensembled all!</p>\n<p>For making networks from B0-B7, 256~1024, I used Chris's stratified notebook and I left the network original that Chris had made<br>\n(Efficient net -&gt; GAP -&gt; dense(1, smoothing = 0.05))<br>\nAll the same network (From B0-B7, SIZE 256 to 1024)<br>\nAll 2019 ext data and coarse dropout was applied</p>\n<p>as I had noticed so many strong competitors reported that their CV was around 0.95,<br>\nI guessed that if I could make my cv converge to around 0.95, and if my cv score &gt; 0.95 then I thought that I supposedly could be good at this competition</p>\n<p><strong>This is how I ensembled 95 deep learning models</strong></p>\n<p><strong>NET BASED ENSEMBLE(EFFICIENT NET 0 ~ EFFICIENT NET 7)</strong><br>\nMODEL 0 : EffcientNet B0 from 256 to 1024, noisy + imagenet CV 0.941 PB 0.9441 PV 0.9350<br>\nMODEL 1 : EffcientNet  B1 from 256 to 1024, noisy + imagenet CV 0.943 PB 0.9491 PV 0.9355<br>\nMODEL 2 : EffcientNet  B2 from 256 to 1024, noisy + imagenet CV 0.946 PB 0.9516 PV 0.9406<br>\nMODEL 3 : EffcientNet  B3 from 256 to 1024, noisy + imagenet CV 0.950 PB 0.9515 PV 0.9411<br>\nMODEL 4 : EffcientNet  B4 from 256 to 1024, noisy + imagenet CV 0.953 PB 0.9543 PV 0.9422<br>\nMODEL 5 : EffcientNet  B5 from 256 to 1024, noisy + imagenet CV 0.953 PB 0.9540 PV 0.9451<br>\nMODEL 6 : EffcientNet  B6 from 256 to 1024, noisy + imagenet CV 0.953  PB 0.9536 PV 0.9472<br>\nMODEL 7 : EffcientNet  B7 from 256 to 1024, noisy + imagenet CV 0.951 PB 0.9549 PV 0.9434<br>\n<strong>-&gt; It was CV 0.9551, PUBLIC LB : 9555 PRIVATE LB 0.9461</strong></p>\n<p><strong>SIZE BASED ENSEMBLE(256 to 1024)</strong><br>\nMODEL 8 : ALL 256 SIZE, EffcientNet  0 ~ 7  CV : 0.939<br>\nMODEL 9 : ALL 384 SIZE, EffcientNet  0 ~ 7  CV : 0.953<br>\nMODEL 10 : ALL 512 SIZE, EffcientNet  0 ~ 7  CV : 0.954<br>\nMODEL 11 : ALL 768 SIZE, EffcientNet  0 ~ 7  CV : 0.949<br>\nMODEL 12 : ALL 1024 SIZE, EffcientNet  0 ~ 7  CV : 0.938</p>\n<p>and I adopted Power ensemble (square 2) that's because AUC is summarized to draw a line between malignant ones and benign ones.<br>\nIf my model is robust, then the resultant power ensemble can be not much different from original ones as AUC is intrinsically just orders</p>\n<p><strong>SIZE BASED ENSEMBLE (with square 2, prediction^2)</strong><br>\nMODEL 13 : ALL 256 SIZE, EffcientNet  0 ~ 7  CV : 0.938<br>\nMODEL 14 : ALL 384 SIZE, EffcientNet  0 ~ 7  CV : 0.947<br>\nMODEL 15 : ALL 512 SIZE, EffcientNet  0 ~ 7  CV : 0.952<br>\nMODEL 16 : ALL 768 SIZE, EffcientNet  0 ~ 7  CV : 0.952<br>\nMODEL 17 : ALL 1024 SIZE, EffcientNet  0 ~ 7  CV : 0.947</p>\n<p><strong>FINAL MODEL = NET BASED MODEL + SIZE BASED MODEL + Square 2 SIZE BASED MODEL</strong></p>\n<p>FINAL MODEL :<strong>TOTAL SIMPLE AVERAGE (MODEL 0 TO 17)  CV -&gt; 0.9538</strong><br>\nTo find appropriate weights, I adopted a differential evolution strategy</p>\n<h1><img src=\"https://i.imgur.com/xRrTIGE.png\" alt=\"https://i.imgur.com/xRrTIGE.png\"></h1>\n<p>FINAL MODEL : OPTIMIZE WEIGHTS WITH SCIPY DIFFERENTIAL_EVOLUTION MODULE -&gt;  <strong>CV 0.9562</strong></p>\n<p>0.8 * FINAL_MODEL + 0.2 * Tabular Meta(FROM XGBOOST) -&gt; <strong>current LB 0.9446</strong><br>\n(If I didn't add meta info to the final model, I could have nearly reached the gold medal)<br>\nbut at Public LB, adding meta info helped improve PUBLIC LB SCORE as much as 0.002, so I guessed that tabular data regularizes overfitting to CV(Cv dropped when added meta), that's why I couldn't abandon it</p>\n<p>By implementing things above described, I found out that my cv is proportional to LB</p>\n<p><strong>When I ensembled all, the resultant CV was around 95.62 and public LB was 95.80\nand Private LB is now 94.46</strong></p>\n<p>I am a Kaggle novice and I am satisfied with the current result!</p>\n<p><strong>Some of you might be curious about scipy.optimize.differential_evolution model</strong></p>\n<p>This is not a special thing.</p>\n<p><a href=\"https://machinelearningmastery.com/weighted-average-ensemble-for-deep-learning-neural-networks/?fbclid=IwAR3GYgj0Fu4Mp3RhTeyacb99H2QyP5uuWJizR7ei6DOOC-NbERKQIGyBB4o\" target=\"_blank\">https://machinelearningmastery.com/weighted-average-ensemble-for-deep-learning-neural-networks/?fbclid=IwAR3GYgj0Fu4Mp3RhTeyacb99H2QyP5uuWJizR7ei6DOOC-NbERKQIGyBB4o</a></p>\n<p>I found the method here</p>\n<p>Actually, To optimize AUC ensemble, I tested the Bayesian method, the Powell optimization method(CV 9614, Public LB:9493 it turned out to be not good), and a lot of things. but CV from evolution differential method was so much proportional to Public LB since I guess it produced the optimized sum of weights = 1 and it is scaled to well for probability while bayesian and powell was not scaled to probability</p>\n<p><strong>What worked:</strong><br>\ntrained model solely with 2019 data<br>\nCoarse dropout<br>\nLabel smoothing<br>\nmalignant upsampling<br>\nnoisy student</p>\n<p><strong>What didn't work:</strong><br>\nFocal loss<br>\nDual input with META + CNN(but it dramatically boosted CV)<br>\nCustom head<br>\nMeta tabular info(but it improved Public LB)<br>\nKNN feature bridging from train to test, from test to train</p>",
  "messages": [
    {
      "id": "974582",
      "postDate": "08/18/2020 01:44:10",
      "content": "<p>Since I had been so much afraid of shakedown, I tried to make an <strong>extremely low variance ensemble and it turned out to be a not so bad approach.</strong></p>\n<p><strong>With efficient net B0 ~ B7 imagenet  + B0 ~ B7 noisy-student</strong><br>\n<strong>Size 256 , 384, 512, 768, 1024, it amounts to total 95 MODEL!</strong></p>\n<p>I ensembled all!</p>\n<p>For making networks from B0-B7, 256~1024, I used Chris's stratified notebook and I left the network original that Chris had made<br>\n(Efficient net -&gt; GAP -&gt; dense(1, smoothing = 0.05))<br>\nAll the same network (From B0-B7, SIZE 256 to 1024)<br>\nAll 2019 ext data and coarse dropout was applied</p>\n<p>as I had noticed so many strong competitors reported that their CV was around 0.95,<br>\nI guessed that if I could make my cv converge to around 0.95, and if my cv score &gt; 0.95 then I thought that I supposedly could be good at this competition</p>\n<p><strong>This is how I ensembled 95 deep learning models</strong></p>\n<p><strong>NET BASED ENSEMBLE(EFFICIENT NET 0 ~ EFFICIENT NET 7)</strong><br>\nMODEL 0 : EffcientNet B0 from 256 to 1024, noisy + imagenet CV 0.941 PB 0.9441 PV 0.9350<br>\nMODEL 1 : EffcientNet  B1 from 256 to 1024, noisy + imagenet CV 0.943 PB 0.9491 PV 0.9355<br>\nMODEL 2 : EffcientNet  B2 from 256 to 1024, noisy + imagenet CV 0.946 PB 0.9516 PV 0.9406<br>\nMODEL 3 : EffcientNet  B3 from 256 to 1024, noisy + imagenet CV 0.950 PB 0.9515 PV 0.9411<br>\nMODEL 4 : EffcientNet  B4 from 256 to 1024, noisy + imagenet CV 0.953 PB 0.9543 PV 0.9422<br>\nMODEL 5 : EffcientNet  B5 from 256 to 1024, noisy + imagenet CV 0.953 PB 0.9540 PV 0.9451<br>\nMODEL 6 : EffcientNet  B6 from 256 to 1024, noisy + imagenet CV 0.953  PB 0.9536 PV 0.9472<br>\nMODEL 7 : EffcientNet  B7 from 256 to 1024, noisy + imagenet CV 0.951 PB 0.9549 PV 0.9434<br>\n<strong>-&gt; It was CV 0.9551, PUBLIC LB : 9555 PRIVATE LB 0.9461</strong></p>\n<p><strong>SIZE BASED ENSEMBLE(256 to 1024)</strong><br>\nMODEL 8 : ALL 256 SIZE, EffcientNet  0 ~ 7  CV : 0.939<br>\nMODEL 9 : ALL 384 SIZE, EffcientNet  0 ~ 7  CV : 0.953<br>\nMODEL 10 : ALL 512 SIZE, EffcientNet  0 ~ 7  CV : 0.954<br>\nMODEL 11 : ALL 768 SIZE, EffcientNet  0 ~ 7  CV : 0.949<br>\nMODEL 12 : ALL 1024 SIZE, EffcientNet  0 ~ 7  CV : 0.938</p>\n<p>and I adopted Power ensemble (square 2) that's because AUC is summarized to draw a line between malignant ones and benign ones.<br>\nIf my model is robust, then the resultant power ensemble can be not much different from original ones as AUC is intrinsically just orders</p>\n<p><strong>SIZE BASED ENSEMBLE (with square 2, prediction^2)</strong><br>\nMODEL 13 : ALL 256 SIZE, EffcientNet  0 ~ 7  CV : 0.938<br>\nMODEL 14 : ALL 384 SIZE, EffcientNet  0 ~ 7  CV : 0.947<br>\nMODEL 15 : ALL 512 SIZE, EffcientNet  0 ~ 7  CV : 0.952<br>\nMODEL 16 : ALL 768 SIZE, EffcientNet  0 ~ 7  CV : 0.952<br>\nMODEL 17 : ALL 1024 SIZE, EffcientNet  0 ~ 7  CV : 0.947</p>\n<p><strong>FINAL MODEL = NET BASED MODEL + SIZE BASED MODEL + Square 2 SIZE BASED MODEL</strong></p>\n<p>FINAL MODEL :<strong>TOTAL SIMPLE AVERAGE (MODEL 0 TO 17)  CV -&gt; 0.9538</strong><br>\nTo find appropriate weights, I adopted a differential evolution strategy</p>\n<h1><img src=\"https://i.imgur.com/xRrTIGE.png\" alt=\"https://i.imgur.com/xRrTIGE.png\"></h1>\n<p>FINAL MODEL : OPTIMIZE WEIGHTS WITH SCIPY DIFFERENTIAL_EVOLUTION MODULE -&gt;  <strong>CV 0.9562</strong></p>\n<p>0.8 * FINAL_MODEL + 0.2 * Tabular Meta(FROM XGBOOST) -&gt; <strong>current LB 0.9446</strong><br>\n(If I didn't add meta info to the final model, I could have nearly reached the gold medal)<br>\nbut at Public LB, adding meta info helped improve PUBLIC LB SCORE as much as 0.002, so I guessed that tabular data regularizes overfitting to CV(Cv dropped when added meta), that's why I couldn't abandon it</p>\n<p>By implementing things above described, I found out that my cv is proportional to LB</p>\n<p><strong>When I ensembled all, the resultant CV was around 95.62 and public LB was 95.80\nand Private LB is now 94.46</strong></p>\n<p>I am a Kaggle novice and I am satisfied with the current result!</p>\n<p><strong>Some of you might be curious about scipy.optimize.differential_evolution model</strong></p>\n<p>This is not a special thing.</p>\n<p><a href=\"https://machinelearningmastery.com/weighted-average-ensemble-for-deep-learning-neural-networks/?fbclid=IwAR3GYgj0Fu4Mp3RhTeyacb99H2QyP5uuWJizR7ei6DOOC-NbERKQIGyBB4o\" target=\"_blank\">https://machinelearningmastery.com/weighted-average-ensemble-for-deep-learning-neural-networks/?fbclid=IwAR3GYgj0Fu4Mp3RhTeyacb99H2QyP5uuWJizR7ei6DOOC-NbERKQIGyBB4o</a></p>\n<p>I found the method here</p>\n<p>Actually, To optimize AUC ensemble, I tested the Bayesian method, the Powell optimization method(CV 9614, Public LB:9493 it turned out to be not good), and a lot of things. but CV from evolution differential method was so much proportional to Public LB since I guess it produced the optimized sum of weights = 1 and it is scaled to well for probability while bayesian and powell was not scaled to probability</p>\n<p><strong>What worked:</strong><br>\ntrained model solely with 2019 data<br>\nCoarse dropout<br>\nLabel smoothing<br>\nmalignant upsampling<br>\nnoisy student</p>\n<p><strong>What didn't work:</strong><br>\nFocal loss<br>\nDual input with META + CNN(but it dramatically boosted CV)<br>\nCustom head<br>\nMeta tabular info(but it improved Public LB)<br>\nKNN feature bridging from train to test, from test to train</p>",
      "rawMarkdown": "Since I had been so much afraid of shakedown, I tried to make an **extremely low variance ensemble and it turned out to be a not so bad approach.**\n\n\n**With efficient net B0 ~ B7 imagenet  + B0 ~ B7 noisy-student**\n**Size 256 , 384, 512, 768, 1024, it amounts to total 95 MODEL!**\n\nI ensembled all!\n\nFor making networks from B0-B7, 256~1024, I used Chris's stratified notebook and I left the network original that Chris had made\n(Efficient net -> GAP -> dense(1, smoothing = 0.05))\nAll the same network (From B0-B7, SIZE 256 to 1024)\nAll 2019 ext data and coarse dropout was applied\n\nas I had noticed so many strong competitors reported that their CV was around 0.95,\nI guessed that if I could make my cv converge to around 0.95, and if my cv score > 0.95 then I thought that I supposedly could be good at this competition\n\n**This is how I ensembled 95 deep learning models**\n\n**NET BASED ENSEMBLE(EFFICIENT NET 0 ~ EFFICIENT NET 7)**\nMODEL 0 : EffcientNet B0 from 256 to 1024, noisy + imagenet CV 0.941 PB 0.9441 PV 0.9350\nMODEL 1 : EffcientNet  B1 from 256 to 1024, noisy + imagenet CV 0.943 PB 0.9491 PV 0.9355\nMODEL 2 : EffcientNet  B2 from 256 to 1024, noisy + imagenet CV 0.946 PB 0.9516 PV 0.9406\nMODEL 3 : EffcientNet  B3 from 256 to 1024, noisy + imagenet CV 0.950 PB 0.9515 PV 0.9411\nMODEL 4 : EffcientNet  B4 from 256 to 1024, noisy + imagenet CV 0.953 PB 0.9543 PV 0.9422\nMODEL 5 : EffcientNet  B5 from 256 to 1024, noisy + imagenet CV 0.953 PB 0.9540 PV 0.9451\nMODEL 6 : EffcientNet  B6 from 256 to 1024, noisy + imagenet CV 0.953  PB 0.9536 PV 0.9472\nMODEL 7 : EffcientNet  B7 from 256 to 1024, noisy + imagenet CV 0.951 PB 0.9549 PV 0.9434\n**-> It was CV 0.9551, PUBLIC LB : 9555 PRIVATE LB 0.9461**\n\n**SIZE BASED ENSEMBLE(256 to 1024)**\nMODEL 8 : ALL 256 SIZE, EffcientNet  0 ~ 7  CV : 0.939\nMODEL 9 : ALL 384 SIZE, EffcientNet  0 ~ 7  CV : 0.953\nMODEL 10 : ALL 512 SIZE, EffcientNet  0 ~ 7  CV : 0.954\nMODEL 11 : ALL 768 SIZE, EffcientNet  0 ~ 7  CV : 0.949\nMODEL 12 : ALL 1024 SIZE, EffcientNet  0 ~ 7  CV : 0.938\n\nand I adopted Power ensemble (square 2) that's because AUC is summarized to draw a line between malignant ones and benign ones.\nIf my model is robust, then the resultant power ensemble can be not much different from original ones as AUC is intrinsically just orders\n\n**SIZE BASED ENSEMBLE (with square 2, prediction^2)**\nMODEL 13 : ALL 256 SIZE, EffcientNet  0 ~ 7  CV : 0.938\nMODEL 14 : ALL 384 SIZE, EffcientNet  0 ~ 7  CV : 0.947\nMODEL 15 : ALL 512 SIZE, EffcientNet  0 ~ 7  CV : 0.952\nMODEL 16 : ALL 768 SIZE, EffcientNet  0 ~ 7  CV : 0.952\nMODEL 17 : ALL 1024 SIZE, EffcientNet  0 ~ 7  CV : 0.947\n\n**FINAL MODEL = NET BASED MODEL + SIZE BASED MODEL + Square 2 SIZE BASED MODEL**\n\nFINAL MODEL :**TOTAL SIMPLE AVERAGE (MODEL 0 TO 17)  CV -> 0.9538**\nTo find appropriate weights, I adopted a differential evolution strategy\n\n#![https://i.imgur.com/xRrTIGE.png](https://i.imgur.com/xRrTIGE.png)\n\n\nFINAL MODEL : OPTIMIZE WEIGHTS WITH SCIPY DIFFERENTIAL_EVOLUTION MODULE ->  **CV 0.9562**\n\n0.8 * FINAL_MODEL + 0.2 * Tabular Meta(FROM XGBOOST) -> **current LB 0.9446**\n(If I didn't add meta info to the final model, I could have nearly reached the gold medal)\nbut at Public LB, adding meta info helped improve PUBLIC LB SCORE as much as 0.002, so I guessed that tabular data regularizes overfitting to CV(Cv dropped when added meta), that's why I couldn't abandon it\n\nBy implementing things above described, I found out that my cv is proportional to LB\n\n**When I ensembled all, the resultant CV was around 95.62 and public LB was 95.80\nand Private LB is now 94.46**\n\nI am a Kaggle novice and I am satisfied with the current result!\n\n\n\n**Some of you might be curious about scipy.optimize.differential_evolution model**\n\nThis is not a special thing.\n\nhttps://machinelearningmastery.com/weighted-average-ensemble-for-deep-learning-neural-networks/?fbclid=IwAR3GYgj0Fu4Mp3RhTeyacb99H2QyP5uuWJizR7ei6DOOC-NbERKQIGyBB4o\n\nI found the method here\n\nActually, To optimize AUC ensemble, I tested the Bayesian method, the Powell optimization method(CV 9614, Public LB:9493 it turned out to be not good), and a lot of things. but CV from evolution differential method was so much proportional to Public LB since I guess it produced the optimized sum of weights = 1 and it is scaled to well for probability while bayesian and powell was not scaled to probability\n\n\n**What worked:**\ntrained model solely with 2019 data\nCoarse dropout\nLabel smoothing\nmalignant upsampling\nnoisy student\n\n**What didn't work:**\nFocal loss\nDual input with META + CNN(but it dramatically boosted CV)\nCustom head\nMeta tabular info(but it improved Public LB)\nKNN feature bridging from train to test, from test to train",
      "votes": null
    },
    {
      "id": "974600",
      "postDate": "08/18/2020 01:54:27",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a>!</p>",
      "rawMarkdown": "Congrats @deepkim!",
      "votes": null
    },
    {
      "id": "974610",
      "postDate": "08/18/2020 02:00:14",
      "content": "<p>Interesting. Thanks for sharing. Do you mind sharing more details about your ensemble technique or maybe publish a kernel?</p>",
      "rawMarkdown": "Interesting. Thanks for sharing. Do you mind sharing more details about your ensemble technique or maybe publish a kernel?",
      "votes": null
    },
    {
      "id": "974699",
      "postDate": "08/18/2020 02:39:28",
      "content": "<p>Congratulations Statking. Achieving CV 0.9562 is excellent. Well done.</p>\n<p>I'm curious about this <code>scipy differential evolution method</code>, I plan to read more about this.</p>",
      "rawMarkdown": "Congratulations Statking. Achieving CV 0.9562 is excellent. Well done.\n\nI'm curious about this `scipy differential evolution method`, I plan to read more about this.",
      "votes": null
    },
    {
      "id": "974908",
      "postDate": "08/18/2020 04:50:00",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> , time well spent on creating so many models.</p>",
      "rawMarkdown": "Congratulations @deepkim , time well spent on creating so many models.",
      "votes": null
    },
    {
      "id": "975303",
      "postDate": "08/18/2020 08:32:01",
      "content": "<p>Congrats! I wonder ur ensemble method in detail. I will wait for you leaving work 👍</p>",
      "rawMarkdown": "Congrats! I wonder ur ensemble method in detail. I will wait for you leaving work 👍",
      "votes": null
    },
    {
      "id": "975568",
      "postDate": "08/18/2020 11:03:45",
      "content": "<p>Thanks for sharing!<br>\nIt's a shame that the table models lowered the score. Top solutions in alaska2 seem to do better with stacking and MLP ensemble for models that have extremely lower score than max CV score. maybe it is helpful.<br>\nref : <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168870\" target=\"_blank\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168870</a></p>\n<blockquote>\n  <p>mixing it well improves the accuracy, but since it is considerably weaker than 3-channel models as a single model, it doesn't work well when simply averaging it. So I used MLP.</p>\n</blockquote>\n<p>You've trained many models, have you trained all the folds?<br>\nAlso, did you use the Chris's notebook? (TF with TPU is needed for train many models?)</p>",
      "rawMarkdown": "Thanks for sharing!\nIt's a shame that the table models lowered the score. Top solutions in alaska2 seem to do better with stacking and MLP ensemble for models that have extremely lower score than max CV score. maybe it is helpful.\nref : https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168870\n>mixing it well improves the accuracy, but since it is considerably weaker than 3-channel models as a single model, it doesn't work well when simply averaging it. So I used MLP.\n\nYou've trained many models, have you trained all the folds?\nAlso, did you use the Chris's notebook? (TF with TPU is needed for train many models?)",
      "votes": null
    },
    {
      "id": "975624",
      "postDate": "08/18/2020 11:40:21",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    },
    {
      "id": "975630",
      "postDate": "08/18/2020 11:44:52",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thank you so much. i have detailed it further</p>",
      "rawMarkdown": "cdeotte thank you so much. i have detailed it further",
      "votes": null
    },
    {
      "id": "975632",
      "postDate": "08/18/2020 11:45:08",
      "content": "<p>Thank you so much</p>",
      "rawMarkdown": "Thank you so much",
      "votes": null
    },
    {
      "id": "975634",
      "postDate": "08/18/2020 11:45:36",
      "content": "<p>I have detailed it further.</p>",
      "rawMarkdown": "I have detailed it further.",
      "votes": null
    },
    {
      "id": "975659",
      "postDate": "08/18/2020 12:05:41",
      "content": "<ol>\n<li>I have trained all 5 folds models</li>\n<li>I utilized Chris's notebook. It is really well made and logical with stratified Tfrecords. </li>\n<li>with TPU, I could train models faster.</li>\n</ol>",
      "rawMarkdown": "1. I have trained all 5 folds models\n2. I utilized Chris's notebook. It is really well made and logical with stratified Tfrecords. \n3. with TPU, I could train models faster.",
      "votes": null
    },
    {
      "id": "975906",
      "postDate": "08/18/2020 14:14:50",
      "content": "<p>Thank you for your sincere congratulation! </p>",
      "rawMarkdown": "Thank you for your sincere congratulation!",
      "votes": null
    },
    {
      "id": "976502",
      "postDate": "08/18/2020 22:48:44",
      "content": "<p>Congrats! I was pursuing the same approach as you, but I gave up before I had run all of the models because the admin of dealing with all the kernels was too much, and I had other approaches I wanted to pursue. </p>\n<p>I am curious as to how you dealt with this? Running 95 kernels on kaggle is a time consuming process.</p>\n<p>I am also curious as to why you say that dual input did not work when it dramatically increased your CV?</p>",
      "rawMarkdown": "Congrats! I was pursuing the same approach as you, but I gave up before I had run all of the models because the admin of dealing with all the kernels was too much, and I had other approaches I wanted to pursue. \n\nI am curious as to how you dealt with this? Running 95 kernels on kaggle is a time consuming process.\n\nI am also curious as to why you say that dual input did not work when it dramatically increased your CV?",
      "votes": null
    },
    {
      "id": "976822",
      "postDate": "08/19/2020 05:43:22",
      "content": "<ol>\n<li><p>256, 384, 512, 768 Size didn't take time too much with the usage of kaggle tpu and colab<br>\nbut 1024 size training was a somewhat time-consuming work<br>\nbut we had 2~3 months.</p></li>\n<li><p>I tried meta info + CNN dual inputs training, it doesn't show  a sign of improving Public LB<br>\nwith meta info added to CNN, my cv could reach above 96 just on ISIC 2020 data<br>\nbut it didn't work for <strong>Public LB</strong> so i abandoned it(it was my mistake)<br>\nby an adversarial analysis, we can infer that training datasets and test datasets are somewhat different.<br>\nI guessed using meta info can lead my model to overfitting to just training datasets</p></li>\n</ol>",
      "rawMarkdown": "1. 256, 384, 512, 768 Size didn't take time too much with the usage of kaggle tpu and colab\nbut 1024 size training was a somewhat time-consuming work\nbut we had 2~3 months.\n\n2. I tried meta info + CNN dual inputs training, it doesn't show  a sign of improving Public LB\nwith meta info added to CNN, my cv could reach above 96 just on ISIC 2020 data\nbut it didn't work for **Public LB** so i abandoned it(it was my mistake)\nby an adversarial analysis, we can infer that training datasets and test datasets are somewhat different.\nI guessed using meta info can lead my model to overfitting to just training datasets",
      "votes": null
    },
    {
      "id": "977025",
      "postDate": "08/19/2020 08:40:19",
      "content": "<p>Okay thanks! Sorry though, I wasn't clear about my first question. The time consumption that I was speaking about was the manual task of running all the kernels, it takes a bit of effort to commit lots of different notebooks and then keep track of the separate parts (I assume your 1024 size had to be trained across multiple sessions?) and I was wondering if you have any way of dealing with that?</p>",
      "rawMarkdown": "Okay thanks! Sorry though, I wasn't clear about my first question. The time consumption that I was speaking about was the manual task of running all the kernels, it takes a bit of effort to commit lots of different notebooks and then keep track of the separate parts (I assume your 1024 size had to be trained across multiple sessions?) and I was wondering if you have any way of dealing with that?",
      "votes": null
    },
    {
      "id": "977135",
      "postDate": "08/19/2020 10:08:42",
      "content": "<p>You are right<br>\nKaggle tpu time limit is 3 - hours per training , 30 hours per week (but it is really fast!)<br>\nIf I had to train 5 folds of high-resolution images, then I needed to train 0,1,2 fold and after that, I train 3,4 fold seperately and merge [0,1,2] and [3,4] fold outputs into one output</p>\n<p>since it was sort of tedious jobs, for most of training, I used my local environment</p>",
      "rawMarkdown": "You are right\nKaggle tpu time limit is 3 - hours per training , 30 hours per week (but it is really fast!)\nIf I had to train 5 folds of high-resolution images, then I needed to train 0,1,2 fold and after that, I train 3,4 fold seperately and merge [0,1,2] and [3,4] fold outputs into one output\n\n\nsince it was sort of tedious jobs, for most of training, I used my local environment",
      "votes": null
    },
    {
      "id": "980269",
      "postDate": "08/21/2020 12:57:45",
      "content": "<p>Thanks for this. Go ahead, man :D</p>",
      "rawMarkdown": "Thanks for this. Go ahead, man :D",
      "votes": null
    },
    {
      "id": "982972",
      "postDate": "08/23/2020 22:35:42",
      "content": "<p>Congrats!<br>\nI am proud that you are Korean</p>\n<p>I have a question.</p>\n<ol>\n<li>Why did you apply 'label smooting'? Did you think there would be mislabelling?</li>\n<li>How to apply 'upsampling'?</li>\n</ol>\n<p>Thank you ~</p>",
      "rawMarkdown": "Congrats!\nI am proud that you are Korean\n\nI have a question.\n\n1. Why did you apply 'label smooting'? Did you think there would be mislabelling?\n2. How to apply 'upsampling'?\n\nThank you ~",
      "votes": null
    },
    {
      "id": "984740",
      "postDate": "08/25/2020 08:55:00",
      "content": "<p>Thank you for your congratulation!</p>\n<p>Let me give you an answer about what you asked</p>\n<ol>\n<li>the evaluation metric is AUC. In this competition, you don't need to classify accurately 0 : benign, 1 : malignant. it's somewhat different from measuring accuracy</li>\n</ol>\n<p>you just need to calculate how much it is likely to be,,, benign one or malignant one<br>\nit is expressed by \"a probability\"<br>\nfor example,  an output of benign one: 0.00001, malignant : 0.9495……</p>\n<p>i had tested a lot of networks and found out that my CV auc values were clustered around 95.x</p>\n<p>it means my model performance is 95.x</p>\n<p>in this situation, generalizing output probability around 0.95x can be a good way for maintaining model stability. if you don't apply label smoothing, your model's output can be overconfident about malignant ones or benign ones. </p>\n<p>that's why i applied label smoothing</p>\n<p>2.For upsampling, watch this!<br>\n<a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\" target=\"_blank\">https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout</a></p>",
      "rawMarkdown": "Thank you for your congratulation!\n\nLet me give you an answer about what you asked\n\n1. the evaluation metric is AUC. In this competition, you don't need to classify accurately 0 : benign, 1 : malignant. it's somewhat different from measuring accuracy\n\nyou just need to calculate how much it is likely to be,,, benign one or malignant one\nit is expressed by \"a probability\"\nfor example,  an output of benign one: 0.00001, malignant : 0.9495......\n\ni had tested a lot of networks and found out that my CV auc values were clustered around 95.x\n\nit means my model performance is 95.x\n\nin this situation, generalizing output probability around 0.95x can be a good way for maintaining model stability. if you don't apply label smoothing, your model's output can be overconfident about malignant ones or benign ones. \n\nthat's why i applied label smoothing\n\n2.For upsampling, watch this!\nhttps://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
      "votes": null
    },
    {
      "id": "2917803",
      "postDate": "07/11/2024 19:42:46",
      "content": "<p>Hello there :) In case you see this message… ISIC is now running a new challenge on Kaggle:</p>\n<p><a href=\"https://www.kaggle.com/competitions/isic-2024-challenge/overview\" target=\"_blank\">https://www.kaggle.com/competitions/isic-2024-challenge/overview</a></p>\n<p>Would love to have you take part!</p>",
      "rawMarkdown": "Hello there :) In case you see this message... ISIC is now running a new challenge on Kaggle:\n\nhttps://www.kaggle.com/competitions/isic-2024-challenge/overview\n\nWould love to have you take part!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974600,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "08/18/2020 01:54:27",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a>!</p>",
      "votes": null,
      "replies": [
        {
          "id": 975624,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/18/2020 11:40:21",
          "content": "<p>Thank you so much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974610,
      "author_name": "luohongchen1993",
      "author_url": "",
      "post_date": "08/18/2020 02:00:14",
      "content": "<p>Interesting. Thanks for sharing. Do you mind sharing more details about your ensemble technique or maybe publish a kernel?</p>",
      "votes": null,
      "replies": [
        {
          "id": 975634,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/18/2020 11:45:36",
          "content": "<p>I have detailed it further.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974699,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/18/2020 02:39:28",
      "content": "<p>Congratulations Statking. Achieving CV 0.9562 is excellent. Well done.</p>\n<p>I'm curious about this <code>scipy differential evolution method</code>, I plan to read more about this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 975630,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/18/2020 11:44:52",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thank you so much. i have detailed it further</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974908,
      "author_name": "karrak3256",
      "author_url": "",
      "post_date": "08/18/2020 04:50:00",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> , time well spent on creating so many models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 975632,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/18/2020 11:45:08",
          "content": "<p>Thank you so much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 975303,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "08/18/2020 08:32:01",
      "content": "<p>Congrats! I wonder ur ensemble method in detail. I will wait for you leaving work 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 975906,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/18/2020 14:14:50",
          "content": "<p>Thank you for your sincere congratulation! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 975568,
      "author_name": "ajtryt2",
      "author_url": "",
      "post_date": "08/18/2020 11:03:45",
      "content": "<p>Thanks for sharing!<br>\nIt's a shame that the table models lowered the score. Top solutions in alaska2 seem to do better with stacking and MLP ensemble for models that have extremely lower score than max CV score. maybe it is helpful.<br>\nref : <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168870\" target=\"_blank\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168870</a></p>\n<blockquote>\n  <p>mixing it well improves the accuracy, but since it is considerably weaker than 3-channel models as a single model, it doesn't work well when simply averaging it. So I used MLP.</p>\n</blockquote>\n<p>You've trained many models, have you trained all the folds?<br>\nAlso, did you use the Chris's notebook? (TF with TPU is needed for train many models?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 975659,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/18/2020 12:05:41",
          "content": "<ol>\n<li>I have trained all 5 folds models</li>\n<li>I utilized Chris's notebook. It is really well made and logical with stratified Tfrecords. </li>\n<li>with TPU, I could train models faster.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976502,
      "author_name": "samklein",
      "author_url": "",
      "post_date": "08/18/2020 22:48:44",
      "content": "<p>Congrats! I was pursuing the same approach as you, but I gave up before I had run all of the models because the admin of dealing with all the kernels was too much, and I had other approaches I wanted to pursue. </p>\n<p>I am curious as to how you dealt with this? Running 95 kernels on kaggle is a time consuming process.</p>\n<p>I am also curious as to why you say that dual input did not work when it dramatically increased your CV?</p>",
      "votes": null,
      "replies": [
        {
          "id": 976822,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/19/2020 05:43:22",
          "content": "<ol>\n<li><p>256, 384, 512, 768 Size didn't take time too much with the usage of kaggle tpu and colab<br>\nbut 1024 size training was a somewhat time-consuming work<br>\nbut we had 2~3 months.</p></li>\n<li><p>I tried meta info + CNN dual inputs training, it doesn't show  a sign of improving Public LB<br>\nwith meta info added to CNN, my cv could reach above 96 just on ISIC 2020 data<br>\nbut it didn't work for <strong>Public LB</strong> so i abandoned it(it was my mistake)<br>\nby an adversarial analysis, we can infer that training datasets and test datasets are somewhat different.<br>\nI guessed using meta info can lead my model to overfitting to just training datasets</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 977025,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "08/19/2020 08:40:19",
          "content": "<p>Okay thanks! Sorry though, I wasn't clear about my first question. The time consumption that I was speaking about was the manual task of running all the kernels, it takes a bit of effort to commit lots of different notebooks and then keep track of the separate parts (I assume your 1024 size had to be trained across multiple sessions?) and I was wondering if you have any way of dealing with that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 977135,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/19/2020 10:08:42",
          "content": "<p>You are right<br>\nKaggle tpu time limit is 3 - hours per training , 30 hours per week (but it is really fast!)<br>\nIf I had to train 5 folds of high-resolution images, then I needed to train 0,1,2 fold and after that, I train 3,4 fold seperately and merge [0,1,2] and [3,4] fold outputs into one output</p>\n<p>since it was sort of tedious jobs, for most of training, I used my local environment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 982972,
      "author_name": "hongym7",
      "author_url": "",
      "post_date": "08/23/2020 22:35:42",
      "content": "<p>Congrats!<br>\nI am proud that you are Korean</p>\n<p>I have a question.</p>\n<ol>\n<li>Why did you apply 'label smooting'? Did you think there would be mislabelling?</li>\n<li>How to apply 'upsampling'?</li>\n</ol>\n<p>Thank you ~</p>",
      "votes": null,
      "replies": [
        {
          "id": 984740,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/25/2020 08:55:00",
          "content": "<p>Thank you for your congratulation!</p>\n<p>Let me give you an answer about what you asked</p>\n<ol>\n<li>the evaluation metric is AUC. In this competition, you don't need to classify accurately 0 : benign, 1 : malignant. it's somewhat different from measuring accuracy</li>\n</ol>\n<p>you just need to calculate how much it is likely to be,,, benign one or malignant one<br>\nit is expressed by \"a probability\"<br>\nfor example,  an output of benign one: 0.00001, malignant : 0.9495……</p>\n<p>i had tested a lot of networks and found out that my CV auc values were clustered around 95.x</p>\n<p>it means my model performance is 95.x</p>\n<p>in this situation, generalizing output probability around 0.95x can be a good way for maintaining model stability. if you don't apply label smoothing, your model's output can be overconfident about malignant ones or benign ones. </p>\n<p>that's why i applied label smoothing</p>\n<p>2.For upsampling, watch this!<br>\n<a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\" target=\"_blank\">https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2917803,
      "author_name": "jwebermsk",
      "author_url": "",
      "post_date": "07/11/2024 19:42:46",
      "content": "<p>Hello there :) In case you see this message… ISIC is now running a new challenge on Kaggle:</p>\n<p><a href=\"https://www.kaggle.com/competitions/isic-2024-challenge/overview\" target=\"_blank\">https://www.kaggle.com/competitions/isic-2024-challenge/overview</a></p>\n<p>Would love to have you take part!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 980269,
      "author_name": "utshabkumarghosh",
      "author_url": "",
      "post_date": "08/21/2020 12:57:45",
      "content": "<p>Thanks for this. Go ahead, man :D</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "974582": "Since I had been so much afraid of shakedown, I tried to make an **extremely low variance ensemble and it turned out to be a not so bad approach.**\n\n\n**With efficient net B0 ~ B7 imagenet  + B0 ~ B7 noisy-student**\n**Size 256 , 384, 512, 768, 1024, it amounts to total 95 MODEL!**\n\nI ensembled all!\n\nFor making networks from B0-B7, 256~1024, I used Chris's stratified notebook and I left the network original that Chris had made\n(Efficient net -> GAP -> dense(1, smoothing = 0.05))\nAll the same network (From B0-B7, SIZE 256 to 1024)\nAll 2019 ext data and coarse dropout was applied\n\nas I had noticed so many strong competitors reported that their CV was around 0.95,\nI guessed that if I could make my cv converge to around 0.95, and if my cv score > 0.95 then I thought that I supposedly could be good at this competition\n\n**This is how I ensembled 95 deep learning models**\n\n**NET BASED ENSEMBLE(EFFICIENT NET 0 ~ EFFICIENT NET 7)**\nMODEL 0 : EffcientNet B0 from 256 to 1024, noisy + imagenet CV 0.941 PB 0.9441 PV 0.9350\nMODEL 1 : EffcientNet  B1 from 256 to 1024, noisy + imagenet CV 0.943 PB 0.9491 PV 0.9355\nMODEL 2 : EffcientNet  B2 from 256 to 1024, noisy + imagenet CV 0.946 PB 0.9516 PV 0.9406\nMODEL 3 : EffcientNet  B3 from 256 to 1024, noisy + imagenet CV 0.950 PB 0.9515 PV 0.9411\nMODEL 4 : EffcientNet  B4 from 256 to 1024, noisy + imagenet CV 0.953 PB 0.9543 PV 0.9422\nMODEL 5 : EffcientNet  B5 from 256 to 1024, noisy + imagenet CV 0.953 PB 0.9540 PV 0.9451\nMODEL 6 : EffcientNet  B6 from 256 to 1024, noisy + imagenet CV 0.953  PB 0.9536 PV 0.9472\nMODEL 7 : EffcientNet  B7 from 256 to 1024, noisy + imagenet CV 0.951 PB 0.9549 PV 0.9434\n**-> It was CV 0.9551, PUBLIC LB : 9555 PRIVATE LB 0.9461**\n\n**SIZE BASED ENSEMBLE(256 to 1024)**\nMODEL 8 : ALL 256 SIZE, EffcientNet  0 ~ 7  CV : 0.939\nMODEL 9 : ALL 384 SIZE, EffcientNet  0 ~ 7  CV : 0.953\nMODEL 10 : ALL 512 SIZE, EffcientNet  0 ~ 7  CV : 0.954\nMODEL 11 : ALL 768 SIZE, EffcientNet  0 ~ 7  CV : 0.949\nMODEL 12 : ALL 1024 SIZE, EffcientNet  0 ~ 7  CV : 0.938\n\nand I adopted Power ensemble (square 2) that's because AUC is summarized to draw a line between malignant ones and benign ones.\nIf my model is robust, then the resultant power ensemble can be not much different from original ones as AUC is intrinsically just orders\n\n**SIZE BASED ENSEMBLE (with square 2, prediction^2)**\nMODEL 13 : ALL 256 SIZE, EffcientNet  0 ~ 7  CV : 0.938\nMODEL 14 : ALL 384 SIZE, EffcientNet  0 ~ 7  CV : 0.947\nMODEL 15 : ALL 512 SIZE, EffcientNet  0 ~ 7  CV : 0.952\nMODEL 16 : ALL 768 SIZE, EffcientNet  0 ~ 7  CV : 0.952\nMODEL 17 : ALL 1024 SIZE, EffcientNet  0 ~ 7  CV : 0.947\n\n**FINAL MODEL = NET BASED MODEL + SIZE BASED MODEL + Square 2 SIZE BASED MODEL**\n\nFINAL MODEL :**TOTAL SIMPLE AVERAGE (MODEL 0 TO 17)  CV -> 0.9538**\nTo find appropriate weights, I adopted a differential evolution strategy\n\n#![https://i.imgur.com/xRrTIGE.png](https://i.imgur.com/xRrTIGE.png)\n\n\nFINAL MODEL : OPTIMIZE WEIGHTS WITH SCIPY DIFFERENTIAL_EVOLUTION MODULE ->  **CV 0.9562**\n\n0.8 * FINAL_MODEL + 0.2 * Tabular Meta(FROM XGBOOST) -> **current LB 0.9446**\n(If I didn't add meta info to the final model, I could have nearly reached the gold medal)\nbut at Public LB, adding meta info helped improve PUBLIC LB SCORE as much as 0.002, so I guessed that tabular data regularizes overfitting to CV(Cv dropped when added meta), that's why I couldn't abandon it\n\nBy implementing things above described, I found out that my cv is proportional to LB\n\n**When I ensembled all, the resultant CV was around 95.62 and public LB was 95.80\nand Private LB is now 94.46**\n\nI am a Kaggle novice and I am satisfied with the current result!\n\n\n\n**Some of you might be curious about scipy.optimize.differential_evolution model**\n\nThis is not a special thing.\n\nhttps://machinelearningmastery.com/weighted-average-ensemble-for-deep-learning-neural-networks/?fbclid=IwAR3GYgj0Fu4Mp3RhTeyacb99H2QyP5uuWJizR7ei6DOOC-NbERKQIGyBB4o\n\nI found the method here\n\nActually, To optimize AUC ensemble, I tested the Bayesian method, the Powell optimization method(CV 9614, Public LB:9493 it turned out to be not good), and a lot of things. but CV from evolution differential method was so much proportional to Public LB since I guess it produced the optimized sum of weights = 1 and it is scaled to well for probability while bayesian and powell was not scaled to probability\n\n\n**What worked:**\ntrained model solely with 2019 data\nCoarse dropout\nLabel smoothing\nmalignant upsampling\nnoisy student\n\n**What didn't work:**\nFocal loss\nDual input with META + CNN(but it dramatically boosted CV)\nCustom head\nMeta tabular info(but it improved Public LB)\nKNN feature bridging from train to test, from test to train",
    "974600": "Congrats @deepkim!",
    "974610": "Interesting. Thanks for sharing. Do you mind sharing more details about your ensemble technique or maybe publish a kernel?",
    "974699": "Congratulations Statking. Achieving CV 0.9562 is excellent. Well done.\n\nI'm curious about this `scipy differential evolution method`, I plan to read more about this.",
    "974908": "Congratulations @deepkim , time well spent on creating so many models.",
    "975303": "Congrats! I wonder ur ensemble method in detail. I will wait for you leaving work 👍",
    "975568": "Thanks for sharing!\nIt's a shame that the table models lowered the score. Top solutions in alaska2 seem to do better with stacking and MLP ensemble for models that have extremely lower score than max CV score. maybe it is helpful.\nref : https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168870\n>mixing it well improves the accuracy, but since it is considerably weaker than 3-channel models as a single model, it doesn't work well when simply averaging it. So I used MLP.\n\nYou've trained many models, have you trained all the folds?\nAlso, did you use the Chris's notebook? (TF with TPU is needed for train many models?)",
    "975624": "Thank you so much!",
    "975630": "cdeotte thank you so much. i have detailed it further",
    "975632": "Thank you so much",
    "975634": "I have detailed it further.",
    "975659": "1. I have trained all 5 folds models\n2. I utilized Chris's notebook. It is really well made and logical with stratified Tfrecords. \n3. with TPU, I could train models faster.",
    "975906": "Thank you for your sincere congratulation!",
    "976502": "Congrats! I was pursuing the same approach as you, but I gave up before I had run all of the models because the admin of dealing with all the kernels was too much, and I had other approaches I wanted to pursue. \n\nI am curious as to how you dealt with this? Running 95 kernels on kaggle is a time consuming process.\n\nI am also curious as to why you say that dual input did not work when it dramatically increased your CV?",
    "976822": "1. 256, 384, 512, 768 Size didn't take time too much with the usage of kaggle tpu and colab\nbut 1024 size training was a somewhat time-consuming work\nbut we had 2~3 months.\n\n2. I tried meta info + CNN dual inputs training, it doesn't show  a sign of improving Public LB\nwith meta info added to CNN, my cv could reach above 96 just on ISIC 2020 data\nbut it didn't work for **Public LB** so i abandoned it(it was my mistake)\nby an adversarial analysis, we can infer that training datasets and test datasets are somewhat different.\nI guessed using meta info can lead my model to overfitting to just training datasets",
    "977025": "Okay thanks! Sorry though, I wasn't clear about my first question. The time consumption that I was speaking about was the manual task of running all the kernels, it takes a bit of effort to commit lots of different notebooks and then keep track of the separate parts (I assume your 1024 size had to be trained across multiple sessions?) and I was wondering if you have any way of dealing with that?",
    "977135": "You are right\nKaggle tpu time limit is 3 - hours per training , 30 hours per week (but it is really fast!)\nIf I had to train 5 folds of high-resolution images, then I needed to train 0,1,2 fold and after that, I train 3,4 fold seperately and merge [0,1,2] and [3,4] fold outputs into one output\n\n\nsince it was sort of tedious jobs, for most of training, I used my local environment",
    "980269": "Thanks for this. Go ahead, man :D",
    "982972": "Congrats!\nI am proud that you are Korean\n\nI have a question.\n\n1. Why did you apply 'label smooting'? Did you think there would be mislabelling?\n2. How to apply 'upsampling'?\n\nThank you ~",
    "984740": "Thank you for your congratulation!\n\nLet me give you an answer about what you asked\n\n1. the evaluation metric is AUC. In this competition, you don't need to classify accurately 0 : benign, 1 : malignant. it's somewhat different from measuring accuracy\n\nyou just need to calculate how much it is likely to be,,, benign one or malignant one\nit is expressed by \"a probability\"\nfor example,  an output of benign one: 0.00001, malignant : 0.9495......\n\ni had tested a lot of networks and found out that my CV auc values were clustered around 95.x\n\nit means my model performance is 95.x\n\nin this situation, generalizing output probability around 0.95x can be a good way for maintaining model stability. if you don't apply label smoothing, your model's output can be overconfident about malignant ones or benign ones. \n\nthat's why i applied label smoothing\n\n2.For upsampling, watch this!\nhttps://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
    "2917803": "Hello there :) In case you see this message... ISIC is now running a new challenge on Kaggle:\n\nhttps://www.kaggle.com/competitions/isic-2024-challenge/overview\n\nWould love to have you take part!"
  },
  "source": "meta"
}