{
  "id": 175846,
  "title": " Justification and 4th place Solution Overview",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/atagi-yuya-justification-and-4th-place-solution-ov",
  "author_name": "",
  "post_date": "2020-08-19T15:52:34.337Z",
  "votes": 21,
  "comment_count": 14,
  "views": 0,
  "content": "<p>First of all, thanks to the Kaggle community and the organizers.<br>\nI learned a lot through the competition.</p>\n<p>Some people suspect me to be a bot, so I will provide 4th place solution.<br>\nCertainly, I can't hide my surprise at this result.<br>\nBut I'm a little disappointed with the skepticism of the competition itself.<br>\nI hope the contest is held in our good faith.</p>\n<p>I could not submit it many times, of course, because the participation period was not long.<br>\nIn addition, the model was selected based on the CV results.<br>\nCV was done with multiple resolutions, and 384*384 was the best.<br>\nAlso, When ensemble with the best scoring model,<br>\nI was able to efficiently raise the score by performing the ensemble at the rate that the difference from the result after the ensemble became the minimum.<br>\nBy using this method, we were able to exclude the ones with very bad differences.<br>\nHowever, this method also carries the risk of overfitting public data.<br>\nI think there is also a risk in ignoring CV results and increasing the number of applications and raising the score.<br>\nIn fact, the 4th private score I got wasn't the best model in public.</p>\n<p>And I think the main factors to win are:<br>\nI have found that some of the prediction results of the CV model predict a high positive rate, whereas many kernels predict a low positive rate.<br>\nI thought this was a false negative in many published models.</p>\n<p>I finally submitted the following 3 models.<br>\n(1) Best score<br>\n(2) Ensemble with best score and model considered in CV<br>\n(3) Due to risk of overfitting, ensemble with model other than the best score and model considered in CV </p>\n<p>As a result, Model (3) was the 4th place result. and, (1) was the worst.<br>\nFrom this result, the following can be said.<br>\nLike many kernel predictive models, the best-scoring model are also models with low sensitivity and is more likely to predict false negatives.<br>\nThe highest scoring model fits public data too much.</p>\n<p>I think the following are valid:<br>\nTrust CV as many claims show<br>\nWe also observe trends by comparing with the results of many models that have a low correlation with the results examined in CV.</p>\n<p>Of course, it is very lucky that the model (3) got the 4th place result.</p>\n<p>I am aware of my lack of skills. However, I've found the fun of kaggle.<br>\nI will continue to deepen my learning through other contests.</p>\n<p>Best Regards</p>",
  "messages": [
    {
      "id": "977593",
      "postDate": "08/19/2020 15:29:43",
      "content": "<p>First of all, thanks to the Kaggle community and the organizers.<br>\nI learned a lot through the competition.</p>\n<p>Some people suspect me to be a bot, so I will provide 4th place solution.<br>\nCertainly, I can't hide my surprise at this result.<br>\nBut I'm a little disappointed with the skepticism of the competition itself.<br>\nI hope the contest is held in our good faith.</p>\n<p>I could not submit it many times, of course, because the participation period was not long.<br>\nIn addition, the model was selected based on the CV results.<br>\nCV was done with multiple resolutions, and 384*384 was the best.<br>\nAlso, When ensemble with the best scoring model,<br>\nI was able to efficiently raise the score by performing the ensemble at the rate that the difference from the result after the ensemble became the minimum.<br>\nBy using this method, we were able to exclude the ones with very bad differences.<br>\nHowever, this method also carries the risk of overfitting public data.<br>\nI think there is also a risk in ignoring CV results and increasing the number of applications and raising the score.<br>\nIn fact, the 4th private score I got wasn't the best model in public.</p>\n<p>And I think the main factors to win are:<br>\nI have found that some of the prediction results of the CV model predict a high positive rate, whereas many kernels predict a low positive rate.<br>\nI thought this was a false negative in many published models.</p>\n<p>I finally submitted the following 3 models.<br>\n(1) Best score<br>\n(2) Ensemble with best score and model considered in CV<br>\n(3) Due to risk of overfitting, ensemble with model other than the best score and model considered in CV </p>\n<p>As a result, Model (3) was the 4th place result. and, (1) was the worst.<br>\nFrom this result, the following can be said.<br>\nLike many kernel predictive models, the best-scoring model are also models with low sensitivity and is more likely to predict false negatives.<br>\nThe highest scoring model fits public data too much.</p>\n<p>I think the following are valid:<br>\nTrust CV as many claims show<br>\nWe also observe trends by comparing with the results of many models that have a low correlation with the results examined in CV.</p>\n<p>Of course, it is very lucky that the model (3) got the 4th place result.</p>\n<p>I am aware of my lack of skills. However, I've found the fun of kaggle.<br>\nI will continue to deepen my learning through other contests.</p>\n<p>Best Regards</p>",
      "rawMarkdown": "First of all, thanks to the Kaggle community and the organizers.\nI learned a lot through the competition.\n\nSome people suspect me to be a bot, so I will provide 4th place solution.\nCertainly, I can't hide my surprise at this result.\nBut I'm a little disappointed with the skepticism of the competition itself.\nI hope the contest is held in our good faith.\n\nI could not submit it many times, of course, because the participation period was not long.\nIn addition, the model was selected based on the CV results.\nCV was done with multiple resolutions, and 384*384 was the best.\nAlso, When ensemble with the best scoring model,\nI was able to efficiently raise the score by performing the ensemble at the rate that the difference from the result after the ensemble became the minimum.\nBy using this method, we were able to exclude the ones with very bad differences.\nHowever, this method also carries the risk of overfitting public data.\nI think there is also a risk in ignoring CV results and increasing the number of applications and raising the score.\nIn fact, the 4th private score I got wasn't the best model in public.\n\nAnd I think the main factors to win are:\nI have found that some of the prediction results of the CV model predict a high positive rate, whereas many kernels predict a low positive rate.\nI thought this was a false negative in many published models.\n\nI finally submitted the following 3 models.\n(1) Best score\n(2) Ensemble with best score and model considered in CV\n(3) Due to risk of overfitting, ensemble with model other than the best score and model considered in CV \n\nAs a result, Model (3) was the 4th place result. and, (1) was the worst.\nFrom this result, the following can be said.\nLike many kernel predictive models, the best-scoring model are also models with low sensitivity and is more likely to predict false negatives.\nThe highest scoring model fits public data too much.\n\nI think the following are valid:\nTrust CV as many claims show\nWe also observe trends by comparing with the results of many models that have a low correlation with the results examined in CV.\n\nOf course, it is very lucky that the model (3) got the 4th place result.\n\nI am aware of my lack of skills. However, I've found the fun of kaggle.\nI will continue to deepen my learning through other contests.\n\nBest Regards",
      "votes": null
    },
    {
      "id": "977623",
      "postDate": "08/19/2020 15:46:59",
      "content": "<p>Honestly, you did not convince me not to be a bot! 😂</p>\n<p>Just kidding, did not get your solution but congrats on your gold medal anyway!</p>",
      "rawMarkdown": "Honestly, you did not convince me not to be a bot! 😂\n\nJust kidding, did not get your solution but congrats on your gold medal anyway!",
      "votes": null
    },
    {
      "id": "977641",
      "postDate": "08/19/2020 15:57:47",
      "content": "<p>Omedetou Yuya-san, keep kaggling</p>",
      "rawMarkdown": "Omedetou Yuya-san, keep kaggling",
      "votes": null
    },
    {
      "id": "977674",
      "postDate": "08/19/2020 16:29:23",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/atagiyuya\" target=\"_blank\">@atagiyuya</a> and I hope to see you other competitions.</p>",
      "rawMarkdown": "congrats @atagiyuya and I hope to see you other competitions.",
      "votes": null
    },
    {
      "id": "977714",
      "postDate": "08/19/2020 16:59:45",
      "content": "<p>Congrats, man. I don't think you're a bot anymore. Next time, no one will think you're a bot. Good luck!</p>",
      "rawMarkdown": "Congrats, man. I don't think you're a bot anymore. Next time, no one will think you're a bot. Good luck!",
      "votes": null
    },
    {
      "id": "977825",
      "postDate": "08/19/2020 18:16:05",
      "content": "<p>Congrats! Very impressive with so few submissions! May I ask a question, please?</p>\n<p><code>Also, When ensemble with the best scoring model, I was able to efficiently raise the score by performing the ensemble at the rate that the difference from the result after the ensemble became the minimum. By using this method, we were able to exclude the ones with very bad differences.</code></p>\n<p>Would you mind explained a little bit about this part? How do you measure differences? By variance? Could do provide any formulations or codes? Thanks!</p>",
      "rawMarkdown": "Congrats! Very impressive with so few submissions! May I ask a question, please?\n\n`Also, When ensemble with the best scoring model, I was able to efficiently raise the score by performing the ensemble at the rate that the difference from the result after the ensemble became the minimum. By using this method, we were able to exclude the ones with very bad differences.`\n\nWould you mind explained a little bit about this part? How do you measure differences? By variance? Could do provide any formulations or codes? Thanks!",
      "votes": null
    },
    {
      "id": "978092",
      "postDate": "08/19/2020 23:45:35",
      "content": "<p>Congratulations with the gold! </p>",
      "rawMarkdown": "Congratulations with the gold!",
      "votes": null
    },
    {
      "id": "978788",
      "postDate": "08/20/2020 11:42:59",
      "content": "<p>Thank you everyone for celebrating me.</p>\n<p>To Helen-san<br>\nThank you for your comment.<br>\nI'm sorry to explain a very simple method.<br>\nTo use this method, you first need to find one prediction that contributed to best score model by the ensemble.<br>\nThat prediction is y.Then try to run the following code, where x is the prediction of the model you want to ensemble.<br>\nThat's how the models that helped to best score model are improved and help to score further.</p>\n<p>`import pandas as pd<br>\nimport numpy as np<br>\nfrom statistics import mean</p>\n<p>x=pd.read_csv(\"submission_x.csv\")[\"target\"]<br>\ny=pd.read_csv(\"submission_y.csv\")[\"target\"]<br>\nz=pd.read_csv(\"submission_best.csv\")[\"target\"]</p>\n<h1>x:Prediction you want to ensemble</h1>\n<h1>y:Prediction that contributed to the best prediction</h1>\n<h1>z:Best prediction</h1>\n<p>df=pd.read_csv(\"sample_submission.csv\")</p>\n<p>best=len(x)<br>\nfor a in np.arange(0, 1, 0.1):<br>\n    b=1-a<br>\n    c=(a<em>x+b</em>y)<br>\n    d=abs(z-c)/z<br>\n    res=mean(d)<br>\n    if res &lt;best:<br>\n        best=res<br>\n        beat_a=a<br>\n        beat_b=b<br>\n        best_c=c</p>\n<p>if best&gt;0.5:<br>\n    print(\"Do not recommend ensemble\")<br>\nelse:<br>\n    best_pre=z<em>0.9+best_c</em>0.1<br>\n    df[\"target\"]=best_pre<br>\n    df.to_csv(\"submission.csv\", index=False)<br>\nprint(best)<br>\nprint(f'a: {beat_a}')<br>\nprint(f'b: {beat_b}')`</p>",
      "rawMarkdown": "Thank you everyone for celebrating me.\n\nTo Helen-san\nThank you for your comment.\nI'm sorry to explain a very simple method.\nTo use this method, you first need to find one prediction that contributed to best score model by the ensemble.\nThat prediction is y.Then try to run the following code, where x is the prediction of the model you want to ensemble.\nThat's how the models that helped to best score model are improved and help to score further.\n\n`import pandas as pd\nimport numpy as np\nfrom statistics import mean\n\nx=pd.read_csv(\"submission_x.csv\")[\"target\"]\ny=pd.read_csv(\"submission_y.csv\")[\"target\"]\nz=pd.read_csv(\"submission_best.csv\")[\"target\"]\n#x:Prediction you want to ensemble\n#y:Prediction that contributed to the best prediction\n#z:Best prediction\ndf=pd.read_csv(\"sample_submission.csv\")\n\nbest=len(x)\nfor a in np.arange(0, 1, 0.1):\n    b=1-a\n    c=(a*x+b*y)\n    d=abs(z-c)/z\n    res=mean(d)\n    if res <best:\n        best=res\n        beat_a=a\n        beat_b=b\n        best_c=c\n\nif best>0.5:\n    print(\"Do not recommend ensemble\")\nelse:\n    best_pre=z*0.9+best_c*0.1\n    df[\"target\"]=best_pre\n    df.to_csv(\"submission.csv\", index=False)\nprint(best)\nprint(f'a: {beat_a}')\nprint(f'b: {beat_b}')`",
      "votes": null
    },
    {
      "id": "978796",
      "postDate": "08/20/2020 11:53:01",
      "content": "<pre><code># I think I understand what's going on here.\nx = New Ensemble candidate\ny = Strongest single model's predictions\nz = Highest Scoring Current CV (or LB) Ensemble\n</code></pre>\n<p>Is that accurate?</p>",
      "rawMarkdown": "```\n# I think I understand what's going on here.\nx = New Ensemble candidate\ny = Strongest single model's predictions\nz = Highest Scoring Current CV (or LB) Ensemble\n```\n\nIs that accurate?",
      "votes": null
    },
    {
      "id": "978845",
      "postDate": "08/20/2020 12:53:13",
      "content": "<p>Hi, you don't give much details.  What data do you use?  Images only?  meta data?  Do you use any external data?  What image sizes?</p>\n<p>You don't describe what models you use.  Efficientnets?  What model sizes?  What loss function?</p>\n<p>Etc.</p>\n<p>Frankly, I don't know what one can learn from your writeup.  </p>",
      "rawMarkdown": "Hi, you don't give much details.  What data do you use?  Images only?  meta data?  Do you use any external data?  What image sizes?\n\nYou don't describe what models you use.  Efficientnets?  What model sizes?  What loss function?\n\nEtc.\n\nFrankly, I don't know what one can learn from your writeup.",
      "votes": null
    },
    {
      "id": "979163",
      "postDate": "08/20/2020 17:02:13",
      "content": "<p>You made me notice a lot. Thank you very much.</p>\n<p>Each parameter I used is:</p>\n<p>DataSet:2017-2018-2019 + 2020 TFrecords (thanks to  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>)<br>\nImage_size:384*384<br>\nModel:efficientnet-b6<br>\nLoss:binary_crossentropy<br>\nOptimizer:Adam<br>\nSeed:42<br>\nTTA:15<br>\nModel building:<br>\nWith reference to *<a href=\"https://www.kaggle.com/ajaykumar7778/efficientnet-cv\" target=\"_blank\">https://www.kaggle.com/ajaykumar7778/efficientnet-cv</a> (thanks to <a href=\"https://www.kaggle.com/ajaykumar7778\" target=\"_blank\">@ajaykumar7778</a>)<br>\n  x = base(keras.layers.Input(shape=(384,384,3)))<br>\n  x = keras.layers.GlobalAveragePooling2D()(x)<br>\n  x = keras.layers.Dropout(0.3)(x)<br>\n  x = keras.layers.Dense(1024)(x)<br>\n  x = keras.layers.Dropout(0.2)(x)<br>\n  x = keras.layers.Dense(512)(x)<br>\n  x = keras.layers.Dropout(0.4)(x)<br>\n  x = keras.layers.Dense(256)(x)<br>\n  x = keras.layers.Dropout(0.3)(x)<br>\n  x = keras.layers.Dense(128)(x)  <br>\n  x = keras.layers.Dense(1)(x)<br>\n  x = keras.layers.Activation('sigmoid', dtype='float32')(x)</p>\n<p>By blending with *<a href=\"https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619\" target=\"_blank\">https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619</a>, Scored 0.9596 public LB and 0.9476 private LB. (thanks to@Pak Lau9)<br>\nI believe that one of the reason is probably that the R2 score is less than 0.6 between the ensemble models, the correlation is not large, and the generalization performance is high.</p>",
      "rawMarkdown": "You made me notice a lot. Thank you very much.\n\nEach parameter I used is:\n\nDataSet:2017-2018-2019 + 2020 TFrecords (thanks to  @cdeotte)\nImage_size:384*384\nModel:efficientnet-b6\nLoss:binary_crossentropy\nOptimizer:Adam\nSeed:42\nTTA:15\nModel building:\nWith reference to *https://www.kaggle.com/ajaykumar7778/efficientnet-cv (thanks to @ajaykumar7778)\n  x = base(keras.layers.Input(shape=(384,384,3)))\n  x = keras.layers.GlobalAveragePooling2D()(x)\n  x = keras.layers.Dropout(0.3)(x)\n  x = keras.layers.Dense(1024)(x)\n  x = keras.layers.Dropout(0.2)(x)\n  x = keras.layers.Dense(512)(x)\n  x = keras.layers.Dropout(0.4)(x)\n  x = keras.layers.Dense(256)(x)\n  x = keras.layers.Dropout(0.3)(x)\n  x = keras.layers.Dense(128)(x)  \n  x = keras.layers.Dense(1)(x)\n  x = keras.layers.Activation('sigmoid', dtype='float32')(x)\n\nBy blending with *https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619, Scored 0.9596 public LB and 0.9476 private LB. (thanks to@Pak Lau9)\nI believe that one of the reason is probably that the R2 score is less than 0.6 between the ensemble models, the correlation is not large, and the generalization performance is high.",
      "votes": null
    },
    {
      "id": "979171",
      "postDate": "08/20/2020 17:06:59",
      "content": "<p>Almost the same interpretation<br>\n<em>y</em> does not have to be a single model, but it is not recommended if the R2 score between <em>y</em> and <em>z</em> is high (R2&gt; 0.7).<br>\nOn the other hand, choosing <em>y</em> with a higher R2 may increase the public score, but at the risk of overfitting the public score.<br>\nIf you find a model where R2 between <em>y</em> and <em>z</em> is between 0.5 and 0.7 and can contribute to the public score, your private score will be guaranteed too.</p>",
      "rawMarkdown": "Almost the same interpretation\n*y* does not have to be a single model, but it is not recommended if the R2 score between *y* and *z* is high (R2> 0.7).\nOn the other hand, choosing *y* with a higher R2 may increase the public score, but at the risk of overfitting the public score.\nIf you find a model where R2 between *y* and *z* is between 0.5 and 0.7 and can contribute to the public score, your private score will be guaranteed too.",
      "votes": null
    },
    {
      "id": "979193",
      "postDate": "08/20/2020 17:24:52",
      "content": "<p>thanks, this is much better.</p>",
      "rawMarkdown": "thanks, this is much better.",
      "votes": null
    },
    {
      "id": "982710",
      "postDate": "08/23/2020 15:46:57",
      "content": "<p><a href=\"https://www.kaggle.com/atagiyuya\" target=\"_blank\">@atagiyuya</a> Thanks for your explanation. Sorry for the late reply, i missed it. Now I understand, it's really a good idea. Learned a lot! 👍</p>",
      "rawMarkdown": "atagiyuya Thanks for your explanation. Sorry for the late reply, i missed it. Now I understand, it's really a good idea. Learned a lot! 👍",
      "votes": null
    },
    {
      "id": "2917812",
      "postDate": "07/11/2024 19:49:58",
      "content": "<p>Hello Atagi Yuya :)</p>\n<p>I figured given your excellent performance in the last (2020 SIIM-ISIC) Melanoma challenge, I'd just point out we are running our next challenge with Kaggle:</p>\n<p><a href=\"https://www.kaggle.com/competitions/isic-2024-challenge/overview\" target=\"_blank\">https://www.kaggle.com/competitions/isic-2024-challenge/overview</a></p>\n<p>Would be fun to see you join!</p>",
      "rawMarkdown": "Hello Atagi Yuya :)\n\nI figured given your excellent performance in the last (2020 SIIM-ISIC) Melanoma challenge, I'd just point out we are running our next challenge with Kaggle:\n\nhttps://www.kaggle.com/competitions/isic-2024-challenge/overview\n\nWould be fun to see you join!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 977623,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "08/19/2020 15:46:59",
      "content": "<p>Honestly, you did not convince me not to be a bot! 😂</p>\n<p>Just kidding, did not get your solution but congrats on your gold medal anyway!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977641,
      "author_name": "authman",
      "author_url": "",
      "post_date": "08/19/2020 15:57:47",
      "content": "<p>Omedetou Yuya-san, keep kaggling</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977674,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "08/19/2020 16:29:23",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/atagiyuya\" target=\"_blank\">@atagiyuya</a> and I hope to see you other competitions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977714,
      "author_name": "vadimtimakin",
      "author_url": "",
      "post_date": "08/19/2020 16:59:45",
      "content": "<p>Congrats, man. I don't think you're a bot anymore. Next time, no one will think you're a bot. Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977825,
      "author_name": "yuanlin08",
      "author_url": "",
      "post_date": "08/19/2020 18:16:05",
      "content": "<p>Congrats! Very impressive with so few submissions! May I ask a question, please?</p>\n<p><code>Also, When ensemble with the best scoring model, I was able to efficiently raise the score by performing the ensemble at the rate that the difference from the result after the ensemble became the minimum. By using this method, we were able to exclude the ones with very bad differences.</code></p>\n<p>Would you mind explained a little bit about this part? How do you measure differences? By variance? Could do provide any formulations or codes? Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 978092,
      "author_name": "antaresnyc",
      "author_url": "",
      "post_date": "08/19/2020 23:45:35",
      "content": "<p>Congratulations with the gold! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 978788,
      "author_name": "atagiyuya",
      "author_url": "",
      "post_date": "08/20/2020 11:42:59",
      "content": "<p>Thank you everyone for celebrating me.</p>\n<p>To Helen-san<br>\nThank you for your comment.<br>\nI'm sorry to explain a very simple method.<br>\nTo use this method, you first need to find one prediction that contributed to best score model by the ensemble.<br>\nThat prediction is y.Then try to run the following code, where x is the prediction of the model you want to ensemble.<br>\nThat's how the models that helped to best score model are improved and help to score further.</p>\n<p>`import pandas as pd<br>\nimport numpy as np<br>\nfrom statistics import mean</p>\n<p>x=pd.read_csv(\"submission_x.csv\")[\"target\"]<br>\ny=pd.read_csv(\"submission_y.csv\")[\"target\"]<br>\nz=pd.read_csv(\"submission_best.csv\")[\"target\"]</p>\n<h1>x:Prediction you want to ensemble</h1>\n<h1>y:Prediction that contributed to the best prediction</h1>\n<h1>z:Best prediction</h1>\n<p>df=pd.read_csv(\"sample_submission.csv\")</p>\n<p>best=len(x)<br>\nfor a in np.arange(0, 1, 0.1):<br>\n    b=1-a<br>\n    c=(a<em>x+b</em>y)<br>\n    d=abs(z-c)/z<br>\n    res=mean(d)<br>\n    if res &lt;best:<br>\n        best=res<br>\n        beat_a=a<br>\n        beat_b=b<br>\n        best_c=c</p>\n<p>if best&gt;0.5:<br>\n    print(\"Do not recommend ensemble\")<br>\nelse:<br>\n    best_pre=z<em>0.9+best_c</em>0.1<br>\n    df[\"target\"]=best_pre<br>\n    df.to_csv(\"submission.csv\", index=False)<br>\nprint(best)<br>\nprint(f'a: {beat_a}')<br>\nprint(f'b: {beat_b}')`</p>",
      "votes": null,
      "replies": [
        {
          "id": 978796,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/20/2020 11:53:01",
          "content": "<pre><code># I think I understand what's going on here.\nx = New Ensemble candidate\ny = Strongest single model's predictions\nz = Highest Scoring Current CV (or LB) Ensemble\n</code></pre>\n<p>Is that accurate?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979171,
          "author_name": "atagiyuya",
          "author_url": "",
          "post_date": "08/20/2020 17:06:59",
          "content": "<p>Almost the same interpretation<br>\n<em>y</em> does not have to be a single model, but it is not recommended if the R2 score between <em>y</em> and <em>z</em> is high (R2&gt; 0.7).<br>\nOn the other hand, choosing <em>y</em> with a higher R2 may increase the public score, but at the risk of overfitting the public score.<br>\nIf you find a model where R2 between <em>y</em> and <em>z</em> is between 0.5 and 0.7 and can contribute to the public score, your private score will be guaranteed too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 982710,
          "author_name": "yuanlin08",
          "author_url": "",
          "post_date": "08/23/2020 15:46:57",
          "content": "<p><a href=\"https://www.kaggle.com/atagiyuya\" target=\"_blank\">@atagiyuya</a> Thanks for your explanation. Sorry for the late reply, i missed it. Now I understand, it's really a good idea. Learned a lot! 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 978845,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/20/2020 12:53:13",
      "content": "<p>Hi, you don't give much details.  What data do you use?  Images only?  meta data?  Do you use any external data?  What image sizes?</p>\n<p>You don't describe what models you use.  Efficientnets?  What model sizes?  What loss function?</p>\n<p>Etc.</p>\n<p>Frankly, I don't know what one can learn from your writeup.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 979163,
          "author_name": "atagiyuya",
          "author_url": "",
          "post_date": "08/20/2020 17:02:13",
          "content": "<p>You made me notice a lot. Thank you very much.</p>\n<p>Each parameter I used is:</p>\n<p>DataSet:2017-2018-2019 + 2020 TFrecords (thanks to  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>)<br>\nImage_size:384*384<br>\nModel:efficientnet-b6<br>\nLoss:binary_crossentropy<br>\nOptimizer:Adam<br>\nSeed:42<br>\nTTA:15<br>\nModel building:<br>\nWith reference to *<a href=\"https://www.kaggle.com/ajaykumar7778/efficientnet-cv\" target=\"_blank\">https://www.kaggle.com/ajaykumar7778/efficientnet-cv</a> (thanks to <a href=\"https://www.kaggle.com/ajaykumar7778\" target=\"_blank\">@ajaykumar7778</a>)<br>\n  x = base(keras.layers.Input(shape=(384,384,3)))<br>\n  x = keras.layers.GlobalAveragePooling2D()(x)<br>\n  x = keras.layers.Dropout(0.3)(x)<br>\n  x = keras.layers.Dense(1024)(x)<br>\n  x = keras.layers.Dropout(0.2)(x)<br>\n  x = keras.layers.Dense(512)(x)<br>\n  x = keras.layers.Dropout(0.4)(x)<br>\n  x = keras.layers.Dense(256)(x)<br>\n  x = keras.layers.Dropout(0.3)(x)<br>\n  x = keras.layers.Dense(128)(x)  <br>\n  x = keras.layers.Dense(1)(x)<br>\n  x = keras.layers.Activation('sigmoid', dtype='float32')(x)</p>\n<p>By blending with *<a href=\"https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619\" target=\"_blank\">https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619</a>, Scored 0.9596 public LB and 0.9476 private LB. (thanks to@Pak Lau9)<br>\nI believe that one of the reason is probably that the R2 score is less than 0.6 between the ensemble models, the correlation is not large, and the generalization performance is high.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979193,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/20/2020 17:24:52",
          "content": "<p>thanks, this is much better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2917812,
      "author_name": "jwebermsk",
      "author_url": "",
      "post_date": "07/11/2024 19:49:58",
      "content": "<p>Hello Atagi Yuya :)</p>\n<p>I figured given your excellent performance in the last (2020 SIIM-ISIC) Melanoma challenge, I'd just point out we are running our next challenge with Kaggle:</p>\n<p><a href=\"https://www.kaggle.com/competitions/isic-2024-challenge/overview\" target=\"_blank\">https://www.kaggle.com/competitions/isic-2024-challenge/overview</a></p>\n<p>Would be fun to see you join!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "977593": "First of all, thanks to the Kaggle community and the organizers.\nI learned a lot through the competition.\n\nSome people suspect me to be a bot, so I will provide 4th place solution.\nCertainly, I can't hide my surprise at this result.\nBut I'm a little disappointed with the skepticism of the competition itself.\nI hope the contest is held in our good faith.\n\nI could not submit it many times, of course, because the participation period was not long.\nIn addition, the model was selected based on the CV results.\nCV was done with multiple resolutions, and 384*384 was the best.\nAlso, When ensemble with the best scoring model,\nI was able to efficiently raise the score by performing the ensemble at the rate that the difference from the result after the ensemble became the minimum.\nBy using this method, we were able to exclude the ones with very bad differences.\nHowever, this method also carries the risk of overfitting public data.\nI think there is also a risk in ignoring CV results and increasing the number of applications and raising the score.\nIn fact, the 4th private score I got wasn't the best model in public.\n\nAnd I think the main factors to win are:\nI have found that some of the prediction results of the CV model predict a high positive rate, whereas many kernels predict a low positive rate.\nI thought this was a false negative in many published models.\n\nI finally submitted the following 3 models.\n(1) Best score\n(2) Ensemble with best score and model considered in CV\n(3) Due to risk of overfitting, ensemble with model other than the best score and model considered in CV \n\nAs a result, Model (3) was the 4th place result. and, (1) was the worst.\nFrom this result, the following can be said.\nLike many kernel predictive models, the best-scoring model are also models with low sensitivity and is more likely to predict false negatives.\nThe highest scoring model fits public data too much.\n\nI think the following are valid:\nTrust CV as many claims show\nWe also observe trends by comparing with the results of many models that have a low correlation with the results examined in CV.\n\nOf course, it is very lucky that the model (3) got the 4th place result.\n\nI am aware of my lack of skills. However, I've found the fun of kaggle.\nI will continue to deepen my learning through other contests.\n\nBest Regards",
    "977623": "Honestly, you did not convince me not to be a bot! 😂\n\nJust kidding, did not get your solution but congrats on your gold medal anyway!",
    "977641": "Omedetou Yuya-san, keep kaggling",
    "977674": "congrats @atagiyuya and I hope to see you other competitions.",
    "977714": "Congrats, man. I don't think you're a bot anymore. Next time, no one will think you're a bot. Good luck!",
    "977825": "Congrats! Very impressive with so few submissions! May I ask a question, please?\n\n`Also, When ensemble with the best scoring model, I was able to efficiently raise the score by performing the ensemble at the rate that the difference from the result after the ensemble became the minimum. By using this method, we were able to exclude the ones with very bad differences.`\n\nWould you mind explained a little bit about this part? How do you measure differences? By variance? Could do provide any formulations or codes? Thanks!",
    "978092": "Congratulations with the gold!",
    "978788": "Thank you everyone for celebrating me.\n\nTo Helen-san\nThank you for your comment.\nI'm sorry to explain a very simple method.\nTo use this method, you first need to find one prediction that contributed to best score model by the ensemble.\nThat prediction is y.Then try to run the following code, where x is the prediction of the model you want to ensemble.\nThat's how the models that helped to best score model are improved and help to score further.\n\n`import pandas as pd\nimport numpy as np\nfrom statistics import mean\n\nx=pd.read_csv(\"submission_x.csv\")[\"target\"]\ny=pd.read_csv(\"submission_y.csv\")[\"target\"]\nz=pd.read_csv(\"submission_best.csv\")[\"target\"]\n#x:Prediction you want to ensemble\n#y:Prediction that contributed to the best prediction\n#z:Best prediction\ndf=pd.read_csv(\"sample_submission.csv\")\n\nbest=len(x)\nfor a in np.arange(0, 1, 0.1):\n    b=1-a\n    c=(a*x+b*y)\n    d=abs(z-c)/z\n    res=mean(d)\n    if res <best:\n        best=res\n        beat_a=a\n        beat_b=b\n        best_c=c\n\nif best>0.5:\n    print(\"Do not recommend ensemble\")\nelse:\n    best_pre=z*0.9+best_c*0.1\n    df[\"target\"]=best_pre\n    df.to_csv(\"submission.csv\", index=False)\nprint(best)\nprint(f'a: {beat_a}')\nprint(f'b: {beat_b}')`",
    "978796": "```\n# I think I understand what's going on here.\nx = New Ensemble candidate\ny = Strongest single model's predictions\nz = Highest Scoring Current CV (or LB) Ensemble\n```\n\nIs that accurate?",
    "978845": "Hi, you don't give much details.  What data do you use?  Images only?  meta data?  Do you use any external data?  What image sizes?\n\nYou don't describe what models you use.  Efficientnets?  What model sizes?  What loss function?\n\nEtc.\n\nFrankly, I don't know what one can learn from your writeup.",
    "979163": "You made me notice a lot. Thank you very much.\n\nEach parameter I used is:\n\nDataSet:2017-2018-2019 + 2020 TFrecords (thanks to  @cdeotte)\nImage_size:384*384\nModel:efficientnet-b6\nLoss:binary_crossentropy\nOptimizer:Adam\nSeed:42\nTTA:15\nModel building:\nWith reference to *https://www.kaggle.com/ajaykumar7778/efficientnet-cv (thanks to @ajaykumar7778)\n  x = base(keras.layers.Input(shape=(384,384,3)))\n  x = keras.layers.GlobalAveragePooling2D()(x)\n  x = keras.layers.Dropout(0.3)(x)\n  x = keras.layers.Dense(1024)(x)\n  x = keras.layers.Dropout(0.2)(x)\n  x = keras.layers.Dense(512)(x)\n  x = keras.layers.Dropout(0.4)(x)\n  x = keras.layers.Dense(256)(x)\n  x = keras.layers.Dropout(0.3)(x)\n  x = keras.layers.Dense(128)(x)  \n  x = keras.layers.Dense(1)(x)\n  x = keras.layers.Activation('sigmoid', dtype='float32')(x)\n\nBy blending with *https://www.kaggle.com/paklau9/minmax-highest-public-lb-9619, Scored 0.9596 public LB and 0.9476 private LB. (thanks to@Pak Lau9)\nI believe that one of the reason is probably that the R2 score is less than 0.6 between the ensemble models, the correlation is not large, and the generalization performance is high.",
    "979171": "Almost the same interpretation\n*y* does not have to be a single model, but it is not recommended if the R2 score between *y* and *z* is high (R2> 0.7).\nOn the other hand, choosing *y* with a higher R2 may increase the public score, but at the risk of overfitting the public score.\nIf you find a model where R2 between *y* and *z* is between 0.5 and 0.7 and can contribute to the public score, your private score will be guaranteed too.",
    "979193": "thanks, this is much better.",
    "982710": "atagiyuya Thanks for your explanation. Sorry for the late reply, i missed it. Now I understand, it's really a good idea. Learned a lot! 👍",
    "2917812": "Hello Atagi Yuya :)\n\nI figured given your excellent performance in the last (2020 SIIM-ISIC) Melanoma challenge, I'd just point out we are running our next challenge with Kaggle:\n\nhttps://www.kaggle.com/competitions/isic-2024-challenge/overview\n\nWould be fun to see you join!"
  },
  "source": "meta"
}