{
  "id": 175656,
  "title": "I chose wrongly  !!!, I had 12 submissions in the medal zone",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175656",
  "author_name": "LwKzGonzalez",
  "post_date": "2020-08-18T23:38:32.795000",
  "votes": 0,
  "comment_count": 14,
  "views": 0,
  "content": "<p>This is my second Kaggle competion, and I have really learned alot.</p>\n<p>Many thanks to:</p>\n<p>Chris Deotte <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> (Triple Stratified KFold CV with TFRecords, Coarse Dropout)<br>\nSirish Somanchi <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> (Explanation of Power Averaging techniques)<br>\nGiba <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> (Simple baseline tabular data)</p>\n<p>Also thanks to all the the Kagglers who have posted notebooks and given ideas in the discussions list that have made me think and help me understerstand better this amazing field.</p>\n<p>First time ensembling and I have learned some basics of ensembling in this competition but still do not understand how to choose the good models for the private leaderboard.</p>\n<p>My ensembles were very simple. When trying to ensemble larger number of models my public leaderboard always worsened. Therefore I just did many different experiments but most of my main ensembles just consisted of two efficientnets models and one tabular model.</p>\n<p>I also added more layers to the efficientnets, and also applied 50 TTA.</p>\n<p>For the ensembles I used the power averaging technique.</p>\n<p>I had a whole range of submissions of public leaderboards with scores from 0.910 to 0.9570.</p>\n<p>Now that the competition has finished and saw that I could have chosen 12 better submissions that would have given me my first medal I almost fell down, being 0.9433 my best \"would have been\" private score. Even though everyone said \"Trust Your CV\", I didnt listen, I just chose my best public scores for the final submission.</p>\n<p>But well, I have learned a lesson the hard way.                 </p>",
  "messages": [
    {
      "id": 976546,
      "postDate": "2020-08-19T00:04:38.007Z",
      "content": "<p>Don't feel bad, it happens to everyone.  In my first competition, I chose badly and dropped 1000+ positions. This is part of the learning process.</p>\n<p>\"Trust your CV\" is difficult for everyone. It is so tempting to choose models that do well on public LB. The secret is to use KFold validation and save OOF for every one of your models. Then ensemble the models as explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a>. Lastly have confidence and choose the ensemble with greatest CV AUC OOF.</p>\n<p>You'll get a medal in your next comp!</p>",
      "rawMarkdown": "Don't feel bad, it happens to everyone.  In my first competition, I chose badly and dropped 1000+ positions. This is part of the learning process.\n\n\"Trust your CV\" is difficult for everyone. It is so tempting to choose models that do well on public LB. The secret is to use KFold validation and save OOF for every one of your models. Then ensemble the models as explained [here][1]. Lastly have confidence and choose the ensemble with greatest CV AUC OOF.\n\nYou'll get a medal in your next comp!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614",
      "votes": 12,
      "replies": [
        {
          "id": 976576,
          "postDate": "2020-08-19T00:41:11.147Z",
          "content": "<p>Thanks, you are truly a Grandmaster, you have offered guidance to all during the competion, and now the competition has finished you are continuing giving out more valuable machine learning techniques and advice. Will definetely continue my learning journey with your \"ensemble models using OOF files\".</p>",
          "rawMarkdown": "Thanks, you are truly a Grandmaster, you have offered guidance to all during the competion, and now the competition has finished you are continuing giving out more valuable machine learning techniques and advice. Will definetely continue my learning journey with your \"ensemble models using OOF files\".",
          "votes": 1
        },
        {
          "id": 976657,
          "postDate": "2020-08-19T02:42:44.923Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> how do you keep track of all your experiments? Do you use something like <a href=\"https://www.wandb.com/\" target=\"_blank\">https://www.wandb.com/</a> or just pen and paper? </p>",
          "rawMarkdown": "@cdeotte how do you keep track of all your experiments? Do you use something like https://www.wandb.com/ or just pen and paper? ",
          "votes": 1
        },
        {
          "id": 976663,
          "postDate": "2020-08-19T02:47:03.617Z",
          "content": "<p>I have three ways (1) good memory (2) hand written journal where I write what i do every day (3) disk folders - all my code each day with logs gets saved into a folder named by the day.</p>",
          "rawMarkdown": "I have three ways (1) good memory (2) hand written journal where I write what i do every day (3) disk folders - all my code each day with logs gets saved into a folder named by the day.",
          "votes": 7
        },
        {
          "id": 976665,
          "postDate": "2020-08-19T02:49:57.970Z",
          "content": "<p>Records are very important. To do well in competitions, you must do lots of experiments and you need to record the results somehow.</p>",
          "rawMarkdown": "Records are very important. To do well in competitions, you must do lots of experiments and you need to record the results somehow.",
          "votes": 4
        },
        {
          "id": 976669,
          "postDate": "2020-08-19T02:57:33.610Z",
          "content": "<p>Nice! I rely a lot on my journal too but sometimes my notes are not that well organize so I didn't find it too \"scientific\" haha so I started playing with wandb.com and found it amazing although sometimes pen and paper is easier.</p>",
          "rawMarkdown": "Nice! I rely a lot on my journal too but sometimes my notes are not that well organize so I didn't find it too \"scientific\" haha so I started playing with wandb.com and found it amazing although sometimes pen and paper is easier.",
          "votes": 3
        },
        {
          "id": 976699,
          "postDate": "2020-08-19T03:37:11.543Z",
          "content": "<p>I have similar problem, I was writing down my ideas on my notebook with pen but they are more like scribbles on paper, they look like encrypted writings from distance. I couldn't find my best single model parameters now 😃</p>",
          "rawMarkdown": "I have similar problem, I was writing down my ideas on my notebook with pen but they are more like scribbles on paper, they look like encrypted writings from distance. I couldn't find my best single model parameters now 😃",
          "votes": 3
        },
        {
          "id": 976713,
          "postDate": "2020-08-19T03:58:43.917Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, As you said, \"Trust your CV\" is really difficult. Do Finding out private/public test data samples split has any concern about choosing submission? If so what are the ways to mimic the private/public split on CV?</p>",
          "rawMarkdown": "Hey @cdeotte, As you said, \"Trust your CV\" is really difficult. Do Finding out private/public test data samples split has any concern about choosing submission? If so what are the ways to mimic the private/public split on CV?",
          "votes": 1
        },
        {
          "id": 977077,
          "postDate": "2020-08-19T09:36:24.157Z",
          "content": "<p>I automated the logging process by collating data such as model, image size and most of the hyper-parameter selections on each experiment or cross validation run and logging to a CSV file. The output of both OOF CV and CV with meta together with standard deviations for both were included. That process was embedded in each notebook and automatically updated the CSV. If I ran any predictions through the LB then I would manually include those scores in the CSV. It was invaluable for analysing CV/LB stability and making my final choices on this basis. I don't think I had the strongest models but that discipline definitely helped me get a better private LB score.</p>\n<p>There is a little effort involved in updating your template notebooks and CSV columns every time you want to include additional types of information but its worth it.</p>\n<p>Also notebooks which I ran on Kaggle or Colab would need to throw off an import file which I would then run a small import script into my local CSV file but that was OK. I would typically save locally all those models, weights and predictions in any case.</p>",
          "rawMarkdown": "I automated the logging process by collating data such as model, image size and most of the hyper-parameter selections on each experiment or cross validation run and logging to a CSV file. The output of both OOF CV and CV with meta together with standard deviations for both were included. That process was embedded in each notebook and automatically updated the CSV. If I ran any predictions through the LB then I would manually include those scores in the CSV. It was invaluable for analysing CV/LB stability and making my final choices on this basis. I don't think I had the strongest models but that discipline definitely helped me get a better private LB score.\n\nThere is a little effort involved in updating your template notebooks and CSV columns every time you want to include additional types of information but its worth it.\n\nAlso notebooks which I ran on Kaggle or Colab would need to throw off an import file which I would then run a small import script into my local CSV file but that was OK. I would typically save locally all those models, weights and predictions in any case.",
          "votes": 2
        },
        {
          "id": 978380,
          "postDate": "2020-08-20T06:30:18.617Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 976528,
      "postDate": "2020-08-18T23:38:32.797Z",
      "content": "<p>This is my second Kaggle competion, and I have really learned alot.</p>\n<p>Many thanks to:</p>\n<p>Chris Deotte <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> (Triple Stratified KFold CV with TFRecords, Coarse Dropout)<br>\nSirish Somanchi <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> (Explanation of Power Averaging techniques)<br>\nGiba <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> (Simple baseline tabular data)</p>\n<p>Also thanks to all the the Kagglers who have posted notebooks and given ideas in the discussions list that have made me think and help me understerstand better this amazing field.</p>\n<p>First time ensembling and I have learned some basics of ensembling in this competition but still do not understand how to choose the good models for the private leaderboard.</p>\n<p>My ensembles were very simple. When trying to ensemble larger number of models my public leaderboard always worsened. Therefore I just did many different experiments but most of my main ensembles just consisted of two efficientnets models and one tabular model.</p>\n<p>I also added more layers to the efficientnets, and also applied 50 TTA.</p>\n<p>For the ensembles I used the power averaging technique.</p>\n<p>I had a whole range of submissions of public leaderboards with scores from 0.910 to 0.9570.</p>\n<p>Now that the competition has finished and saw that I could have chosen 12 better submissions that would have given me my first medal I almost fell down, being 0.9433 my best \"would have been\" private score. Even though everyone said \"Trust Your CV\", I didnt listen, I just chose my best public scores for the final submission.</p>\n<p>But well, I have learned a lesson the hard way.                 </p>",
      "rawMarkdown": "This is my second Kaggle competion, and I have really learned alot.\n\nMany thanks to:\n\nChris Deotte @cdeotte (Triple Stratified KFold CV with TFRecords, Coarse Dropout)\nSirish Somanchi @sirishks (Explanation of Power Averaging techniques)\nGiba @titericz (Simple baseline tabular data)\n\nAlso thanks to all the the Kagglers who have posted notebooks and given ideas in the discussions list that have made me think and help me understerstand better this amazing field.\n\nFirst time ensembling and I have learned some basics of ensembling in this competition but still do not understand how to choose the good models for the private leaderboard.\n\nMy ensembles were very simple. When trying to ensemble larger number of models my public leaderboard always worsened. Therefore I just did many different experiments but most of my main ensembles just consisted of two efficientnets models and one tabular model.\n\nI also added more layers to the efficientnets, and also applied 50 TTA.\n\nFor the ensembles I used the power averaging technique.\n\nI had a whole range of submissions of public leaderboards with scores from 0.910 to 0.9570.\n\nNow that the competition has finished and saw that I could have chosen 12 better submissions that would have given me my first medal I almost fell down, being 0.9433 my best \"would have been\" private score. Even though everyone said \"Trust Your CV\", I didnt listen, I just chose my best public scores for the final submission.\n\nBut well, I have learned a lesson the hard way.                 ",
      "votes": 4
    },
    {
      "id": 978372,
      "postDate": "2020-08-20T06:23:51.717Z",
      "content": "<p>We have to use Some Heuristics based on which we should automate this process of chossing our final submissions that can maximise our private LB. it may vary from competition to competition. Like Balancing Bias and Variance, how can we balance our private and public LB? Any ideas</p>",
      "rawMarkdown": "We have to use Some Heuristics based on which we should automate this process of chossing our final submissions that can maximise our private LB. it may vary from competition to competition. Like Balancing Bias and Variance, how can we balance our private and public LB? Any ideas"
    },
    {
      "id": 977053,
      "postDate": "2020-08-19T09:11:49.943Z",
      "content": "<p>I think many kagglers in this competition has better submissions than selected ones. It is because tight leader board scores and high randomness. </p>",
      "rawMarkdown": "I think many kagglers in this competition has better submissions than selected ones. It is because tight leader board scores and high randomness. ",
      "replies": [
        {
          "id": 977902,
          "postDate": "2020-08-19T19:36:22.840Z",
          "content": "<p>So do you mean that in other Kaggle competitions the public leaderboard is similar to the private leaerboard? </p>",
          "rawMarkdown": "So do you mean that in other Kaggle competitions the public leaderboard is similar to the private leaerboard? "
        },
        {
          "id": 978367,
          "postDate": "2020-08-20T06:18:21.587Z",
          "content": "<p>Is some competitions yes in others no. It depends on many factors, such as test size, test/train similarity and other.</p>",
          "rawMarkdown": "Is some competitions yes in others no. It depends on many factors, such as test size, test/train similarity and other.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 976546,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-19T00:04:38.007000",
      "content": "<p>Don't feel bad, it happens to everyone.  In my first competition, I chose badly and dropped 1000+ positions. This is part of the learning process.</p>\n<p>\"Trust your CV\" is difficult for everyone. It is so tempting to choose models that do well on public LB. The secret is to use KFold validation and save OOF for every one of your models. Then ensemble the models as explained <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a>. Lastly have confidence and choose the ensemble with greatest CV AUC OOF.</p>\n<p>You'll get a medal in your next comp!</p>",
      "votes": 12,
      "replies": [
        {
          "id": 976576,
          "author_name": "LwKzGonzalez",
          "author_url": "",
          "post_date": "2020-08-19T00:41:11.147000",
          "content": "<p>Thanks, you are truly a Grandmaster, you have offered guidance to all during the competion, and now the competition has finished you are continuing giving out more valuable machine learning techniques and advice. Will definetely continue my learning journey with your \"ensemble models using OOF files\".</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976657,
          "author_name": "Santiago Viquez",
          "author_url": "",
          "post_date": "2020-08-19T02:42:44.923000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> how do you keep track of all your experiments? Do you use something like <a href=\"https://www.wandb.com/\" target=\"_blank\">https://www.wandb.com/</a> or just pen and paper? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976663,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T02:47:03.617000",
          "content": "<p>I have three ways (1) good memory (2) hand written journal where I write what i do every day (3) disk folders - all my code each day with logs gets saved into a folder named by the day.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 976665,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T02:49:57.970000",
          "content": "<p>Records are very important. To do well in competitions, you must do lots of experiments and you need to record the results somehow.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 976669,
          "author_name": "Santiago Viquez",
          "author_url": "",
          "post_date": "2020-08-19T02:57:33.610000",
          "content": "<p>Nice! I rely a lot on my journal too but sometimes my notes are not that well organize so I didn't find it too \"scientific\" haha so I started playing with wandb.com and found it amazing although sometimes pen and paper is easier.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 976699,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-08-19T03:37:11.543000",
          "content": "<p>I have similar problem, I was writing down my ideas on my notebook with pen but they are more like scribbles on paper, they look like encrypted writings from distance. I couldn't find my best single model parameters now 😃</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 976713,
          "author_name": "Aakash V",
          "author_url": "",
          "post_date": "2020-08-19T03:58:43.917000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, As you said, \"Trust your CV\" is really difficult. Do Finding out private/public test data samples split has any concern about choosing submission? If so what are the ways to mimic the private/public split on CV?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 977077,
          "author_name": "jsyphil",
          "author_url": "",
          "post_date": "2020-08-19T09:36:24.157000",
          "content": "<p>I automated the logging process by collating data such as model, image size and most of the hyper-parameter selections on each experiment or cross validation run and logging to a CSV file. The output of both OOF CV and CV with meta together with standard deviations for both were included. That process was embedded in each notebook and automatically updated the CSV. If I ran any predictions through the LB then I would manually include those scores in the CSV. It was invaluable for analysing CV/LB stability and making my final choices on this basis. I don't think I had the strongest models but that discipline definitely helped me get a better private LB score.</p>\n<p>There is a little effort involved in updating your template notebooks and CSV columns every time you want to include additional types of information but its worth it.</p>\n<p>Also notebooks which I ran on Kaggle or Colab would need to throw off an import file which I would then run a small import script into my local CSV file but that was OK. I would typically save locally all those models, weights and predictions in any case.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 978380,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-20T06:30:18.617000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 978372,
      "author_name": "Gokagglers ",
      "author_url": "",
      "post_date": "2020-08-20T06:23:51.717000",
      "content": "<p>We have to use Some Heuristics based on which we should automate this process of chossing our final submissions that can maximise our private LB. it may vary from competition to competition. Like Balancing Bias and Variance, how can we balance our private and public LB? Any ideas</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 977053,
      "author_name": "Nickolay Safronov",
      "author_url": "",
      "post_date": "2020-08-19T09:11:49.943000",
      "content": "<p>I think many kagglers in this competition has better submissions than selected ones. It is because tight leader board scores and high randomness. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 977902,
          "author_name": "LwKzGonzalez",
          "author_url": "",
          "post_date": "2020-08-19T19:36:22.840000",
          "content": "<p>So do you mean that in other Kaggle competitions the public leaderboard is similar to the private leaerboard? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 978367,
          "author_name": "Nickolay Safronov",
          "author_url": "",
          "post_date": "2020-08-20T06:18:21.587000",
          "content": "<p>Is some competitions yes in others no. It depends on many factors, such as test size, test/train similarity and other.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "976546": "Don't feel bad, it happens to everyone.  In my first competition, I chose badly and dropped 1000+ positions. This is part of the learning process.\n\n\"Trust your CV\" is difficult for everyone. It is so tempting to choose models that do well on public LB. The secret is to use KFold validation and save OOF for every one of your models. Then ensemble the models as explained [here][1]. Lastly have confidence and choose the ensemble with greatest CV AUC OOF.\n\nYou'll get a medal in your next comp!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614",
    "976528": "This is my second Kaggle competion, and I have really learned alot.\n\nMany thanks to:\n\nChris Deotte @cdeotte (Triple Stratified KFold CV with TFRecords, Coarse Dropout)\nSirish Somanchi @sirishks (Explanation of Power Averaging techniques)\nGiba @titericz (Simple baseline tabular data)\n\nAlso thanks to all the the Kagglers who have posted notebooks and given ideas in the discussions list that have made me think and help me understerstand better this amazing field.\n\nFirst time ensembling and I have learned some basics of ensembling in this competition but still do not understand how to choose the good models for the private leaderboard.\n\nMy ensembles were very simple. When trying to ensemble larger number of models my public leaderboard always worsened. Therefore I just did many different experiments but most of my main ensembles just consisted of two efficientnets models and one tabular model.\n\nI also added more layers to the efficientnets, and also applied 50 TTA.\n\nFor the ensembles I used the power averaging technique.\n\nI had a whole range of submissions of public leaderboards with scores from 0.910 to 0.9570.\n\nNow that the competition has finished and saw that I could have chosen 12 better submissions that would have given me my first medal I almost fell down, being 0.9433 my best \"would have been\" private score. Even though everyone said \"Trust Your CV\", I didnt listen, I just chose my best public scores for the final submission.\n\nBut well, I have learned a lesson the hard way.                 ",
    "978372": "We have to use Some Heuristics based on which we should automate this process of chossing our final submissions that can maximise our private LB. it may vary from competition to competition. Like Balancing Bias and Variance, how can we balance our private and public LB? Any ideas",
    "977053": "I think many kagglers in this competition has better submissions than selected ones. It is because tight leader board scores and high randomness. "
  }
}