{
  "id": 160884,
  "title": "Two bronze medals gone to dust. Guide Needed from Kagglers.",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/160884",
  "author_name": "",
  "post_date": "2020-06-23T02:57:26.015012700Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone. I just joined Kaggle and after working very hard \n[1]: I was on Bronze List in \"Tweet Sentiment Extraction\" and the shake up took it away.\n[2]: My one submission on Public LB in \"Jigsaw Multilingual Toxic comment classification\" was on Bronze in Private LB but I missed it while selecting the best ones.</p>\n\n<p>What could have been done to avoid the above ? Any guidance would be appreciated. </p>",
  "messages": [
    {
      "id": "897660",
      "postDate": "06/23/2020 02:57:26",
      "content": "<p>Hi everyone. I just joined Kaggle and after working very hard \n[1]: I was on Bronze List in \"Tweet Sentiment Extraction\" and the shake up took it away.\n[2]: My one submission on Public LB in \"Jigsaw Multilingual Toxic comment classification\" was on Bronze in Private LB but I missed it while selecting the best ones.</p>\n\n<p>What could have been done to avoid the above ? Any guidance would be appreciated. </p>",
      "rawMarkdown": "Hi everyone. I just joined Kaggle and after working very hard \n[1]: I was on Bronze List in \"Tweet Sentiment Extraction\" and the shake up took it away.\n[2]: My one submission on Public LB in \"Jigsaw Multilingual Toxic comment classification\" was on Bronze in Private LB but I missed it while selecting the best ones.\n\nWhat could have been done to avoid the above ? Any guidance would be appreciated.",
      "votes": null
    },
    {
      "id": "897845",
      "postDate": "06/23/2020 06:18:33",
      "content": "<p>Selecting the best submission it's sometimes a matter of luck, especially if they are close in score. However, one can typically limit the effect of shakeup by choosing 2 diverse submissions, possibly addressing different scenarios for private LB. People often choose best public LB + best local CV. But sometimes the public LB is so not reliable that you may just want to choose based on the CV. It depends. \nIn general, one should really not rely entirely on the public LB.</p>",
      "rawMarkdown": "Selecting the best submission it's sometimes a matter of luck, especially if they are close in score. However, one can typically limit the effect of shakeup by choosing 2 diverse submissions, possibly addressing different scenarios for private LB. People often choose best public LB + best local CV. But sometimes the public LB is so not reliable that you may just want to choose based on the CV. It depends. \nIn general, one should really not rely entirely on the public LB.",
      "votes": null
    },
    {
      "id": "897902",
      "postDate": "06/23/2020 07:09:00",
      "content": "<p>Thank you. That is informative. Definitely going to consider this in future.</p>",
      "rawMarkdown": "Thank you. That is informative. Definitely going to consider this in future.",
      "votes": null
    },
    {
      "id": "897929",
      "postDate": "06/23/2020 07:23:20",
      "content": "<p>What I did was this while selecting the submission. \n1. I had many submission in the range of 0.9477 to 0.9479. So, either I could choose one with a high public LB score or a little bit less Public LB score. So, I looked at count of positive and negative prediction. Based on previous experience with the submissions I could tell that there can not be more than 9K positive samples in whole test data. The reason is that only those submission score 0.947x whose positive sample count was between 9K to 10K. So, I choose the one with a lower count of positive sample and hence submitted 0.9477. This was my hunch and theory based on a number of samples in validation and train data. \n2. For the second submission, I did the same. I choose the submission with the lowest count of positive samples. There was one with public score 0.9475 and positive samples around 8K only. So I choose that one for the second final submission. \nFortunately, the one with 0.9477 scored more on private LB more compared to 0.9479 one and 0.9475 one. So, it is like understanding of total test set and overfitting on the public LB.</p>",
      "rawMarkdown": "What I did was this while selecting the submission. \n1. I had many submission in the range of 0.9477 to 0.9479. So, either I could choose one with a high public LB score or a little bit less Public LB score. So, I looked at count of positive and negative prediction. Based on previous experience with the submissions I could tell that there can not be more than 9K positive samples in whole test data. The reason is that only those submission score 0.947x whose positive sample count was between 9K to 10K. So, I choose the one with a lower count of positive sample and hence submitted 0.9477. This was my hunch and theory based on a number of samples in validation and train data. \n2. For the second submission, I did the same. I choose the submission with the lowest count of positive samples. There was one with public score 0.9475 and positive samples around 8K only. So I choose that one for the second final submission. \nFortunately, the one with 0.9477 scored more on private LB more compared to 0.9479 one and 0.9475 one. So, it is like understanding of total test set and overfitting on the public LB.",
      "votes": null
    },
    {
      "id": "898028",
      "postDate": "06/23/2020 08:33:42",
      "content": "<p>That's the crux of Kaggle: being able to asses model performance as correctly as possible so that you select the right one.  I would recommend you try to understand what cold have made you select the right one here.  How can you modify your evaluation procedure (aka CV or LB probing, or  a combination) so that you asses your model performance correctly?</p>\n\n<p>When reading top teams writeups, make sure you understand how they validated model performance.  Often this is not described enough for my tatste, but sometimes people do it.  For instance I describe my CV scheme in Tweet.</p>",
      "rawMarkdown": "That's the crux of Kaggle: being able to asses model performance as correctly as possible so that you select the right one.  I would recommend you try to understand what cold have made you select the right one here.  How can you modify your evaluation procedure (aka CV or LB probing, or  a combination) so that you asses your model performance correctly?\n\nWhen reading top teams writeups, make sure you understand how they validated model performance.  Often this is not described enough for my tatste, but sometimes people do it.  For instance I describe my CV scheme in Tweet.",
      "votes": null
    },
    {
      "id": "898219",
      "postDate": "06/23/2020 11:31:24",
      "content": "<p>Thank you <a href=\"/cpmpml\">@cpmpml</a> . I have been following you everywhere. Keep inspiring. That was informative.</p>",
      "rawMarkdown": "Thank you @cpmpml . I have been following you everywhere. Keep inspiring. That was informative.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 897845,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "06/23/2020 06:18:33",
      "content": "<p>Selecting the best submission it's sometimes a matter of luck, especially if they are close in score. However, one can typically limit the effect of shakeup by choosing 2 diverse submissions, possibly addressing different scenarios for private LB. People often choose best public LB + best local CV. But sometimes the public LB is so not reliable that you may just want to choose based on the CV. It depends. \nIn general, one should really not rely entirely on the public LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 897902,
          "author_name": "jamshaidsohail5",
          "author_url": "",
          "post_date": "06/23/2020 07:09:00",
          "content": "<p>Thank you. That is informative. Definitely going to consider this in future.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897929,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "06/23/2020 07:23:20",
      "content": "<p>What I did was this while selecting the submission. \n1. I had many submission in the range of 0.9477 to 0.9479. So, either I could choose one with a high public LB score or a little bit less Public LB score. So, I looked at count of positive and negative prediction. Based on previous experience with the submissions I could tell that there can not be more than 9K positive samples in whole test data. The reason is that only those submission score 0.947x whose positive sample count was between 9K to 10K. So, I choose the one with a lower count of positive sample and hence submitted 0.9477. This was my hunch and theory based on a number of samples in validation and train data. \n2. For the second submission, I did the same. I choose the submission with the lowest count of positive samples. There was one with public score 0.9475 and positive samples around 8K only. So I choose that one for the second final submission. \nFortunately, the one with 0.9477 scored more on private LB more compared to 0.9479 one and 0.9475 one. So, it is like understanding of total test set and overfitting on the public LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898028,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/23/2020 08:33:42",
      "content": "<p>That's the crux of Kaggle: being able to asses model performance as correctly as possible so that you select the right one.  I would recommend you try to understand what cold have made you select the right one here.  How can you modify your evaluation procedure (aka CV or LB probing, or  a combination) so that you asses your model performance correctly?</p>\n\n<p>When reading top teams writeups, make sure you understand how they validated model performance.  Often this is not described enough for my tatste, but sometimes people do it.  For instance I describe my CV scheme in Tweet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898219,
          "author_name": "jamshaidsohail5",
          "author_url": "",
          "post_date": "06/23/2020 11:31:24",
          "content": "<p>Thank you <a href=\"/cpmpml\">@cpmpml</a> . I have been following you everywhere. Keep inspiring. That was informative.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "897660": "Hi everyone. I just joined Kaggle and after working very hard \n[1]: I was on Bronze List in \"Tweet Sentiment Extraction\" and the shake up took it away.\n[2]: My one submission on Public LB in \"Jigsaw Multilingual Toxic comment classification\" was on Bronze in Private LB but I missed it while selecting the best ones.\n\nWhat could have been done to avoid the above ? Any guidance would be appreciated.",
    "897845": "Selecting the best submission it's sometimes a matter of luck, especially if they are close in score. However, one can typically limit the effect of shakeup by choosing 2 diverse submissions, possibly addressing different scenarios for private LB. People often choose best public LB + best local CV. But sometimes the public LB is so not reliable that you may just want to choose based on the CV. It depends. \nIn general, one should really not rely entirely on the public LB.",
    "897902": "Thank you. That is informative. Definitely going to consider this in future.",
    "897929": "What I did was this while selecting the submission. \n1. I had many submission in the range of 0.9477 to 0.9479. So, either I could choose one with a high public LB score or a little bit less Public LB score. So, I looked at count of positive and negative prediction. Based on previous experience with the submissions I could tell that there can not be more than 9K positive samples in whole test data. The reason is that only those submission score 0.947x whose positive sample count was between 9K to 10K. So, I choose the one with a lower count of positive sample and hence submitted 0.9477. This was my hunch and theory based on a number of samples in validation and train data. \n2. For the second submission, I did the same. I choose the submission with the lowest count of positive samples. There was one with public score 0.9475 and positive samples around 8K only. So I choose that one for the second final submission. \nFortunately, the one with 0.9477 scored more on private LB more compared to 0.9479 one and 0.9475 one. So, it is like understanding of total test set and overfitting on the public LB.",
    "898028": "That's the crux of Kaggle: being able to asses model performance as correctly as possible so that you select the right one.  I would recommend you try to understand what cold have made you select the right one here.  How can you modify your evaluation procedure (aka CV or LB probing, or  a combination) so that you asses your model performance correctly?\n\nWhen reading top teams writeups, make sure you understand how they validated model performance.  Often this is not described enough for my tatste, but sometimes people do it.  For instance I describe my CV scheme in Tweet.",
    "898219": "Thank you @cpmpml . I have been following you everywhere. Keep inspiring. That was informative."
  },
  "source": "meta"
}