{
  "id": 75381,
  "title": "What do you think about the public test set?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75381",
  "author_name": "",
  "post_date": "2018-12-21T01:21:42.301233400Z",
  "votes": 15,
  "comment_count": 19,
  "views": 0,
  "content": "<p>For example, I took this kernel which scored 0.674 locally and 0.690-0.692 LB, <a href=\"https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr\">https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr</a>, made some changes so it got to 0.681 local. And then whooosh, 0.681 LB. Then I made some other changes, local stayed the same, but this time 0.690 LB.</p>\n\n<p>This has happened way too often, in fact my best scored attempt has the most stupid ideas I could think of. Local and LB scores just seem to have no correlation at all.</p>\n\n<p>I know the set is very small but if it was fairly drawn from the original data I don't think the results can be that unpredictable. I saw some people having the same problems, I wonder how many of us there are and what the top top people on the LB think about this.</p>",
  "messages": [
    {
      "id": "443078",
      "postDate": "12/21/2018 01:21:42",
      "content": "<p>For example, I took this kernel which scored 0.674 locally and 0.690-0.692 LB, <a href=\"https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr\">https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr</a>, made some changes so it got to 0.681 local. And then whooosh, 0.681 LB. Then I made some other changes, local stayed the same, but this time 0.690 LB.</p>\n\n<p>This has happened way too often, in fact my best scored attempt has the most stupid ideas I could think of. Local and LB scores just seem to have no correlation at all.</p>\n\n<p>I know the set is very small but if it was fairly drawn from the original data I don't think the results can be that unpredictable. I saw some people having the same problems, I wonder how many of us there are and what the top top people on the LB think about this.</p>",
      "rawMarkdown": "For example, I took this kernel which scored 0.674 locally and 0.690-0.692 LB, https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr, made some changes so it got to 0.681 local. And then whooosh, 0.681 LB. Then I made some other changes, local stayed the same, but this time 0.690 LB.\n\nThis has happened way too often, in fact my best scored attempt has the most stupid ideas I could think of. Local and LB scores just seem to have no correlation at all.\n\nI know the set is very small but if it was fairly drawn from the original data I don't think the results can be that unpredictable. I saw some people having the same problems, I wonder how many of us there are and what the top top people on the LB think about this.",
      "votes": null
    },
    {
      "id": "443096",
      "postDate": "12/21/2018 02:29:13",
      "content": "<p>Maybe the distribution of the two sets has a little bit difference. \nBTW, Who knows the test set size in Stage 2 ?</p>",
      "rawMarkdown": "Maybe the distribution of the two sets has a little bit difference. \nBTW, Who knows the test set size in Stage 2 ?",
      "votes": null
    },
    {
      "id": "443101",
      "postDate": "12/21/2018 02:37:06",
      "content": "<p>That is another thing I'm worrying about, if the stage 2 test size is on par with the training set many public kernels will break.</p>",
      "rawMarkdown": "That is another thing I'm worrying about, if the stage 2 test size is on par with the training set many public kernels will break.",
      "votes": null
    },
    {
      "id": "443107",
      "postDate": "12/21/2018 02:56:16",
      "content": "<p>I don't think that will happen, but push your model run though within 7000s will be indispensable.</p>",
      "rawMarkdown": "I don't think that will happen, but push your model run though within 7000s will be indispensable.",
      "votes": null
    },
    {
      "id": "443188",
      "postDate": "12/21/2018 07:24:49",
      "content": "<p>train: 1306122\n4-fold val: 1306122/4=326530\ntest: 56370</p>\n\n<p>What do we trust? local cv or LB？</p>",
      "rawMarkdown": "train: 1306122\n4-fold val: 1306122/4=326530\ntest: 56370\n\nWhat do we trust? local cv or LB？",
      "votes": null
    },
    {
      "id": "443190",
      "postDate": "12/21/2018 07:30:47",
      "content": "<p>Here are some of my thoughts：\n1、The data of Public LB is part of the Pricate LB data, so it has a certain degree of credibility.\n2、If both val and LB improve, I will trust this change very much.\n3、If val is lowered, LB is improved, and I don't really think that this is an increase in generalization ability.\n4、This is an NLP task, so thinking about model generalization is more worthwhile than thinking about data distribution.</p>\n\n<p>Thank you for your contribution to this competition.</p>",
      "rawMarkdown": "Here are some of my thoughts：\n1、The data of Public LB is part of the Pricate LB data, so it has a certain degree of credibility.\n2、If both val and LB improve, I will trust this change very much.\n3、If val is lowered, LB is improved, and I don't really think that this is an increase in generalization ability.\n4、This is an NLP task, so thinking about model generalization is more worthwhile than thinking about data distribution.\n\nThank you for your contribution to this competition.",
      "votes": null
    },
    {
      "id": "443192",
      "postDate": "12/21/2018 07:33:44",
      "content": "<p>So we should check the metric about (CV+LB)/2 ?</p>",
      "rawMarkdown": "So we should check the metric about (CV+LB)/2 ?",
      "votes": null
    },
    {
      "id": "443193",
      "postDate": "12/21/2018 07:38:56",
      "content": "<p>Private LB test size will be 7 times as large, it is stated in the competition description. Your kernel needs to also account for the additional runtime.</p>",
      "rawMarkdown": "Private LB test size will be 7 times as large, it is stated in the competition description. Your kernel needs to also account for the additional runtime.",
      "votes": null
    },
    {
      "id": "443201",
      "postDate": "12/21/2018 07:48:44",
      "content": "<p>I don't think it is meaningless to do this. since a higher score and a lower score will give a meddle score.</p>",
      "rawMarkdown": "I don't think it is meaningless to do this. since a higher score and a lower score will give a meddle score.",
      "votes": null
    },
    {
      "id": "443244",
      "postDate": "12/21/2018 09:46:33",
      "content": "<p>Ahhh, I really want to public another kernel but without this correlation it's hard to be sure that my ideas worked.</p>",
      "rawMarkdown": "Ahhh, I really want to public another kernel but without this correlation it's hard to be sure that my ideas worked.",
      "votes": null
    },
    {
      "id": "443250",
      "postDate": "12/21/2018 09:53:11",
      "content": "<p>Listen to what you mean, the local is very high, but the LB is very low.</p>",
      "rawMarkdown": "Listen to what you mean, the local is very high, but the LB is very low.",
      "votes": null
    },
    {
      "id": "443269",
      "postDate": "12/21/2018 10:20:14",
      "content": "<blockquote>\n  <p>sample_submission.csv - similar to test.csv, this will be changed from\n  ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain\n  the same.</p>\n</blockquote>\n\n<p>So it'd be ~376k samples in Stage 2 test set.</p>",
      "rawMarkdown": "&gt; sample_submission.csv - similar to test.csv, this will be changed from\n&gt; ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain\n&gt; the same.\n\nSo it'd be ~376k samples in Stage 2 test set.",
      "votes": null
    },
    {
      "id": "443271",
      "postDate": "12/21/2018 10:25:02",
      "content": "<blockquote>\n  <p>1、The data of Public LB is part of the Pricate LB data, so it has a\n  certain degree of credibility.</p>\n</blockquote>\n\n<p>Hi <a href=\"/xiaobai1123q\">@xiaobai1123q</a>,\nCould you point where it's stated, please? I just thought that Private test set would be brand new and wouldn't contain Public test set.</p>",
      "rawMarkdown": "&gt; 1、The data of Public LB is part of the Pricate LB data, so it has a\n&gt; certain degree of credibility.\n\nHi @xiaobai1123q,\nCould you point where it's stated, please? I just thought that Private test set would be brand new and wouldn't contain Public test set.",
      "votes": null
    },
    {
      "id": "443274",
      "postDate": "12/21/2018 10:30:12",
      "content": "<p>I think <a href=\"/xiaobai1123q\">@xiaobai1123q</a> misunderstood something, the private test sets are always different from the public unless explicitly mentioned otherwise.</p>",
      "rawMarkdown": "I think @xiaobai1123q misunderstood something, the private test sets are always different from the public unless explicitly mentioned otherwise.",
      "votes": null
    },
    {
      "id": "443298",
      "postDate": "12/21/2018 11:39:06",
      "content": "<p>I have the same doubt. It is mentioned in the data page.\nCan you refer to that here : <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75308#443297\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75308#443297</a></p>",
      "rawMarkdown": "I have the same doubt. It is mentioned in the data page.\nCan you refer to that here : https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75308#443297",
      "votes": null
    },
    {
      "id": "443536",
      "postDate": "12/21/2018 20:28:46",
      "content": "<p>I think one of the reasons why the noise is so strong in this competition is because of the binary predictions we are supposed to submit. That way, the errors are very fickle. If we could submit probabilities and get a smooth loss as in most other competitions, the errors and noise would not be so dramatic.</p>",
      "rawMarkdown": "I think one of the reasons why the noise is so strong in this competition is because of the binary predictions we are supposed to submit. That way, the errors are very fickle. If we could submit probabilities and get a smooth loss as in most other competitions, the errors and noise would not be so dramatic.",
      "votes": null
    },
    {
      "id": "443679",
      "postDate": "12/22/2018 05:56:41",
      "content": "<hr>\n\n<p>The private leaderboard is calculated over the same rows as the public leaderboard in this competition.</p>\n\n<hr>\n\n<p><a href=\"/thinline72\">@thinline72</a>\n<a href=\"/suicaokhoailang\">@suicaokhoailang</a>\nI feel this is because my English is not good. What does the meaning of same rows mean here?</p>",
      "rawMarkdown": "The private leaderboard is calculated over the same rows as the public leaderboard in this competition.\n\n\n----------\n\n\n@thinline72\n@suicaokhoailang\nI feel this is because my English is not good. What does the meaning of same rows mean here?",
      "votes": null
    },
    {
      "id": "443717",
      "postDate": "12/22/2018 08:13:09",
      "content": "<p>Ah, I think the text here is misleading, usually it is: \"The private leaderboard is calculated over x% of total test data and the public leaderboard on the (100 - x)%\".  (You see all the test data, but don't know what samples will be calculated for which leaderboard)</p>\n\n<p>In some competitions, usually the playgrounds where leaderboard don't really matter, they won't do that split and just calculate everything on all test data and thus you'll see that description, those competitions <strong>do not</strong> have a private leaderboard. This one for example <a href=\"https://www.kaggle.com/c/whale-categorization-playground/leaderboard\">https://www.kaggle.com/c/whale-categorization-playground/leaderboard</a> .</p>\n\n<p>For this competition it's a bit special since your model won't see all the test data in stage 1, so the x% description won't apply and I guess they defaulted to the second one as we saw.  </p>",
      "rawMarkdown": "Ah, I think the text here is misleading, usually it is: \"The private leaderboard is calculated over x% of total test data and the public leaderboard on the (100 - x)%\".  (You see all the test data, but don't know what samples will be calculated for which leaderboard)\n\nIn some competitions, usually the playgrounds where leaderboard don't really matter, they won't do that split and just calculate everything on all test data and thus you'll see that description, those competitions **do not** have a private leaderboard. This one for example https://www.kaggle.com/c/whale-categorization-playground/leaderboard .\n\nFor this competition it's a bit special since your model won't see all the test data in stage 1, so the x% description won't apply and I guess they defaulted to the second one as we saw.",
      "votes": null
    },
    {
      "id": "444049",
      "postDate": "12/23/2018 02:37:55",
      "content": "<p>What if local cv improved, but pb droped? </p>",
      "rawMarkdown": "What if local cv improved, but pb droped?",
      "votes": null
    },
    {
      "id": "444114",
      "postDate": "12/23/2018 08:17:35",
      "content": "<p><a href=\"/suicaokhoailang\">@suicaokhoailang</a>\nSo do you think the data in private lb is brand new?</p>",
      "rawMarkdown": "suicaokhoailang\nSo do you think the data in private lb is brand new?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 443096,
      "author_name": "mr007rin",
      "author_url": "",
      "post_date": "12/21/2018 02:29:13",
      "content": "<p>Maybe the distribution of the two sets has a little bit difference. \nBTW, Who knows the test set size in Stage 2 ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 443101,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "12/21/2018 02:37:06",
          "content": "<p>That is another thing I'm worrying about, if the stage 2 test size is on par with the training set many public kernels will break.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443107,
          "author_name": "mr007rin",
          "author_url": "",
          "post_date": "12/21/2018 02:56:16",
          "content": "<p>I don't think that will happen, but push your model run though within 7000s will be indispensable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443193,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/21/2018 07:38:56",
          "content": "<p>Private LB test size will be 7 times as large, it is stated in the competition description. Your kernel needs to also account for the additional runtime.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443269,
          "author_name": "thinline72",
          "author_url": "",
          "post_date": "12/21/2018 10:20:14",
          "content": "<blockquote>\n  <p>sample_submission.csv - similar to test.csv, this will be changed from\n  ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain\n  the same.</p>\n</blockquote>\n\n<p>So it'd be ~376k samples in Stage 2 test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443188,
      "author_name": "hengzheng",
      "author_url": "",
      "post_date": "12/21/2018 07:24:49",
      "content": "<p>train: 1306122\n4-fold val: 1306122/4=326530\ntest: 56370</p>\n\n<p>What do we trust? local cv or LB？</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443190,
      "author_name": "xiaobai1123q",
      "author_url": "",
      "post_date": "12/21/2018 07:30:47",
      "content": "<p>Here are some of my thoughts：\n1、The data of Public LB is part of the Pricate LB data, so it has a certain degree of credibility.\n2、If both val and LB improve, I will trust this change very much.\n3、If val is lowered, LB is improved, and I don't really think that this is an increase in generalization ability.\n4、This is an NLP task, so thinking about model generalization is more worthwhile than thinking about data distribution.</p>\n\n<p>Thank you for your contribution to this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 443192,
          "author_name": "hengzheng",
          "author_url": "",
          "post_date": "12/21/2018 07:33:44",
          "content": "<p>So we should check the metric about (CV+LB)/2 ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443201,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/21/2018 07:48:44",
          "content": "<p>I don't think it is meaningless to do this. since a higher score and a lower score will give a meddle score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443244,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "12/21/2018 09:46:33",
          "content": "<p>Ahhh, I really want to public another kernel but without this correlation it's hard to be sure that my ideas worked.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443250,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/21/2018 09:53:11",
          "content": "<p>Listen to what you mean, the local is very high, but the LB is very low.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443271,
          "author_name": "thinline72",
          "author_url": "",
          "post_date": "12/21/2018 10:25:02",
          "content": "<blockquote>\n  <p>1、The data of Public LB is part of the Pricate LB data, so it has a\n  certain degree of credibility.</p>\n</blockquote>\n\n<p>Hi <a href=\"/xiaobai1123q\">@xiaobai1123q</a>,\nCould you point where it's stated, please? I just thought that Private test set would be brand new and wouldn't contain Public test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443274,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "12/21/2018 10:30:12",
          "content": "<p>I think <a href=\"/xiaobai1123q\">@xiaobai1123q</a> misunderstood something, the private test sets are always different from the public unless explicitly mentioned otherwise.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443298,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "12/21/2018 11:39:06",
          "content": "<p>I have the same doubt. It is mentioned in the data page.\nCan you refer to that here : <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75308#443297\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75308#443297</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443679,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/22/2018 05:56:41",
          "content": "<hr>\n\n<p>The private leaderboard is calculated over the same rows as the public leaderboard in this competition.</p>\n\n<hr>\n\n<p><a href=\"/thinline72\">@thinline72</a>\n<a href=\"/suicaokhoailang\">@suicaokhoailang</a>\nI feel this is because my English is not good. What does the meaning of same rows mean here?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443717,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "12/22/2018 08:13:09",
          "content": "<p>Ah, I think the text here is misleading, usually it is: \"The private leaderboard is calculated over x% of total test data and the public leaderboard on the (100 - x)%\".  (You see all the test data, but don't know what samples will be calculated for which leaderboard)</p>\n\n<p>In some competitions, usually the playgrounds where leaderboard don't really matter, they won't do that split and just calculate everything on all test data and thus you'll see that description, those competitions <strong>do not</strong> have a private leaderboard. This one for example <a href=\"https://www.kaggle.com/c/whale-categorization-playground/leaderboard\">https://www.kaggle.com/c/whale-categorization-playground/leaderboard</a> .</p>\n\n<p>For this competition it's a bit special since your model won't see all the test data in stage 1, so the x% description won't apply and I guess they defaulted to the second one as we saw.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444049,
          "author_name": "laevatein",
          "author_url": "",
          "post_date": "12/23/2018 02:37:55",
          "content": "<p>What if local cv improved, but pb droped? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444114,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/23/2018 08:17:35",
          "content": "<p><a href=\"/suicaokhoailang\">@suicaokhoailang</a>\nSo do you think the data in private lb is brand new?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443536,
      "author_name": "mschumacher",
      "author_url": "",
      "post_date": "12/21/2018 20:28:46",
      "content": "<p>I think one of the reasons why the noise is so strong in this competition is because of the binary predictions we are supposed to submit. That way, the errors are very fickle. If we could submit probabilities and get a smooth loss as in most other competitions, the errors and noise would not be so dramatic.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "443078": "For example, I took this kernel which scored 0.674 locally and 0.690-0.692 LB, https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr, made some changes so it got to 0.681 local. And then whooosh, 0.681 LB. Then I made some other changes, local stayed the same, but this time 0.690 LB.\n\nThis has happened way too often, in fact my best scored attempt has the most stupid ideas I could think of. Local and LB scores just seem to have no correlation at all.\n\nI know the set is very small but if it was fairly drawn from the original data I don't think the results can be that unpredictable. I saw some people having the same problems, I wonder how many of us there are and what the top top people on the LB think about this.",
    "443096": "Maybe the distribution of the two sets has a little bit difference. \nBTW, Who knows the test set size in Stage 2 ?",
    "443101": "That is another thing I'm worrying about, if the stage 2 test size is on par with the training set many public kernels will break.",
    "443107": "I don't think that will happen, but push your model run though within 7000s will be indispensable.",
    "443188": "train: 1306122\n4-fold val: 1306122/4=326530\ntest: 56370\n\nWhat do we trust? local cv or LB？",
    "443190": "Here are some of my thoughts：\n1、The data of Public LB is part of the Pricate LB data, so it has a certain degree of credibility.\n2、If both val and LB improve, I will trust this change very much.\n3、If val is lowered, LB is improved, and I don't really think that this is an increase in generalization ability.\n4、This is an NLP task, so thinking about model generalization is more worthwhile than thinking about data distribution.\n\nThank you for your contribution to this competition.",
    "443192": "So we should check the metric about (CV+LB)/2 ?",
    "443193": "Private LB test size will be 7 times as large, it is stated in the competition description. Your kernel needs to also account for the additional runtime.",
    "443201": "I don't think it is meaningless to do this. since a higher score and a lower score will give a meddle score.",
    "443244": "Ahhh, I really want to public another kernel but without this correlation it's hard to be sure that my ideas worked.",
    "443250": "Listen to what you mean, the local is very high, but the LB is very low.",
    "443269": "&gt; sample_submission.csv - similar to test.csv, this will be changed from\n&gt; ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain\n&gt; the same.\n\nSo it'd be ~376k samples in Stage 2 test set.",
    "443271": "&gt; 1、The data of Public LB is part of the Pricate LB data, so it has a\n&gt; certain degree of credibility.\n\nHi @xiaobai1123q,\nCould you point where it's stated, please? I just thought that Private test set would be brand new and wouldn't contain Public test set.",
    "443274": "I think @xiaobai1123q misunderstood something, the private test sets are always different from the public unless explicitly mentioned otherwise.",
    "443298": "I have the same doubt. It is mentioned in the data page.\nCan you refer to that here : https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75308#443297",
    "443536": "I think one of the reasons why the noise is so strong in this competition is because of the binary predictions we are supposed to submit. That way, the errors are very fickle. If we could submit probabilities and get a smooth loss as in most other competitions, the errors and noise would not be so dramatic.",
    "443679": "The private leaderboard is calculated over the same rows as the public leaderboard in this competition.\n\n\n----------\n\n\n@thinline72\n@suicaokhoailang\nI feel this is because my English is not good. What does the meaning of same rows mean here?",
    "443717": "Ah, I think the text here is misleading, usually it is: \"The private leaderboard is calculated over x% of total test data and the public leaderboard on the (100 - x)%\".  (You see all the test data, but don't know what samples will be calculated for which leaderboard)\n\nIn some competitions, usually the playgrounds where leaderboard don't really matter, they won't do that split and just calculate everything on all test data and thus you'll see that description, those competitions **do not** have a private leaderboard. This one for example https://www.kaggle.com/c/whale-categorization-playground/leaderboard .\n\nFor this competition it's a bit special since your model won't see all the test data in stage 1, so the x% description won't apply and I guess they defaulted to the second one as we saw.",
    "444049": "What if local cv improved, but pb droped?",
    "444114": "suicaokhoailang\nSo do you think the data in private lb is brand new?"
  },
  "source": "meta"
}