{
  "id": 75281,
  "title": "Is the training set and test set distributed differently?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75281",
  "author_name": "",
  "post_date": "2018-12-20T07:04:13.314315100Z",
  "votes": null,
  "comment_count": 23,
  "views": 0,
  "content": "<p>my f1(0.10 valid) is 0.7062, but lb is 0.697\nwhy is there such a big difference between the two results.\ndistribution between validation data and test data is so different?</p>",
  "messages": [
    {
      "id": "442591",
      "postDate": "12/20/2018 07:04:13",
      "content": "<p>my f1(0.10 valid) is 0.7062, but lb is 0.697\nwhy is there such a big difference between the two results.\ndistribution between validation data and test data is so different?</p>",
      "rawMarkdown": "my f1(0.10 valid) is 0.7062, but lb is 0.697\nwhy is there such a big difference between the two results.\ndistribution between validation data and test data is so different?",
      "votes": null
    },
    {
      "id": "442612",
      "postDate": "12/20/2018 07:43:04",
      "content": "<p>We're all sharing that headache! :( My theory is that the current test set is just too tiny. Let's wait for the second test set to arrive!</p>",
      "rawMarkdown": "We're all sharing that headache! :( My theory is that the current test set is just too tiny. Let's wait for the second test set to arrive!",
      "votes": null
    },
    {
      "id": "442711",
      "postDate": "12/20/2018 11:18:19",
      "content": "<p>I think they are the same. just see the kernel <a href=\"https://www.kaggle.com/tunguz/quora-adversarial-validation\">https://www.kaggle.com/tunguz/quora-adversarial-validation</a> </p>",
      "rawMarkdown": "I think they are the same. just see the kernel https://www.kaggle.com/tunguz/quora-adversarial-validation",
      "votes": null
    },
    {
      "id": "442726",
      "postDate": "12/20/2018 11:47:33",
      "content": "<p>Are you sure there will be second test set in this competition? I didn't find any such information in its description nor timeline.</p>",
      "rawMarkdown": "Are you sure there will be second test set in this competition? I didn't find any such information in its description nor timeline.",
      "votes": null
    },
    {
      "id": "442727",
      "postDate": "12/20/2018 11:50:55",
      "content": "<p>I agree.\nAnd the difference between 0.7062 and 0.697 isn't that big though. How does the score differ between folds?</p>",
      "rawMarkdown": "I agree.\nAnd the difference between 0.7062 and 0.697 isn't that big though. How does the score differ between folds?",
      "votes": null
    },
    {
      "id": "442733",
      "postDate": "12/20/2018 11:59:59",
      "content": "<p>Check under the \"Data\" tab, then look for the <code>What will be available in the 2nd stage of the competition?</code> heading. Here's a quote:</p>\n\n<blockquote>\n  <p>In the second stage of the competition, we will re-run your selected Kernels. [...] This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.</p>\n</blockquote>\n\n<p>Hmm ok maybe I'm misinterpreting it, though. <code>The public leaderboard data remains the same for both versions.</code> probably means we will not get more information through the second test set.</p>",
      "rawMarkdown": "Check under the \"Data\" tab, then look for the `What will be available in the 2nd stage of the competition?` heading. Here's a quote:\n\n&gt; In the second stage of the competition, we will re-run your selected Kernels. [...] This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\n\nHmm ok maybe I'm misinterpreting it, though. `The public leaderboard data remains the same for both versions.` probably means we will not get more information through the second test set.",
      "votes": null
    },
    {
      "id": "442762",
      "postDate": "12/20/2018 13:10:27",
      "content": "<p>I think it is normal. I have also achieved such results before. There is no relationship between data distribution and model generalization capabilities. After all, this is an NLP task.</p>",
      "rawMarkdown": "I think it is normal. I have also achieved such results before. There is no relationship between data distribution and model generalization capabilities. After all, this is an NLP task.",
      "votes": null
    },
    {
      "id": "442799",
      "postDate": "12/20/2018 14:19:27",
      "content": "<p>You are right, it seems that 2nd stage is just the evaluation on the final day:</p>\n\n<blockquote>\n  <p>Stage 2 files will only be available in Kernels and not available for download.</p>\n</blockquote>\n\n<p>But then - as they will rerun my two best kernels - will they take the versions which lead to best results or the latest versions of them?</p>",
      "rawMarkdown": "You are right, it seems that 2nd stage is just the evaluation on the final day:\n\n&gt; Stage 2 files will only be available in Kernels and not available for download.\n\nBut then - as they will rerun my two best kernels - will they take the versions which lead to best results or the latest versions of them?",
      "votes": null
    },
    {
      "id": "442878",
      "postDate": "12/20/2018 16:22:08",
      "content": "<p>latest version it seems</p>",
      "rawMarkdown": "latest version it seems",
      "votes": null
    },
    {
      "id": "442902",
      "postDate": "12/20/2018 16:58:10",
      "content": "<p>no. you can choose two kernels for stage 2. and the kernels contains version code</p>",
      "rawMarkdown": "no. you can choose two kernels for stage 2. and the kernels contains version code",
      "votes": null
    },
    {
      "id": "443175",
      "postDate": "12/21/2018 06:39:07",
      "content": "<p>Isn't your score 0.7062 simply a (r.v.) sample?</p>\n\n<p>That being said, you can easily find the distribution of you CV scores and predict how likely it falls to 0.697. If your modeling and CV are done correctly, I bet you would find a pretty high probability to get a score like 0.697 -- even if your CV/validation, or simply a testing sample score was around 0.7062.</p>\n\n<p>(BTW, this is one of the typical/fundamental concepts that are used building up the general learning theory.)</p>",
      "rawMarkdown": "Isn't your score 0.7062 simply a (r.v.) sample?\n\nThat being said, you can easily find the distribution of you CV scores and predict how likely it falls to 0.697. If your modeling and CV are done correctly, I bet you would find a pretty high probability to get a score like 0.697 -- even if your CV/validation, or simply a testing sample score was around 0.7062.\n\n(BTW, this is one of the typical/fundamental concepts that are used building up the general learning theory.)",
      "votes": null
    },
    {
      "id": "443299",
      "postDate": "12/21/2018 11:41:55",
      "content": "<p>thanks, but sorry, I don't understand you...</p>",
      "rawMarkdown": "thanks, but sorry, I don't understand you...",
      "votes": null
    },
    {
      "id": "443301",
      "postDate": "12/21/2018 11:46:33",
      "content": "<p>why? is there any special about NLP task?\nI see your lb is 0.706, what is your val score, if it is convenient. </p>",
      "rawMarkdown": "why? is there any special about NLP task?\nI see your lb is 0.706, what is your val score, if it is convenient.",
      "votes": null
    },
    {
      "id": "443302",
      "postDate": "12/21/2018 11:47:09",
      "content": "<p>I agree with you</p>",
      "rawMarkdown": "I agree with you",
      "votes": null
    },
    {
      "id": "445254",
      "postDate": "12/26/2018 03:56:01",
      "content": "<p>how you achieve 0.7062 local cv , k-fold or train-val split?</p>",
      "rawMarkdown": "how you achieve 0.7062 local cv , k-fold or train-val split?",
      "votes": null
    },
    {
      "id": "445338",
      "postDate": "12/26/2018 08:38:46",
      "content": "<p>I don't use cv, k-fold.\nI randomly split training data(9:1), and get train, val data </p>",
      "rawMarkdown": "I don't use cv, k-fold.\nI randomly split training data(9:1), and get train, val data",
      "votes": null
    },
    {
      "id": "445399",
      "postDate": "12/26/2018 11:34:03",
      "content": "<p><a href=\"/xyzhang09\">@xyzhang09</a>\nour local score is 0.705, when run four times, we achieved 0.706, 0.705, 0.706, 0.702. I think this score has some mistakes.\nI think the generalization ability of NLP tasks is crucial.</p>",
      "rawMarkdown": "xyzhang09\nour local score is 0.705, when run four times, we achieved 0.706, 0.705, 0.706, 0.702. I think this score has some mistakes.\nI think the generalization ability of NLP tasks is crucial.",
      "votes": null
    },
    {
      "id": "445400",
      "postDate": "12/26/2018 11:38:06",
      "content": "<p><a href=\"/mschumacher\">@mschumacher</a>\nFrom my experience, the 50k test data is not small for NLP tasks. There may be a lot of important tricks we didn't find.</p>",
      "rawMarkdown": "mschumacher\nFrom my experience, the 50k test data is not small for NLP tasks. There may be a lot of important tricks we didn't find.",
      "votes": null
    },
    {
      "id": "445627",
      "postDate": "12/26/2018 20:38:07",
      "content": "<p>@iaobai1123q how many insincere questions did your 0.705 model predict?</p>",
      "rawMarkdown": "iaobai1123q how many insincere questions did your 0.705 model predict?",
      "votes": null
    },
    {
      "id": "445730",
      "postDate": "12/27/2018 01:57:59",
      "content": "<p><a href=\"/outflow\">@outflow</a>\n0.1</p>",
      "rawMarkdown": "outflow\n0.1",
      "votes": null
    },
    {
      "id": "445993",
      "postDate": "12/27/2018 10:24:39",
      "content": "<p>@iaobai1123q nice I think the test set have around 6370 insincere questions. Are you doing any Pseudo-Labelling?</p>",
      "rawMarkdown": "iaobai1123q nice I think the test set have around 6370 insincere questions. Are you doing any Pseudo-Labelling?",
      "votes": null
    },
    {
      "id": "446067",
      "postDate": "12/27/2018 12:58:41",
      "content": "<p>hi, maybe you are wrong\nI think test set have 3375 insincere questions</p>",
      "rawMarkdown": "hi, maybe you are wrong\nI think test set have 3375 insincere questions",
      "votes": null
    },
    {
      "id": "447761",
      "postDate": "12/30/2018 13:40:50",
      "content": "<p>How do you know the number of insincere questions?</p>",
      "rawMarkdown": "How do you know the number of insincere questions?",
      "votes": null
    },
    {
      "id": "448192",
      "postDate": "12/31/2018 12:10:54",
      "content": "<p>Hi <a href=\"/xyzhang09\">@xyzhang09</a> why do you think it have 3375?\n<a href=\"/david26694\">@david26694</a> submitting a csv with all predictions true got me 0.113 ish I am not F1 score expert I just assumed that the test set have 1.33% True positive which is  6370 </p>",
      "rawMarkdown": "Hi @xyzhang09 why do you think it have 3375?\n@david26694 submitting a csv with all predictions true got me 0.113 ish I am not F1 score expert I just assumed that the test set have 1.33% True positive which is  6370",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442612,
      "author_name": "mschumacher",
      "author_url": "",
      "post_date": "12/20/2018 07:43:04",
      "content": "<p>We're all sharing that headache! :( My theory is that the current test set is just too tiny. Let's wait for the second test set to arrive!</p>",
      "votes": null,
      "replies": [
        {
          "id": 442726,
          "author_name": "akuropatwinski",
          "author_url": "",
          "post_date": "12/20/2018 11:47:33",
          "content": "<p>Are you sure there will be second test set in this competition? I didn't find any such information in its description nor timeline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442733,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/20/2018 11:59:59",
          "content": "<p>Check under the \"Data\" tab, then look for the <code>What will be available in the 2nd stage of the competition?</code> heading. Here's a quote:</p>\n\n<blockquote>\n  <p>In the second stage of the competition, we will re-run your selected Kernels. [...] This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.</p>\n</blockquote>\n\n<p>Hmm ok maybe I'm misinterpreting it, though. <code>The public leaderboard data remains the same for both versions.</code> probably means we will not get more information through the second test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442799,
          "author_name": "akuropatwinski",
          "author_url": "",
          "post_date": "12/20/2018 14:19:27",
          "content": "<p>You are right, it seems that 2nd stage is just the evaluation on the final day:</p>\n\n<blockquote>\n  <p>Stage 2 files will only be available in Kernels and not available for download.</p>\n</blockquote>\n\n<p>But then - as they will rerun my two best kernels - will they take the versions which lead to best results or the latest versions of them?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442878,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "12/20/2018 16:22:08",
          "content": "<p>latest version it seems</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442902,
          "author_name": "syhens",
          "author_url": "",
          "post_date": "12/20/2018 16:58:10",
          "content": "<p>no. you can choose two kernels for stage 2. and the kernels contains version code</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443302,
          "author_name": "xyzhang09",
          "author_url": "",
          "post_date": "12/21/2018 11:47:09",
          "content": "<p>I agree with you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445400,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/26/2018 11:38:06",
          "content": "<p><a href=\"/mschumacher\">@mschumacher</a>\nFrom my experience, the 50k test data is not small for NLP tasks. There may be a lot of important tricks we didn't find.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442711,
      "author_name": "suchith0312",
      "author_url": "",
      "post_date": "12/20/2018 11:18:19",
      "content": "<p>I think they are the same. just see the kernel <a href=\"https://www.kaggle.com/tunguz/quora-adversarial-validation\">https://www.kaggle.com/tunguz/quora-adversarial-validation</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 442727,
          "author_name": "akuropatwinski",
          "author_url": "",
          "post_date": "12/20/2018 11:50:55",
          "content": "<p>I agree.\nAnd the difference between 0.7062 and 0.697 isn't that big though. How does the score differ between folds?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442762,
      "author_name": "xiaobai1123q",
      "author_url": "",
      "post_date": "12/20/2018 13:10:27",
      "content": "<p>I think it is normal. I have also achieved such results before. There is no relationship between data distribution and model generalization capabilities. After all, this is an NLP task.</p>",
      "votes": null,
      "replies": [
        {
          "id": 443301,
          "author_name": "xyzhang09",
          "author_url": "",
          "post_date": "12/21/2018 11:46:33",
          "content": "<p>why? is there any special about NLP task?\nI see your lb is 0.706, what is your val score, if it is convenient. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445399,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/26/2018 11:34:03",
          "content": "<p><a href=\"/xyzhang09\">@xyzhang09</a>\nour local score is 0.705, when run four times, we achieved 0.706, 0.705, 0.706, 0.702. I think this score has some mistakes.\nI think the generalization ability of NLP tasks is crucial.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445627,
          "author_name": "outflow",
          "author_url": "",
          "post_date": "12/26/2018 20:38:07",
          "content": "<p>@iaobai1123q how many insincere questions did your 0.705 model predict?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445730,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/27/2018 01:57:59",
          "content": "<p><a href=\"/outflow\">@outflow</a>\n0.1</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445993,
          "author_name": "outflow",
          "author_url": "",
          "post_date": "12/27/2018 10:24:39",
          "content": "<p>@iaobai1123q nice I think the test set have around 6370 insincere questions. Are you doing any Pseudo-Labelling?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446067,
          "author_name": "xyzhang09",
          "author_url": "",
          "post_date": "12/27/2018 12:58:41",
          "content": "<p>hi, maybe you are wrong\nI think test set have 3375 insincere questions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447761,
          "author_name": "david26694",
          "author_url": "",
          "post_date": "12/30/2018 13:40:50",
          "content": "<p>How do you know the number of insincere questions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448192,
          "author_name": "outflow",
          "author_url": "",
          "post_date": "12/31/2018 12:10:54",
          "content": "<p>Hi <a href=\"/xyzhang09\">@xyzhang09</a> why do you think it have 3375?\n<a href=\"/david26694\">@david26694</a> submitting a csv with all predictions true got me 0.113 ish I am not F1 score expert I just assumed that the test set have 1.33% True positive which is  6370 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443175,
      "author_name": "feynmann",
      "author_url": "",
      "post_date": "12/21/2018 06:39:07",
      "content": "<p>Isn't your score 0.7062 simply a (r.v.) sample?</p>\n\n<p>That being said, you can easily find the distribution of you CV scores and predict how likely it falls to 0.697. If your modeling and CV are done correctly, I bet you would find a pretty high probability to get a score like 0.697 -- even if your CV/validation, or simply a testing sample score was around 0.7062.</p>\n\n<p>(BTW, this is one of the typical/fundamental concepts that are used building up the general learning theory.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 443299,
          "author_name": "xyzhang09",
          "author_url": "",
          "post_date": "12/21/2018 11:41:55",
          "content": "<p>thanks, but sorry, I don't understand you...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 445254,
      "author_name": "luckyboyde",
      "author_url": "",
      "post_date": "12/26/2018 03:56:01",
      "content": "<p>how you achieve 0.7062 local cv , k-fold or train-val split?</p>",
      "votes": null,
      "replies": [
        {
          "id": 445338,
          "author_name": "xyzhang09",
          "author_url": "",
          "post_date": "12/26/2018 08:38:46",
          "content": "<p>I don't use cv, k-fold.\nI randomly split training data(9:1), and get train, val data </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "442591": "my f1(0.10 valid) is 0.7062, but lb is 0.697\nwhy is there such a big difference between the two results.\ndistribution between validation data and test data is so different?",
    "442612": "We're all sharing that headache! :( My theory is that the current test set is just too tiny. Let's wait for the second test set to arrive!",
    "442711": "I think they are the same. just see the kernel https://www.kaggle.com/tunguz/quora-adversarial-validation",
    "442726": "Are you sure there will be second test set in this competition? I didn't find any such information in its description nor timeline.",
    "442727": "I agree.\nAnd the difference between 0.7062 and 0.697 isn't that big though. How does the score differ between folds?",
    "442733": "Check under the \"Data\" tab, then look for the `What will be available in the 2nd stage of the competition?` heading. Here's a quote:\n\n&gt; In the second stage of the competition, we will re-run your selected Kernels. [...] This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\n\nHmm ok maybe I'm misinterpreting it, though. `The public leaderboard data remains the same for both versions.` probably means we will not get more information through the second test set.",
    "442762": "I think it is normal. I have also achieved such results before. There is no relationship between data distribution and model generalization capabilities. After all, this is an NLP task.",
    "442799": "You are right, it seems that 2nd stage is just the evaluation on the final day:\n\n&gt; Stage 2 files will only be available in Kernels and not available for download.\n\nBut then - as they will rerun my two best kernels - will they take the versions which lead to best results or the latest versions of them?",
    "442878": "latest version it seems",
    "442902": "no. you can choose two kernels for stage 2. and the kernels contains version code",
    "443175": "Isn't your score 0.7062 simply a (r.v.) sample?\n\nThat being said, you can easily find the distribution of you CV scores and predict how likely it falls to 0.697. If your modeling and CV are done correctly, I bet you would find a pretty high probability to get a score like 0.697 -- even if your CV/validation, or simply a testing sample score was around 0.7062.\n\n(BTW, this is one of the typical/fundamental concepts that are used building up the general learning theory.)",
    "443299": "thanks, but sorry, I don't understand you...",
    "443301": "why? is there any special about NLP task?\nI see your lb is 0.706, what is your val score, if it is convenient.",
    "443302": "I agree with you",
    "445254": "how you achieve 0.7062 local cv , k-fold or train-val split?",
    "445338": "I don't use cv, k-fold.\nI randomly split training data(9:1), and get train, val data",
    "445399": "xyzhang09\nour local score is 0.705, when run four times, we achieved 0.706, 0.705, 0.706, 0.702. I think this score has some mistakes.\nI think the generalization ability of NLP tasks is crucial.",
    "445400": "mschumacher\nFrom my experience, the 50k test data is not small for NLP tasks. There may be a lot of important tricks we didn't find.",
    "445627": "iaobai1123q how many insincere questions did your 0.705 model predict?",
    "445730": "outflow\n0.1",
    "445993": "iaobai1123q nice I think the test set have around 6370 insincere questions. Are you doing any Pseudo-Labelling?",
    "446067": "hi, maybe you are wrong\nI think test set have 3375 insincere questions",
    "447761": "How do you know the number of insincere questions?",
    "448192": "Hi @xyzhang09 why do you think it have 3375?\n@david26694 submitting a csv with all predictions true got me 0.113 ish I am not F1 score expert I just assumed that the test set have 1.33% True positive which is  6370"
  },
  "source": "meta"
}