{
  "id": 194022,
  "title": "Questions about difference score between Local & Public",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194022",
  "author_name": "",
  "post_date": "2020-10-30T08:54:43.719715100Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>For my experiment,  I have settled with 13M dataset with about 10M train_set and 2M test_set.</p>\n<p>I have got a matrix AUC score by test_set <strong>0.7683</strong> . <br>\nHowever, my submissions score (public score) is just <strong>0.690</strong>. Who have met such situation before? </p>\n<p><strong>I guess some of you might have had such puzzle.</strong></p>\n<p>Solved or not, could you please share your analysis or experience in the comment</p>",
  "messages": [
    {
      "id": "1064563",
      "postDate": "10/30/2020 08:54:43",
      "content": "<p>For my experiment,  I have settled with 13M dataset with about 10M train_set and 2M test_set.</p>\n<p>I have got a matrix AUC score by test_set <strong>0.7683</strong> . <br>\nHowever, my submissions score (public score) is just <strong>0.690</strong>. Who have met such situation before? </p>\n<p><strong>I guess some of you might have had such puzzle.</strong></p>\n<p>Solved or not, could you please share your analysis or experience in the comment</p>",
      "rawMarkdown": "For my experiment,  I have settled with 13M dataset with about 10M train_set and 2M test_set.\n\nI have got a matrix AUC score by test_set **0.7683** . \nHowever, my submissions score (public score) is just **0.690**. Who have met such situation before? \n\n**I guess some of you might have had such puzzle.**\n\nSolved or not, could you please share your analysis or experience in the comment",
      "votes": null
    },
    {
      "id": "1064585",
      "postDate": "10/30/2020 09:36:30",
      "content": "<p>Almost everyone who shared their CV-LB scores <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919\" target=\"_blank\">here</a> have reported very stable correlation and a small absolute difference. In some cases LB is a bit higher than CV.</p>\n<p>Local <strong>0.768</strong> vs LB <strong>0.690</strong> implies something is definitely off. I would agree with <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>'s <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919#1063896\" target=\"_blank\">comment</a> that it is most likely due to some difference in the feature preparation for local test vs private test. Maybe use the public test iterator to check / verify the feature values but beyond that it is up to you to investigate your pipeline in detail.</p>",
      "rawMarkdown": "Almost everyone who shared their CV-LB scores [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919) have reported very stable correlation and a small absolute difference. In some cases LB is a bit higher than CV.\n\nLocal **0.768** vs LB **0.690** implies something is definitely off. I would agree with @aquatic's [comment](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919#1063896) that it is most likely due to some difference in the feature preparation for local test vs private test. Maybe use the public test iterator to check / verify the feature values but beyond that it is up to you to investigate your pipeline in detail.",
      "votes": null
    },
    {
      "id": "1069066",
      "postDate": "11/04/2020 03:07:27",
      "content": "<p>Pretty good, thanks. I have just found some outliers null value in test cases which might cause a low performance in public testing. </p>",
      "rawMarkdown": "Pretty good, thanks. I have just found some outliers null value in test cases which might cause a low performance in public testing.",
      "votes": null
    },
    {
      "id": "1069070",
      "postDate": "11/04/2020 03:16:59",
      "content": "<p>Btw, sir, I want to ask you an additional question that could I divide one dataset into several pieces and predict them in different model then combining the result, that is, in some columns there are nulls in part of features, and I could not find an apropos substitution. Thus, I want to ignore such features while predicting the relevant columns.  Is this idea legal？</p>",
      "rawMarkdown": "Btw, sir, I want to ask you an additional question that could I divide one dataset into several pieces and predict them in different model then combining the result, that is, in some columns there are nulls in part of features, and I could not find an apropos substitution. Thus, I want to ignore such features while predicting the relevant columns.  Is this idea legal？",
      "votes": null
    },
    {
      "id": "1069080",
      "postDate": "11/04/2020 03:33:12",
      "content": "<p>You can use any approach for predicting as per the test API workflow. It's usually a good idea to test any approach on a validation dataset from train data before going to the test data and let the validation decide whether a particular idea is good or not.</p>",
      "rawMarkdown": "You can use any approach for predicting as per the test API workflow. It's usually a good idea to test any approach on a validation dataset from train data before going to the test data and let the validation decide whether a particular idea is good or not.",
      "votes": null
    },
    {
      "id": "1069161",
      "postDate": "11/04/2020 06:52:44",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1064585,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/30/2020 09:36:30",
      "content": "<p>Almost everyone who shared their CV-LB scores <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919\" target=\"_blank\">here</a> have reported very stable correlation and a small absolute difference. In some cases LB is a bit higher than CV.</p>\n<p>Local <strong>0.768</strong> vs LB <strong>0.690</strong> implies something is definitely off. I would agree with <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>'s <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919#1063896\" target=\"_blank\">comment</a> that it is most likely due to some difference in the feature preparation for local test vs private test. Maybe use the public test iterator to check / verify the feature values but beyond that it is up to you to investigate your pipeline in detail.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1069066,
          "author_name": "puyuzhou",
          "author_url": "",
          "post_date": "11/04/2020 03:07:27",
          "content": "<p>Pretty good, thanks. I have just found some outliers null value in test cases which might cause a low performance in public testing. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069070,
          "author_name": "puyuzhou",
          "author_url": "",
          "post_date": "11/04/2020 03:16:59",
          "content": "<p>Btw, sir, I want to ask you an additional question that could I divide one dataset into several pieces and predict them in different model then combining the result, that is, in some columns there are nulls in part of features, and I could not find an apropos substitution. Thus, I want to ignore such features while predicting the relevant columns.  Is this idea legal？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069080,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "11/04/2020 03:33:12",
          "content": "<p>You can use any approach for predicting as per the test API workflow. It's usually a good idea to test any approach on a validation dataset from train data before going to the test data and let the validation decide whether a particular idea is good or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069161,
          "author_name": "puyuzhou",
          "author_url": "",
          "post_date": "11/04/2020 06:52:44",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1064563": "For my experiment,  I have settled with 13M dataset with about 10M train_set and 2M test_set.\n\nI have got a matrix AUC score by test_set **0.7683** . \nHowever, my submissions score (public score) is just **0.690**. Who have met such situation before? \n\n**I guess some of you might have had such puzzle.**\n\nSolved or not, could you please share your analysis or experience in the comment",
    "1064585": "Almost everyone who shared their CV-LB scores [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919) have reported very stable correlation and a small absolute difference. In some cases LB is a bit higher than CV.\n\nLocal **0.768** vs LB **0.690** implies something is definitely off. I would agree with @aquatic's [comment](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919#1063896) that it is most likely due to some difference in the feature preparation for local test vs private test. Maybe use the public test iterator to check / verify the feature values but beyond that it is up to you to investigate your pipeline in detail.",
    "1069066": "Pretty good, thanks. I have just found some outliers null value in test cases which might cause a low performance in public testing.",
    "1069070": "Btw, sir, I want to ask you an additional question that could I divide one dataset into several pieces and predict them in different model then combining the result, that is, in some columns there are nulls in part of features, and I could not find an apropos substitution. Thus, I want to ignore such features while predicting the relevant columns.  Is this idea legal？",
    "1069080": "You can use any approach for predicting as per the test API workflow. It's usually a good idea to test any approach on a validation dataset from train data before going to the test data and let the validation decide whether a particular idea is good or not.",
    "1069161": "Thank you!"
  },
  "source": "meta"
}