{
  "id": 552481,
  "title": "Are the data distribution different between public test data and private test data?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552481",
  "author_name": "",
  "post_date": "2024-12-20T00:37:14.651438300Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Almost all the private scores of my submissions were lower than the public scores. Even the notebooks which I made without looking at the public scores and using only simple scikit-learn models (like SGDRegressor) caused the same problem. So I don't think the reason is just leakage. Maybe the private test data was harder than the public test data?</p>",
  "messages": [
    {
      "id": "3076401",
      "postDate": "12/20/2024 00:37:14",
      "content": "<p>Almost all the private scores of my submissions were lower than the public scores. Even the notebooks which I made without looking at the public scores and using only simple scikit-learn models (like SGDRegressor) caused the same problem. So I don't think the reason is just leakage. Maybe the private test data was harder than the public test data?</p>",
      "rawMarkdown": "Almost all the private scores of my submissions were lower than the public scores. Even the notebooks which I made without looking at the public scores and using only simple scikit-learn models (like SGDRegressor) caused the same problem. So I don't think the reason is just leakage. Maybe the private test data was harder than the public test data?",
      "votes": null
    },
    {
      "id": "3076465",
      "postDate": "12/20/2024 01:28:43",
      "content": "<p>Yes of course, the private test set was harder and contained most of the class 3 sii samples <a href=\"https://www.kaggle.com/ykawakita\" target=\"_blank\">@ykawakita</a> <br>\nThe CV-LB relations indicated this through the competition!</p>",
      "rawMarkdown": "Yes of course, the private test set was harder and contained most of the class 3 sii samples @ykawakita \nThe CV-LB relations indicated this through the competition!",
      "votes": null
    },
    {
      "id": "3076496",
      "postDate": "12/20/2024 02:14:44",
      "content": "<p>Wow, I didn't realize it at all. I assumed that the train data and the test data were completely randomly taken from the data which Child Mind Institute has. Thank you for the lesson!</p>",
      "rawMarkdown": "Wow, I didn't realize it at all. I assumed that the train data and the test data were completely randomly taken from the data which Child Mind Institute has. Thank you for the lesson!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076465,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/20/2024 01:28:43",
      "content": "<p>Yes of course, the private test set was harder and contained most of the class 3 sii samples <a href=\"https://www.kaggle.com/ykawakita\" target=\"_blank\">@ykawakita</a> <br>\nThe CV-LB relations indicated this through the competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3076496,
          "author_name": "ykawakita",
          "author_url": "",
          "post_date": "12/20/2024 02:14:44",
          "content": "<p>Wow, I didn't realize it at all. I assumed that the train data and the test data were completely randomly taken from the data which Child Mind Institute has. Thank you for the lesson!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3076401": "Almost all the private scores of my submissions were lower than the public scores. Even the notebooks which I made without looking at the public scores and using only simple scikit-learn models (like SGDRegressor) caused the same problem. So I don't think the reason is just leakage. Maybe the private test data was harder than the public test data?",
    "3076465": "Yes of course, the private test set was harder and contained most of the class 3 sii samples @ykawakita \nThe CV-LB relations indicated this through the competition!",
    "3076496": "Wow, I didn't realize it at all. I assumed that the train data and the test data were completely randomly taken from the data which Child Mind Institute has. Thank you for the lesson!"
  },
  "source": "meta"
}