{
  "id": 357241,
  "title": "The test dataset is as imbalanced as the train dataset",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/357241",
  "author_name": "",
  "post_date": "2022-10-03T16:54:55.707231900Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>As we know, the train dataset is highly imbalanced, with over 85% samples scoring no goals.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4206802%2F378b1c8e5fe2f05aa8c1dd21d11f110f%2FScreenshot%20from%202022-10-03%2021-40-26.png?generation=1664813452105951&amp;alt=media\" alt=\"\"></p>\n<p>But I was curious about the test dataset, how were the outcomes distributed there? So I submitted the sample submission file, with all 0s, and received a score of 2.05. </p>\n<p>Each sample in the test dataset that actually has someone scoring (only one will score, and either way, per-sample loss will be the same), will contribute a loss of <br>\n$$ \\frac{-1}{2} [\\log(10^{-15}) + \\log(1 - 10^{-15})] \\approx 17.27 $$</p>\n<p>Then we iterate over the number of all possible % of the test dataset that could be of the Class 1 - No one scoring, and see where we get this score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4206802%2Fe68c7e09eb421c47b00f4630819f69a5%2FScreenshot%20from%202022-10-03%2021-53-52.png?generation=1664814254664632&amp;alt=media\" alt=\"\"></p>\n<p>And there we have it. Even in the test dataset, about 88% of the samples are of snapshots where the probability of either teams scoring is 0. You could go on and make another submission with all of another class, then again reverse engineer from the loss to find out approximate distributions of the other two classes. That wouldn't be of much help though, so we stop here.</p>\n<p>Simply being able to precisely predict the 88% of the samples that <strong>don't</strong> score a goal, and then giving a random (0.5, 0.5) for other remaining samples, will approximately bring your score to 0.08</p>",
  "messages": [
    {
      "id": "1969737",
      "postDate": "10/03/2022 16:54:55",
      "content": "<p>As we know, the train dataset is highly imbalanced, with over 85% samples scoring no goals.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4206802%2F378b1c8e5fe2f05aa8c1dd21d11f110f%2FScreenshot%20from%202022-10-03%2021-40-26.png?generation=1664813452105951&amp;alt=media\" alt=\"\"></p>\n<p>But I was curious about the test dataset, how were the outcomes distributed there? So I submitted the sample submission file, with all 0s, and received a score of 2.05. </p>\n<p>Each sample in the test dataset that actually has someone scoring (only one will score, and either way, per-sample loss will be the same), will contribute a loss of <br>\n$$ \\frac{-1}{2} [\\log(10^{-15}) + \\log(1 - 10^{-15})] \\approx 17.27 $$</p>\n<p>Then we iterate over the number of all possible % of the test dataset that could be of the Class 1 - No one scoring, and see where we get this score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4206802%2Fe68c7e09eb421c47b00f4630819f69a5%2FScreenshot%20from%202022-10-03%2021-53-52.png?generation=1664814254664632&amp;alt=media\" alt=\"\"></p>\n<p>And there we have it. Even in the test dataset, about 88% of the samples are of snapshots where the probability of either teams scoring is 0. You could go on and make another submission with all of another class, then again reverse engineer from the loss to find out approximate distributions of the other two classes. That wouldn't be of much help though, so we stop here.</p>\n<p>Simply being able to precisely predict the 88% of the samples that <strong>don't</strong> score a goal, and then giving a random (0.5, 0.5) for other remaining samples, will approximately bring your score to 0.08</p>",
      "rawMarkdown": "As we know, the train dataset is highly imbalanced, with over 85% samples scoring no goals.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4206802%2F378b1c8e5fe2f05aa8c1dd21d11f110f%2FScreenshot%20from%202022-10-03%2021-40-26.png?generation=1664813452105951&alt=media)\n\nBut I was curious about the test dataset, how were the outcomes distributed there? So I submitted the sample submission file, with all 0s, and received a score of 2.05. \n\nEach sample in the test dataset that actually has someone scoring (only one will score, and either way, per-sample loss will be the same), will contribute a loss of \n$$ \\frac{-1}{2} [\\log(10^{-15}) + \\log(1 - 10^{-15})] \\approx 17.27 $$\n\nThen we iterate over the number of all possible % of the test dataset that could be of the Class 1 - No one scoring, and see where we get this score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4206802%2Fe68c7e09eb421c47b00f4630819f69a5%2FScreenshot%20from%202022-10-03%2021-53-52.png?generation=1664814254664632&alt=media)\n\nAnd there we have it. Even in the test dataset, about 88% of the samples are of snapshots where the probability of either teams scoring is 0. You could go on and make another submission with all of another class, then again reverse engineer from the loss to find out approximate distributions of the other two classes. That wouldn't be of much help though, so we stop here.\n\nSimply being able to precisely predict the 88% of the samples that **don't** score a goal, and then giving a random (0.5, 0.5) for other remaining samples, will approximately bring your score to 0.08",
      "votes": null
    },
    {
      "id": "1970167",
      "postDate": "10/03/2022 22:31:24",
      "content": "<p>That the data is imbalanced is to be expected. Even if every event is a goaling event, the average time to goal  is likely much higher than 10 sec. Say the average time to goal is 90 sec. Then 89% of the samples would have neither team scoring  within the next 10 sec.</p>",
      "rawMarkdown": "That the data is imbalanced is to be expected. Even if every event is a goaling event, the average time to goal  is likely much higher than 10 sec. Say the average time to goal is 90 sec. Then 89% of the samples would have neither team scoring  within the next 10 sec.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1970167,
      "author_name": "siukeitin",
      "author_url": "",
      "post_date": "10/03/2022 22:31:24",
      "content": "<p>That the data is imbalanced is to be expected. Even if every event is a goaling event, the average time to goal  is likely much higher than 10 sec. Say the average time to goal is 90 sec. Then 89% of the samples would have neither team scoring  within the next 10 sec.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1969737": "As we know, the train dataset is highly imbalanced, with over 85% samples scoring no goals.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4206802%2F378b1c8e5fe2f05aa8c1dd21d11f110f%2FScreenshot%20from%202022-10-03%2021-40-26.png?generation=1664813452105951&alt=media)\n\nBut I was curious about the test dataset, how were the outcomes distributed there? So I submitted the sample submission file, with all 0s, and received a score of 2.05. \n\nEach sample in the test dataset that actually has someone scoring (only one will score, and either way, per-sample loss will be the same), will contribute a loss of \n$$ \\frac{-1}{2} [\\log(10^{-15}) + \\log(1 - 10^{-15})] \\approx 17.27 $$\n\nThen we iterate over the number of all possible % of the test dataset that could be of the Class 1 - No one scoring, and see where we get this score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4206802%2Fe68c7e09eb421c47b00f4630819f69a5%2FScreenshot%20from%202022-10-03%2021-53-52.png?generation=1664814254664632&alt=media)\n\nAnd there we have it. Even in the test dataset, about 88% of the samples are of snapshots where the probability of either teams scoring is 0. You could go on and make another submission with all of another class, then again reverse engineer from the loss to find out approximate distributions of the other two classes. That wouldn't be of much help though, so we stop here.\n\nSimply being able to precisely predict the 88% of the samples that **don't** score a goal, and then giving a random (0.5, 0.5) for other remaining samples, will approximately bring your score to 0.08",
    "1970167": "That the data is imbalanced is to be expected. Even if every event is a goaling event, the average time to goal  is likely much higher than 10 sec. Say the average time to goal is 90 sec. Then 89% of the samples would have neither team scoring  within the next 10 sec."
  },
  "source": "meta"
}