{
  "id": 207594,
  "title": "Suggesting for New Private Dateset",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207594",
  "author_name": "",
  "post_date": "2020-12-30T12:48:50.098674900Z",
  "votes": null,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi Kaggle Team/Host<br>\nI think it is fairly known that owing to notorious bug many people were seeing till date (not sure since when)  their Private LB  score/Ranking  which would have given unfair advantage in terms of chosing  right  CV strategy, chosing  right model Hyper parameters ,basically all means to prevent overfiting to private set .<br>\nWe must set the new private set for fair competition .</p>",
  "messages": [
    {
      "id": "1132493",
      "postDate": "12/30/2020 12:48:50",
      "content": "<p>Hi Kaggle Team/Host<br>\nI think it is fairly known that owing to notorious bug many people were seeing till date (not sure since when)  their Private LB  score/Ranking  which would have given unfair advantage in terms of chosing  right  CV strategy, chosing  right model Hyper parameters ,basically all means to prevent overfiting to private set .<br>\nWe must set the new private set for fair competition .</p>",
      "rawMarkdown": "Hi Kaggle Team/Host\nI think it is fairly known that owing to notorious bug many people were seeing till date (not sure since when)  their Private LB  score/Ranking  which would have given unfair advantage in terms of chosing  right  CV strategy, chosing  right model Hyper parameters ,basically all means to prevent overfiting to private set .\nWe must set the new private set for fair competition .",
      "votes": null
    },
    {
      "id": "1132505",
      "postDate": "12/30/2020 12:58:45",
      "content": "<p>Yes.<br>\nAnd kaggle needs to choose data that doesn't shake.</p>",
      "rawMarkdown": "Yes.\nAnd kaggle needs to choose data that doesn't shake.",
      "votes": null
    },
    {
      "id": "1132508",
      "postDate": "12/30/2020 13:00:11",
      "content": "<p>yes. I think kaggle admins and the host have a discussion about this now.</p>",
      "rawMarkdown": "yes. I think kaggle admins and the host have a discussion about this now.",
      "votes": null
    },
    {
      "id": "1132593",
      "postDate": "12/30/2020 14:20:21",
      "content": "<p>it might come up at the cost of extended dead line.. but then what is the option else this competition will become like a league Match series person with highest standing  based on  final submission towards a unified leader board ,no private no public.</p>",
      "rawMarkdown": "it might come up at the cost of extended dead line.. but then what is the option else this competition will become like a league Match series person with highest standing  based on  final submission towards a unified leader board ,no private no public.",
      "votes": null
    },
    {
      "id": "1132631",
      "postDate": "12/30/2020 14:53:40",
      "content": "<p>I fear there might not be more data to evaluate with. From the onset, it looks like they went \"all-in\" with their dataset.</p>",
      "rawMarkdown": "I fear there might not be more data to evaluate with. From the onset, it looks like they went \"all-in\" with their dataset.",
      "votes": null
    },
    {
      "id": "1132660",
      "postDate": "12/30/2020 15:28:49",
      "content": "<p>I think they have enough data that contains the history of users that are not in train/test, because it's clear they have more than 400000 customers. However, I think maybe they don't have enough data that contains the history of users in training set, i.e. \"all-in\" here. So, I think this competition will possibly become completely different from now. I mean, maybe it will be unnecessary to handle the users that are in training set.</p>",
      "rawMarkdown": "I think they have enough data that contains the history of users that are not in train/test, because it's clear they have more than 400000 customers. However, I think maybe they don't have enough data that contains the history of users in training set, i.e. \"all-in\" here. So, I think this competition will possibly become completely different from now. I mean, maybe it will be unnecessary to handle the users that are in training set.",
      "votes": null
    },
    {
      "id": "1132663",
      "postDate": "12/30/2020 15:33:22",
      "content": "<p>Totally agree with you <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, and it would be very unfortunate they take that change. For me, handling the users in the training set is one of the important key in this competition, it would be really terrible for a lot of people to remove this aspect.</p>",
      "rawMarkdown": "Totally agree with you @mamasinkgs, and it would be very unfortunate they take that change. For me, handling the users in the training set is one of the important key in this competition, it would be really terrible for a lot of people to remove this aspect.",
      "votes": null
    },
    {
      "id": "1132670",
      "postDate": "12/30/2020 15:42:14",
      "content": "<p>Maybe there will be a solution tomorrow，let's wait</p>",
      "rawMarkdown": "Maybe there will be a solution tomorrow，let's wait",
      "votes": null
    },
    {
      "id": "1132677",
      "postDate": "12/30/2020 15:54:16",
      "content": "<p>If they change the private dataset, I hope it doesn't break our pipelines. Any slight format change can potentially breaks something. Imagine that they decide to increase the size of each group (from test dataset iterator), and suddently you get GPU OOM issue without knowing it … Or some assumptions about the questions / lectures / timestamp order change ….</p>",
      "rawMarkdown": "If they change the private dataset, I hope it doesn't break our pipelines. Any slight format change can potentially breaks something. Imagine that they decide to increase the size of each group (from test dataset iterator), and suddently you get GPU OOM issue without knowing it ... Or some assumptions about the questions / lectures / timestamp order change ....",
      "votes": null
    },
    {
      "id": "1132878",
      "postDate": "12/30/2020 18:56:24",
      "content": "<p>We probably have ~600K new users in test set - we are given ~400k and the title says \"Track knowledge states of 1M+ students in the wild\" </p>",
      "rawMarkdown": "We probably have ~600K new users in test set - we are given ~400k and the title says \"Track knowledge states of 1M+ students in the wild\"",
      "votes": null
    },
    {
      "id": "1133232",
      "postDate": "12/31/2020 03:34:14",
      "content": "<p>no, I checked there are not more than 100K new users in test set.</p>",
      "rawMarkdown": "no, I checked there are not more than 100K new users in test set.",
      "votes": null
    },
    {
      "id": "1133239",
      "postDate": "12/31/2020 03:43:02",
      "content": "<p><a href=\"https://www.kaggle.com/rashmibanthia\" target=\"_blank\">@rashmibanthia</a> train set +100M rows, ~350k students. No way test set 2.5M rows will have 600k new users.</p>",
      "rawMarkdown": "rashmibanthia train set +100M rows, ~350k students. No way test set 2.5M rows will have 600k new users.",
      "votes": null
    },
    {
      "id": "1133243",
      "postDate": "12/31/2020 03:51:05",
      "content": "<p>ok thanks, <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>  I wonder what's this for - \"Track knowledge states of 1M+ students in the wild\"  🤔</p>",
      "rawMarkdown": "ok thanks, @mamasinkgs @authman  I wonder what's this for - \"Track knowledge states of 1M+ students in the wild\"  🤔",
      "votes": null
    },
    {
      "id": "1133255",
      "postDate": "12/31/2020 04:01:44",
      "content": "<p>It would not be a fair thing to do when you just have a week's time. If this was disclosed 2/3 weeks back, it would have been fine to modify the test set with an extension grace as well. Unless it's identified that (top?) people in this comp have used it, i don't think the test data should be reset.</p>\n<p>As of now, People are desperately waiting to look at each other solutions and discuss the myriad of ideas they tried, what worked, what didn't etc. So care should be taken if they are going to swap the test set. Either ways, the cv-lb is pretty tight in this comp from the very beginning, so let's hope for the best.</p>\n<p>Happy New Year In Advance!</p>",
      "rawMarkdown": "It would not be a fair thing to do when you just have a week's time. If this was disclosed 2/3 weeks back, it would have been fine to modify the test set with an extension grace as well. Unless it's identified that (top?) people in this comp have used it, i don't think the test data should be reset.\n\nAs of now, People are desperately waiting to look at each other solutions and discuss the myriad of ideas they tried, what worked, what didn't etc. So care should be taken if they are going to swap the test set. Either ways, the cv-lb is pretty tight in this comp from the very beginning, so let's hope for the best.\n\nHappy New Year In Advance!",
      "votes": null
    },
    {
      "id": "1133276",
      "postDate": "12/31/2020 04:33:57",
      "content": "<p>I completely agree with you.<br>\nHowever, the current dataset has very high quality, as we did not get large CV/LB gap.<br>\nTest dataset must includes both new and old users, while temporal gap between current training and test dataset would be very small. It would be very difficult to create such a high quality datasets from new raw data without making any leakage. For example, we can guess new private dataset would not include the users in old training dataset, because they would all be appeared in old test dataset.</p>",
      "rawMarkdown": "I completely agree with you.\nHowever, the current dataset has very high quality, as we did not get large CV/LB gap.\nTest dataset must includes both new and old users, while temporal gap between current training and test dataset would be very small. It would be very difficult to create such a high quality datasets from new raw data without making any leakage. For example, we can guess new private dataset would not include the users in old training dataset, because they would all be appeared in old test dataset.",
      "votes": null
    },
    {
      "id": "1133298",
      "postDate": "12/31/2020 05:15:54",
      "content": "<p>Yes, to make such a great test set is very hard, because we can guess they have already used the whole history of the users in the old training dataset. If the test set is replaced, The competition will become somewhat fair, but this competition can become meaningless, because of the poor quality test set.</p>",
      "rawMarkdown": "Yes, to make such a great test set is very hard, because we can guess they have already used the whole history of the users in the old training dataset. If the test set is replaced, The competition will become somewhat fair, but this competition can become meaningless, because of the poor quality test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1132505,
      "author_name": "zakopur0",
      "author_url": "",
      "post_date": "12/30/2020 12:58:45",
      "content": "<p>Yes.<br>\nAnd kaggle needs to choose data that doesn't shake.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1132508,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/30/2020 13:00:11",
      "content": "<p>yes. I think kaggle admins and the host have a discussion about this now.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1132593,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/30/2020 14:20:21",
          "content": "<p>it might come up at the cost of extended dead line.. but then what is the option else this competition will become like a league Match series person with highest standing  based on  final submission towards a unified leader board ,no private no public.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1132631,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/30/2020 14:53:40",
      "content": "<p>I fear there might not be more data to evaluate with. From the onset, it looks like they went \"all-in\" with their dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1132660,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/30/2020 15:28:49",
          "content": "<p>I think they have enough data that contains the history of users that are not in train/test, because it's clear they have more than 400000 customers. However, I think maybe they don't have enough data that contains the history of users in training set, i.e. \"all-in\" here. So, I think this competition will possibly become completely different from now. I mean, maybe it will be unnecessary to handle the users that are in training set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132663,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/30/2020 15:33:22",
          "content": "<p>Totally agree with you <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, and it would be very unfortunate they take that change. For me, handling the users in the training set is one of the important key in this competition, it would be really terrible for a lot of people to remove this aspect.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132878,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "12/30/2020 18:56:24",
          "content": "<p>We probably have ~600K new users in test set - we are given ~400k and the title says \"Track knowledge states of 1M+ students in the wild\" </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133232,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/31/2020 03:34:14",
          "content": "<p>no, I checked there are not more than 100K new users in test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133239,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/31/2020 03:43:02",
          "content": "<p><a href=\"https://www.kaggle.com/rashmibanthia\" target=\"_blank\">@rashmibanthia</a> train set +100M rows, ~350k students. No way test set 2.5M rows will have 600k new users.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133243,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "12/31/2020 03:51:05",
          "content": "<p>ok thanks, <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>  I wonder what's this for - \"Track knowledge states of 1M+ students in the wild\"  🤔</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1132670,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "12/30/2020 15:42:14",
      "content": "<p>Maybe there will be a solution tomorrow，let's wait</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1132677,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "12/30/2020 15:54:16",
      "content": "<p>If they change the private dataset, I hope it doesn't break our pipelines. Any slight format change can potentially breaks something. Imagine that they decide to increase the size of each group (from test dataset iterator), and suddently you get GPU OOM issue without knowing it … Or some assumptions about the questions / lectures / timestamp order change ….</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1133255,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "12/31/2020 04:01:44",
      "content": "<p>It would not be a fair thing to do when you just have a week's time. If this was disclosed 2/3 weeks back, it would have been fine to modify the test set with an extension grace as well. Unless it's identified that (top?) people in this comp have used it, i don't think the test data should be reset.</p>\n<p>As of now, People are desperately waiting to look at each other solutions and discuss the myriad of ideas they tried, what worked, what didn't etc. So care should be taken if they are going to swap the test set. Either ways, the cv-lb is pretty tight in this comp from the very beginning, so let's hope for the best.</p>\n<p>Happy New Year In Advance!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1133276,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "12/31/2020 04:33:57",
      "content": "<p>I completely agree with you.<br>\nHowever, the current dataset has very high quality, as we did not get large CV/LB gap.<br>\nTest dataset must includes both new and old users, while temporal gap between current training and test dataset would be very small. It would be very difficult to create such a high quality datasets from new raw data without making any leakage. For example, we can guess new private dataset would not include the users in old training dataset, because they would all be appeared in old test dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1133298,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/31/2020 05:15:54",
          "content": "<p>Yes, to make such a great test set is very hard, because we can guess they have already used the whole history of the users in the old training dataset. If the test set is replaced, The competition will become somewhat fair, but this competition can become meaningless, because of the poor quality test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1132493": "Hi Kaggle Team/Host\nI think it is fairly known that owing to notorious bug many people were seeing till date (not sure since when)  their Private LB  score/Ranking  which would have given unfair advantage in terms of chosing  right  CV strategy, chosing  right model Hyper parameters ,basically all means to prevent overfiting to private set .\nWe must set the new private set for fair competition .",
    "1132505": "Yes.\nAnd kaggle needs to choose data that doesn't shake.",
    "1132508": "yes. I think kaggle admins and the host have a discussion about this now.",
    "1132593": "it might come up at the cost of extended dead line.. but then what is the option else this competition will become like a league Match series person with highest standing  based on  final submission towards a unified leader board ,no private no public.",
    "1132631": "I fear there might not be more data to evaluate with. From the onset, it looks like they went \"all-in\" with their dataset.",
    "1132660": "I think they have enough data that contains the history of users that are not in train/test, because it's clear they have more than 400000 customers. However, I think maybe they don't have enough data that contains the history of users in training set, i.e. \"all-in\" here. So, I think this competition will possibly become completely different from now. I mean, maybe it will be unnecessary to handle the users that are in training set.",
    "1132663": "Totally agree with you @mamasinkgs, and it would be very unfortunate they take that change. For me, handling the users in the training set is one of the important key in this competition, it would be really terrible for a lot of people to remove this aspect.",
    "1132670": "Maybe there will be a solution tomorrow，let's wait",
    "1132677": "If they change the private dataset, I hope it doesn't break our pipelines. Any slight format change can potentially breaks something. Imagine that they decide to increase the size of each group (from test dataset iterator), and suddently you get GPU OOM issue without knowing it ... Or some assumptions about the questions / lectures / timestamp order change ....",
    "1132878": "We probably have ~600K new users in test set - we are given ~400k and the title says \"Track knowledge states of 1M+ students in the wild\"",
    "1133232": "no, I checked there are not more than 100K new users in test set.",
    "1133239": "rashmibanthia train set +100M rows, ~350k students. No way test set 2.5M rows will have 600k new users.",
    "1133243": "ok thanks, @mamasinkgs @authman  I wonder what's this for - \"Track knowledge states of 1M+ students in the wild\"  🤔",
    "1133255": "It would not be a fair thing to do when you just have a week's time. If this was disclosed 2/3 weeks back, it would have been fine to modify the test set with an extension grace as well. Unless it's identified that (top?) people in this comp have used it, i don't think the test data should be reset.\n\nAs of now, People are desperately waiting to look at each other solutions and discuss the myriad of ideas they tried, what worked, what didn't etc. So care should be taken if they are going to swap the test set. Either ways, the cv-lb is pretty tight in this comp from the very beginning, so let's hope for the best.\n\nHappy New Year In Advance!",
    "1133276": "I completely agree with you.\nHowever, the current dataset has very high quality, as we did not get large CV/LB gap.\nTest dataset must includes both new and old users, while temporal gap between current training and test dataset would be very small. It would be very difficult to create such a high quality datasets from new raw data without making any leakage. For example, we can guess new private dataset would not include the users in old training dataset, because they would all be appeared in old test dataset.",
    "1133298": "Yes, to make such a great test set is very hard, because we can guess they have already used the whole history of the users in the old training dataset. If the test set is replaced, The competition will become somewhat fair, but this competition can become meaningless, because of the poor quality test set."
  },
  "source": "meta"
}