{
  "id": 339033,
  "title": "\"Fun fact\":All competitions on kaggle exist using future data.",
  "url": "/competitions/amex-default-prediction/discussion/339033",
  "author_name": "",
  "post_date": "2022-07-23T03:17:53.257050500Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Because kaggle finally takes the best score in the two submissions as the final score, but this is meaningless for the real scene.<br>\nIn other words, we know which plan is good in the future and take that plan as our real plan, which is using future data.</p>\n<p>In the case of large amount of data and consistent online and offline distribution, it will not have a great impact.<br>\nHowever, if the distribution is different, it will have a great impact, or determine the success or failure.<br>\nSome examples:<br>\n1.Positive and negative sample ratio in classification.<br>\n2.The mean value of the test set in the regression problem.</p>\n<p>If we have one prediction and have two submission, we will generate a prediction *1.05 and a prediction *0.95. This may lead to a higher winning rate in the end in shaking competion.<br>\nBut in the actual scenario, if we generate two strategies, the final result should be the average of the two strategies, not the maximum.</p>\n<p>So I think the number of final selection evaluation of kaggle should be once, rather than twice in the past and now. I think this is a mistake left over by history</p>",
  "messages": [
    {
      "id": "1867147",
      "postDate": "07/23/2022 03:17:53",
      "content": "<p>Because kaggle finally takes the best score in the two submissions as the final score, but this is meaningless for the real scene.<br>\nIn other words, we know which plan is good in the future and take that plan as our real plan, which is using future data.</p>\n<p>In the case of large amount of data and consistent online and offline distribution, it will not have a great impact.<br>\nHowever, if the distribution is different, it will have a great impact, or determine the success or failure.<br>\nSome examples:<br>\n1.Positive and negative sample ratio in classification.<br>\n2.The mean value of the test set in the regression problem.</p>\n<p>If we have one prediction and have two submission, we will generate a prediction *1.05 and a prediction *0.95. This may lead to a higher winning rate in the end in shaking competion.<br>\nBut in the actual scenario, if we generate two strategies, the final result should be the average of the two strategies, not the maximum.</p>\n<p>So I think the number of final selection evaluation of kaggle should be once, rather than twice in the past and now. I think this is a mistake left over by history</p>",
      "rawMarkdown": "Because kaggle finally takes the best score in the two submissions as the final score, but this is meaningless for the real scene.\nIn other words, we know which plan is good in the future and take that plan as our real plan, which is using future data.\n\nIn the case of large amount of data and consistent online and offline distribution, it will not have a great impact.\nHowever, if the distribution is different, it will have a great impact, or determine the success or failure.\nSome examples:\n1.Positive and negative sample ratio in classification.\n2.The mean value of the test set in the regression problem.\n\nIf we have one prediction and have two submission, we will generate a prediction *1.05 and a prediction *0.95. This may lead to a higher winning rate in the end in shaking competion.\nBut in the actual scenario, if we generate two strategies, the final result should be the average of the two strategies, not the maximum.\n\nSo I think the number of final selection evaluation of kaggle should be once, rather than twice in the past and now. I think this is a mistake left over by history",
      "votes": null
    },
    {
      "id": "1867210",
      "postDate": "07/23/2022 04:41:02",
      "content": "<p>I've thought about this before and I half agree and disagree. Looking at most competitions and how they are constrained by satisfying some arbitrary loss function, they don't have much to do with how you would do things in real life and for many competitions in the past I've looked at to learn from good kagglers this leads to solutions which have questionable real life application.</p>\n<p>But for your point with the two submissions I think to the contrary that for most competitions you should be given even more submissions, because imagine if you were only given only one submission, then you would have to guess which strategy to follow because you would be at a disadvantage if you took some \"safe average\" compared to the others taking a risk of following the either *1.05 or the *0.95 example you outlined in your post.</p>",
      "rawMarkdown": "I've thought about this before and I half agree and disagree. Looking at most competitions and how they are constrained by satisfying some arbitrary loss function, they don't have much to do with how you would do things in real life and for many competitions in the past I've looked at to learn from good kagglers this leads to solutions which have questionable real life application.\n\nBut for your point with the two submissions I think to the contrary that for most competitions you should be given even more submissions, because imagine if you were only given only one submission, then you would have to guess which strategy to follow because you would be at a disadvantage if you took some \"safe average\" compared to the others taking a risk of following the either *1.05 or the *0.95 example you outlined in your post.",
      "votes": null
    },
    {
      "id": "1867254",
      "postDate": "07/23/2022 05:30:14",
      "content": "<p>Yes, I also think that more submissions with large differences will be more stable, but this cannot prevent others from making the same copy of all submissions.<br>\nSo this scheme is more difficult to implement.</p>",
      "rawMarkdown": "Yes, I also think that more submissions with large differences will be more stable, but this cannot prevent others from making the same copy of all submissions.\nSo this scheme is more difficult to implement.",
      "votes": null
    },
    {
      "id": "1867569",
      "postDate": "07/23/2022 10:44:02",
      "content": "<p>Hi, interesting point.</p>\n<p>This might be a trick choose the final submission, but most people will not do that. They choose the submissions generarally made by two completely different methods. Even it is just an optimization for cv and an optimization for LB would still be a relevant. </p>\n<p>For example, I have chosen the final two submission in one of the competitions: one is the prediction made by LGB, and the other is the prediction of LGB+NN. These two results are not related to maximum and minimum predictions, but completely different scenarios.</p>\n<p>At the end, the competition host needs a solution to solve a problem, not a one-time correct selection. Therefore I personally think selecting 2 submissions as the final should be fine.</p>",
      "rawMarkdown": "Hi, interesting point.\n\nThis might be a trick choose the final submission, but most people will not do that. They choose the submissions generarally made by two completely different methods. Even it is just an optimization for cv and an optimization for LB would still be a relevant. \n\nFor example, I have chosen the final two submission in one of the competitions: one is the prediction made by LGB, and the other is the prediction of LGB+NN. These two results are not related to maximum and minimum predictions, but completely different scenarios.\n\nAt the end, the competition host needs a solution to solve a problem, not a one-time correct selection. Therefore I personally think selecting 2 submissions as the final should be fine.",
      "votes": null
    },
    {
      "id": "1867617",
      "postDate": "07/23/2022 11:35:39",
      "content": "<p>We don't get 2 submissions using future data. We get 6,464 submissions using future data. In this current competition there are 3232 teams and each team has 2 submissions. Therefore the best model is chosen from 6,464 submissions in the future.</p>\n<p>The choice whether each team gets 1 or 2 submissions doesn't make much difference since 3,232 total submissions or 6,464 total submissions are basically the same.</p>",
      "rawMarkdown": "We don't get 2 submissions using future data. We get 6,464 submissions using future data. In this current competition there are 3232 teams and each team has 2 submissions. Therefore the best model is chosen from 6,464 submissions in the future.\n\nThe choice whether each team gets 1 or 2 submissions doesn't make much difference since 3,232 total submissions or 6,464 total submissions are basically the same.",
      "votes": null
    },
    {
      "id": "1867649",
      "postDate": "07/23/2022 12:18:57",
      "content": "<p>Haha,you are right!</p>",
      "rawMarkdown": "Haha,you are right!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1867210,
      "author_name": "kaggledoer",
      "author_url": "",
      "post_date": "07/23/2022 04:41:02",
      "content": "<p>I've thought about this before and I half agree and disagree. Looking at most competitions and how they are constrained by satisfying some arbitrary loss function, they don't have much to do with how you would do things in real life and for many competitions in the past I've looked at to learn from good kagglers this leads to solutions which have questionable real life application.</p>\n<p>But for your point with the two submissions I think to the contrary that for most competitions you should be given even more submissions, because imagine if you were only given only one submission, then you would have to guess which strategy to follow because you would be at a disadvantage if you took some \"safe average\" compared to the others taking a risk of following the either *1.05 or the *0.95 example you outlined in your post.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1867254,
          "author_name": "dc5e964768ef56302a32",
          "author_url": "",
          "post_date": "07/23/2022 05:30:14",
          "content": "<p>Yes, I also think that more submissions with large differences will be more stable, but this cannot prevent others from making the same copy of all submissions.<br>\nSo this scheme is more difficult to implement.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1867569,
      "author_name": "meli19",
      "author_url": "",
      "post_date": "07/23/2022 10:44:02",
      "content": "<p>Hi, interesting point.</p>\n<p>This might be a trick choose the final submission, but most people will not do that. They choose the submissions generarally made by two completely different methods. Even it is just an optimization for cv and an optimization for LB would still be a relevant. </p>\n<p>For example, I have chosen the final two submission in one of the competitions: one is the prediction made by LGB, and the other is the prediction of LGB+NN. These two results are not related to maximum and minimum predictions, but completely different scenarios.</p>\n<p>At the end, the competition host needs a solution to solve a problem, not a one-time correct selection. Therefore I personally think selecting 2 submissions as the final should be fine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1867617,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/23/2022 11:35:39",
      "content": "<p>We don't get 2 submissions using future data. We get 6,464 submissions using future data. In this current competition there are 3232 teams and each team has 2 submissions. Therefore the best model is chosen from 6,464 submissions in the future.</p>\n<p>The choice whether each team gets 1 or 2 submissions doesn't make much difference since 3,232 total submissions or 6,464 total submissions are basically the same.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1867649,
          "author_name": "dc5e964768ef56302a32",
          "author_url": "",
          "post_date": "07/23/2022 12:18:57",
          "content": "<p>Haha,you are right!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1867147": "Because kaggle finally takes the best score in the two submissions as the final score, but this is meaningless for the real scene.\nIn other words, we know which plan is good in the future and take that plan as our real plan, which is using future data.\n\nIn the case of large amount of data and consistent online and offline distribution, it will not have a great impact.\nHowever, if the distribution is different, it will have a great impact, or determine the success or failure.\nSome examples:\n1.Positive and negative sample ratio in classification.\n2.The mean value of the test set in the regression problem.\n\nIf we have one prediction and have two submission, we will generate a prediction *1.05 and a prediction *0.95. This may lead to a higher winning rate in the end in shaking competion.\nBut in the actual scenario, if we generate two strategies, the final result should be the average of the two strategies, not the maximum.\n\nSo I think the number of final selection evaluation of kaggle should be once, rather than twice in the past and now. I think this is a mistake left over by history",
    "1867210": "I've thought about this before and I half agree and disagree. Looking at most competitions and how they are constrained by satisfying some arbitrary loss function, they don't have much to do with how you would do things in real life and for many competitions in the past I've looked at to learn from good kagglers this leads to solutions which have questionable real life application.\n\nBut for your point with the two submissions I think to the contrary that for most competitions you should be given even more submissions, because imagine if you were only given only one submission, then you would have to guess which strategy to follow because you would be at a disadvantage if you took some \"safe average\" compared to the others taking a risk of following the either *1.05 or the *0.95 example you outlined in your post.",
    "1867254": "Yes, I also think that more submissions with large differences will be more stable, but this cannot prevent others from making the same copy of all submissions.\nSo this scheme is more difficult to implement.",
    "1867569": "Hi, interesting point.\n\nThis might be a trick choose the final submission, but most people will not do that. They choose the submissions generarally made by two completely different methods. Even it is just an optimization for cv and an optimization for LB would still be a relevant. \n\nFor example, I have chosen the final two submission in one of the competitions: one is the prediction made by LGB, and the other is the prediction of LGB+NN. These two results are not related to maximum and minimum predictions, but completely different scenarios.\n\nAt the end, the competition host needs a solution to solve a problem, not a one-time correct selection. Therefore I personally think selecting 2 submissions as the final should be fine.",
    "1867617": "We don't get 2 submissions using future data. We get 6,464 submissions using future data. In this current competition there are 3232 teams and each team has 2 submissions. Therefore the best model is chosen from 6,464 submissions in the future.\n\nThe choice whether each team gets 1 or 2 submissions doesn't make much difference since 3,232 total submissions or 6,464 total submissions are basically the same.",
    "1867649": "Haha,you are right!"
  },
  "source": "meta"
}