{
  "id": 199829,
  "title": "One nice thing I observed in this competition ... Thanks to @iglovikov",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199829",
  "author_name": "",
  "post_date": "2020-11-27T14:48:45.804253600Z",
  "votes": 22,
  "comment_count": 10,
  "views": 0,
  "content": "<p>In this competition, How the External Data Disclosure Thread is maintained is great  … <br> <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> replied on each of the post and let the competitors know in advance whether the code/data they use is acceptable or not ..  <br>Great work and not sure if this can be possible with all competitions but I think its a new good standard .. Thank you <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> </p>",
  "messages": [
    {
      "id": "1093222",
      "postDate": "11/27/2020 14:48:45",
      "content": "<p>In this competition, How the External Data Disclosure Thread is maintained is great  … <br> <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> replied on each of the post and let the competitors know in advance whether the code/data they use is acceptable or not ..  <br>Great work and not sure if this can be possible with all competitions but I think its a new good standard .. Thank you <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> </p>",
      "rawMarkdown": "In this competition, How the External Data Disclosure Thread is maintained is great  ... </br> @iglovikov replied on each of the post and let the competitors know in advance whether the code/data they use is acceptable or not ..  </br>Great work and not sure if this can be possible with all competitions but I think its a new good standard .. Thank you @iglovikov",
      "votes": null
    },
    {
      "id": "1093294",
      "postDate": "11/27/2020 15:45:11",
      "content": "<p>The hosts did an outstanding job during this competiton quickly anwering any questions. Also, there seem to be no leaks, clear validation methods, clear method of test data creation, etc.<br>\nIt was a great experience without drama. Thanks!</p>",
      "rawMarkdown": "The hosts did an outstanding job during this competiton quickly anwering any questions. Also, there seem to be no leaks, clear validation methods, clear method of test data creation, etc.\nIt was a great experience without drama. Thanks!",
      "votes": null
    },
    {
      "id": "1093427",
      "postDate": "11/27/2020 17:49:02",
      "content": "<p>I count <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199498\" target=\"_blank\">this</a> as a leak. Don't you?</p>",
      "rawMarkdown": "I count [this](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199498) as a leak. Don't you?",
      "votes": null
    },
    {
      "id": "1093481",
      "postDate": "11/27/2020 19:03:46",
      "content": "<p>Indeed that could be one, but in my opinion it's argueable, as those are the \"easy\" cases anyway. Very low, or zero velocity in heavy traffic. In those cases even very simple models already predicted the correct trajectory. <br>\nIn all cases you have 15 seconds of uncertainty and the known state 15 seconds into the future is still 10 seconds away from the last needed prediction. I am argueing that the improvement one can get from this knowledge may be neglectable.</p>\n<p>By chance, did you ran any experiments on the impact of this future peeking?</p>",
      "rawMarkdown": "Indeed that could be one, but in my opinion it's argueable, as those are the \"easy\" cases anyway. Very low, or zero velocity in heavy traffic. In those cases even very simple models already predicted the correct trajectory. \nIn all cases you have 15 seconds of uncertainty and the known state 15 seconds into the future is still 10 seconds away from the last needed prediction. I am argueing that the improvement one can get from this knowledge may be neglectable.\n\nBy chance, did you ran any experiments on the impact of this future peeking?",
      "votes": null
    },
    {
      "id": "1093583",
      "postDate": "11/27/2020 21:03:55",
      "content": "<p>Only the one that I reported there. I agree with your assessment, the possible improvement is small.</p>",
      "rawMarkdown": "Only the one that I reported there. I agree with your assessment, the possible improvement is small.",
      "votes": null
    },
    {
      "id": "1093695",
      "postDate": "11/27/2020 23:55:48",
      "content": "<p>The external data especially and overall the host feedback has been handled great, should be the new Kaggle standard.</p>\n<p>The only improvement I'd really like we had - to have the training, validation and test datasets separated in space. We would have to build solution which generalizes to the new unseen areas, would be much more challenging competition. With the kernel competition, the test dataset can be hidden from us to avoid leaks.</p>",
      "rawMarkdown": "The external data especially and overall the host feedback has been handled great, should be the new Kaggle standard.\n\nThe only improvement I'd really like we had - to have the training, validation and test datasets separated in space. We would have to build solution which generalizes to the new unseen areas, would be much more challenging competition. With the kernel competition, the test dataset can be hidden from us to avoid leaks.",
      "votes": null
    },
    {
      "id": "1094384",
      "postDate": "11/28/2020 15:17:40",
      "content": "<p>I have suggested this to kaggle  <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158747\" target=\"_blank\">here</a> 5 months ago after seeing  what happened on deepfake.<br>\nHere is what was replied:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/valanm\" target=\"_blank\">@valanm</a> I appreciate your recommendation, but realistically we (Kaggle) are not in a position to be vetting every external dataset. The evaluation and enforcement of external dataset eligibility is ultimately in the hands of the host for each competition. The extent of scrutiny of external data or any competition-specific rules a host opts to add is very host-specific (some, frankly, don't care), so the proposed approach will not be scalable or practical for all hosts.</p>\n  <p>So what does that mean for this competition? In this competition, the host has opened up the opportunity for any external data to be used (including private/non-public datasets), due to their interests in inclusivity of all approaches. However, for the purposes of prize-eligibility, if you opt to use external datasets (which I must emphasize is not a requirement), then those datasets must be public (with licensing that minimally includes academic/research use) with dataset provenance clearly shared on the external data thread.</p>\n  <p>In the interest of mitigating any issues around uncertainty in external data use, we are urging hosts (this competition included) to monitor closely the external data thread, in the event questions or concerns are raised around the use of certain datasets. If you check that thread, you'll see many of these detailed questions have already been addressed, so I encourage that to be the method for effective escalation.</p>\n</blockquote>",
      "rawMarkdown": "I have suggested this to kaggle  [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158747) 5 months ago after seeing  what happened on deepfake.\nHere is what was replied:\n> @valanm I appreciate your recommendation, but realistically we (Kaggle) are not in a position to be vetting every external dataset. The evaluation and enforcement of external dataset eligibility is ultimately in the hands of the host for each competition. The extent of scrutiny of external data or any competition-specific rules a host opts to add is very host-specific (some, frankly, don't care), so the proposed approach will not be scalable or practical for all hosts.\n\n> So what does that mean for this competition? In this competition, the host has opened up the opportunity for any external data to be used (including private/non-public datasets), due to their interests in inclusivity of all approaches. However, for the purposes of prize-eligibility, if you opt to use external datasets (which I must emphasize is not a requirement), then those datasets must be public (with licensing that minimally includes academic/research use) with dataset provenance clearly shared on the external data thread.\n\n> In the interest of mitigating any issues around uncertainty in external data use, we are urging hosts (this competition included) to monitor closely the external data thread, in the event questions or concerns are raised around the use of certain datasets. If you check that thread, you'll see many of these detailed questions have already been addressed, so I encourage that to be the method for effective escalation.",
      "votes": null
    },
    {
      "id": "1094389",
      "postDate": "11/28/2020 15:19:42",
      "content": "<p>So it all comes down to a host. Some are great (i.e. <a href=\"https://www.kaggle.com/iglovikovl\" target=\"_blank\">Vladimir</a>), some not so much</p>",
      "rawMarkdown": "So it all comes down to a host. Some are great (i.e. [Vladimir](https://www.kaggle.com/iglovikovl)), some not so much",
      "votes": null
    },
    {
      "id": "1094735",
      "postDate": "11/28/2020 21:27:12",
      "content": "<p>I agree, generalizing to new areas would be very interesting to study. But then, this goal should be made transparent to participants, I really do not like the idea of having a complete different private dataset without prior knowledge.</p>",
      "rawMarkdown": "I agree, generalizing to new areas would be very interesting to study. But then, this goal should be made transparent to participants, I really do not like the idea of having a complete different private dataset without prior knowledge.",
      "votes": null
    },
    {
      "id": "1094906",
      "postDate": "11/29/2020 04:33:35",
      "content": "<p>I'd completely agree with this, it should be clearly communicated the private test set is different (but with the similar distribution as the train one, not from the country with the different direction of travel etc) to the public test train sets.</p>",
      "rawMarkdown": "I'd completely agree with this, it should be clearly communicated the private test set is different (but with the similar distribution as the train one, not from the country with the different direction of travel etc) to the public test train sets.",
      "votes": null
    },
    {
      "id": "1104179",
      "postDate": "12/06/2020 17:20:06",
      "content": "<p>I had some troubles with the task and the data understanding, but later found that most of the the answers were given by the hosts. Just had to start with the post \"Dataset Tutorial Recording and Q&amp;A Resources\".</p>",
      "rawMarkdown": "I had some troubles with the task and the data understanding, but later found that most of the the answers were given by the hosts. Just had to start with the post \"Dataset Tutorial Recording and Q&A Resources\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1093294,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "11/27/2020 15:45:11",
      "content": "<p>The hosts did an outstanding job during this competiton quickly anwering any questions. Also, there seem to be no leaks, clear validation methods, clear method of test data creation, etc.<br>\nIt was a great experience without drama. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1093427,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/27/2020 17:49:02",
          "content": "<p>I count <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199498\" target=\"_blank\">this</a> as a leak. Don't you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093481,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "11/27/2020 19:03:46",
          "content": "<p>Indeed that could be one, but in my opinion it's argueable, as those are the \"easy\" cases anyway. Very low, or zero velocity in heavy traffic. In those cases even very simple models already predicted the correct trajectory. <br>\nIn all cases you have 15 seconds of uncertainty and the known state 15 seconds into the future is still 10 seconds away from the last needed prediction. I am argueing that the improvement one can get from this knowledge may be neglectable.</p>\n<p>By chance, did you ran any experiments on the impact of this future peeking?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093583,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "11/27/2020 21:03:55",
          "content": "<p>Only the one that I reported there. I agree with your assessment, the possible improvement is small.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1093695,
      "author_name": "dmytropoplavskiy",
      "author_url": "",
      "post_date": "11/27/2020 23:55:48",
      "content": "<p>The external data especially and overall the host feedback has been handled great, should be the new Kaggle standard.</p>\n<p>The only improvement I'd really like we had - to have the training, validation and test datasets separated in space. We would have to build solution which generalizes to the new unseen areas, would be much more challenging competition. With the kernel competition, the test dataset can be hidden from us to avoid leaks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1094735,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/28/2020 21:27:12",
          "content": "<p>I agree, generalizing to new areas would be very interesting to study. But then, this goal should be made transparent to participants, I really do not like the idea of having a complete different private dataset without prior knowledge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094906,
          "author_name": "dmytropoplavskiy",
          "author_url": "",
          "post_date": "11/29/2020 04:33:35",
          "content": "<p>I'd completely agree with this, it should be clearly communicated the private test set is different (but with the similar distribution as the train one, not from the country with the different direction of travel etc) to the public test train sets.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1094384,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "11/28/2020 15:17:40",
      "content": "<p>I have suggested this to kaggle  <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158747\" target=\"_blank\">here</a> 5 months ago after seeing  what happened on deepfake.<br>\nHere is what was replied:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/valanm\" target=\"_blank\">@valanm</a> I appreciate your recommendation, but realistically we (Kaggle) are not in a position to be vetting every external dataset. The evaluation and enforcement of external dataset eligibility is ultimately in the hands of the host for each competition. The extent of scrutiny of external data or any competition-specific rules a host opts to add is very host-specific (some, frankly, don't care), so the proposed approach will not be scalable or practical for all hosts.</p>\n  <p>So what does that mean for this competition? In this competition, the host has opened up the opportunity for any external data to be used (including private/non-public datasets), due to their interests in inclusivity of all approaches. However, for the purposes of prize-eligibility, if you opt to use external datasets (which I must emphasize is not a requirement), then those datasets must be public (with licensing that minimally includes academic/research use) with dataset provenance clearly shared on the external data thread.</p>\n  <p>In the interest of mitigating any issues around uncertainty in external data use, we are urging hosts (this competition included) to monitor closely the external data thread, in the event questions or concerns are raised around the use of certain datasets. If you check that thread, you'll see many of these detailed questions have already been addressed, so I encourage that to be the method for effective escalation.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1094389,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "11/28/2020 15:19:42",
          "content": "<p>So it all comes down to a host. Some are great (i.e. <a href=\"https://www.kaggle.com/iglovikovl\" target=\"_blank\">Vladimir</a>), some not so much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1104179,
      "author_name": "valvex",
      "author_url": "",
      "post_date": "12/06/2020 17:20:06",
      "content": "<p>I had some troubles with the task and the data understanding, but later found that most of the the answers were given by the hosts. Just had to start with the post \"Dataset Tutorial Recording and Q&amp;A Resources\".</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1093222": "In this competition, How the External Data Disclosure Thread is maintained is great  ... </br> @iglovikov replied on each of the post and let the competitors know in advance whether the code/data they use is acceptable or not ..  </br>Great work and not sure if this can be possible with all competitions but I think its a new good standard .. Thank you @iglovikov",
    "1093294": "The hosts did an outstanding job during this competiton quickly anwering any questions. Also, there seem to be no leaks, clear validation methods, clear method of test data creation, etc.\nIt was a great experience without drama. Thanks!",
    "1093427": "I count [this](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199498) as a leak. Don't you?",
    "1093481": "Indeed that could be one, but in my opinion it's argueable, as those are the \"easy\" cases anyway. Very low, or zero velocity in heavy traffic. In those cases even very simple models already predicted the correct trajectory. \nIn all cases you have 15 seconds of uncertainty and the known state 15 seconds into the future is still 10 seconds away from the last needed prediction. I am argueing that the improvement one can get from this knowledge may be neglectable.\n\nBy chance, did you ran any experiments on the impact of this future peeking?",
    "1093583": "Only the one that I reported there. I agree with your assessment, the possible improvement is small.",
    "1093695": "The external data especially and overall the host feedback has been handled great, should be the new Kaggle standard.\n\nThe only improvement I'd really like we had - to have the training, validation and test datasets separated in space. We would have to build solution which generalizes to the new unseen areas, would be much more challenging competition. With the kernel competition, the test dataset can be hidden from us to avoid leaks.",
    "1094384": "I have suggested this to kaggle  [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158747) 5 months ago after seeing  what happened on deepfake.\nHere is what was replied:\n> @valanm I appreciate your recommendation, but realistically we (Kaggle) are not in a position to be vetting every external dataset. The evaluation and enforcement of external dataset eligibility is ultimately in the hands of the host for each competition. The extent of scrutiny of external data or any competition-specific rules a host opts to add is very host-specific (some, frankly, don't care), so the proposed approach will not be scalable or practical for all hosts.\n\n> So what does that mean for this competition? In this competition, the host has opened up the opportunity for any external data to be used (including private/non-public datasets), due to their interests in inclusivity of all approaches. However, for the purposes of prize-eligibility, if you opt to use external datasets (which I must emphasize is not a requirement), then those datasets must be public (with licensing that minimally includes academic/research use) with dataset provenance clearly shared on the external data thread.\n\n> In the interest of mitigating any issues around uncertainty in external data use, we are urging hosts (this competition included) to monitor closely the external data thread, in the event questions or concerns are raised around the use of certain datasets. If you check that thread, you'll see many of these detailed questions have already been addressed, so I encourage that to be the method for effective escalation.",
    "1094389": "So it all comes down to a host. Some are great (i.e. [Vladimir](https://www.kaggle.com/iglovikovl)), some not so much",
    "1094735": "I agree, generalizing to new areas would be very interesting to study. But then, this goal should be made transparent to participants, I really do not like the idea of having a complete different private dataset without prior knowledge.",
    "1094906": "I'd completely agree with this, it should be clearly communicated the private test set is different (but with the similar distribution as the train one, not from the country with the different direction of travel etc) to the public test train sets.",
    "1104179": "I had some troubles with the task and the data understanding, but later found that most of the the answers were given by the hosts. Just had to start with the post \"Dataset Tutorial Recording and Q&A Resources\"."
  },
  "source": "meta"
}