{
  "id": 37186,
  "title": "Should we forbid using test data to retrain model in stage 2?",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/37186",
  "author_name": "",
  "post_date": "2017-07-28T15:43:31.339253200Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have a feeling that some tricks used in <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection\">https://www.kaggle.com/c/state-farm-distracted-driver-detection</a> could be used in this competition too. Specifically, if a person appears multiple times in the test data, then those unlabelled test data could still somehow be exploited (e.g. it is possible that algorithms could identify that, say, 30 test cases involves a same person, and one could make some interesting assumptions about the threat distribution in these 30 test case). </p>\n\n<p>Note that in the real TSA application, one could never obtain such information. I would propose that the model cannot be changed using any stage 2 data. The model should read and predict test cases one by one in a streaming fashion.</p>",
  "messages": [
    {
      "id": "208109",
      "postDate": "07/28/2017 15:43:31",
      "content": "<p>I have a feeling that some tricks used in <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection\">https://www.kaggle.com/c/state-farm-distracted-driver-detection</a> could be used in this competition too. Specifically, if a person appears multiple times in the test data, then those unlabelled test data could still somehow be exploited (e.g. it is possible that algorithms could identify that, say, 30 test cases involves a same person, and one could make some interesting assumptions about the threat distribution in these 30 test case). </p>\n\n<p>Note that in the real TSA application, one could never obtain such information. I would propose that the model cannot be changed using any stage 2 data. The model should read and predict test cases one by one in a streaming fashion.</p>",
      "rawMarkdown": "I have a feeling that some tricks used in https://www.kaggle.com/c/state-farm-distracted-driver-detection could be used in this competition too. Specifically, if a person appears multiple times in the test data, then those unlabelled test data could still somehow be exploited (e.g. it is possible that algorithms could identify that, say, 30 test cases involves a same person, and one could make some interesting assumptions about the threat distribution in these 30 test case). \n\nNote that in the real TSA application, one could never obtain such information. I would propose that the model cannot be changed using any stage 2 data. The model should read and predict test cases one by one in a streaming fashion.",
      "votes": null
    },
    {
      "id": "219522",
      "postDate": "09/08/2017 13:15:59",
      "content": "<p>You have to upload the model before the stage 2 data is released.</p>",
      "rawMarkdown": "You have to upload the model before the stage 2 data is released.",
      "votes": null
    },
    {
      "id": "235077",
      "postDate": "10/24/2017 17:39:51",
      "content": "<p>how to upload model? I can not find \"More-&gt;Team-&gt;Your Model\".</p>",
      "rawMarkdown": "how to upload model? I can not find \"More-&gt;Team-&gt;Your Model\".",
      "votes": null
    },
    {
      "id": "235078",
      "postDate": "10/24/2017 17:42:52",
      "content": "<p>Hi Zjp2origin,</p>\n\n<p>The Model Upload will be enabled at the conclusion of Phase 1.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi Zjp2origin,\n\nThe Model Upload will be enabled at the conclusion of Phase 1.\n\nThanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 219522,
      "author_name": "jduffy",
      "author_url": "",
      "post_date": "09/08/2017 13:15:59",
      "content": "<p>You have to upload the model before the stage 2 data is released.</p>",
      "votes": null,
      "replies": [
        {
          "id": 235077,
          "author_name": "zjp2origin",
          "author_url": "",
          "post_date": "10/24/2017 17:39:51",
          "content": "<p>how to upload model? I can not find \"More-&gt;Team-&gt;Your Model\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 235078,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "10/24/2017 17:42:52",
          "content": "<p>Hi Zjp2origin,</p>\n\n<p>The Model Upload will be enabled at the conclusion of Phase 1.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "208109": "I have a feeling that some tricks used in https://www.kaggle.com/c/state-farm-distracted-driver-detection could be used in this competition too. Specifically, if a person appears multiple times in the test data, then those unlabelled test data could still somehow be exploited (e.g. it is possible that algorithms could identify that, say, 30 test cases involves a same person, and one could make some interesting assumptions about the threat distribution in these 30 test case). \n\nNote that in the real TSA application, one could never obtain such information. I would propose that the model cannot be changed using any stage 2 data. The model should read and predict test cases one by one in a streaming fashion.",
    "219522": "You have to upload the model before the stage 2 data is released.",
    "235077": "how to upload model? I can not find \"More-&gt;Team-&gt;Your Model\".",
    "235078": "Hi Zjp2origin,\n\nThe Model Upload will be enabled at the conclusion of Phase 1.\n\nThanks!"
  },
  "source": "meta"
}