{
  "id": 42791,
  "title": "stage1 and stage2 training dataset have the same names?",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/42791",
  "author_name": "",
  "post_date": "2017-11-05T04:52:39.317948Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I guess we'll have stage2_labels.csv for the 2nd stage training dataset.\nDoes this file have the union of stage1_labels.csv and stage1_sample_submission.csv with exactly the same namings? I mean a ct scan named \"00360f79fd6e02781457eda48f85da90\" in stgae1 has the same name \"00360f79fd6e02781457eda48f85da90\" in stage2?</p>",
  "messages": [
    {
      "id": "239956",
      "postDate": "11/05/2017 04:52:39",
      "content": "<p>I guess we'll have stage2_labels.csv for the 2nd stage training dataset.\nDoes this file have the union of stage1_labels.csv and stage1_sample_submission.csv with exactly the same namings? I mean a ct scan named \"00360f79fd6e02781457eda48f85da90\" in stgae1 has the same name \"00360f79fd6e02781457eda48f85da90\" in stage2?</p>",
      "rawMarkdown": "I guess we'll have stage2_labels.csv for the 2nd stage training dataset.\nDoes this file have the union of stage1_labels.csv and stage1_sample_submission.csv with exactly the same namings? I mean a ct scan named \"00360f79fd6e02781457eda48f85da90\" in stgae1 has the same name \"00360f79fd6e02781457eda48f85da90\" in stage2?",
      "votes": null
    },
    {
      "id": "240761",
      "postDate": "11/07/2017 09:51:53",
      "content": "<p>Take this with a grain of salt, but my understanding is that there will be no training for stage 2.  The ranking will be a result of how well the model you trained from stage 1 performs on unseen, unreleased data: if you change your model (e.g. train it on the stage 2 data), you're disqualified entirely.**</p>\n\n<p>I believe they're flexible when it comes to changing a line here or there to allow for your model being used with different filenames, but that's about it.</p>\n\n<p>** Source: random information found here on Kaggle, and common sense.</p>",
      "rawMarkdown": "Take this with a grain of salt, but my understanding is that there will be no training for stage 2.  The ranking will be a result of how well the model you trained from stage 1 performs on unseen, unreleased data: if you change your model (e.g. train it on the stage 2 data), you're disqualified entirely.**\n\nI believe they're flexible when it comes to changing a line here or there to allow for your model being used with different filenames, but that's about it.\n\n** Source: random information found here on Kaggle, and common sense.",
      "votes": null
    },
    {
      "id": "241070",
      "postDate": "11/08/2017 01:27:24",
      "content": "<p>@wcukierski</p>\n\n<ul>\n<li>What files will I get for the 2nd stage? (Do we get labels for stage1 test dataset?)</li>\n<li>Is semi-supervised learning with 2nd stage data allowed?</li>\n</ul>\n\n<p>@Murray\nThank you for your answer, I think I was biased by a competition where we got a new training dataset (combination of stage1 training and test dataset). And, now I have definitely no idea what the stage2 dataset will look like...\n<a href=\"https://www.kaggle.com/c/web-traffic-time-series-forecasting/data\">https://www.kaggle.com/c/web-traffic-time-series-forecasting/data</a></p>",
      "rawMarkdown": "wcukierski\n\n* What files will I get for the 2nd stage? (Do we get labels for stage1 test dataset?)\n* Is semi-supervised learning with 2nd stage data allowed?\n\n@Murray\nThank you for your answer, I think I was biased by a competition where we got a new training dataset (combination of stage1 training and test dataset). And, now I have definitely no idea what the stage2 dataset will look like...\nhttps://www.kaggle.com/c/web-traffic-time-series-forecasting/data",
      "votes": null
    },
    {
      "id": "243269",
      "postDate": "11/13/2017 17:22:20",
      "content": "<p>The 2nd stage test set will be a set of scans similar to the current test set (with different file names). You will get the stage 1 labels. Semi supervised learning is allowed, but has to be automated.</p>",
      "rawMarkdown": "The 2nd stage test set will be a set of scans similar to the current test set (with different file names). You will get the stage 1 labels. Semi supervised learning is allowed, but has to be automated.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 240761,
      "author_name": "mmiron",
      "author_url": "",
      "post_date": "11/07/2017 09:51:53",
      "content": "<p>Take this with a grain of salt, but my understanding is that there will be no training for stage 2.  The ranking will be a result of how well the model you trained from stage 1 performs on unseen, unreleased data: if you change your model (e.g. train it on the stage 2 data), you're disqualified entirely.**</p>\n\n<p>I believe they're flexible when it comes to changing a line here or there to allow for your model being used with different filenames, but that's about it.</p>\n\n<p>** Source: random information found here on Kaggle, and common sense.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 241070,
      "author_name": "ckomaki",
      "author_url": "",
      "post_date": "11/08/2017 01:27:24",
      "content": "<p>@wcukierski</p>\n\n<ul>\n<li>What files will I get for the 2nd stage? (Do we get labels for stage1 test dataset?)</li>\n<li>Is semi-supervised learning with 2nd stage data allowed?</li>\n</ul>\n\n<p>@Murray\nThank you for your answer, I think I was biased by a competition where we got a new training dataset (combination of stage1 training and test dataset). And, now I have definitely no idea what the stage2 dataset will look like...\n<a href=\"https://www.kaggle.com/c/web-traffic-time-series-forecasting/data\">https://www.kaggle.com/c/web-traffic-time-series-forecasting/data</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 243269,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "11/13/2017 17:22:20",
          "content": "<p>The 2nd stage test set will be a set of scans similar to the current test set (with different file names). You will get the stage 1 labels. Semi supervised learning is allowed, but has to be automated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "239956": "I guess we'll have stage2_labels.csv for the 2nd stage training dataset.\nDoes this file have the union of stage1_labels.csv and stage1_sample_submission.csv with exactly the same namings? I mean a ct scan named \"00360f79fd6e02781457eda48f85da90\" in stgae1 has the same name \"00360f79fd6e02781457eda48f85da90\" in stage2?",
    "240761": "Take this with a grain of salt, but my understanding is that there will be no training for stage 2.  The ranking will be a result of how well the model you trained from stage 1 performs on unseen, unreleased data: if you change your model (e.g. train it on the stage 2 data), you're disqualified entirely.**\n\nI believe they're flexible when it comes to changing a line here or there to allow for your model being used with different filenames, but that's about it.\n\n** Source: random information found here on Kaggle, and common sense.",
    "241070": "wcukierski\n\n* What files will I get for the 2nd stage? (Do we get labels for stage1 test dataset?)\n* Is semi-supervised learning with 2nd stage data allowed?\n\n@Murray\nThank you for your answer, I think I was biased by a competition where we got a new training dataset (combination of stage1 training and test dataset). And, now I have definitely no idea what the stage2 dataset will look like...\nhttps://www.kaggle.com/c/web-traffic-time-series-forecasting/data",
    "243269": "The 2nd stage test set will be a set of scans similar to the current test set (with different file names). You will get the stage 1 labels. Semi supervised learning is allowed, but has to be automated."
  },
  "source": "meta"
}