{
  "id": 39484,
  "title": "Will the same volunteers appear in the stage 2 images multiple times?",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/39484",
  "author_name": "",
  "post_date": "2017-09-15T00:58:13.724635500Z",
  "votes": 13,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I understand that the volunteers used in Stage 2 will be different than those used in Stage 1, but will volunteers used in Stage 2 appear multiple times?</p>\n\n<p>I ask because if the same volunteers appear multiple times in Stage 2 one could conceivably:</p>\n\n<ol>\n<li>Use a clustering algorithm to separate the images into clusters (ideally one cluster for each volunteer)</li>\n<li>For each volunteer you could combine all of their images to create a \"normal\" image i.e. an image of what each volunteer would look like without any threats. Since each zone probably won't have a threat in the majority of images this could be achieved through some kind of averaging. Admittedly, this step is dependent on threats appearing about as frequently as they do in the training set. If every volunteer is loaded head to toe with threats in every image then this bit might not work.</li>\n<li>Use the difference of each image and the \"normal\" image to detect a threat. </li>\n</ol>\n\n<p>Since in a real life situation there will only be 1 scan of each person passing through security this method will be completely unusable. If volunteers appear multiple times in the test set is such a method allowed in this competition?</p>",
  "messages": [
    {
      "id": "221365",
      "postDate": "09/15/2017 00:58:13",
      "content": "<p>I understand that the volunteers used in Stage 2 will be different than those used in Stage 1, but will volunteers used in Stage 2 appear multiple times?</p>\n\n<p>I ask because if the same volunteers appear multiple times in Stage 2 one could conceivably:</p>\n\n<ol>\n<li>Use a clustering algorithm to separate the images into clusters (ideally one cluster for each volunteer)</li>\n<li>For each volunteer you could combine all of their images to create a \"normal\" image i.e. an image of what each volunteer would look like without any threats. Since each zone probably won't have a threat in the majority of images this could be achieved through some kind of averaging. Admittedly, this step is dependent on threats appearing about as frequently as they do in the training set. If every volunteer is loaded head to toe with threats in every image then this bit might not work.</li>\n<li>Use the difference of each image and the \"normal\" image to detect a threat. </li>\n</ol>\n\n<p>Since in a real life situation there will only be 1 scan of each person passing through security this method will be completely unusable. If volunteers appear multiple times in the test set is such a method allowed in this competition?</p>",
      "rawMarkdown": "I understand that the volunteers used in Stage 2 will be different than those used in Stage 1, but will volunteers used in Stage 2 appear multiple times?\n\nI ask because if the same volunteers appear multiple times in Stage 2 one could conceivably:\n\n1. Use a clustering algorithm to separate the images into clusters (ideally one cluster for each volunteer)\n2. For each volunteer you could combine all of their images to create a \"normal\" image i.e. an image of what each volunteer would look like without any threats. Since each zone probably won't have a threat in the majority of images this could be achieved through some kind of averaging. Admittedly, this step is dependent on threats appearing about as frequently as they do in the training set. If every volunteer is loaded head to toe with threats in every image then this bit might not work.\n3. Use the difference of each image and the \"normal\" image to detect a threat. \n\nSince in a real life situation there will only be 1 scan of each person passing through security this method will be completely unusable. If volunteers appear multiple times in the test set is such a method allowed in this competition?",
      "votes": null
    },
    {
      "id": "223050",
      "postDate": "09/21/2017 01:03:54",
      "content": "<p>Can we get an answer from an admin on this please? Are solutions that exploit the fact that the same volunteers may appear multiple times in Stage 2 admissible? </p>",
      "rawMarkdown": "Can we get an answer from an admin on this please? Are solutions that exploit the fact that the same volunteers may appear multiple times in Stage 2 admissible?",
      "votes": null
    },
    {
      "id": "223150",
      "postDate": "09/21/2017 08:47:57",
      "content": "<p>Branden,</p>\n\n<blockquote>\n  <p>is such a method allowed in this competition?</p>\n</blockquote>\n\n<p>It's my understanding that this is allowed, because you are allowed to use test set scans other than\nthe scan you are making predictions for. (Clustering is but a special case of unsupervised learning)</p>\n\n<p><a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/35116#199421\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/35116#199421</a></p>\n\n<blockquote>\n  <p>Since in a real life situation there will only be 1 scan of each person passing through security this method will be completely unusable. </p>\n</blockquote>\n\n<p>There are other factors to consider, such as that, in practice, they could create a much larger unlabeled dataset than we'll have here, which is a boon to semi-supervised methods. Additionally, semi-supervised methods may generalize better to previously unseen contraband.</p>",
      "rawMarkdown": "Branden,\n\n&gt; is such a method allowed in this competition?\n\nIt's my understanding that this is allowed, because you are allowed to use test set scans other than\nthe scan you are making predictions for. (Clustering is but a special case of unsupervised learning)\n\nhttps://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/35116#199421\n\n&gt; Since in a real life situation there will only be 1 scan of each person passing through security this method will be completely unusable. \n\nThere are other factors to consider, such as that, in practice, they could create a much larger unlabeled dataset than we'll have here, which is a boon to semi-supervised methods. Additionally, semi-supervised methods may generalize better to previously unseen contraband.",
      "votes": null
    },
    {
      "id": "224059",
      "postDate": "09/25/2017 00:39:45",
      "content": "<p>I agree this could use some clarification from the admins.  I don't see anything in the rules that explicitly prohibits this, but I doubt this is the kind of solution the sponsor would want to see.  In real world deployment, each person is scanned only once.  Algorithms that compare multiple scans from the same person to figure out when they have objects on them would be pretty useless.   </p>",
      "rawMarkdown": "I agree this could use some clarification from the admins.  I don't see anything in the rules that explicitly prohibits this, but I doubt this is the kind of solution the sponsor would want to see.  In real world deployment, each person is scanned only once.  Algorithms that compare multiple scans from the same person to figure out when they have objects on them would be pretty useless.",
      "votes": null
    },
    {
      "id": "224074",
      "postDate": "09/25/2017 02:59:48",
      "content": "<p>TLDR: I think it's a feature, rather than a bug.</p>\n\n<p>I've done some clustering, and in my opinion, simple clustering won't predict contraband well anyway. The reason for that is that people have different postures in different scans, wear different clothing, exhale/inhale and move their limbs considerably. If you were to subtract one scan from another, the innocuous differences will often be greater than some of the contraband you are trying to find.</p>\n\n<p>That is not to say that some more sophisticated unsupervised learning algorithms might not be potentially useful in this competition.</p>\n\n<p>As far as I can tell, this competition is different from real life in three relevant aspects:</p>\n\n<ol>\n<li>There is only one scan per person in real life, as others pointed out.</li>\n<li>The contraband in the training dataset is not very diverse. It's impossible to think, in advance, of all possible types of shapes and materials someone might try to smuggle, let alone include such contraband in the dataset.</li>\n<li>The unlabeled dataset will be small. We have a few dozen subjects here, but realistically, one could find thousands, and if the DHS gave passengers the option to allow storage of their scans to help airport security, they could build a huge unlabeled dataset, with perhaps on the order of a million subjects.</li>\n</ol>\n\n<p>Since only unsupervised/semi-supervised learning takes advantage of unlabeled data, the unrealistically small size of the unlabeled dataset here already \"penalizes\" unsupervised/semi-supervised learning in this competition relative to its real-world applicability.</p>\n\n<p>Additionally, unsupervised learning would learn what a normal scan looks like, in order to spot deviations from the norm. It doesn't focus on the contraband, and therefore it should generalize better to contraband that's not in the dataset. This is obviously beneficial in real life.</p>\n\n<p>What about multiple scans of the same subject in the test set? You can look at it as effectively a proxy for a large unlabeled dataset (which we don't have access to here): If you had a large unlabeled dataset, you could, for example, easily find doppelgangers for passengers.</p>\n\n<p>If, despite the aforementioned innocuous differences between scans of the same person, some algorithm does manage to take advantage of such unlabeled data, it should also do well in practice, when the unlabeled dataset is large (and contraband is not what was in the dataset).</p>\n\n<p>In other words, I see this as a feature, rather than a bug. (That said, I agree that if Kaggle/DHS wish to change the rules during the competition, they need to announce this sooner rather than later)</p>\n\n<p>By the way, I asked about unsupervised/semi-supervised learning in the \"Welcome\" thread, and it was explicitly allowed.</p>",
      "rawMarkdown": "TLDR: I think it's a feature, rather than a bug.\n\nI've done some clustering, and in my opinion, simple clustering won't predict contraband well anyway. The reason for that is that people have different postures in different scans, wear different clothing, exhale/inhale and move their limbs considerably. If you were to subtract one scan from another, the innocuous differences will often be greater than some of the contraband you are trying to find.\n\nThat is not to say that some more sophisticated unsupervised learning algorithms might not be potentially useful in this competition.\n\nAs far as I can tell, this competition is different from real life in three relevant aspects:\n\n 1. There is only one scan per person in real life, as others pointed out.\n 2. The contraband in the training dataset is not very diverse. It's impossible to think, in advance, of all possible types of shapes and materials someone might try to smuggle, let alone include such contraband in the dataset.\n 3. The unlabeled dataset will be small. We have a few dozen subjects here, but realistically, one could find thousands, and if the DHS gave passengers the option to allow storage of their scans to help airport security, they could build a huge unlabeled dataset, with perhaps on the order of a million subjects.\n\nSince only unsupervised/semi-supervised learning takes advantage of unlabeled data, the unrealistically small size of the unlabeled dataset here already \"penalizes\" unsupervised/semi-supervised learning in this competition relative to its real-world applicability.\n\nAdditionally, unsupervised learning would learn what a normal scan looks like, in order to spot deviations from the norm. It doesn't focus on the contraband, and therefore it should generalize better to contraband that's not in the dataset. This is obviously beneficial in real life.\n\nWhat about multiple scans of the same subject in the test set? You can look at it as effectively a proxy for a large unlabeled dataset (which we don't have access to here): If you had a large unlabeled dataset, you could, for example, easily find doppelgangers for passengers.\n\nIf, despite the aforementioned innocuous differences between scans of the same person, some algorithm does manage to take advantage of such unlabeled data, it should also do well in practice, when the unlabeled dataset is large (and contraband is not what was in the dataset).\n\nIn other words, I see this as a feature, rather than a bug. (That said, I agree that if Kaggle/DHS wish to change the rules during the competition, they need to announce this sooner rather than later)\n\nBy the way, I asked about unsupervised/semi-supervised learning in the \"Welcome\" thread, and it was explicitly allowed.",
      "votes": null
    },
    {
      "id": "232940",
      "postDate": "10/18/2017 20:39:57",
      "content": "<p>We would consider such an approach to be semi-supervised learning, which is generally allowed in Kaggle competitions (including here). Note that your method has to be fully automated and locked in before the 2nd-stage test set is released (i.e. you may not manually draw clusters, manually set the number of clusters, manually register scans, etc.)</p>",
      "rawMarkdown": "We would consider such an approach to be semi-supervised learning, which is generally allowed in Kaggle competitions (including here). Note that your method has to be fully automated and locked in before the 2nd-stage test set is released (i.e. you may not manually draw clusters, manually set the number of clusters, manually register scans, etc.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 223050,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "09/21/2017 01:03:54",
      "content": "<p>Can we get an answer from an admin on this please? Are solutions that exploit the fact that the same volunteers may appear multiple times in Stage 2 admissible? </p>",
      "votes": null,
      "replies": [
        {
          "id": 232940,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "10/18/2017 20:39:57",
          "content": "<p>We would consider such an approach to be semi-supervised learning, which is generally allowed in Kaggle competitions (including here). Note that your method has to be fully automated and locked in before the 2nd-stage test set is released (i.e. you may not manually draw clusters, manually set the number of clusters, manually register scans, etc.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 223150,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "09/21/2017 08:47:57",
      "content": "<p>Branden,</p>\n\n<blockquote>\n  <p>is such a method allowed in this competition?</p>\n</blockquote>\n\n<p>It's my understanding that this is allowed, because you are allowed to use test set scans other than\nthe scan you are making predictions for. (Clustering is but a special case of unsupervised learning)</p>\n\n<p><a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/35116#199421\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/35116#199421</a></p>\n\n<blockquote>\n  <p>Since in a real life situation there will only be 1 scan of each person passing through security this method will be completely unusable. </p>\n</blockquote>\n\n<p>There are other factors to consider, such as that, in practice, they could create a much larger unlabeled dataset than we'll have here, which is a boon to semi-supervised methods. Additionally, semi-supervised methods may generalize better to previously unseen contraband.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224059,
      "author_name": "stella09",
      "author_url": "",
      "post_date": "09/25/2017 00:39:45",
      "content": "<p>I agree this could use some clarification from the admins.  I don't see anything in the rules that explicitly prohibits this, but I doubt this is the kind of solution the sponsor would want to see.  In real world deployment, each person is scanned only once.  Algorithms that compare multiple scans from the same person to figure out when they have objects on them would be pretty useless.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 224074,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "09/25/2017 02:59:48",
          "content": "<p>TLDR: I think it's a feature, rather than a bug.</p>\n\n<p>I've done some clustering, and in my opinion, simple clustering won't predict contraband well anyway. The reason for that is that people have different postures in different scans, wear different clothing, exhale/inhale and move their limbs considerably. If you were to subtract one scan from another, the innocuous differences will often be greater than some of the contraband you are trying to find.</p>\n\n<p>That is not to say that some more sophisticated unsupervised learning algorithms might not be potentially useful in this competition.</p>\n\n<p>As far as I can tell, this competition is different from real life in three relevant aspects:</p>\n\n<ol>\n<li>There is only one scan per person in real life, as others pointed out.</li>\n<li>The contraband in the training dataset is not very diverse. It's impossible to think, in advance, of all possible types of shapes and materials someone might try to smuggle, let alone include such contraband in the dataset.</li>\n<li>The unlabeled dataset will be small. We have a few dozen subjects here, but realistically, one could find thousands, and if the DHS gave passengers the option to allow storage of their scans to help airport security, they could build a huge unlabeled dataset, with perhaps on the order of a million subjects.</li>\n</ol>\n\n<p>Since only unsupervised/semi-supervised learning takes advantage of unlabeled data, the unrealistically small size of the unlabeled dataset here already \"penalizes\" unsupervised/semi-supervised learning in this competition relative to its real-world applicability.</p>\n\n<p>Additionally, unsupervised learning would learn what a normal scan looks like, in order to spot deviations from the norm. It doesn't focus on the contraband, and therefore it should generalize better to contraband that's not in the dataset. This is obviously beneficial in real life.</p>\n\n<p>What about multiple scans of the same subject in the test set? You can look at it as effectively a proxy for a large unlabeled dataset (which we don't have access to here): If you had a large unlabeled dataset, you could, for example, easily find doppelgangers for passengers.</p>\n\n<p>If, despite the aforementioned innocuous differences between scans of the same person, some algorithm does manage to take advantage of such unlabeled data, it should also do well in practice, when the unlabeled dataset is large (and contraband is not what was in the dataset).</p>\n\n<p>In other words, I see this as a feature, rather than a bug. (That said, I agree that if Kaggle/DHS wish to change the rules during the competition, they need to announce this sooner rather than later)</p>\n\n<p>By the way, I asked about unsupervised/semi-supervised learning in the \"Welcome\" thread, and it was explicitly allowed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "221365": "I understand that the volunteers used in Stage 2 will be different than those used in Stage 1, but will volunteers used in Stage 2 appear multiple times?\n\nI ask because if the same volunteers appear multiple times in Stage 2 one could conceivably:\n\n1. Use a clustering algorithm to separate the images into clusters (ideally one cluster for each volunteer)\n2. For each volunteer you could combine all of their images to create a \"normal\" image i.e. an image of what each volunteer would look like without any threats. Since each zone probably won't have a threat in the majority of images this could be achieved through some kind of averaging. Admittedly, this step is dependent on threats appearing about as frequently as they do in the training set. If every volunteer is loaded head to toe with threats in every image then this bit might not work.\n3. Use the difference of each image and the \"normal\" image to detect a threat. \n\nSince in a real life situation there will only be 1 scan of each person passing through security this method will be completely unusable. If volunteers appear multiple times in the test set is such a method allowed in this competition?",
    "223050": "Can we get an answer from an admin on this please? Are solutions that exploit the fact that the same volunteers may appear multiple times in Stage 2 admissible?",
    "223150": "Branden,\n\n&gt; is such a method allowed in this competition?\n\nIt's my understanding that this is allowed, because you are allowed to use test set scans other than\nthe scan you are making predictions for. (Clustering is but a special case of unsupervised learning)\n\nhttps://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/35116#199421\n\n&gt; Since in a real life situation there will only be 1 scan of each person passing through security this method will be completely unusable. \n\nThere are other factors to consider, such as that, in practice, they could create a much larger unlabeled dataset than we'll have here, which is a boon to semi-supervised methods. Additionally, semi-supervised methods may generalize better to previously unseen contraband.",
    "224059": "I agree this could use some clarification from the admins.  I don't see anything in the rules that explicitly prohibits this, but I doubt this is the kind of solution the sponsor would want to see.  In real world deployment, each person is scanned only once.  Algorithms that compare multiple scans from the same person to figure out when they have objects on them would be pretty useless.",
    "224074": "TLDR: I think it's a feature, rather than a bug.\n\nI've done some clustering, and in my opinion, simple clustering won't predict contraband well anyway. The reason for that is that people have different postures in different scans, wear different clothing, exhale/inhale and move their limbs considerably. If you were to subtract one scan from another, the innocuous differences will often be greater than some of the contraband you are trying to find.\n\nThat is not to say that some more sophisticated unsupervised learning algorithms might not be potentially useful in this competition.\n\nAs far as I can tell, this competition is different from real life in three relevant aspects:\n\n 1. There is only one scan per person in real life, as others pointed out.\n 2. The contraband in the training dataset is not very diverse. It's impossible to think, in advance, of all possible types of shapes and materials someone might try to smuggle, let alone include such contraband in the dataset.\n 3. The unlabeled dataset will be small. We have a few dozen subjects here, but realistically, one could find thousands, and if the DHS gave passengers the option to allow storage of their scans to help airport security, they could build a huge unlabeled dataset, with perhaps on the order of a million subjects.\n\nSince only unsupervised/semi-supervised learning takes advantage of unlabeled data, the unrealistically small size of the unlabeled dataset here already \"penalizes\" unsupervised/semi-supervised learning in this competition relative to its real-world applicability.\n\nAdditionally, unsupervised learning would learn what a normal scan looks like, in order to spot deviations from the norm. It doesn't focus on the contraband, and therefore it should generalize better to contraband that's not in the dataset. This is obviously beneficial in real life.\n\nWhat about multiple scans of the same subject in the test set? You can look at it as effectively a proxy for a large unlabeled dataset (which we don't have access to here): If you had a large unlabeled dataset, you could, for example, easily find doppelgangers for passengers.\n\nIf, despite the aforementioned innocuous differences between scans of the same person, some algorithm does manage to take advantage of such unlabeled data, it should also do well in practice, when the unlabeled dataset is large (and contraband is not what was in the dataset).\n\nIn other words, I see this as a feature, rather than a bug. (That said, I agree that if Kaggle/DHS wish to change the rules during the competition, they need to announce this sooner rather than later)\n\nBy the way, I asked about unsupervised/semi-supervised learning in the \"Welcome\" thread, and it was explicitly allowed.",
    "232940": "We would consider such an approach to be semi-supervised learning, which is generally allowed in Kaggle competitions (including here). Note that your method has to be fully automated and locked in before the 2nd-stage test set is released (i.e. you may not manually draw clusters, manually set the number of clusters, manually register scans, etc.)"
  },
  "source": "meta"
}