{
  "id": 40021,
  "title": "Psychology, OpenData and Image Processing",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/40021",
  "author_name": "",
  "post_date": "2017-09-26T11:20:32.370056100Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I am wondering how Psychology, Open datasets and Image processing can help in this problem.  I have few thoughts:</p>\n\n<ol>\n<li><p>One thing is clear, that Stage 2 real data will be different from what we have for now. The prominence of Zone 1 can only be possible due to hand-picked or formatted data and getting a small (real) sample from a very large dataset.</p></li>\n<li><p>All zones are not equal. Can we try to remove few zones all together from the calculation and see how much it would effect on overall prediction. or I can simply duplicate the results for Zones (14, 8, 2 and 1) and Zones (13, 11, 6 and 15) without loosing much.</p></li>\n<li><p>How many attacks are there and can we sort them based on body parts being used to carry these attacks? here is a Wikipedia article on <a href=\"https://en.wikipedia.org/wiki/Timeline_of_airliner_bombing_attacks\">number of airline attacks.</a></p></li>\n<li><p>How about comparing the given image with local images and getting the difference to find out if there exists any threats. I really like this paper from <a href=\"https://www.spiedigitallibrary.org/conference-proceedings-of-spie/9993/1/New-algorithm-for-detection-of-dangerous-objects-hidden-on-a/10.1117/12.2242720.short?SSO=1\">SPIE terahertz conference</a>.  </p></li>\n<li><p>Why we need security scanning images. The same passengers can be captured by dozen other cameras at the airport from taxi curb-dropping to local coffee shop. Can we compare the same person's image with non-scanning image and find the difference? or are they actually quite different when it comes to detecting threats?</p></li>\n<li><p>The actual possible ways someone \"should or can\" hide something is quite different from the actual ways people used to hide things. How do we make sure we are not overly biased in our thinking and get as much help from the data as can?</p></li>\n<li><p>Can weather plays an important part where you are wearing more in cold season, does it effect the images in hand?</p></li>\n<li><p>Do we have different patterns from different airports? Or do we know if all the data is coming from the same machines, brands, airports, population?</p></li>\n<li><p>How do we scale the same detection algorithm across the country and all over the world, are we assuming that we will find similar patterns everywhere?</p></li>\n<li><p>Any idea of current False Positives and False Negatives? What is the ideal % we like?</p></li>\n</ol>\n\n<p>While, I am exploring the data like anyone else, please help me if anyone can answer some of these questions or shed some light to make problem more understandable for us.</p>",
  "messages": [
    {
      "id": "224449",
      "postDate": "09/26/2017 11:20:32",
      "content": "<p>I am wondering how Psychology, Open datasets and Image processing can help in this problem.  I have few thoughts:</p>\n\n<ol>\n<li><p>One thing is clear, that Stage 2 real data will be different from what we have for now. The prominence of Zone 1 can only be possible due to hand-picked or formatted data and getting a small (real) sample from a very large dataset.</p></li>\n<li><p>All zones are not equal. Can we try to remove few zones all together from the calculation and see how much it would effect on overall prediction. or I can simply duplicate the results for Zones (14, 8, 2 and 1) and Zones (13, 11, 6 and 15) without loosing much.</p></li>\n<li><p>How many attacks are there and can we sort them based on body parts being used to carry these attacks? here is a Wikipedia article on <a href=\"https://en.wikipedia.org/wiki/Timeline_of_airliner_bombing_attacks\">number of airline attacks.</a></p></li>\n<li><p>How about comparing the given image with local images and getting the difference to find out if there exists any threats. I really like this paper from <a href=\"https://www.spiedigitallibrary.org/conference-proceedings-of-spie/9993/1/New-algorithm-for-detection-of-dangerous-objects-hidden-on-a/10.1117/12.2242720.short?SSO=1\">SPIE terahertz conference</a>.  </p></li>\n<li><p>Why we need security scanning images. The same passengers can be captured by dozen other cameras at the airport from taxi curb-dropping to local coffee shop. Can we compare the same person's image with non-scanning image and find the difference? or are they actually quite different when it comes to detecting threats?</p></li>\n<li><p>The actual possible ways someone \"should or can\" hide something is quite different from the actual ways people used to hide things. How do we make sure we are not overly biased in our thinking and get as much help from the data as can?</p></li>\n<li><p>Can weather plays an important part where you are wearing more in cold season, does it effect the images in hand?</p></li>\n<li><p>Do we have different patterns from different airports? Or do we know if all the data is coming from the same machines, brands, airports, population?</p></li>\n<li><p>How do we scale the same detection algorithm across the country and all over the world, are we assuming that we will find similar patterns everywhere?</p></li>\n<li><p>Any idea of current False Positives and False Negatives? What is the ideal % we like?</p></li>\n</ol>\n\n<p>While, I am exploring the data like anyone else, please help me if anyone can answer some of these questions or shed some light to make problem more understandable for us.</p>",
      "rawMarkdown": "I am wondering how Psychology, Open datasets and Image processing can help in this problem.  I have few thoughts:\n\n1. One thing is clear, that Stage 2 real data will be different from what we have for now. The prominence of Zone 1 can only be possible due to hand-picked or formatted data and getting a small (real) sample from a very large dataset.\n\n2. All zones are not equal. Can we try to remove few zones all together from the calculation and see how much it would effect on overall prediction. or I can simply duplicate the results for Zones (14, 8, 2 and 1) and Zones (13, 11, 6 and 15) without loosing much.\n\n3. How many attacks are there and can we sort them based on body parts being used to carry these attacks? here is a Wikipedia article on [number of airline attacks.][1]\n\n4. How about comparing the given image with local images and getting the difference to find out if there exists any threats. I really like this paper from [SPIE terahertz conference][2].  \n\n5. Why we need security scanning images. The same passengers can be captured by dozen other cameras at the airport from taxi curb-dropping to local coffee shop. Can we compare the same person's image with non-scanning image and find the difference? or are they actually quite different when it comes to detecting threats?\n\n6.  The actual possible ways someone \"should or can\" hide something is quite different from the actual ways people used to hide things. How do we make sure we are not overly biased in our thinking and get as much help from the data as can?\n\n7. Can weather plays an important part where you are wearing more in cold season, does it effect the images in hand?\n\n8. Do we have different patterns from different airports? Or do we know if all the data is coming from the same machines, brands, airports, population?\n\n9. How do we scale the same detection algorithm across the country and all over the world, are we assuming that we will find similar patterns everywhere?\n\n10. Any idea of current False Positives and False Negatives? What is the ideal % we like?\n\n\nWhile, I am exploring the data like anyone else, please help me if anyone can answer some of these questions or shed some light to make problem more understandable for us.\n \n\n\n  [1]: https://en.wikipedia.org/wiki/Timeline_of_airliner_bombing_attacks\n  [2]: https://www.spiedigitallibrary.org/conference-proceedings-of-spie/9993/1/New-algorithm-for-detection-of-dangerous-objects-hidden-on-a/10.1117/12.2242720.short?SSO=1",
      "votes": null
    },
    {
      "id": "231426",
      "postDate": "10/14/2017 19:06:21",
      "content": "<p>Forgive me if I'm misunderstanding what your true point is -- I easily could be -- but since nobody else has responded, I don't mind speaking up.</p>\n\n<p>I have no idea whatsoever why they spend so much money on such cool toys.  Honestly, regardless of their usefulness without improved software, those machines are quite a nice example of engineering.  Moreover, I applaud any agency that's willing to seek out improvements (or even just verify that their primary vendor is still the best one for the job.)</p>\n\n<p>That said, I doubt the rules are going to change.  To use a metaphor: personally, I think requiring zones at all is a catastrophic error that's basically hobbling us before the race even starts.  The labels for the training data we have not being 100% reliable can be tested for by training a solution, trying to identify incorrect labels, and then throwing it away to train a new one on the corrected data.  That's just the sort of thing that happens outside of class rooms, you know.</p>",
      "rawMarkdown": "Forgive me if I'm misunderstanding what your true point is -- I easily could be -- but since nobody else has responded, I don't mind speaking up.\n\nI have no idea whatsoever why they spend so much money on such cool toys.  Honestly, regardless of their usefulness without improved software, those machines are quite a nice example of engineering.  Moreover, I applaud any agency that's willing to seek out improvements (or even just verify that their primary vendor is still the best one for the job.)\n\nThat said, I doubt the rules are going to change.  To use a metaphor: personally, I think requiring zones at all is a catastrophic error that's basically hobbling us before the race even starts.  The labels for the training data we have not being 100% reliable can be tested for by training a solution, trying to identify incorrect labels, and then throwing it away to train a new one on the corrected data.  That's just the sort of thing that happens outside of class rooms, you know.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 231426,
      "author_name": "mmiron",
      "author_url": "",
      "post_date": "10/14/2017 19:06:21",
      "content": "<p>Forgive me if I'm misunderstanding what your true point is -- I easily could be -- but since nobody else has responded, I don't mind speaking up.</p>\n\n<p>I have no idea whatsoever why they spend so much money on such cool toys.  Honestly, regardless of their usefulness without improved software, those machines are quite a nice example of engineering.  Moreover, I applaud any agency that's willing to seek out improvements (or even just verify that their primary vendor is still the best one for the job.)</p>\n\n<p>That said, I doubt the rules are going to change.  To use a metaphor: personally, I think requiring zones at all is a catastrophic error that's basically hobbling us before the race even starts.  The labels for the training data we have not being 100% reliable can be tested for by training a solution, trying to identify incorrect labels, and then throwing it away to train a new one on the corrected data.  That's just the sort of thing that happens outside of class rooms, you know.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "224449": "I am wondering how Psychology, Open datasets and Image processing can help in this problem.  I have few thoughts:\n\n1. One thing is clear, that Stage 2 real data will be different from what we have for now. The prominence of Zone 1 can only be possible due to hand-picked or formatted data and getting a small (real) sample from a very large dataset.\n\n2. All zones are not equal. Can we try to remove few zones all together from the calculation and see how much it would effect on overall prediction. or I can simply duplicate the results for Zones (14, 8, 2 and 1) and Zones (13, 11, 6 and 15) without loosing much.\n\n3. How many attacks are there and can we sort them based on body parts being used to carry these attacks? here is a Wikipedia article on [number of airline attacks.][1]\n\n4. How about comparing the given image with local images and getting the difference to find out if there exists any threats. I really like this paper from [SPIE terahertz conference][2].  \n\n5. Why we need security scanning images. The same passengers can be captured by dozen other cameras at the airport from taxi curb-dropping to local coffee shop. Can we compare the same person's image with non-scanning image and find the difference? or are they actually quite different when it comes to detecting threats?\n\n6.  The actual possible ways someone \"should or can\" hide something is quite different from the actual ways people used to hide things. How do we make sure we are not overly biased in our thinking and get as much help from the data as can?\n\n7. Can weather plays an important part where you are wearing more in cold season, does it effect the images in hand?\n\n8. Do we have different patterns from different airports? Or do we know if all the data is coming from the same machines, brands, airports, population?\n\n9. How do we scale the same detection algorithm across the country and all over the world, are we assuming that we will find similar patterns everywhere?\n\n10. Any idea of current False Positives and False Negatives? What is the ideal % we like?\n\n\nWhile, I am exploring the data like anyone else, please help me if anyone can answer some of these questions or shed some light to make problem more understandable for us.\n \n\n\n  [1]: https://en.wikipedia.org/wiki/Timeline_of_airliner_bombing_attacks\n  [2]: https://www.spiedigitallibrary.org/conference-proceedings-of-spie/9993/1/New-algorithm-for-detection-of-dangerous-objects-hidden-on-a/10.1117/12.2242720.short?SSO=1",
    "231426": "Forgive me if I'm misunderstanding what your true point is -- I easily could be -- but since nobody else has responded, I don't mind speaking up.\n\nI have no idea whatsoever why they spend so much money on such cool toys.  Honestly, regardless of their usefulness without improved software, those machines are quite a nice example of engineering.  Moreover, I applaud any agency that's willing to seek out improvements (or even just verify that their primary vendor is still the best one for the job.)\n\nThat said, I doubt the rules are going to change.  To use a metaphor: personally, I think requiring zones at all is a catastrophic error that's basically hobbling us before the race even starts.  The labels for the training data we have not being 100% reliable can be tested for by training a solution, trying to identify incorrect labels, and then throwing it away to train a new one on the corrected data.  That's just the sort of thing that happens outside of class rooms, you know."
  },
  "source": "meta"
}