{
  "id": 45795,
  "title": "Any success using a3d or a3daps images?",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/45795",
  "author_name": "Branden Murray",
  "post_date": "2017-12-15T23:50:20.750000",
  "votes": 2,
  "comment_count": 32,
  "views": 0,
  "content": "<p>Did anyone have success using the .a3d or .a3daps (or even .ahi) images? I spent a decent amount of time trying to get good scores out of them, but I think the best I got out of .a3daps was ~0.12 (using 'proper' validation, i.e. separating folds by subject) and for .a3d my best was probably around 0.17. Was anyone able to get less than &lt;0.10 or 0.05 on the Public LB using these?</p>",
  "messages": [
    {
      "id": 258787,
      "postDate": "2017-12-17T03:00:43.937Z",
      "content": "<p>We observed very quickly that there was a big difference in the score we could achieve with the .a3daps/a3d compared with the .aps data trained using the same model, and we could not figure out why. Fortunately we started out with the .aps data and our first submission achieved a score of 0.023 putting us in 12th place in stage 1. Knowing the aps data had only 16 view points and the .a3daps had 64 viewpoints we expected the same model trained using the same approach on the a3daps data with more view points should achieve a much better score but this was not the case.  We checked mean, standard deviation, checked if subjects were mixed up  btw aps and a3daps e.t.c. We tried multiple times  but no success, our single best model with .aps achieves a stage 1 score of ~0.012 and stage 2 score of ~0.035 while the same model with the .a3daps  achieves a score of ~0.03 stage1  and ~0.08  stage2.  Another observation looking at the images is that the a3daps images look much noisier than the aps data. I think there are some issues with the processing of the a3d, and .a3daps, and knowing the .a3daps was derived from the a3d while the aps was acquired separately gave us comfort in pursuing models that used the .aps data know that logically these models should not be outperforming the  .a3d and a3daps models.</p>",
      "rawMarkdown": "We observed very quickly that there was a big difference in the score we could achieve with the .a3daps/a3d compared with the .aps data trained using the same model, and we could not figure out why. Fortunately we started out with the .aps data and our first submission achieved a score of 0.023 putting us in 12th place in stage 1. Knowing the aps data had only 16 view points and the .a3daps had 64 viewpoints we expected the same model trained using the same approach on the a3daps data with more view points should achieve a much better score but this was not the case.  We checked mean, standard deviation, checked if subjects were mixed up  btw aps and a3daps e.t.c. We tried multiple times  but no success, our single best model with .aps achieves a stage 1 score of ~0.012 and stage 2 score of ~0.035 while the same model with the .a3daps  achieves a score of ~0.03 stage1  and ~0.08  stage2.  Another observation looking at the images is that the a3daps images look much noisier than the aps data. I think there are some issues with the processing of the a3d, and .a3daps, and knowing the .a3daps was derived from the a3d while the aps was acquired separately gave us comfort in pursuing models that used the .aps data know that logically these models should not be outperforming the  .a3d and a3daps models.",
      "votes": 4,
      "replies": [
        {
          "id": 258824,
          "postDate": "2017-12-17T05:05:26.043Z",
          "content": "<p>Thanks, David, and congrats! I'm curious if you, and anyone who did well in this contest, used</p>\n\n<ol>\n<li>pretrained networks or external data</li>\n<li>did \"semi-supervised\" learning, including any kind of subject identity discovery on Stage2 data?</li>\n</ol>\n\n<p>I mentioned in another thread one heuristic that I think would have worked very well:  \"If the same subject has a bomb in the same zone all the time, it's probably a false positive\". I think I should have pursued \"semi-supervised\" approaches more, in retrospect.</p>",
          "rawMarkdown": "Thanks, David, and congrats! I'm curious if you, and anyone who did well in this contest, used\n\n 1. pretrained networks or external data\n 2. did \"semi-supervised\" learning, including any kind of subject identity discovery on Stage2 data?\n\nI mentioned in another thread one heuristic that I think would have worked very well:  \"If the same subject has a bomb in the same zone all the time, it's probably a false positive\". I think I should have pursued \"semi-supervised\" approaches more, in retrospect."
        },
        {
          "id": 258838,
          "postDate": "2017-12-17T05:31:12.540Z",
          "content": "<p>Hi Oleg, We did not use external data but we did use keras imagenet pretrained networks, Xception, Inceptionv3,  Resnet50, Inception_Resnet, and VGG19 as components of our model pipeline, our model with Resnet50 was the best while as expected VGG19 was the worst but added value in our ensemble, we did not do semi-supervised or any kind of subject identity discovery on stage 2 data.</p>",
          "rawMarkdown": "Hi Oleg, We did not use external data but we did use keras imagenet pretrained networks, Xception, Inceptionv3,  Resnet50, Inception_Resnet, and VGG19 as components of our model pipeline, our model with Resnet50 was the best while as expected VGG19 was the worst but added value in our ensemble, we did not do semi-supervised or any kind of subject identity discovery on stage 2 data.\n",
          "votes": 2
        },
        {
          "id": 258847,
          "postDate": "2017-12-17T06:00:11.463Z",
          "content": "<p>Congrats David and thank you for providing all of your valuable insights! </p>\n\n<p>I still find it quite shocking/unbelievable that using only aps images could produce such a high score/low loss rate. I too started using aps images but could never get below .03 on stage 1 while getting .005 with a3d only. Just based on my own observations I never came across a threat that was not visible with a3d yet I oftentimes came across threats that I simply could not find visibly in the aps images from any of the 16 views. </p>\n\n<p>For this reason, while aps may have \"won\" in this competition with this specific mix of training set and test sets for stage 1/2, but I'm not sure if the aps images would be very reliable in practice since one could argue that there is a greater chance of false negatives which would be a lot worse than a few more false positives (for good reason). Interested to hear your thoughts on that as well!</p>",
          "rawMarkdown": "Congrats David and thank you for providing all of your valuable insights! \n\nI still find it quite shocking/unbelievable that using only aps images could produce such a high score/low loss rate. I too started using aps images but could never get below .03 on stage 1 while getting .005 with a3d only. Just based on my own observations I never came across a threat that was not visible with a3d yet I oftentimes came across threats that I simply could not find visibly in the aps images from any of the 16 views. \n\nFor this reason, while aps may have \"won\" in this competition with this specific mix of training set and test sets for stage 1/2, but I'm not sure if the aps images would be very reliable in practice since one could argue that there is a greater chance of false negatives which would be a lot worse than a few more false positives (for good reason). Interested to hear your thoughts on that as well!"
        },
        {
          "id": 258855,
          "postDate": "2017-12-17T06:49:37.960Z",
          "content": "<p>I agree that using just the .aps increases the chances of false negatives, the threats on the forearms inner ankles and groin are particularly difficult, some of the threats are only visible in 2 frames of the aps data, however we had 2 models that used the  a3daps data in our ensemble to compensate for potential false negatives. I think a robust system will need much more training samples and subject variety. As far as not seeing the threats in the .aps we created frame level labels for all the images in the stage 1 aps data, we found some label errors but did find threats that were not visible.</p>",
          "rawMarkdown": "I agree that using just the .aps increases the chances of false negatives, the threats on the forearms inner ankles and groin are particularly difficult, some of the threats are only visible in 2 frames of the aps data, however we had 2 models that used the  a3daps data in our ensemble to compensate for potential false negatives. I think a robust system will need much more training samples and subject variety. As far as not seeing the threats in the .aps we created frame level labels for all the images in the stage 1 aps data, we found some label errors but did find threats that were not visible.",
          "votes": 1
        }
      ]
    },
    {
      "id": 258371,
      "postDate": "2017-12-16T01:56:08.710Z",
      "content": "<p>I trained (from scratch) InceptionResnetV2 on a set of 299x299 mosaic-images each containing aps/a3daps crops based on multiple views on a zone. Did not use the a3d-images.</p>\n\n<p>In hindsight I should have gone for a more specialized net like MVCNN or Conv3D I guess. </p>",
      "rawMarkdown": "I trained (from scratch) InceptionResnetV2 on a set of 299x299 mosaic-images each containing aps/a3daps crops based on multiple views on a zone. Did not use the a3d-images.\n\nIn hindsight I should have gone for a more specialized net like MVCNN or Conv3D I guess. ",
      "votes": 1
    },
    {
      "id": 258338,
      "postDate": "2017-12-16T00:13:43.853Z",
      "content": "<p>I also used only a3d. </p>",
      "rawMarkdown": "I also used only a3d. ",
      "votes": 1
    },
    {
      "id": 258759,
      "postDate": "2017-12-17T00:53:18.017Z",
      "content": "<p>I did something a bit different than most.  Using .a3d files as input, I \"unwrapped\" a passenger's body surface into a 2 dimensional representation using a cylindrical coordinate system for each body-part.  The 7 body-parts were the 2 legs, 1 trunk, 2 biceps, and 2 forearms.  Cylindrical coordinate systems were registered to body-parts using estimates of the positions of wrists, elbows, shoulders, feet, the buttock/leg meeting point, and the center of mass.  A cylinder's axis was allowed to curve but slices were parallel.  Surface reflectivity, radius and thickness became 3 image channels.  All 7 body-part images were joined into one. Variability due to differing body types was subtracted away using a nearest neighbors approach.  </p>\n\n<p>Contraband really stands out in the resulting images as areas of saturated color on a gray background.  Using only color histograms and logistic regression, I placed in the top 20% on the initial public leaderboard.  So even without using shape and texture information in the processed images, the technique already performed well.</p>\n\n<p>I ran out of time for a few reasons:  this is my first Python project, I started with only 2 months to go, I have no CNN experience, yada yada. The next thing I was going to do was try the simplest sort of CNN transfer learning.  I may still do it, and make a late submission.</p>",
      "rawMarkdown": "I did something a bit different than most.  Using .a3d files as input, I \"unwrapped\" a passenger's body surface into a 2 dimensional representation using a cylindrical coordinate system for each body-part.  The 7 body-parts were the 2 legs, 1 trunk, 2 biceps, and 2 forearms.  Cylindrical coordinate systems were registered to body-parts using estimates of the positions of wrists, elbows, shoulders, feet, the buttock/leg meeting point, and the center of mass.  A cylinder's axis was allowed to curve but slices were parallel.  Surface reflectivity, radius and thickness became 3 image channels.  All 7 body-part images were joined into one. Variability due to differing body types was subtracted away using a nearest neighbors approach.  \n\nContraband really stands out in the resulting images as areas of saturated color on a gray background.  Using only color histograms and logistic regression, I placed in the top 20% on the initial public leaderboard.  So even without using shape and texture information in the processed images, the technique already performed well.\n\nI ran out of time for a few reasons:  this is my first Python project, I started with only 2 months to go, I have no CNN experience, yada yada. The next thing I was going to do was try the simplest sort of CNN transfer learning.  I may still do it, and make a late submission.",
      "votes": 2,
      "replies": [
        {
          "id": 258765,
          "postDate": "2017-12-17T01:21:36.927Z",
          "content": "<p>Cool! Can you explain in more detail how you got the reflectivity, surface height and thickness out of you unwrapped body parts?</p>",
          "rawMarkdown": "Cool! Can you explain in more detail how you got the reflectivity, surface height and thickness out of you unwrapped body parts?",
          "votes": 1
        },
        {
          "id": 258780,
          "postDate": "2017-12-17T02:18:24.083Z",
          "content": "<p>Thank you!  Once a cylindrical coordinate transform is performed for a body-part, I have an array with dimensions (r, \\theta, w), where r is radius, \\theta is the angular coordinate,  and w is the axial coordinate.  For each (\\theta, w) coordinate, I get the peaks along the r direction.  The tallest peak's height, position, and width are the surface reflectivity, radius (AKA surface height), and thickness, respectively, for coordinate (\\theta, w).  That's not a precise description, by the way, since the determination of peak height, position and width (I.e. zeroeth, first and second moments) needs to be tweeked to give good results.  Anyway, I end up with surface reflectivity, radius and thickness for each (\\theta, w) coordinate.</p>",
          "rawMarkdown": "Thank you!  Once a cylindrical coordinate transform is performed for a body-part, I have an array with dimensions (r, \\theta, w), where r is radius, \\theta is the angular coordinate,  and w is the axial coordinate.  For each (\\theta, w) coordinate, I get the peaks along the r direction.  The tallest peak's height, position, and width are the surface reflectivity, radius (AKA surface height), and thickness, respectively, for coordinate (\\theta, w).  That's not a precise description, by the way, since the determination of peak height, position and width (I.e. zeroeth, first and second moments) needs to be tweeked to give good results.  Anyway, I end up with surface reflectivity, radius and thickness for each (\\theta, w) coordinate."
        },
        {
          "id": 258850,
          "postDate": "2017-12-17T06:19:12.597Z",
          "content": "<p>@Nathaniel This is a cool idea!</p>\n\n<p>I too used the a3d images not as 3d images but as 2d images, however I treated them instead like frames from a video so my model basically is traveling up through the body one \"frame\" at a time. If you flip through the 3d image slices you can pretty easily spot the threats in their evolution as they appear to be \"growing\". This also had added benefits of making the training set exponentially bigger, using large batch sizes, and also allowed for pre-trained nets which wouldn't otherwise be possible with 3d convnets.</p>",
          "rawMarkdown": "@Nathaniel This is a cool idea!\n\nI too used the a3d images not as 3d images but as 2d images, however I treated them instead like frames from a video so my model basically is traveling up through the body one \"frame\" at a time. If you flip through the 3d image slices you can pretty easily spot the threats in their evolution as they appear to be \"growing\". This also had added benefits of making the training set exponentially bigger, using large batch sizes, and also allowed for pre-trained nets which wouldn't otherwise be possible with 3d convnets.",
          "votes": 1
        }
      ]
    },
    {
      "id": 258376,
      "postDate": "2017-12-16T02:25:30.883Z",
      "content": "<p>In my opinion, you need both. Sometimes you see the threat on APS, but not A3DAPS. Less commonly, it's the other way around. I fed both of them to the network, as separate channels. No need to choose.</p>",
      "rawMarkdown": "In my opinion, you need both. Sometimes you see the threat on APS, but not A3DAPS. Less commonly, it's the other way around. I fed both of them to the network, as separate channels. No need to choose.",
      "votes": 2,
      "replies": [
        {
          "id": 258383,
          "postDate": "2017-12-16T02:33:28.023Z",
          "content": "<p>I observed the same, and also passed both aps and a3daps data as separate channels, but strangely my final loss didn't change. Did you see improvement over just using aps?</p>",
          "rawMarkdown": "I observed the same, and also passed both aps and a3daps data as separate channels, but strangely my final loss didn't change. Did you see improvement over just using aps?"
        },
        {
          "id": 258384,
          "postDate": "2017-12-16T02:38:46.640Z",
          "content": "<p>I did not experiment with just using one. I decided early on that I'd use both.</p>",
          "rawMarkdown": "I did not experiment with just using one. I decided early on that I'd use both."
        },
        {
          "id": 258454,
          "postDate": "2017-12-16T05:39:11.343Z",
          "content": "<p>I concatenated aps &amp; a3daps into one image and got higher accuracy than when I used aps or a3daps separately (0.03 logloss using this strategy on stage 1).</p>",
          "rawMarkdown": "I concatenated aps &amp; a3daps into one image and got higher accuracy than when I used aps or a3daps separately (0.03 logloss using this strategy on stage 1)."
        }
      ]
    },
    {
      "id": 258326,
      "postDate": "2017-12-15T23:50:20.750Z",
      "content": "<p>Did anyone have success using the .a3d or .a3daps (or even .ahi) images? I spent a decent amount of time trying to get good scores out of them, but I think the best I got out of .a3daps was ~0.12 (using 'proper' validation, i.e. separating folds by subject) and for .a3d my best was probably around 0.17. Was anyone able to get less than &lt;0.10 or 0.05 on the Public LB using these?</p>",
      "rawMarkdown": "Did anyone have success using the .a3d or .a3daps (or even .ahi) images? I spent a decent amount of time trying to get good scores out of them, but I think the best I got out of .a3daps was ~0.12 (using 'proper' validation, i.e. separating folds by subject) and for .a3d my best was probably around 0.17. Was anyone able to get less than &lt;0.10 or 0.05 on the Public LB using these?",
      "votes": 2
    },
    {
      "id": 258435,
      "postDate": "2017-12-16T04:33:49.090Z",
      "content": "<p>I used only APS at full scale(512x660 I think). However, my model consisted of two separate stages: the first stage detected the contrabands and the second figured out which zone the contraband was in. I did try to use a3daps but that was too slow in the first phase for me to make meaningful progress so I gave up on it even though I noticed my contraband detection model was missing around 2-3% of the contrabands. Infact those 2-3% were not even visible visually when I looked at the images.</p>",
      "rawMarkdown": "I used only APS at full scale(512x660 I think). However, my model consisted of two separate stages: the first stage detected the contrabands and the second figured out which zone the contraband was in. I did try to use a3daps but that was too slow in the first phase for me to make meaningful progress so I gave up on it even though I noticed my contraband detection model was missing around 2-3% of the contrabands. Infact those 2-3% were not even visible visually when I looked at the images."
    },
    {
      "id": 258420,
      "postDate": "2017-12-16T03:45:18.660Z",
      "content": "<p>I have a theory regarding why models based on APS images performed better. See <a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/45814\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/45814</a>. In short, I suspect that labels were generated by looking at APS images, :-).</p>",
      "rawMarkdown": "I have a theory regarding why models based on APS images performed better. See https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/45814. In short, I suspect that labels were generated by looking at APS images, :-)."
    },
    {
      "id": 258403,
      "postDate": "2017-12-16T03:14:16.707Z",
      "content": "<p>Interesting! We only used A3DAPS with 2D CNN models like resnet/densenet. Should have used APS!</p>",
      "rawMarkdown": "Interesting! We only used A3DAPS with 2D CNN models like resnet/densenet. Should have used APS!"
    },
    {
      "id": 258347,
      "postDate": "2017-12-16T00:35:35.333Z",
      "content": "<p>My model overfit more easily when a3daps was involved.</p>",
      "rawMarkdown": "My model overfit more easily when a3daps was involved.",
      "replies": [
        {
          "id": 258350,
          "postDate": "2017-12-16T00:38:37.577Z",
          "content": "<p>I thought aps data was too coarse, and didn't even bother.</p>",
          "rawMarkdown": "I thought aps data was too coarse, and didn't even bother."
        }
      ]
    },
    {
      "id": 258343,
      "postDate": "2017-12-16T00:24:15.023Z",
      "content": "<p>I only used a3d.</p>",
      "rawMarkdown": "I only used a3d."
    },
    {
      "id": 258334,
      "postDate": "2017-12-16T00:11:03.080Z",
      "content": "<p>I was only using .a3d data. With the hinder sight, looks that this was a wrong decision, :-)</p>",
      "rawMarkdown": "I was only using .a3d data. With the hinder sight, looks that this was a wrong decision, :-)",
      "replies": [
        {
          "id": 258455,
          "postDate": "2017-12-16T05:41:20.033Z",
          "content": "<p>Yep I made this same mistake. My first model used both aps/a3daps and would have done better in stage 2 but since my a3d only model was doing much better in stage 1 so I abandoned it. </p>",
          "rawMarkdown": "Yep I made this same mistake. My first model used both aps/a3daps and would have done better in stage 2 but since my a3d only model was doing much better in stage 1 so I abandoned it. "
        }
      ]
    },
    {
      "id": 258332,
      "postDate": "2017-12-16T00:06:06.260Z",
      "content": "<p>I'm also curious about this. I trained a few threat segmentation models using both aps and a3daps images as input, but didn't see any improvement over segmenting from just aps images.</p>",
      "rawMarkdown": "I'm also curious about this. I trained a few threat segmentation models using both aps and a3daps images as input, but didn't see any improvement over segmenting from just aps images.",
      "replies": [
        {
          "id": 258339,
          "postDate": "2017-12-16T00:15:51.243Z",
          "content": "<p>I only used a3d (scaled down by 1/4 and simple Conv3D) and just one model. I wish I added another model with aps :(</p>",
          "rawMarkdown": "I only used a3d (scaled down by 1/4 and simple Conv3D) and just one model. I wish I added another model with aps :("
        },
        {
          "id": 258344,
          "postDate": "2017-12-16T00:30:04.303Z",
          "content": "<p>Just for clarification, do you mean you used images that were 3/4 the size of the original or 1/4 the size of the original? I think I tried 3/4, but that didn't work. My architectures were probably crap.</p>",
          "rawMarkdown": "Just for clarification, do you mean you used images that were 3/4 the size of the original or 1/4 the size of the original? I think I tried 3/4, but that didn't work. My architectures were probably crap."
        },
        {
          "id": 258352,
          "postDate": "2017-12-16T00:39:14.613Z",
          "content": "<p>Scaled down to 1/4 on each dimension - so it becomes 128x128x165</p>",
          "rawMarkdown": "Scaled down to 1/4 on each dimension - so it becomes 128x128x165",
          "votes": 2
        },
        {
          "id": 258353,
          "postDate": "2017-12-16T00:41:21.443Z",
          "content": "<p>oh wow.  Congrats on the results.  I tried Conv3D and was not converging so I went with MVCNN.</p>",
          "rawMarkdown": "oh wow.  Congrats on the results.  I tried Conv3D and was not converging so I went with MVCNN."
        },
        {
          "id": 258358,
          "postDate": "2017-12-16T00:54:36.083Z",
          "content": "<p>I found keeping the original resolution increased my performance considerably.</p>",
          "rawMarkdown": "I found keeping the original resolution increased my performance considerably.",
          "votes": 1
        },
        {
          "id": 258361,
          "postDate": "2017-12-16T00:59:44.273Z",
          "content": "<p>Interesting, I thought the scans were just too big so I didn't even try Conv3D -- what did your architectures look like? It seems like it would be pretty easy to overfit, which was why I mostly stuck to MVCNN-like models.</p>",
          "rawMarkdown": "Interesting, I thought the scans were just too big so I didn't even try Conv3D -- what did your architectures look like? It seems like it would be pretty easy to overfit, which was why I mostly stuck to MVCNN-like models.",
          "votes": 1
        },
        {
          "id": 258372,
          "postDate": "2017-12-16T01:59:37.967Z",
          "content": "<p>It's quite simple architecture, so each epoch just takes 10 minutes on single 1080 Ti. I used 5 blocks of resnet like conv3d. First 3 blocks are shared for all zones, and last 2 blocks are per zone. Basic rotation/zoom/shift augmentation/flip was used. </p>",
          "rawMarkdown": "It's quite simple architecture, so each epoch just takes 10 minutes on single 1080 Ti. I used 5 blocks of resnet like conv3d. First 3 blocks are shared for all zones, and last 2 blocks are per zone. Basic rotation/zoom/shift augmentation/flip was used. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 258359,
      "postDate": "2017-12-16T00:55:56.367Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 258787,
      "author_name": "DavidGbodiOdaibo",
      "author_url": "",
      "post_date": "2017-12-17T03:00:43.937000",
      "content": "<p>We observed very quickly that there was a big difference in the score we could achieve with the .a3daps/a3d compared with the .aps data trained using the same model, and we could not figure out why. Fortunately we started out with the .aps data and our first submission achieved a score of 0.023 putting us in 12th place in stage 1. Knowing the aps data had only 16 view points and the .a3daps had 64 viewpoints we expected the same model trained using the same approach on the a3daps data with more view points should achieve a much better score but this was not the case.  We checked mean, standard deviation, checked if subjects were mixed up  btw aps and a3daps e.t.c. We tried multiple times  but no success, our single best model with .aps achieves a stage 1 score of ~0.012 and stage 2 score of ~0.035 while the same model with the .a3daps  achieves a score of ~0.03 stage1  and ~0.08  stage2.  Another observation looking at the images is that the a3daps images look much noisier than the aps data. I think there are some issues with the processing of the a3d, and .a3daps, and knowing the .a3daps was derived from the a3d while the aps was acquired separately gave us comfort in pursuing models that used the .aps data know that logically these models should not be outperforming the  .a3d and a3daps models.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 258824,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2017-12-17T05:05:26.043000",
          "content": "<p>Thanks, David, and congrats! I'm curious if you, and anyone who did well in this contest, used</p>\n\n<ol>\n<li>pretrained networks or external data</li>\n<li>did \"semi-supervised\" learning, including any kind of subject identity discovery on Stage2 data?</li>\n</ol>\n\n<p>I mentioned in another thread one heuristic that I think would have worked very well:  \"If the same subject has a bomb in the same zone all the time, it's probably a false positive\". I think I should have pursued \"semi-supervised\" approaches more, in retrospect.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258838,
          "author_name": "DavidGbodiOdaibo",
          "author_url": "",
          "post_date": "2017-12-17T05:31:12.540000",
          "content": "<p>Hi Oleg, We did not use external data but we did use keras imagenet pretrained networks, Xception, Inceptionv3,  Resnet50, Inception_Resnet, and VGG19 as components of our model pipeline, our model with Resnet50 was the best while as expected VGG19 was the worst but added value in our ensemble, we did not do semi-supervised or any kind of subject identity discovery on stage 2 data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 258847,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2017-12-17T06:00:11.463000",
          "content": "<p>Congrats David and thank you for providing all of your valuable insights! </p>\n\n<p>I still find it quite shocking/unbelievable that using only aps images could produce such a high score/low loss rate. I too started using aps images but could never get below .03 on stage 1 while getting .005 with a3d only. Just based on my own observations I never came across a threat that was not visible with a3d yet I oftentimes came across threats that I simply could not find visibly in the aps images from any of the 16 views. </p>\n\n<p>For this reason, while aps may have \"won\" in this competition with this specific mix of training set and test sets for stage 1/2, but I'm not sure if the aps images would be very reliable in practice since one could argue that there is a greater chance of false negatives which would be a lot worse than a few more false positives (for good reason). Interested to hear your thoughts on that as well!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258855,
          "author_name": "DavidGbodiOdaibo",
          "author_url": "",
          "post_date": "2017-12-17T06:49:37.960000",
          "content": "<p>I agree that using just the .aps increases the chances of false negatives, the threats on the forearms inner ankles and groin are particularly difficult, some of the threats are only visible in 2 frames of the aps data, however we had 2 models that used the  a3daps data in our ensemble to compensate for potential false negatives. I think a robust system will need much more training samples and subject variety. As far as not seeing the threats in the .aps we created frame level labels for all the images in the stage 1 aps data, we found some label errors but did find threats that were not visible.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 258371,
      "author_name": "Hans Bouwmeester",
      "author_url": "",
      "post_date": "2017-12-16T01:56:08.710000",
      "content": "<p>I trained (from scratch) InceptionResnetV2 on a set of 299x299 mosaic-images each containing aps/a3daps crops based on multiple views on a zone. Did not use the a3d-images.</p>\n\n<p>In hindsight I should have gone for a more specialized net like MVCNN or Conv3D I guess. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 258338,
      "author_name": "dhammack",
      "author_url": "",
      "post_date": "2017-12-16T00:13:43.853000",
      "content": "<p>I also used only a3d. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 258759,
      "author_name": "Nathaniel Maddux",
      "author_url": "",
      "post_date": "2017-12-17T00:53:18.017000",
      "content": "<p>I did something a bit different than most.  Using .a3d files as input, I \"unwrapped\" a passenger's body surface into a 2 dimensional representation using a cylindrical coordinate system for each body-part.  The 7 body-parts were the 2 legs, 1 trunk, 2 biceps, and 2 forearms.  Cylindrical coordinate systems were registered to body-parts using estimates of the positions of wrists, elbows, shoulders, feet, the buttock/leg meeting point, and the center of mass.  A cylinder's axis was allowed to curve but slices were parallel.  Surface reflectivity, radius and thickness became 3 image channels.  All 7 body-part images were joined into one. Variability due to differing body types was subtracted away using a nearest neighbors approach.  </p>\n\n<p>Contraband really stands out in the resulting images as areas of saturated color on a gray background.  Using only color histograms and logistic regression, I placed in the top 20% on the initial public leaderboard.  So even without using shape and texture information in the processed images, the technique already performed well.</p>\n\n<p>I ran out of time for a few reasons:  this is my first Python project, I started with only 2 months to go, I have no CNN experience, yada yada. The next thing I was going to do was try the simplest sort of CNN transfer learning.  I may still do it, and make a late submission.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 258765,
          "author_name": "Bastiaan Bergman",
          "author_url": "",
          "post_date": "2017-12-17T01:21:36.927000",
          "content": "<p>Cool! Can you explain in more detail how you got the reflectivity, surface height and thickness out of you unwrapped body parts?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 258780,
          "author_name": "Nathaniel Maddux",
          "author_url": "",
          "post_date": "2017-12-17T02:18:24.083000",
          "content": "<p>Thank you!  Once a cylindrical coordinate transform is performed for a body-part, I have an array with dimensions (r, \\theta, w), where r is radius, \\theta is the angular coordinate,  and w is the axial coordinate.  For each (\\theta, w) coordinate, I get the peaks along the r direction.  The tallest peak's height, position, and width are the surface reflectivity, radius (AKA surface height), and thickness, respectively, for coordinate (\\theta, w).  That's not a precise description, by the way, since the determination of peak height, position and width (I.e. zeroeth, first and second moments) needs to be tweeked to give good results.  Anyway, I end up with surface reflectivity, radius and thickness for each (\\theta, w) coordinate.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258850,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2017-12-17T06:19:12.597000",
          "content": "<p>@Nathaniel This is a cool idea!</p>\n\n<p>I too used the a3d images not as 3d images but as 2d images, however I treated them instead like frames from a video so my model basically is traveling up through the body one \"frame\" at a time. If you flip through the 3d image slices you can pretty easily spot the threats in their evolution as they appear to be \"growing\". This also had added benefits of making the training set exponentially bigger, using large batch sizes, and also allowed for pre-trained nets which wouldn't otherwise be possible with 3d convnets.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 258376,
      "author_name": "Oleg Trott",
      "author_url": "",
      "post_date": "2017-12-16T02:25:30.883000",
      "content": "<p>In my opinion, you need both. Sometimes you see the threat on APS, but not A3DAPS. Less commonly, it's the other way around. I fed both of them to the network, as separate channels. No need to choose.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 258383,
          "author_name": "Suchir Balaji",
          "author_url": "",
          "post_date": "2017-12-16T02:33:28.023000",
          "content": "<p>I observed the same, and also passed both aps and a3daps data as separate channels, but strangely my final loss didn't change. Did you see improvement over just using aps?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258384,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2017-12-16T02:38:46.640000",
          "content": "<p>I did not experiment with just using one. I decided early on that I'd use both.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258454,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2017-12-16T05:39:11.343000",
          "content": "<p>I concatenated aps &amp; a3daps into one image and got higher accuracy than when I used aps or a3daps separately (0.03 logloss using this strategy on stage 1).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 258435,
      "author_name": "Shiv Gowda",
      "author_url": "",
      "post_date": "2017-12-16T04:33:49.090000",
      "content": "<p>I used only APS at full scale(512x660 I think). However, my model consisted of two separate stages: the first stage detected the contrabands and the second figured out which zone the contraband was in. I did try to use a3daps but that was too slow in the first phase for me to make meaningful progress so I gave up on it even though I noticed my contraband detection model was missing around 2-3% of the contrabands. Infact those 2-3% were not even visible visually when I looked at the images.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 258420,
      "author_name": "Hillview",
      "author_url": "",
      "post_date": "2017-12-16T03:45:18.660000",
      "content": "<p>I have a theory regarding why models based on APS images performed better. See <a href=\"https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/45814\">https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/45814</a>. In short, I suspect that labels were generated by looking at APS images, :-).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 258403,
      "author_name": "Naoto Usuyama",
      "author_url": "",
      "post_date": "2017-12-16T03:14:16.707000",
      "content": "<p>Interesting! We only used A3DAPS with 2D CNN models like resnet/densenet. Should have used APS!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 258347,
      "author_name": "Moejoe",
      "author_url": "",
      "post_date": "2017-12-16T00:35:35.333000",
      "content": "<p>My model overfit more easily when a3daps was involved.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 258350,
          "author_name": "Hillview",
          "author_url": "",
          "post_date": "2017-12-16T00:38:37.577000",
          "content": "<p>I thought aps data was too coarse, and didn't even bother.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 258343,
      "author_name": "Yusaku Sako",
      "author_url": "",
      "post_date": "2017-12-16T00:24:15.023000",
      "content": "<p>I only used a3d.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 258334,
      "author_name": "Hillview",
      "author_url": "",
      "post_date": "2017-12-16T00:11:03.080000",
      "content": "<p>I was only using .a3d data. With the hinder sight, looks that this was a wrong decision, :-)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 258455,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2017-12-16T05:41:20.033000",
          "content": "<p>Yep I made this same mistake. My first model used both aps/a3daps and would have done better in stage 2 but since my a3d only model was doing much better in stage 1 so I abandoned it. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 258332,
      "author_name": "Suchir Balaji",
      "author_url": "",
      "post_date": "2017-12-16T00:06:06.260000",
      "content": "<p>I'm also curious about this. I trained a few threat segmentation models using both aps and a3daps images as input, but didn't see any improvement over segmenting from just aps images.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 258339,
          "author_name": "Sukjae Cho",
          "author_url": "",
          "post_date": "2017-12-16T00:15:51.243000",
          "content": "<p>I only used a3d (scaled down by 1/4 and simple Conv3D) and just one model. I wish I added another model with aps :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258344,
          "author_name": "Branden Murray",
          "author_url": "",
          "post_date": "2017-12-16T00:30:04.303000",
          "content": "<p>Just for clarification, do you mean you used images that were 3/4 the size of the original or 1/4 the size of the original? I think I tried 3/4, but that didn't work. My architectures were probably crap.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258352,
          "author_name": "Sukjae Cho",
          "author_url": "",
          "post_date": "2017-12-16T00:39:14.613000",
          "content": "<p>Scaled down to 1/4 on each dimension - so it becomes 128x128x165</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 258353,
          "author_name": "Yusaku Sako",
          "author_url": "",
          "post_date": "2017-12-16T00:41:21.443000",
          "content": "<p>oh wow.  Congrats on the results.  I tried Conv3D and was not converging so I went with MVCNN.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 258358,
          "author_name": "Moejoe",
          "author_url": "",
          "post_date": "2017-12-16T00:54:36.083000",
          "content": "<p>I found keeping the original resolution increased my performance considerably.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 258361,
          "author_name": "Suchir Balaji",
          "author_url": "",
          "post_date": "2017-12-16T00:59:44.273000",
          "content": "<p>Interesting, I thought the scans were just too big so I didn't even try Conv3D -- what did your architectures look like? It seems like it would be pretty easy to overfit, which was why I mostly stuck to MVCNN-like models.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 258372,
          "author_name": "Sukjae Cho",
          "author_url": "",
          "post_date": "2017-12-16T01:59:37.967000",
          "content": "<p>It's quite simple architecture, so each epoch just takes 10 minutes on single 1080 Ti. I used 5 blocks of resnet like conv3d. First 3 blocks are shared for all zones, and last 2 blocks are per zone. Basic rotation/zoom/shift augmentation/flip was used. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 258359,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-12-16T00:55:56.367000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "258787": "We observed very quickly that there was a big difference in the score we could achieve with the .a3daps/a3d compared with the .aps data trained using the same model, and we could not figure out why. Fortunately we started out with the .aps data and our first submission achieved a score of 0.023 putting us in 12th place in stage 1. Knowing the aps data had only 16 view points and the .a3daps had 64 viewpoints we expected the same model trained using the same approach on the a3daps data with more view points should achieve a much better score but this was not the case.  We checked mean, standard deviation, checked if subjects were mixed up  btw aps and a3daps e.t.c. We tried multiple times  but no success, our single best model with .aps achieves a stage 1 score of ~0.012 and stage 2 score of ~0.035 while the same model with the .a3daps  achieves a score of ~0.03 stage1  and ~0.08  stage2.  Another observation looking at the images is that the a3daps images look much noisier than the aps data. I think there are some issues with the processing of the a3d, and .a3daps, and knowing the .a3daps was derived from the a3d while the aps was acquired separately gave us comfort in pursuing models that used the .aps data know that logically these models should not be outperforming the  .a3d and a3daps models.",
    "258371": "I trained (from scratch) InceptionResnetV2 on a set of 299x299 mosaic-images each containing aps/a3daps crops based on multiple views on a zone. Did not use the a3d-images.\n\nIn hindsight I should have gone for a more specialized net like MVCNN or Conv3D I guess. ",
    "258338": "I also used only a3d. ",
    "258759": "I did something a bit different than most.  Using .a3d files as input, I \"unwrapped\" a passenger's body surface into a 2 dimensional representation using a cylindrical coordinate system for each body-part.  The 7 body-parts were the 2 legs, 1 trunk, 2 biceps, and 2 forearms.  Cylindrical coordinate systems were registered to body-parts using estimates of the positions of wrists, elbows, shoulders, feet, the buttock/leg meeting point, and the center of mass.  A cylinder's axis was allowed to curve but slices were parallel.  Surface reflectivity, radius and thickness became 3 image channels.  All 7 body-part images were joined into one. Variability due to differing body types was subtracted away using a nearest neighbors approach.  \n\nContraband really stands out in the resulting images as areas of saturated color on a gray background.  Using only color histograms and logistic regression, I placed in the top 20% on the initial public leaderboard.  So even without using shape and texture information in the processed images, the technique already performed well.\n\nI ran out of time for a few reasons:  this is my first Python project, I started with only 2 months to go, I have no CNN experience, yada yada. The next thing I was going to do was try the simplest sort of CNN transfer learning.  I may still do it, and make a late submission.",
    "258376": "In my opinion, you need both. Sometimes you see the threat on APS, but not A3DAPS. Less commonly, it's the other way around. I fed both of them to the network, as separate channels. No need to choose.",
    "258326": "Did anyone have success using the .a3d or .a3daps (or even .ahi) images? I spent a decent amount of time trying to get good scores out of them, but I think the best I got out of .a3daps was ~0.12 (using 'proper' validation, i.e. separating folds by subject) and for .a3d my best was probably around 0.17. Was anyone able to get less than &lt;0.10 or 0.05 on the Public LB using these?",
    "258435": "I used only APS at full scale(512x660 I think). However, my model consisted of two separate stages: the first stage detected the contrabands and the second figured out which zone the contraband was in. I did try to use a3daps but that was too slow in the first phase for me to make meaningful progress so I gave up on it even though I noticed my contraband detection model was missing around 2-3% of the contrabands. Infact those 2-3% were not even visible visually when I looked at the images.",
    "258420": "I have a theory regarding why models based on APS images performed better. See https://www.kaggle.com/c/passenger-screening-algorithm-challenge/discussion/45814. In short, I suspect that labels were generated by looking at APS images, :-).",
    "258403": "Interesting! We only used A3DAPS with 2D CNN models like resnet/densenet. Should have used APS!",
    "258347": "My model overfit more easily when a3daps was involved.",
    "258343": "I only used a3d.",
    "258334": "I was only using .a3d data. With the hinder sight, looks that this was a wrong decision, :-)",
    "258332": "I'm also curious about this. I trained a few threat segmentation models using both aps and a3daps images as input, but didn't see any improvement over segmenting from just aps images.",
    "258359": ""
  }
}