{
  "id": 45805,
  "title": "Post Competition Architecture Discussion",
  "url": "/competitions/passenger-screening-algorithm-challenge/writeups/moejoe-post-competition-architecture-discussion",
  "author_name": "",
  "post_date": "2018-04-10T21:53:28.130Z",
  "votes": 19,
  "comment_count": 85,
  "views": 0,
  "content": "<p>Now that the competition is over let's have a place to discuss our architecture, findings, and hindsight?</p>\n\n<p>Edit: Here's my 10th place solution: <a href=\"https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\">https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution</a></p>",
  "messages": [
    {
      "id": "258374",
      "postDate": "12/16/2017 02:14:05",
      "content": "<p>Now that the competition is over let's have a place to discuss our architecture, findings, and hindsight?</p>\n\n<p>Edit: Here's my 10th place solution: <a href=\"https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\">https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution</a></p>",
      "rawMarkdown": "Now that the competition is over let's have a place to discuss our architecture, findings, and hindsight?\n\nEdit: Here's my 10th place solution: https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution",
      "votes": null
    },
    {
      "id": "258391",
      "postDate": "12/16/2017 02:44:34",
      "content": "<p>Edit: <a href=\"https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\">https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution</a>\nHere's my code.</p>\n\n<p>In my first attempts I tried to use MVCNN on a3daps files and downsampled inputs to a 3rd of the original resolution. This was getting about 0.10 on the original test data. I never tried 3D convolutions as I thought it'd be too slow and memory hungry.</p>\n\n<p>I found I got better results by using a3d files and maintaining the original resolution (and lowering the batch size considerably so it would fit in memory). I tried combining a3d and a3daps by throwing them in separate channels but I found it caused overfitting.</p>\n\n<p>Later on I dropped MVCNN and went with a custom architecture where each view is fed through a pretrained ResNet-50 CNN and each feature map is pyramid pooled (to find objects/features of differing size) and fed to a LSTM with attention. My pyramid pooling doesn't actually use pooling layers but instead passes the last feature map through a 1x1, 3x3, and 5x5 convolution with stride 1, 2, and 3 respectively. I use a form of attention for CNNs that I apply to the feature map before the pyramid pooling. To make it run on my 1080 ti I increased the stride of the first couple layers from 2 to 3, which probably has a similar affect as decreasing the resolution but loses less information. I trained with SGD with a cosine annealing schedule. This was getting around 0.015 on the original test data.</p>\n\n<p>I didn't use any segmentation and had my model predict on all 17 zones at once. In hindsight maybe I should have segmented as it could allow me to train with higher resolution and focus attention on relevant zones at the cost of speed, but I was also worried that the people in the stage 2 data would be drastically different and the segmentations would be wrong. I also probably shouldn't have done experimental things like CNN attention and cosine annealing but then again maybe the risk paid off and they helped.</p>\n\n<p>Regrettably I didn't use ensembling. Just didn't have the time to train more models or think it would be too important. But looks like it could have made a difference.</p>\n\n<p>But in the end I'm still very satisfied with my score, and congratulations to everyone!</p>",
      "rawMarkdown": "Edit: https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\nHere's my code.\n\nIn my first attempts I tried to use MVCNN on a3daps files and downsampled inputs to a 3rd of the original resolution. This was getting about 0.10 on the original test data. I never tried 3D convolutions as I thought it'd be too slow and memory hungry.\n\nI found I got better results by using a3d files and maintaining the original resolution (and lowering the batch size considerably so it would fit in memory). I tried combining a3d and a3daps by throwing them in separate channels but I found it caused overfitting.\n\nLater on I dropped MVCNN and went with a custom architecture where each view is fed through a pretrained ResNet-50 CNN and each feature map is pyramid pooled (to find objects/features of differing size) and fed to a LSTM with attention. My pyramid pooling doesn't actually use pooling layers but instead passes the last feature map through a 1x1, 3x3, and 5x5 convolution with stride 1, 2, and 3 respectively. I use a form of attention for CNNs that I apply to the feature map before the pyramid pooling. To make it run on my 1080 ti I increased the stride of the first couple layers from 2 to 3, which probably has a similar affect as decreasing the resolution but loses less information. I trained with SGD with a cosine annealing schedule. This was getting around 0.015 on the original test data.\n\nI didn't use any segmentation and had my model predict on all 17 zones at once. In hindsight maybe I should have segmented as it could allow me to train with higher resolution and focus attention on relevant zones at the cost of speed, but I was also worried that the people in the stage 2 data would be drastically different and the segmentations would be wrong. I also probably shouldn't have done experimental things like CNN attention and cosine annealing but then again maybe the risk paid off and they helped.\n\nRegrettably I didn't use ensembling. Just didn't have the time to train more models or think it would be too important. But looks like it could have made a difference.\n\nBut in the end I'm still very satisfied with my score, and congratulations to everyone!",
      "votes": null
    },
    {
      "id": "258408",
      "postDate": "12/16/2017 03:19:11",
      "content": "<p>Moejoe,\nThanks for sharing your experience and thoughts.  I wonder how did you decide which zone the threats locate If you did not use any segmentation. Thanks.</p>",
      "rawMarkdown": "Moejoe,\nThanks for sharing your experience and thoughts.  I wonder how did you decide which zone the threats locate If you did not use any segmentation. Thanks.",
      "votes": null
    },
    {
      "id": "258419",
      "postDate": "12/16/2017 03:43:19",
      "content": "<p>Hi Wensu, the CNN learns on its own which locations correspond to which labels without any human guidance. It outputs a size 17 vector and is trained against size 17 targets. Since the classes aren't mutually exclusive I don't use softmax and I train with binary cross entropy loss.</p>\n\n<p>The average pooling layer at the end of ResNet loses all spatial information so one would imagine this would be a problem. Instead I use a pyramid pooling scheme which pools with multiple kernel sizes and feed the output (which maintains information about location) directly to the LSTM. However, although pyramid pooling helped noticeably, I found that it wasn't absolutely necessary and even with the regular average pooling layer it still gave decent results (at least on the stage 1 data).</p>",
      "rawMarkdown": "Hi Wensu, the CNN learns on its own which locations correspond to which labels without any human guidance. It outputs a size 17 vector and is trained against size 17 targets. Since the classes aren't mutually exclusive I don't use softmax and I train with binary cross entropy loss.\n\nThe average pooling layer at the end of ResNet loses all spatial information so one would imagine this would be a problem. Instead I use a pyramid pooling scheme which pools with multiple kernel sizes and feed the output (which maintains information about location) directly to the LSTM. However, although pyramid pooling helped noticeably, I found that it wasn't absolutely necessary and even with the regular average pooling layer it still gave decent results (at least on the stage 1 data).",
      "votes": null
    },
    {
      "id": "258430",
      "postDate": "12/16/2017 04:12:53",
      "content": "<p>Congratulations to everyone who participated in this challenging and suspenseful competition! </p>\n\n<p>I used a multi-view CNN custom ResNet that looks at the entire image (without any pre-segmentation) and outputs 17 probabilities. I used the .aps dataset downsampled to half-resolution. </p>\n\n<p>My stage1 score (around 0.025) was much better than my stage2 score (0.13589). This must be because the subjects in the stage1 training set were the same as those in the stage1 test set, but the subjects in the stage2 set were completely different people. </p>\n\n<p>During the optimization process, I used regular cross-validation. After reading comments from others, it has become clear to me that the correct way to do cross validation was to first group all of the instances depicting the same person, and then randomly assign each group of instances to a cross validation set. </p>\n\n<p>I had suspected that this would be an issue, but could not have imagined that it would make such a huge difference between stage1 and stage2 performance. I think there are two possible mechanisms that could explain this:</p>\n\n<ol>\n<li><p>Since my net does not use pre-segmentation, it probably determines the threat zones by identifying landmark features. For example, it might identify a threat between the hand and elbow, and therefore determine that the threat lies in the forearm region. However, different people have different-looking hands and elbows. This may have prevented my net from properly identifying landmarks in stage2. </p></li>\n<li><p>It is also possible that the net learns to whitelist certain objects that would otherwise be identified as threats. The Bob Marley lookalike from stage2 is a perfect example. If the stage1 training set had included someone with this type of hair, the net would most likely have learned to specifically avoid labeling dreadlocks as a threat. </p></li>\n</ol>",
      "rawMarkdown": "Congratulations to everyone who participated in this challenging and suspenseful competition! \n\nI used a multi-view CNN custom ResNet that looks at the entire image (without any pre-segmentation) and outputs 17 probabilities. I used the .aps dataset downsampled to half-resolution. \n\nMy stage1 score (around 0.025) was much better than my stage2 score (0.13589). This must be because the subjects in the stage1 training set were the same as those in the stage1 test set, but the subjects in the stage2 set were completely different people. \n\nDuring the optimization process, I used regular cross-validation. After reading comments from others, it has become clear to me that the correct way to do cross validation was to first group all of the instances depicting the same person, and then randomly assign each group of instances to a cross validation set. \n\nI had suspected that this would be an issue, but could not have imagined that it would make such a huge difference between stage1 and stage2 performance. I think there are two possible mechanisms that could explain this:\n\n1. Since my net does not use pre-segmentation, it probably determines the threat zones by identifying landmark features. For example, it might identify a threat between the hand and elbow, and therefore determine that the threat lies in the forearm region. However, different people have different-looking hands and elbows. This may have prevented my net from properly identifying landmarks in stage2. \n\n2. It is also possible that the net learns to whitelist certain objects that would otherwise be identified as threats. The Bob Marley lookalike from stage2 is a perfect example. If the stage1 training set had included someone with this type of hair, the net would most likely have learned to specifically avoid labeling dreadlocks as a threat.",
      "votes": null
    },
    {
      "id": "258445",
      "postDate": "12/16/2017 05:23:01",
      "content": "<p>Initially, I started out by using a combination of both the aps and a3daps files. Rather than use all 16 and 64 \"views\" from each body scan format, I instead selected 8 prominent views i.e. front-center, front-left, etc. from both files and then further segmented and cropped out all 17 zones into equal 100x100 square images. I then concatenated these into a 400x400 \"mosaic\" for each passenger/zone combo. This was fed to my model that generated a binary classification. Using this strategy I was able to achieve .03 logloss on stage 1 test set. Like others have echoed on the forum threads, I too found that if I left out the aps images I would end up with lower accuracy yet I did observe that I got better accuracy when having a3daps images included in the mix rather than just using aps by itself since not all threats were visible in all cases. </p>\n\n<p>With around 2 weeks left in the competition I tried using the a3d images to see if I could improve my score further. I found that I couldn't get good results using a 3d convnet, so instead I tried a different strategy which was to treat all image slices as separate images to feed to the model in a sequence. So for each passenger/subject there was a total of 310 slices (step-size of 2 pixels and clipping everything over 620 which was mostly noise). I also segmented by zone. With this strategy I was surprised that I was able to get as low as .005 logloss on stage 1 test set (this also corresponded with my local validation loss). This was using just a single model (DenseNet169) and one 15% validation hold out. Not only was it much easier for my model to locate and classify the threats, using this strategy also really helped me with segmentation because I could also see patterns in the data like exactly how long a threat was and which image slices were most confident etc. This high score in stage 1 gave me enough confidence that I should abandon my first strategy and only use a3d images for stage 2. In hindsight, I can now see this was not such a good idea as my score ended up dropping to 0.17 on stage 2!! After reviewing my predictions, I can see that some new features of the passengers introduced in stage 2 (that didn't exist in stage 1 training set or test set) caused my model to generate a large number of false positives in a few isolated zones. Had my model at least seen examples of something similar I think it probably would have scored much higher as it would have learned to whitelist some of those new features in stage 2 (i.e. dreadlocks, suspender clips, etc.) that looked very similar to threats in stage 1 training set. </p>\n\n<p>While I am quite disappointed in my final result, but overall I think this was a really interesting challenge and was a great learning experience for me!! Hopefully someone here can also learn from my mistakes!</p>",
      "rawMarkdown": "Initially, I started out by using a combination of both the aps and a3daps files. Rather than use all 16 and 64 \"views\" from each body scan format, I instead selected 8 prominent views i.e. front-center, front-left, etc. from both files and then further segmented and cropped out all 17 zones into equal 100x100 square images. I then concatenated these into a 400x400 \"mosaic\" for each passenger/zone combo. This was fed to my model that generated a binary classification. Using this strategy I was able to achieve .03 logloss on stage 1 test set. Like others have echoed on the forum threads, I too found that if I left out the aps images I would end up with lower accuracy yet I did observe that I got better accuracy when having a3daps images included in the mix rather than just using aps by itself since not all threats were visible in all cases. \n\nWith around 2 weeks left in the competition I tried using the a3d images to see if I could improve my score further. I found that I couldn't get good results using a 3d convnet, so instead I tried a different strategy which was to treat all image slices as separate images to feed to the model in a sequence. So for each passenger/subject there was a total of 310 slices (step-size of 2 pixels and clipping everything over 620 which was mostly noise). I also segmented by zone. With this strategy I was surprised that I was able to get as low as .005 logloss on stage 1 test set (this also corresponded with my local validation loss). This was using just a single model (DenseNet169) and one 15% validation hold out. Not only was it much easier for my model to locate and classify the threats, using this strategy also really helped me with segmentation because I could also see patterns in the data like exactly how long a threat was and which image slices were most confident etc. This high score in stage 1 gave me enough confidence that I should abandon my first strategy and only use a3d images for stage 2. In hindsight, I can now see this was not such a good idea as my score ended up dropping to 0.17 on stage 2!! After reviewing my predictions, I can see that some new features of the passengers introduced in stage 2 (that didn't exist in stage 1 training set or test set) caused my model to generate a large number of false positives in a few isolated zones. Had my model at least seen examples of something similar I think it probably would have scored much higher as it would have learned to whitelist some of those new features in stage 2 (i.e. dreadlocks, suspender clips, etc.) that looked very similar to threats in stage 1 training set. \n\nWhile I am quite disappointed in my final result, but overall I think this was a really interesting challenge and was a great learning experience for me!! Hopefully someone here can also learn from my mistakes!",
      "votes": null
    },
    {
      "id": "258447",
      "postDate": "12/16/2017 05:26:58",
      "content": "<p>My CV was 0.032-ish (No subject was in multiple CV folds, of course). </p>\n\n<p>The big surprise in Stage2 was the guy with massive dreadlocks, like no one had in the training dataset. They looked a lot like some of the bombs in the training  scans instead, except for the fact that they adjoined his head. My model was understandably suspicious. I had calculated that that would cost me 0.01 on the score, and it seems like it did.</p>\n\n<p>If I could have seen the Stage2 data for just 10 minutes, before finalizing the model, I would have reduced the confidence of zone 6 &amp; 7 predictions, and negated most of the 0.01 damage, but the predictions are supposed to be fully automatic (or such was my understanding).</p>",
      "rawMarkdown": "My CV was 0.032-ish (No subject was in multiple CV folds, of course). \n\nThe big surprise in Stage2 was the guy with massive dreadlocks, like no one had in the training dataset. They looked a lot like some of the bombs in the training  scans instead, except for the fact that they adjoined his head. My model was understandably suspicious. I had calculated that that would cost me 0.01 on the score, and it seems like it did.\n\nIf I could have seen the Stage2 data for just 10 minutes, before finalizing the model, I would have reduced the confidence of zone 6 &amp; 7 predictions, and negated most of the 0.01 damage, but the predictions are supposed to be fully automatic (or such was my understanding).",
      "votes": null
    },
    {
      "id": "258451",
      "postDate": "12/16/2017 05:34:09",
      "content": "<p>Congrats Oleg on placing 5th in this comp, thats very impressive!</p>\n\n<p>Glad you mentioned about this guy and that I'm not the only one who had issues with it. The dreadlocks were the main culprit for pretty much all of my false positives in zones 17, 6 and 7. My model also had problems with the guy with the suspenders near the metal clips (zones 8 &amp; 10). Like you say, if this wasn't a two stage comp I would have made some adjustments to my model's sensitivity/threat thresholds to mitigate this but since I didn't anticipate needing to do this in the model upload so I couldn't make those changes. </p>\n\n<p>&gt; <strong>Oleg Trott wrote</strong>\n&gt; \n&gt; &gt; My CV was 0.032-ish (No subject was in multiple CV folds, of course). \n&gt; \n&gt; The big surprise in Stage2 was the guy with massive dreadlocks, like no one had in the training dataset. They looked a lot like some of the bombs in the training  scans instead, except for the fact that they adjoined his head. My model was understandably suspicious. I had calculated that that would cost me 0.01 on the score, and it seems like it did.\n&gt; \n&gt; If I could have seen the Stage2 data for just 10 minutes, before finalizing the model, I would have reduced the confidence of zone 6 &amp; 7 predictions, and negated most of the 0.01 damage, but the predictions are supposed to be fully automatic (or such was my understanding).\n&gt; \n&gt; \n&gt; \n&gt; </p>",
      "rawMarkdown": "Congrats Oleg on placing 5th in this comp, thats very impressive!\n\nGlad you mentioned about this guy and that I'm not the only one who had issues with it. The dreadlocks were the main culprit for pretty much all of my false positives in zones 17, 6 and 7. My model also had problems with the guy with the suspenders near the metal clips (zones 8 &amp; 10). Like you say, if this wasn't a two stage comp I would have made some adjustments to my model's sensitivity/threat thresholds to mitigate this but since I didn't anticipate needing to do this in the model upload so I couldn't make those changes. \n\n&gt; **Oleg Trott wrote**\n&gt; \n&gt; &gt; My CV was 0.032-ish (No subject was in multiple CV folds, of course). \n&gt; \n&gt; The big surprise in Stage2 was the guy with massive dreadlocks, like no one had in the training dataset. They looked a lot like some of the bombs in the training  scans instead, except for the fact that they adjoined his head. My model was understandably suspicious. I had calculated that that would cost me 0.01 on the score, and it seems like it did.\n&gt; \n&gt; If I could have seen the Stage2 data for just 10 minutes, before finalizing the model, I would have reduced the confidence of zone 6 &amp; 7 predictions, and negated most of the 0.01 damage, but the predictions are supposed to be fully automatic (or such was my understanding).\n&gt; \n&gt; \n&gt; \n&gt;",
      "votes": null
    },
    {
      "id": "258581",
      "postDate": "12/16/2017 14:00:10",
      "content": "<p>Hi Moejoe, Thanks for your response.  Since I spend more than a half time on segmentation, I'd like to learn how others did for this issue.  you said your model learns on its own correspond to which labels, it is great and is what I like to do, but did not figure out how.  you mentioned your model is trained against 17 targets, how did you get the 17 targets without segmentation,  could you explain further how your model did?  Thanks a lot.</p>",
      "rawMarkdown": "Hi Moejoe, Thanks for your response.  Since I spend more than a half time on segmentation, I'd like to learn how others did for this issue.  you said your model learns on its own correspond to which labels, it is great and is what I like to do, but did not figure out how.  you mentioned your model is trained against 17 targets, how did you get the 17 targets without segmentation,  could you explain further how your model did?  Thanks a lot.",
      "votes": null
    },
    {
      "id": "258613",
      "postDate": "12/16/2017 15:47:18",
      "content": "<p>oh boy, I see a future where we data scientist are to blame for people with dreadlocks going through \"special screening\" at the airport :)</p>",
      "rawMarkdown": "oh boy, I see a future where we data scientist are to blame for people with dreadlocks going through \"special screening\" at the airport :)",
      "votes": null
    },
    {
      "id": "258618",
      "postDate": "12/16/2017 16:10:19",
      "content": "<p>we used exclusively A3DAPS (generally 7 of the 64 views), training with a 7 layer model on individual views and then combining the  views using a separate model.</p>\n\n<p>Ensemble: our logistic regression also incorporated some home-brewed algorithms (using morphology and transforms)  which added predictivity in some cases over the ML results.</p>\n\n<p>VGG16 was hard to use on my desktop, it crashed often unless we throttled back the batch size significantly\nwe had som e success training multiple zones together using left right symmetry</p>\n\n<p>feeding all views at once also crashed my desktop (16GB of RAM), I would appreciate feedback how people made that work.</p>\n\n<p>I am surprised that some contestants had success with 100 x100 crops.  we generaly needed to go 225-300 wide depending on the zone and view</p>\n\n<p>Our model seemed to work very well on 11 zones and poorly on 6 zones: arms(4), groin, and upperchest. The arm regions were noisy. I would be interested in strategies for those zones.  We tried to rotate all arms to vertical but saw no improvement. Did anyone use gender recognition to facilitate groin and upperchest? </p>\n\n<p>Did anyone use cross-zone models? In the training set there was a slight inverse correlation among zones  (ie a passenger with contraband in 2 other zones was less likely to have it in a third zone) </p>",
      "rawMarkdown": "we used exclusively A3DAPS (generally 7 of the 64 views), training with a 7 layer model on individual views and then combining the  views using a separate model.\n\nEnsemble: our logistic regression also incorporated some home-brewed algorithms (using morphology and transforms)  which added predictivity in some cases over the ML results.\n\nVGG16 was hard to use on my desktop, it crashed often unless we throttled back the batch size significantly\nwe had som e success training multiple zones together using left right symmetry\n\nfeeding all views at once also crashed my desktop (16GB of RAM), I would appreciate feedback how people made that work.\n\nI am surprised that some contestants had success with 100 x100 crops.  we generaly needed to go 225-300 wide depending on the zone and view\n\nOur model seemed to work very well on 11 zones and poorly on 6 zones: arms(4), groin, and upperchest. The arm regions were noisy. I would be interested in strategies for those zones.  We tried to rotate all arms to vertical but saw no improvement. Did anyone use gender recognition to facilitate groin and upperchest? \n\nDid anyone use cross-zone models? In the training set there was a slight inverse correlation among zones  (ie a passenger with contraband in 2 other zones was less likely to have it in a third zone)",
      "votes": null
    },
    {
      "id": "258673",
      "postDate": "12/16/2017 18:55:17",
      "content": "<p>I started out with APS images first.\nI used various CNNs (VGG, Resnet, Densenet, etc) that have been pretrained with ImageNet weights as featurizers by stripping out their classification layers.\nFor each scan, forward passes were made on the CNN to create feature maps for all 16 views, and those feature maps were flattened and concatenated into a single vector to be fed into LightGBM for classification.\nAs LightGBM does not support multi-label output, each of the 17 zones were trained/eval'd separately.\nThis got me to ~0.19 on the public LB when I used Densenet 121 as the featurizer.  I did not expect this to work too well as ImageNet and TSA images are very different.  Also, some threats were very difficult to see in APS due to occlusion and how projections seem to be computed.</p>\n\n<p>Then I started looking at A3D and thought to myself that segmenting out into various zone groups would help with classification, because:</p>\n\n<ul>\n<li>Cassifiers that solely focus on its assigned specific zone group can be created, rather than having the network figure out what to focus on</li>\n<li>More samples can be fed because mirroring can be performed independently on the body part in question</li>\n<li>By eliminating other parts, threats can be seen more easily (you get a clear shot at the zone(s) in question.)</li>\n</ul>\n\n<p>I ended up with 6 classifiers for the following zone groups: arms, legs, chest-back, torso, waist, and crotch.\nThe \"arms\" classifier will get trained on right arm images, left arm images, and their horizontal mirrors to output 2 labels (i.e., whether a threat exists in the upper arm and/or lower arm.)\nThe \"legs\" classifier will be trained in a similar fashion to output 3 labels for detecting threats on the thigh, knee, and foot zones.  Others were trained to output a single label (i.e., binary classifiers.)</p>\n\n<p>I experimented by training Conv3D-based classifiers on the segmented zone groups, but this did not work so well, and I ended up creating a variant of the Multi-View CNN (<a href=\"https://arxiv.org/pdf/1505.00880.pdf\">https://arxiv.org/pdf/1505.00880.pdf</a>) except I used a learnable convolution layer instead of a view pooling layer.  Train-time augmentation was limited to horizontal flips, and test-time augmentation was limited to 2 combinations of horizontal flips and averaging out the predictions.  No trimming of predictions were performed (boy, the models gave very confident predictions so perhaps I should have looked into this a little more to minimize the penalty on mistakes.)\nThis approach got me to ~0.019 on the public LB, and ~0.089 on the private LB.</p>\n\n<p>My final approach was preprocessing heavy due to all segmentation done based on the center of mass and assumed\nbody proportions, then rotating the zone groups in 3D to create 2D projections (I actually ended up implementing multiple methods of 2D projections, as simply doing \"max\" on the collapsed axis would sometimes hide certain threats that appear on one side only.)  Also, training was very compute intensive since there were 6 separate networks to cover all the zone groups.  And I actually trained multiple networks for each zone group for ensembling, so a total of 38 networks were trained when it was all said and done.</p>\n\n<p>Looking at other competitors posts, I feel a bit shocked and a little embarrassed to find that some did better with a much simpler and elegant pipeline of predicting all 17 labels simultaneously without any segmentation or even ensembling!  I didn't even try that because I had strong convictions in my intuitions (that I unfortunately did not bother to validate.)</p>\n\n<p>I learned a lot in this competition!</p>\n\n<ul>\n<li>Explore wider before settling down on a method (I feel that I turned down my \"learning rate\" or \"temperature\" too quickly and did not get good coverage on the solution space.)</li>\n<li>Before building a complex pipeline, try out and validate simpler solutions first.</li>\n<li>Be very thoughtful in setting up local CV.  I did not other to separate out subjects when creating folds, unlike other wise competitors did.</li>\n</ul>",
      "rawMarkdown": "I started out with APS images first.\nI used various CNNs (VGG, Resnet, Densenet, etc) that have been pretrained with ImageNet weights as featurizers by stripping out their classification layers.\nFor each scan, forward passes were made on the CNN to create feature maps for all 16 views, and those feature maps were flattened and concatenated into a single vector to be fed into LightGBM for classification.\nAs LightGBM does not support multi-label output, each of the 17 zones were trained/eval'd separately.\nThis got me to ~0.19 on the public LB when I used Densenet 121 as the featurizer.  I did not expect this to work too well as ImageNet and TSA images are very different.  Also, some threats were very difficult to see in APS due to occlusion and how projections seem to be computed.\n\nThen I started looking at A3D and thought to myself that segmenting out into various zone groups would help with classification, because:\n\n- Cassifiers that solely focus on its assigned specific zone group can be created, rather than having the network figure out what to focus on\n- More samples can be fed because mirroring can be performed independently on the body part in question\n- By eliminating other parts, threats can be seen more easily (you get a clear shot at the zone(s) in question.)\n\nI ended up with 6 classifiers for the following zone groups: arms, legs, chest-back, torso, waist, and crotch.\nThe \"arms\" classifier will get trained on right arm images, left arm images, and their horizontal mirrors to output 2 labels (i.e., whether a threat exists in the upper arm and/or lower arm.)\nThe \"legs\" classifier will be trained in a similar fashion to output 3 labels for detecting threats on the thigh, knee, and foot zones.  Others were trained to output a single label (i.e., binary classifiers.)\n\nI experimented by training Conv3D-based classifiers on the segmented zone groups, but this did not work so well, and I ended up creating a variant of the Multi-View CNN (https://arxiv.org/pdf/1505.00880.pdf) except I used a learnable convolution layer instead of a view pooling layer.  Train-time augmentation was limited to horizontal flips, and test-time augmentation was limited to 2 combinations of horizontal flips and averaging out the predictions.  No trimming of predictions were performed (boy, the models gave very confident predictions so perhaps I should have looked into this a little more to minimize the penalty on mistakes.)\nThis approach got me to ~0.019 on the public LB, and ~0.089 on the private LB.\n\nMy final approach was preprocessing heavy due to all segmentation done based on the center of mass and assumed\nbody proportions, then rotating the zone groups in 3D to create 2D projections (I actually ended up implementing multiple methods of 2D projections, as simply doing \"max\" on the collapsed axis would sometimes hide certain threats that appear on one side only.)  Also, training was very compute intensive since there were 6 separate networks to cover all the zone groups.  And I actually trained multiple networks for each zone group for ensembling, so a total of 38 networks were trained when it was all said and done.\n\nLooking at other competitors posts, I feel a bit shocked and a little embarrassed to find that some did better with a much simpler and elegant pipeline of predicting all 17 labels simultaneously without any segmentation or even ensembling!  I didn't even try that because I had strong convictions in my intuitions (that I unfortunately did not bother to validate.)\n\nI learned a lot in this competition!\n\n- Explore wider before settling down on a method (I feel that I turned down my \"learning rate\" or \"temperature\" too quickly and did not get good coverage on the solution space.)\n- Before building a complex pipeline, try out and validate simpler solutions first.\n- Be very thoughtful in setting up local CV.  I did not other to separate out subjects when creating folds, unlike other wise competitors did.",
      "votes": null
    },
    {
      "id": "258677",
      "postDate": "12/16/2017 19:18:14",
      "content": "<p>Separating subjects when creating training/validation folds could make huge difference on the stage 2 results. I didn’t do it as well :(</p>",
      "rawMarkdown": "Separating subjects when creating training/validation folds could make huge difference on the stage 2 results. I didn’t do it as well :(",
      "votes": null
    },
    {
      "id": "258684",
      "postDate": "12/16/2017 19:38:17",
      "content": "<p>The targets come from the training data. Consider the following sample from the training data:</p>\n\n<pre><code>0050492f92e22eed3474ae3a6fc907fa_Zone1,0\n0050492f92e22eed3474ae3a6fc907fa_Zone10,0\n0050492f92e22eed3474ae3a6fc907fa_Zone11,0\n0050492f92e22eed3474ae3a6fc907fa_Zone12,0\n0050492f92e22eed3474ae3a6fc907fa_Zone13,0\n0050492f92e22eed3474ae3a6fc907fa_Zone14,0\n0050492f92e22eed3474ae3a6fc907fa_Zone15,0\n0050492f92e22eed3474ae3a6fc907fa_Zone16,1\n0050492f92e22eed3474ae3a6fc907fa_Zone17,0\n0050492f92e22eed3474ae3a6fc907fa_Zone2,0\n0050492f92e22eed3474ae3a6fc907fa_Zone3,0\n0050492f92e22eed3474ae3a6fc907fa_Zone4,1\n0050492f92e22eed3474ae3a6fc907fa_Zone5,0\n0050492f92e22eed3474ae3a6fc907fa_Zone6,0\n0050492f92e22eed3474ae3a6fc907fa_Zone7,0\n0050492f92e22eed3474ae3a6fc907fa_Zone8,1\n0050492f92e22eed3474ae3a6fc907fa_Zone9,0\n</code></pre>\n\n<p>This is treated as an input target pair where 0050492f92e22eed3474ae3a6fc907fa.aps is the input and [0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0] is the target and what I train the model to output. The model is fed no other information.</p>\n\n<p>It may feel like black magic at first but it works and it works pretty well. The architecture of the model is designed so that it's more spatially aware than regular convnets to make it easier to learn.</p>",
      "rawMarkdown": "The targets come from the training data. Consider the following sample from the training data:\n\n    0050492f92e22eed3474ae3a6fc907fa_Zone1,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone10,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone11,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone12,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone13,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone14,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone15,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone16,1\n    0050492f92e22eed3474ae3a6fc907fa_Zone17,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone2,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone3,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone4,1\n    0050492f92e22eed3474ae3a6fc907fa_Zone5,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone6,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone7,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone8,1\n    0050492f92e22eed3474ae3a6fc907fa_Zone9,0\n\nThis is treated as an input target pair where 0050492f92e22eed3474ae3a6fc907fa.aps is the input and [0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0] is the target and what I train the model to output. The model is fed no other information.\n\nIt may feel like black magic at first but it works and it works pretty well. The architecture of the model is designed so that it's more spatially aware than regular convnets to make it easier to learn.",
      "votes": null
    },
    {
      "id": "258716",
      "postDate": "12/16/2017 21:23:39",
      "content": "<p>At first, I went through aps images and I felt that many threats seem to be hard to find, so I decided to go with a3d.\nI used APS to identify each person and make train/val set - I manually annotated a few persons at first and let NN to classify the person. But later I found that PCA can do same thing easily from discussion board :(</p>\n\n<p>When I tried to feed a3d data, the data size was too big - so I scaled down to 1/4 on each dimension and cut down mostly empty areas on front/back/side/top). I could see the threats even after 1/4 scale, so I assumed that 1/4 dowscaling doesn't reduce the accuracy much. And I adjusted vertical scale little bit for each image so that the head sits on similar position(no ML, just simple image processing). Even after scaling down, only 2-3 batches fit into 1080Ti.</p>\n\n<p>Next, I fed the data into resnet-like 3D convolution. All data is fed to 3 blocks of identify block. Then output is splitted into 17 zones(with some overlap) and connected to per-zone 3D CNN(two blocks of identify block). It is quite simple network with not much parameters, so it just takes 10 minutes per epoch and the loss stablized around 80-130 epochs. Interesting thing is that I tried to share same per-zone network for symmetric zones(like left arm and right arm), but it came out worse. I think keeping separate network for two sides made some ensemble effects. </p>\n\n<p>I used basic augmentation like rotation, zoom, shift, flip left-right.</p>\n\n<p>I tried to separate each body part and put each part into the network, but it did much worse. I guess it's because NN losses context of body and have trouble finding whether the thread is the part of body or not. </p>\n\n<p>I didn't try much other experiments due to lack of time/resources. In hindsight, I wish I put little bit more time.</p>",
      "rawMarkdown": "At first, I went through aps images and I felt that many threats seem to be hard to find, so I decided to go with a3d.\nI used APS to identify each person and make train/val set - I manually annotated a few persons at first and let NN to classify the person. But later I found that PCA can do same thing easily from discussion board :(\n\nWhen I tried to feed a3d data, the data size was too big - so I scaled down to 1/4 on each dimension and cut down mostly empty areas on front/back/side/top). I could see the threats even after 1/4 scale, so I assumed that 1/4 dowscaling doesn't reduce the accuracy much. And I adjusted vertical scale little bit for each image so that the head sits on similar position(no ML, just simple image processing). Even after scaling down, only 2-3 batches fit into 1080Ti.\n\nNext, I fed the data into resnet-like 3D convolution. All data is fed to 3 blocks of identify block. Then output is splitted into 17 zones(with some overlap) and connected to per-zone 3D CNN(two blocks of identify block). It is quite simple network with not much parameters, so it just takes 10 minutes per epoch and the loss stablized around 80-130 epochs. Interesting thing is that I tried to share same per-zone network for symmetric zones(like left arm and right arm), but it came out worse. I think keeping separate network for two sides made some ensemble effects. \n\nI used basic augmentation like rotation, zoom, shift, flip left-right.\n\nI tried to separate each body part and put each part into the network, but it did much worse. I guess it's because NN losses context of body and have trouble finding whether the thread is the part of body or not. \n\nI didn't try much other experiments due to lack of time/resources. In hindsight, I wish I put little bit more time.",
      "votes": null
    },
    {
      "id": "258721",
      "postDate": "12/16/2017 22:02:30",
      "content": "<p>Congratulations to winners, \nthanks to all sharing their approach!</p>\n\n<p>My approach  was binary classification  with a separate set of  CNN models for each zone-view.\nClassifier models were fed with zone segments cropped from a3daps views. \nCropping coordinates were created by Unet segmentation models which I had to train. \nFor this I annotated 0,16,32 and 48-th views from 625 images.  </p>\n\n<p>To  combine models for different views of same zone segment I trained Xgboost logistic regression models.</p>\n\n<p>I used only  0,3,29,16,32,35,48 and 61 views, not all of them in each zone.\nClassifiers were Inception V3, Densenet161 and Resnet152.</p>\n\n<p>I also experimented with mosaic of 2 and 3 views for some zones. <br>\nNo ensembling of final models nor multi-fold validations.</p>\n\n<p>Stage 1 best score was 0.1074 with corresponding Stage2 of 0.1578.</p>\n\n<p>Segmentation worked pretty well on both stages. \nClassification was a real challenge.\nThe hardest zones were 2 and 4 with local validation scores of 0.156 and  0.171 respectively.\nThe easiest  zones were 5 and 17 with 0.045 and 0.012 scores.</p>\n\n<p>If I have time I would like to try approach close to Moejoe's and make some late submissions.\nI very much appreciate the ideas behind this much more elegant and effective approach.</p>",
      "rawMarkdown": "Congratulations to winners, \nthanks to all sharing their approach!\n\nMy approach  was binary classification  with a separate set of  CNN models for each zone-view.\nClassifier models were fed with zone segments cropped from a3daps views. \nCropping coordinates were created by Unet segmentation models which I had to train. \nFor this I annotated 0,16,32 and 48-th views from 625 images.  \n\nTo  combine models for different views of same zone segment I trained Xgboost logistic regression models.\n\nI used only  0,3,29,16,32,35,48 and 61 views, not all of them in each zone.\nClassifiers were Inception V3, Densenet161 and Resnet152.\n\nI also experimented with mosaic of 2 and 3 views for some zones.  \nNo ensembling of final models nor multi-fold validations.\n\nStage 1 best score was 0.1074 with corresponding Stage2 of 0.1578.\n\nSegmentation worked pretty well on both stages. \nClassification was a real challenge.\nThe hardest zones were 2 and 4 with local validation scores of 0.156 and  0.171 respectively.\nThe easiest  zones were 5 and 17 with 0.045 and 0.012 scores.\n\nIf I have time I would like to try approach close to Moejoe's and make some late submissions.\nI very much appreciate the ideas behind this much more elegant and effective approach.",
      "votes": null
    },
    {
      "id": "258734",
      "postDate": "12/16/2017 23:10:31",
      "content": "<p>... Training separate models, one for each zone is very suboptimal. It's easy to see that with the following argument: </p>\n\n<p>If you have ankle-strapped guns in the training dataset, the model that looks at ankles will learn to recognize them. However, if forearm-strapped guns only occur in the test dataset, they will be gratuitously novel to the model that looks at forearms. In other words, you won't transfer knowledge between zones as well as you should.</p>",
      "rawMarkdown": "... Training separate models, one for each zone is very suboptimal. It's easy to see that with the following argument: \n\nIf you have ankle-strapped guns in the training dataset, the model that looks at ankles will learn to recognize them. However, if forearm-strapped guns only occur in the test dataset, they will be gratuitously novel to the model that looks at forearms. In other words, you won't transfer knowledge between zones as well as you should.",
      "votes": null
    },
    {
      "id": "258742",
      "postDate": "12/16/2017 23:36:58",
      "content": "<p>Hi Moejoe,\nThank you very much for your explanation, it sounds abstract but really a clever approach. </p>",
      "rawMarkdown": "Hi Moejoe,\nThank you very much for your explanation, it sounds abstract but really a clever approach.",
      "votes": null
    },
    {
      "id": "258744",
      "postDate": "12/16/2017 23:52:32",
      "content": "<p>Hi Alexander, Thanks for sharing your experience.  I also used a segmentation approach with opencv which is not an effective way but works.  I'd like learn from others to improve my approach.  You used Unet segmentation models, could you explain a  little further about this model and if you can provide  any references?  Thank you.</p>",
      "rawMarkdown": "Hi Alexander, Thanks for sharing your experience.  I also used a segmentation approach with opencv which is not an effective way but works.  I'd like learn from others to improve my approach.  You used Unet segmentation models, could you explain a  little further about this model and if you can provide  any references?  Thank you.",
      "votes": null
    },
    {
      "id": "258750",
      "postDate": "12/17/2017 00:10:12",
      "content": "<p>My architecture:\nI worked with APS only, an MVCNN on all 16 views concurrently while also multi-task learning, in hopes of better generalization. I had 6 resnet-like layers on each view, then a average pooling to put all 16 views together and another 2 resnet-like layers then it split in two: 1) 4 layer resnet to train on the 17 threat labels and 2) a 4 layer resnet to train on the gender of the subject (which I manually labeled on the training data using PCA). I reduced the input images resolution with a factor 2 and used random crop and zoom to compensate for that. I also had the first two conv layers with a stride 2 to reduce the data.</p>\n\n<p>An interesting trick I found useful: You can <code>revolve</code> the weights from each of the view pillars. Take the weights from the 6 resnet layers of view 0 and put them in the layers of view 1, from view 1 to 2 and so on, finally put view 16 layers into view 0. If you do this every 20 epochs you'd expect you'll have to learn pretty much from start every time, but somewhat to my surprice the network doesn't lose it's ability much. Instead it prevents overfitting as  shown by the test-cross entropy which never surpassed my training cross entropy.</p>\n\n<p>Another interesting trick, I suppose most people have figured out is image mirroring for data augmentation. You have to be somewhat careful as the labels change too, e.g. a left arm threat becomes a right arm threat. Moreover the order of the APS views needs to be reversed to keep the subject as a left rotating subject. If you do that, mirror the image, reverse the views and mirror the labels you basically double the dataset. Sort of a standard trick, but with a little more to it in the case here, due to the nature of the data.</p>\n\n<p>I worked on a single GPU and couldn't get my batch size above 4, nor could I increase my network size, both due to memory constraints. Increasing the network size didn't seem to help much anyway. I'm not sure how much the gender training helped for model generalization. I did not test on proper test splits (proper as in separate subjects and threats for different folds), but my stage 1 score was 0.12 to 0.16 in stage 2.</p>\n\n<p>Things I would add next time:\n1) 5 fold training with proper separation of subjects and some boosted tree method to combine them\n2) Also separate threats, but I have no idea how\n3) Reduce image resolution more in the beginning\n4) More GPU's or different hardware to get a bigger model</p>\n\n<p>I don't think any of this will get me the winning model, wonder what the big difference is, what I'm missing. Hope to learn it here!</p>",
      "rawMarkdown": "My architecture:\nI worked with APS only, an MVCNN on all 16 views concurrently while also multi-task learning, in hopes of better generalization. I had 6 resnet-like layers on each view, then a average pooling to put all 16 views together and another 2 resnet-like layers then it split in two: 1) 4 layer resnet to train on the 17 threat labels and 2) a 4 layer resnet to train on the gender of the subject (which I manually labeled on the training data using PCA). I reduced the input images resolution with a factor 2 and used random crop and zoom to compensate for that. I also had the first two conv layers with a stride 2 to reduce the data.\n\nAn interesting trick I found useful: You can `revolve` the weights from each of the view pillars. Take the weights from the 6 resnet layers of view 0 and put them in the layers of view 1, from view 1 to 2 and so on, finally put view 16 layers into view 0. If you do this every 20 epochs you'd expect you'll have to learn pretty much from start every time, but somewhat to my surprice the network doesn't lose it's ability much. Instead it prevents overfitting as  shown by the test-cross entropy which never surpassed my training cross entropy.\n\nAnother interesting trick, I suppose most people have figured out is image mirroring for data augmentation. You have to be somewhat careful as the labels change too, e.g. a left arm threat becomes a right arm threat. Moreover the order of the APS views needs to be reversed to keep the subject as a left rotating subject. If you do that, mirror the image, reverse the views and mirror the labels you basically double the dataset. Sort of a standard trick, but with a little more to it in the case here, due to the nature of the data.\n\nI worked on a single GPU and couldn't get my batch size above 4, nor could I increase my network size, both due to memory constraints. Increasing the network size didn't seem to help much anyway. I'm not sure how much the gender training helped for model generalization. I did not test on proper test splits (proper as in separate subjects and threats for different folds), but my stage 1 score was 0.12 to 0.16 in stage 2.\n\nThings I would add next time:\n1) 5 fold training with proper separation of subjects and some boosted tree method to combine them\n2) Also separate threats, but I have no idea how\n3) Reduce image resolution more in the beginning\n4) More GPU's or different hardware to get a bigger model\n\nI don't think any of this will get me the winning model, wonder what the big difference is, what I'm missing. Hope to learn it here!",
      "votes": null
    },
    {
      "id": "258760",
      "postDate": "12/17/2017 00:56:53",
      "content": "<p>I have really banged my head against walls for how other got so much better than I, with or without overfitting I don't care. Can you describe some more details how you got to 0.025? I never got past 0.11,.. See my writeup elsewhere in this thread fo rmy architecture.</p>",
      "rawMarkdown": "I have really banged my head against walls for how other got so much better than I, with or without overfitting I don't care. Can you describe some more details how you got to 0.025? I never got past 0.11,.. See my writeup elsewhere in this thread fo rmy architecture.",
      "votes": null
    },
    {
      "id": "258761",
      "postDate": "12/17/2017 01:06:06",
      "content": "<p>I didn't use segmentation either. I felt it wasn't possible to proper segment the image and sometimes the threat was more easily visible from a strange angle. For example the ankle strapped threat on the inside of the right leg may be best viewed from the left (with the left leg obstructing most of the view).\nI'm not too worried about \"landmarks\" as other have mentioned as the pixel position remains \"known\" throughout your convolutional layers until the resolution gets very small at which point it supposedly has figured out towards which label we're going.</p>",
      "rawMarkdown": "I didn't use segmentation either. I felt it wasn't possible to proper segment the image and sometimes the threat was more easily visible from a strange angle. For example the ankle strapped threat on the inside of the right leg may be best viewed from the left (with the left leg obstructing most of the view).\nI'm not too worried about \"landmarks\" as other have mentioned as the pixel position remains \"known\" throughout your convolutional layers until the resolution gets very small at which point it supposedly has figured out towards which label we're going.",
      "votes": null
    },
    {
      "id": "258768",
      "postDate": "12/17/2017 01:37:16",
      "content": "<p>That argument only applies if what is going on is that the network learns to detect threats and not that the network learns to detect normalcy.  It is my interpretation of my own models that they were in fact doing the later and not the former.</p>",
      "rawMarkdown": "That argument only applies if what is going on is that the network learns to detect threats and not that the network learns to detect normalcy.  It is my interpretation of my own models that they were in fact doing the later and not the former.",
      "votes": null
    },
    {
      "id": "258773",
      "postDate": "12/17/2017 01:43:26",
      "content": "<p>I found one paper in particular that was very helpful: \n<a href=\"https://arxiv.org/pdf/1703.07047.pdf\">High-Resolution Breast Cancer Screening with Multi-View Deep Convolutional Neural Networks</a></p>\n\n<p>My net was very similar to the one described in this paper. I added a convolutional layer with 17 output channels right before global average pooling. I optimized the layer parameters, but as described above it seems that the optimization was performed under unrealistic test conditions. </p>\n\n<p>I trained the net using SGD + Momentum + Nesterov with lr=0.4 and momentum=0.32. I settled on these unusual values after running an extensive random search, and was very surprised that the optimal momentum turned out to be so low. I also got a lot of unusual \"spikes\" in the train and test error during training. The graph below shows the log loss on one of the cross validation folds, with exponential lr decay. I will be very grateful if someone can offer insights into this unusual behavior. </p>\n\n<p><img src=\"http://seansoleyman.com/wp-content/uploads/eval.png\" alt=\"Learning Curve\" title=\"\"></p>",
      "rawMarkdown": "I found one paper in particular that was very helpful: \n[High-Resolution Breast Cancer Screening with Multi-View Deep Convolutional Neural Networks][1]\n\nMy net was very similar to the one described in this paper. I added a convolutional layer with 17 output channels right before global average pooling. I optimized the layer parameters, but as described above it seems that the optimization was performed under unrealistic test conditions. \n\nI trained the net using SGD + Momentum + Nesterov with lr=0.4 and momentum=0.32. I settled on these unusual values after running an extensive random search, and was very surprised that the optimal momentum turned out to be so low. I also got a lot of unusual \"spikes\" in the train and test error during training. The graph below shows the log loss on one of the cross validation folds, with exponential lr decay. I will be very grateful if someone can offer insights into this unusual behavior. \n\n![Learning Curve][2]\n\n  [1]: https://arxiv.org/pdf/1703.07047.pdf\n  [2]: http://seansoleyman.com/wp-content/uploads/eval.png",
      "votes": null
    },
    {
      "id": "258775",
      "postDate": "12/17/2017 01:52:26",
      "content": "<p>I used two networks, one using about 15% of training data to classify threats, and the other had about 1% of training data to classify categories.  I used an IoU metric to match the location of the detection. So in total 2 resnet networks trained on a gtx 1070. Only data augmentation was horizontal flips. and only a3daps images.  I think the approach could yield a much better result with more of the training data used and more time to refine the model.</p>",
      "rawMarkdown": "I used two networks, one using about 15% of training data to classify threats, and the other had about 1% of training data to classify categories.  I used an IoU metric to match the location of the detection. So in total 2 resnet networks trained on a gtx 1070. Only data augmentation was horizontal flips. and only a3daps images.  I think the approach could yield a much better result with more of the training data used and more time to refine the model.",
      "votes": null
    },
    {
      "id": "258776",
      "postDate": "12/17/2017 02:02:26",
      "content": "<p>I don't think it's an either-or type of situation. It helps to know what normal bodies look like and it helps to know what the threats look like if you aim to tell them apart. Quantitatively though, the image patches containing threats were much more scarce. Squandering them seems to me like a serious mistake.</p>",
      "rawMarkdown": "I don't think it's an either-or type of situation. It helps to know what normal bodies look like and it helps to know what the threats look like if you aim to tell them apart. Quantitatively though, the image patches containing threats were much more scarce. Squandering them seems to me like a serious mistake.",
      "votes": null
    },
    {
      "id": "258778",
      "postDate": "12/17/2017 02:14:16",
      "content": "<p>Hi Wensu,\nThis is the original paper regarding Unet:\n<a href=\"https://arxiv.org/pdf/1505.04597.pdf\">https://arxiv.org/pdf/1505.04597.pdf</a></p>\n\n<p>As for me, I adapted Unet segmentation code from Carvana Image Masking challenge posted here:\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208#latest-222649\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208#latest-222649</a></p>\n\n<p>I created very simple masks  from my annotations. </p>",
      "rawMarkdown": "Hi Wensu,\nThis is the original paper regarding Unet:\nhttps://arxiv.org/pdf/1505.04597.pdf\n\nAs for me, I adapted Unet segmentation code from Carvana Image Masking challenge posted here:\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208#latest-222649\n\nI created very simple masks  from my annotations.",
      "votes": null
    },
    {
      "id": "258790",
      "postDate": "12/17/2017 03:05:05",
      "content": "<p>Very valuable points., Oleg. Besides, segmentation models rely on annotations and their quality. Plus complexity and extra work... Much better when single net cares about zone boundaries and threats. Still I reserve my right to make serious mistakes.</p>",
      "rawMarkdown": "Very valuable points., Oleg. Besides, segmentation models rely on annotations and their quality. Plus complexity and extra work... Much better when single net cares about zone boundaries and threats. Still I reserve my right to make serious mistakes.",
      "votes": null
    },
    {
      "id": "258820",
      "postDate": "12/17/2017 04:56:07",
      "content": "<p>Wohaa! You cut straight through to 0.015, you don't want to know how long I have been watching my charts slowly inching down from 0.25 to 0.18 on the test set, never getting further than that. The last bit, from 0.18 to 0.12 I did by repeating the infer at 72 different crops and resizes + mirroring. But then again, you heavily overfit making your final score not all that much better than mine ;-). I wonder if the number of feature maps in the conv layers may have to do with that. I have been very conservative with a maximum of 80 feature maps, the paper you refer to has 256. My rationale was, these standard nets (Resnet, VGG etc) are designed and tested for millions of images with thousands of classes, here we have a lot less of both so the network needs to be scaled down accordingly. Maybe I have been to aggressive with that.</p>\n\n<p>The peak in your training, is that not caused by the faulty image in the APS series? There is one that is half blacked out. Or otherwise maybe the crazy momentum?</p>\n\n<p>Did you do any tricks at infer time? Did that improve matters?</p>",
      "rawMarkdown": "Wohaa! You cut straight through to 0.015, you don't want to know how long I have been watching my charts slowly inching down from 0.25 to 0.18 on the test set, never getting further than that. The last bit, from 0.18 to 0.12 I did by repeating the infer at 72 different crops and resizes + mirroring. But then again, you heavily overfit making your final score not all that much better than mine ;-). I wonder if the number of feature maps in the conv layers may have to do with that. I have been very conservative with a maximum of 80 feature maps, the paper you refer to has 256. My rationale was, these standard nets (Resnet, VGG etc) are designed and tested for millions of images with thousands of classes, here we have a lot less of both so the network needs to be scaled down accordingly. Maybe I have been to aggressive with that.\n\nThe peak in your training, is that not caused by the faulty image in the APS series? There is one that is half blacked out. Or otherwise maybe the crazy momentum?\n\nDid you do any tricks at infer time? Did that improve matters?",
      "votes": null
    },
    {
      "id": "258825",
      "postDate": "12/17/2017 05:07:33",
      "content": "<p>This seems interesting, thanks for sharing! What do you mean with classify threats versus classify categories? What is the difference? And what did you do with the other 84% ot the training data?</p>\n\n<p>I got a great improvement by using random crop and zoom, which I believe has to do with the image reduction I applied. You may see a similar gain with that.</p>",
      "rawMarkdown": "This seems interesting, thanks for sharing! What do you mean with classify threats versus classify categories? What is the difference? And what did you do with the other 84% ot the training data?\n\nI got a great improvement by using random crop and zoom, which I believe has to do with the image reduction I applied. You may see a similar gain with that.",
      "votes": null
    },
    {
      "id": "258830",
      "postDate": "12/17/2017 05:14:46",
      "content": "<p>Thanks for sharing! How did you \"split the data into 17 zones with some overlap\"? You're not referring to segmentation, do you?</p>",
      "rawMarkdown": "Thanks for sharing! How did you \"split the data into 17 zones with some overlap\"? You're not referring to segmentation, do you?",
      "votes": null
    },
    {
      "id": "258842",
      "postDate": "12/17/2017 05:37:45",
      "content": "<p>Actually segmentation wasn't very difficult to get pretty accurate if working with a3d images. While there are 16 &amp; 64 views with aps/a3daps there is really just one aerial view for a3d so all you needed to get right was the z-axis coordinates and the rest you could rely on basic symmetry of the human body. I used a simple CNN on a3daps images to predict the height range coordinates for each mirror zone pair (6/7, 2/4 etc.) and then  later any threat that my classifier found in that slice range was matched up with the corresponding zone. </p>\n\n<p>Although I definitely think Moejoe's \"automated\" solution is much more elegant and wish I had thought of it :)</p>",
      "rawMarkdown": "Actually segmentation wasn't very difficult to get pretty accurate if working with a3d images. While there are 16 &amp; 64 views with aps/a3daps there is really just one aerial view for a3d so all you needed to get right was the z-axis coordinates and the rest you could rely on basic symmetry of the human body. I used a simple CNN on a3daps images to predict the height range coordinates for each mirror zone pair (6/7, 2/4 etc.) and then  later any threat that my classifier found in that slice range was matched up with the corresponding zone. \n\nAlthough I definitely think Moejoe's \"automated\" solution is much more elegant and wish I had thought of it :)",
      "votes": null
    },
    {
      "id": "258845",
      "postDate": "12/17/2017 05:50:12",
      "content": "<p>So how did you do it? Used and MVCNN? ;-)</p>",
      "rawMarkdown": "So how did you do it? Used and MVCNN? ;-)",
      "votes": null
    },
    {
      "id": "258848",
      "postDate": "12/17/2017 06:09:11",
      "content": "<p>It's not segmentation, but just trick to reduce computation since 3D convolution requires a lot even at  1/4 scale.\nAfter 3 resnet block for the whole body, output dimensions are like 10(front-back) x 12(left-right) x 18(foot-hands) x # of last layer filters. I just slice this output to 17 zones - for example, I sliced it like x[2:6, 2:4, 0:1, :] for left feet, so this per-zone CNN only see small area only on center-left-bottom part.</p>\n\n<p>I think the network will eventually figure out the area of interest even if I didn't use this per-zone network, but it will require network with much more capacities and more training time. </p>",
      "rawMarkdown": "It's not segmentation, but just trick to reduce computation since 3D convolution requires a lot even at  1/4 scale.\nAfter 3 resnet block for the whole body, output dimensions are like 10(front-back) x 12(left-right) x 18(foot-hands) x # of last layer filters. I just slice this output to 17 zones - for example, I sliced it like x[2:6, 2:4, 0:1, :] for left feet, so this per-zone CNN only see small area only on center-left-bottom part.\n\nI think the network will eventually figure out the area of interest even if I didn't use this per-zone network, but it will require network with much more capacities and more training time.",
      "votes": null
    },
    {
      "id": "258870",
      "postDate": "12/17/2017 07:37:43",
      "content": "<blockquote>\n  <p><strong>Bastiaan Bergman wrote</strong></p>\n  \n  <blockquote>\n    <p>Wohaa! You cut straight through to 0.015, you don't want to know how long I have been watching my charts slowly inching down from 0.25 to 0.18 on the test set, never getting further than that. </p>\n  </blockquote>\n</blockquote>\n\n<p>A watched pot never boils. The secret is to ignore them. Do what you would normally do through the day. When you finally remember that you are training something and look at the log, then you'll find yourself thinking \"whoa, awesome results!\"</p>",
      "rawMarkdown": "&gt; **Bastiaan Bergman wrote**\n&gt; \n&gt; &gt; Wohaa! You cut straight through to 0.015, you don't want to know how long I have been watching my charts slowly inching down from 0.25 to 0.18 on the test set, never getting further than that. \n\nA watched pot never boils. The secret is to ignore them. Do what you would normally do through the day. When you finally remember that you are training something and look at the log, then you'll find yourself thinking \"whoa, awesome results!\"",
      "votes": null
    },
    {
      "id": "259089",
      "postDate": "12/17/2017 17:11:49",
      "content": "<p>This was after I had removed the bad instance and switched the two sets of labels. This was one of the better CV folds - they averaged out to around 0.025. </p>\n\n<p>I did use a large number of feature maps (somewhere around 256 for the upper layers) because this helped speed up training and reduced the stage1 error. Looking back, I probably would have done better without such a large number of feature maps. Dropout may also have improved stage2 results by preventing the net from memorizing specific features, although it did not help much with stage1. </p>\n\n<p>Much of this is just speculation - I wouldn't read too much into these details!</p>",
      "rawMarkdown": "This was after I had removed the bad instance and switched the two sets of labels. This was one of the better CV folds - they averaged out to around 0.025. \n\nI did use a large number of feature maps (somewhere around 256 for the upper layers) because this helped speed up training and reduced the stage1 error. Looking back, I probably would have done better without such a large number of feature maps. Dropout may also have improved stage2 results by preventing the net from memorizing specific features, although it did not help much with stage1. \n\nMuch of this is just speculation - I wouldn't read too much into these details!",
      "votes": null
    },
    {
      "id": "259134",
      "postDate": "12/17/2017 18:57:44",
      "content": "<p>Hi Alexander, thank you very much for the references you provided.</p>",
      "rawMarkdown": "Hi Alexander, thank you very much for the references you provided.",
      "votes": null
    },
    {
      "id": "259183",
      "postDate": "12/17/2017 21:27:33",
      "content": "<p>I started with zone segmentation from the very beginning. The main reason for this is I've been an engineer for a long time but have only recently gotten into ML. I don't have a lot of tools in the ML toolbox but I can work through an engineering problem like segmenting an image (though a lot of that was new too!). Also I thought I might have had a puncher's chance at doing well because I figured not a lot of teams would spend time on meticulously segmenting the zones. I really didn't know how else to proceed but my solution certainly evolved as I went along.</p>\n\n<p>TL;DR - I did a reasonable job at segmenting the zones</p>\n\n<p>For my first segmentation attempt I took the max point of all slices in the dataset and tried to hand-curate the zones. I didn't like those results so then I decided if I could determine a few critical points, namely where the torso began (at the groin) and ended (at the neck) along with the torso width I could assume a universal body proportion and accurately segment the zones. I was able to determine the points pretty easily with the skimage library and it did seem like most people had similar relative proportions. Several caveats: 1) an individual could be off center on either floor axis which required an adjustment (this could be determined by looking at the center point of slices 0,8 and 4,12), 2) some subjects had excessive...um...girth so that would require an adjustment across slices and finally 3) by observing scanners at airports, I concluded that the scanner would start slow, speed up, then slow down again which required a further adjustment. To all of this, I added some padding for each zone and it seemed I had a pretty robust zone segmenter - it even looked good in the new stage 2 dataset.</p>\n\n<p>TL;DR - Used Keras' TimeDistributed layer with InceptionV3 freezing the first 172 layers and adding an LSTM and dense layer at the end. One classifier for zones 6,7,8,10-16, one for the arm zones and one each for 5, 17 and 9. Weighted threats 90/10.</p>\n\n<p>My model was relatively simple and I never experimented with anything else. I essentially viewed the input as images across time (ie. video). So I used Keras' TimeDistributed layer with a pre-trained Imagenet model (VGG16 and InceptionV3) fed into an LSTM to a dense layer and finally to a binary result. At this point it was all manual experimentation. I tried different slices in each zone, different image input sizes, different sizes for LSTM and dense layer and playing with un-freezing various layers in the pre-trained models. At first I did not fully segment the arm and leg zones (for ex. zones 11, 13, 15 were one classifier) but ultimately I achieved a nice jump on the leaderboard by segmenting the legs so that reinforced the idea, perhaps incorrectly, that segmenting was the way to go. In the end I used 150x150 size images, with 9 time slices (save for 5, 17 and 9 which had less) and then used InceptionV3 in the TimeDistributed layer with the first 172 layers frozen and the rest trainable.  I weighted threats 90/10 in gradient updates since they were so much less prevalent. I discovered the symmetry trick so I started pairing zones (like [6,7], [8,10], [11,12], etc.) by flipping horizontally. To additionally augment, I flipped the temporal axis. Then I thought why not combine more zones? So I did all the leg zones in one classifier and all the torso zones. I kept moving up in the LB so finally I just did every 9-slice zone in one giant classifier and started flipping vertically to additionally augment. The end result: one classifier for zones 6,7,8,10-16,  one for the arm zones and one each for zones 5, 17 and 9. I reached 0.108 on the public LB. I discovered all this combining at the end so I think I could've improved that if I had more time to train.</p>\n\n<p>I thought I had something with the mega-classifier. It allowed me to greatly augment the data and I figured I was relatively immune to new people being introduced in stage 2 since zone segments I would think lack landmark features and by using all the zones together and flipping them around the network could focus on the threats. One glaring problem with segmenting the zones is many of the threats seemed to span more than one zone even though they were labelled for one zone. So I'm sure that confused the network in that it maybe only identified a threat when it was in the center.</p>\n\n<p>My stage 2 results more than doubled. Bob Marley and suspender-dude didn't help. Also in my brief manual inspection, it seemed like there were new, subtle-looking threats that weren't in the stage 1 data. I also didn't account for mislabels by pulling back any 1.0 predictions to 0.999. Live and learn. This competition was a ton of fun though and looking forward to the next one!</p>",
      "rawMarkdown": "I started with zone segmentation from the very beginning. The main reason for this is I've been an engineer for a long time but have only recently gotten into ML. I don't have a lot of tools in the ML toolbox but I can work through an engineering problem like segmenting an image (though a lot of that was new too!). Also I thought I might have had a puncher's chance at doing well because I figured not a lot of teams would spend time on meticulously segmenting the zones. I really didn't know how else to proceed but my solution certainly evolved as I went along.\n\nTL;DR - I did a reasonable job at segmenting the zones\n\nFor my first segmentation attempt I took the max point of all slices in the dataset and tried to hand-curate the zones. I didn't like those results so then I decided if I could determine a few critical points, namely where the torso began (at the groin) and ended (at the neck) along with the torso width I could assume a universal body proportion and accurately segment the zones. I was able to determine the points pretty easily with the skimage library and it did seem like most people had similar relative proportions. Several caveats: 1) an individual could be off center on either floor axis which required an adjustment (this could be determined by looking at the center point of slices 0,8 and 4,12), 2) some subjects had excessive...um...girth so that would require an adjustment across slices and finally 3) by observing scanners at airports, I concluded that the scanner would start slow, speed up, then slow down again which required a further adjustment. To all of this, I added some padding for each zone and it seemed I had a pretty robust zone segmenter - it even looked good in the new stage 2 dataset.\n\nTL;DR - Used Keras' TimeDistributed layer with InceptionV3 freezing the first 172 layers and adding an LSTM and dense layer at the end. One classifier for zones 6,7,8,10-16, one for the arm zones and one each for 5, 17 and 9. Weighted threats 90/10.\n\nMy model was relatively simple and I never experimented with anything else. I essentially viewed the input as images across time (ie. video). So I used Keras' TimeDistributed layer with a pre-trained Imagenet model (VGG16 and InceptionV3) fed into an LSTM to a dense layer and finally to a binary result. At this point it was all manual experimentation. I tried different slices in each zone, different image input sizes, different sizes for LSTM and dense layer and playing with un-freezing various layers in the pre-trained models. At first I did not fully segment the arm and leg zones (for ex. zones 11, 13, 15 were one classifier) but ultimately I achieved a nice jump on the leaderboard by segmenting the legs so that reinforced the idea, perhaps incorrectly, that segmenting was the way to go. In the end I used 150x150 size images, with 9 time slices (save for 5, 17 and 9 which had less) and then used InceptionV3 in the TimeDistributed layer with the first 172 layers frozen and the rest trainable.  I weighted threats 90/10 in gradient updates since they were so much less prevalent. I discovered the symmetry trick so I started pairing zones (like [6,7], [8,10], [11,12], etc.) by flipping horizontally. To additionally augment, I flipped the temporal axis. Then I thought why not combine more zones? So I did all the leg zones in one classifier and all the torso zones. I kept moving up in the LB so finally I just did every 9-slice zone in one giant classifier and started flipping vertically to additionally augment. The end result: one classifier for zones 6,7,8,10-16,  one for the arm zones and one each for zones 5, 17 and 9. I reached 0.108 on the public LB. I discovered all this combining at the end so I think I could've improved that if I had more time to train.\n\nI thought I had something with the mega-classifier. It allowed me to greatly augment the data and I figured I was relatively immune to new people being introduced in stage 2 since zone segments I would think lack landmark features and by using all the zones together and flipping them around the network could focus on the threats. One glaring problem with segmenting the zones is many of the threats seemed to span more than one zone even though they were labelled for one zone. So I'm sure that confused the network in that it maybe only identified a threat when it was in the center.\n\nMy stage 2 results more than doubled. Bob Marley and suspender-dude didn't help. Also in my brief manual inspection, it seemed like there were new, subtle-looking threats that weren't in the stage 1 data. I also didn't account for mislabels by pulling back any 1.0 predictions to 0.999. Live and learn. This competition was a ton of fun though and looking forward to the next one!",
      "votes": null
    },
    {
      "id": "259200",
      "postDate": "12/17/2017 23:01:46",
      "content": "<p>As promised, I'll post my architecture.  Obviously it must have over-fit to the stage 1 participants.  In hindsight I should have split my validation by individual people.  But anyway I guess it was my first real competition and I learned a valuable lesson for next time I do a 2 stager :P</p>\n\n<h2>Normalization</h2>\n\n<p>I used only aps samples.  Height was the most variable feature, so I stretched each person to be the same height.  I chose the height based on the highest pixel above a certain threshold, it actually worked extremely well and reliably.  Once normalized, I created crop locations for each zone in each rotation, skipping occluded channels.</p>\n\n<h2>Augmentation</h2>\n\n<p>I doubled the sample size by including mirror images.  My crop areas were 160x160 pixels, but in the actual training, I randomly sampled 128x128 crops from that larger area to sample random translations.</p>\n\n<h2>Architecture</h2>\n\n<p>Each zone had a different fully convolutional network.  I couldn't use vanilla convolutions since there is no spatial locality between rotations, so I created an 'Independent Convolution' layer that partitions the rotations and applies the same convolutions independently to each rotation.  I had 5 layers of convolutions and pooling before a single dense layer that combines the partitions.</p>\n\n<p>None of my networks had more than 14,000 parameters, usually a lot less.  Keeping a low parameter count was how I avoided over-fitting (at least to the stage 1 validation I guess).  Although it was different for each zone, here's how a typical network looks:</p>\n\n<pre><code>_________________________________________________________________\nLayer (type)                 Output Shape              Param #\n=================================================================\ninput                        (None, 14, 128, 128)      0\n_________________________________________________________________\nind_conv2d_1 (IndConv2D)     (None, 168, 124, 124)     312\n_________________________________________________________________\nmax_pooling2d_1 (MaxPooling2 (None, 168, 62, 62)       0\n_________________________________________________________________\nactivation_1 (Activation)    (None, 168, 62, 62)       0\n_________________________________________________________________\nind_conv2d_2 (IndConv2D)     (None, 168, 60, 60)       1308\n_________________________________________________________________\nmax_pooling2d_2 (MaxPooling2 (None, 168, 30, 30)       0\n_________________________________________________________________\nactivation_2 (Activation)    (None, 168, 30, 30)       0\n_________________________________________________________________\nind_conv2d_3 (IndConv2D)     (None, 168, 28, 28)       1308\n_________________________________________________________________\nmax_pooling2d_3 (MaxPooling2 (None, 168, 14, 14)       0\n_________________________________________________________________\nactivation_3 (Activation)    (None, 168, 14, 14)       0\n_________________________________________________________________\nind_conv2d_4 (IndConv2D)     (None, 168, 12, 12)       1308\n_________________________________________________________________\nmax_pooling2d_4 (MaxPooling2 (None, 168, 6, 6)         0\n_________________________________________________________________\nactivation_4 (Activation)    (None, 168, 6, 6)         0\n_________________________________________________________________\nind_conv2d_5 (IndConv2D)     (None, 168, 4, 4)         1308\n_________________________________________________________________\nflatten_1 (Flatten)          (None, 2688)              0\n_________________________________________________________________\ndropout_1 (Dropout)          (None, 2688)              0\n_________________________________________________________________\ndense_1 (Dense)              (None, 1)                 2689\n_________________________________________________________________\nactivation_5 (Activation)    (None, 1)                 0\n=================================================================\nTotal params: 8,233\n</code></pre>\n\n<h2>Hyperparameter Optimization</h2>\n\n<p>The only hyper-parameters I would tune were the amount of dropout at the dense layer and the number of kernels in the convolutions, which was the same for every layer.</p>\n\n<h2>Ensembleling</h2>\n\n<p>I didn't use ensembling in the traditional sense.  Instead of using multiple models, I just also ran the mirror image through the same model (mirroring the results) as well as 9 different translations for each.  Then I averaged the results, which had to be done pre-sigmoid for statistical correctness.</p>\n\n<h2>Conclusion</h2>\n\n<p>The low parameter count and low stage 1 score gave me too much confidence that \"it couldn't possibly be over-fitting\" but obviously in hindsight I should have tested how well it generalized to an unseen body by re-splitting the data.</p>",
      "rawMarkdown": "As promised, I'll post my architecture.  Obviously it must have over-fit to the stage 1 participants.  In hindsight I should have split my validation by individual people.  But anyway I guess it was my first real competition and I learned a valuable lesson for next time I do a 2 stager :P\n\nNormalization\n-------------\nI used only aps samples.  Height was the most variable feature, so I stretched each person to be the same height.  I chose the height based on the highest pixel above a certain threshold, it actually worked extremely well and reliably.  Once normalized, I created crop locations for each zone in each rotation, skipping occluded channels.\n\nAugmentation\n-------------\nI doubled the sample size by including mirror images.  My crop areas were 160x160 pixels, but in the actual training, I randomly sampled 128x128 crops from that larger area to sample random translations.\n\nArchitecture\n-------------\nEach zone had a different fully convolutional network.  I couldn't use vanilla convolutions since there is no spatial locality between rotations, so I created an 'Independent Convolution' layer that partitions the rotations and applies the same convolutions independently to each rotation.  I had 5 layers of convolutions and pooling before a single dense layer that combines the partitions.\n\nNone of my networks had more than 14,000 parameters, usually a lot less.  Keeping a low parameter count was how I avoided over-fitting (at least to the stage 1 validation I guess).  Although it was different for each zone, here's how a typical network looks:\n\n    _________________________________________________________________\n    Layer (type)                 Output Shape              Param #\n    =================================================================\n    input                        (None, 14, 128, 128)      0\n    _________________________________________________________________\n    ind_conv2d_1 (IndConv2D)     (None, 168, 124, 124)     312\n    _________________________________________________________________\n    max_pooling2d_1 (MaxPooling2 (None, 168, 62, 62)       0\n    _________________________________________________________________\n    activation_1 (Activation)    (None, 168, 62, 62)       0\n    _________________________________________________________________\n    ind_conv2d_2 (IndConv2D)     (None, 168, 60, 60)       1308\n    _________________________________________________________________\n    max_pooling2d_2 (MaxPooling2 (None, 168, 30, 30)       0\n    _________________________________________________________________\n    activation_2 (Activation)    (None, 168, 30, 30)       0\n    _________________________________________________________________\n    ind_conv2d_3 (IndConv2D)     (None, 168, 28, 28)       1308\n    _________________________________________________________________\n    max_pooling2d_3 (MaxPooling2 (None, 168, 14, 14)       0\n    _________________________________________________________________\n    activation_3 (Activation)    (None, 168, 14, 14)       0\n    _________________________________________________________________\n    ind_conv2d_4 (IndConv2D)     (None, 168, 12, 12)       1308\n    _________________________________________________________________\n    max_pooling2d_4 (MaxPooling2 (None, 168, 6, 6)         0\n    _________________________________________________________________\n    activation_4 (Activation)    (None, 168, 6, 6)         0\n    _________________________________________________________________\n    ind_conv2d_5 (IndConv2D)     (None, 168, 4, 4)         1308\n    _________________________________________________________________\n    flatten_1 (Flatten)          (None, 2688)              0\n    _________________________________________________________________\n    dropout_1 (Dropout)          (None, 2688)              0\n    _________________________________________________________________\n    dense_1 (Dense)              (None, 1)                 2689\n    _________________________________________________________________\n    activation_5 (Activation)    (None, 1)                 0\n    =================================================================\n    Total params: 8,233\n\nHyperparameter Optimization\n-------------\nThe only hyper-parameters I would tune were the amount of dropout at the dense layer and the number of kernels in the convolutions, which was the same for every layer.\n\nEnsembleling\n-------------\nI didn't use ensembling in the traditional sense.  Instead of using multiple models, I just also ran the mirror image through the same model (mirroring the results) as well as 9 different translations for each.  Then I averaged the results, which had to be done pre-sigmoid for statistical correctness.\n\nConclusion\n-------------\nThe low parameter count and low stage 1 score gave me too much confidence that \"it couldn't possibly be over-fitting\" but obviously in hindsight I should have tested how well it generalized to an unseen body by re-splitting the data.",
      "votes": null
    },
    {
      "id": "259294",
      "postDate": "12/18/2017 04:41:13",
      "content": "<p>Hi Kevin,\nCongratulations, a great score nevertheless! And thanks for the write up.\nI wonder what exactly your 'Independent Convolution' is about. Is it convolution on the image section of one particular threat zone? Which is different for each of the 16 views?</p>",
      "rawMarkdown": "Hi Kevin,\nCongratulations, a great score nevertheless! And thanks for the write up.\nI wonder what exactly your 'Independent Convolution' is about. Is it convolution on the image section of one particular threat zone? Which is different for each of the 16 views?",
      "votes": null
    },
    {
      "id": "259297",
      "postDate": "12/18/2017 04:48:58",
      "content": "<blockquote>\n  <p>A watched pot never boils.</p>\n</blockquote>\n\n<p>Yeah, I tried that too,.. something tells me you had more in your secret sauce ;-)</p>",
      "rawMarkdown": "&gt; A watched pot never boils.\n\nYeah, I tried that too,.. something tells me you had more in your secret sauce ;-)",
      "votes": null
    },
    {
      "id": "259299",
      "postDate": "12/18/2017 04:55:54",
      "content": "<p>Sort of yeah, I think its similar to a MVCNN, but instead of pooling to combine the views, I just use a dense layer.  The weights are shared for each of the 12-15 views in the zone.  I don't think any zone actually used 16 views since there was occlusion in at least some of them.</p>",
      "rawMarkdown": "Sort of yeah, I think its similar to a MVCNN, but instead of pooling to combine the views, I just use a dense layer.  The weights are shared for each of the 12-15 views in the zone.  I don't think any zone actually used 16 views since there was occlusion in at least some of them.",
      "votes": null
    },
    {
      "id": "259512",
      "postDate": "12/18/2017 14:40:17",
      "content": "<p>Hi, sorry if my question is dumb, but what is this pre-sigmoid averaging? Why would this bring statistical correctness?</p>",
      "rawMarkdown": "Hi, sorry if my question is dumb, but what is this pre-sigmoid averaging? Why would this bring statistical correctness?",
      "votes": null
    },
    {
      "id": "259574",
      "postDate": "12/18/2017 16:45:43",
      "content": "<p>I was wondering the same ;-)</p>",
      "rawMarkdown": "I was wondering the same ;-)",
      "votes": null
    },
    {
      "id": "259594",
      "postDate": "12/18/2017 17:25:55",
      "content": "<p>So when using log-loss, the resulting probability should be interpreted as \"Confidence\".  When you have one network with 50% confidence, and one with 99.999% you definitely wouldn't want to say that it's only 75% confident!  The other network is exponentially more confident.  The difference between 99% and 99.9% is much bigger than 50% and 50.9%.  This whole argument is of course specific to log-loss scoring.  </p>\n\n<p>Because the space is stretched like this, the correct way to combine the values is to evaluate the network all the way down to right before the final sigmoid, average, and then finally apply the sigmoid once everything's been averaged together.  This transforms the weird 'sigmoidy space' to a nice linear space where averaging makes more sense.</p>\n\n<p>Overall this gave me a 30% improvement on the LB score compared to just averaging after the sigmoid.</p>",
      "rawMarkdown": "So when using log-loss, the resulting probability should be interpreted as \"Confidence\".  When you have one network with 50% confidence, and one with 99.999% you definitely wouldn't want to say that it's only 75% confident!  The other network is exponentially more confident.  The difference between 99% and 99.9% is much bigger than 50% and 50.9%.  This whole argument is of course specific to log-loss scoring.  \n\nBecause the space is stretched like this, the correct way to combine the values is to evaluate the network all the way down to right before the final sigmoid, average, and then finally apply the sigmoid once everything's been averaged together.  This transforms the weird 'sigmoidy space' to a nice linear space where averaging makes more sense.\n\nOverall this gave me a 30% improvement on the LB score compared to just averaging after the sigmoid.",
      "votes": null
    },
    {
      "id": "259684",
      "postDate": "12/18/2017 20:56:56",
      "content": "<p>Thanks for posting this, Kevin; sorry that I was wrong about you being the obvious winner, but congratulations all the same -- hey, you beat me at least!  :)</p>",
      "rawMarkdown": "Thanks for posting this, Kevin; sorry that I was wrong about you being the obvious winner, but congratulations all the same -- hey, you beat me at least!  :)",
      "votes": null
    },
    {
      "id": "259789",
      "postDate": "12/19/2017 01:50:46",
      "content": "<p>Overall, I concatenated threat zone crops from the aps files into \"threat plates\" and trained an ensemble of ResNet50 (Imagenet weights) using those. I normalized the data by cropping each of the 16 images around a certain threshold and resizing that to be the full image, so that threat zone locations would be more consistent. Beyond that, there were two key insights behind my solution:</p>\n\n<ol>\n<li><p>Exploit the symmetry of the human body. By switching and horizontally flipping certain images I could use the same crops for right and left forearm, right and left shin, etc. Thus I doubled the size of the dataset and only needed 9 networks for a full segmentation approach.</p></li>\n<li><p>Clipping. Based on the math of the log loss error function, on an inaccurate prediction, overconfidence will be penalized really harshly. So I clipped my predictions to the [0.015, 0.985] interval and this took me from bronze to silver.</p></li>\n</ol>\n\n<p>I really appreciate all the help I got when I asked! I think the community spirit is one of the greatest aspects of Kaggle. Here's a repo with more info and my code - <a href=\"https://github.com/jamespeterthornton/DHS\">https://github.com/jamespeterthornton/DHS</a></p>",
      "rawMarkdown": "Overall, I concatenated threat zone crops from the aps files into \"threat plates\" and trained an ensemble of ResNet50 (Imagenet weights) using those. I normalized the data by cropping each of the 16 images around a certain threshold and resizing that to be the full image, so that threat zone locations would be more consistent. Beyond that, there were two key insights behind my solution:\n\n1. Exploit the symmetry of the human body. By switching and horizontally flipping certain images I could use the same crops for right and left forearm, right and left shin, etc. Thus I doubled the size of the dataset and only needed 9 networks for a full segmentation approach.\n\n2. Clipping. Based on the math of the log loss error function, on an inaccurate prediction, overconfidence will be penalized really harshly. So I clipped my predictions to the [0.015, 0.985] interval and this took me from bronze to silver.\n\nI really appreciate all the help I got when I asked! I think the community spirit is one of the greatest aspects of Kaggle. Here's a repo with more info and my code - https://github.com/jamespeterthornton/DHS",
      "votes": null
    },
    {
      "id": "260002",
      "postDate": "12/19/2017 11:46:36",
      "content": "<p>Hello, Thanks for sharing. I don't have enough hardware to retrain the ResNet50, so can you share the re-trained ResNet50 model weights? </p>",
      "rawMarkdown": "Hello, Thanks for sharing. I don't have enough hardware to retrain the ResNet50, so can you share the re-trained ResNet50 model weights?",
      "votes": null
    },
    {
      "id": "260123",
      "postDate": "12/19/2017 16:19:56",
      "content": "<p>I used both the .aps and .a3daps files for a public leaderboard score of 0.09559, and a private score of 0.2 -- I tried retraining with the additional 100 images that were used as the test set for stage 1 (I had some complications that demanded my attention, so I didn't get a chance to do that before stage 1 ended), but only hit 0.189 as a late submission with them.</p>\n\n<p>Like JamesThornton, I segmented the input images into \"zone plates.\"  I ended up discarding most of the input data though, which in hindsight was a major mistake I think; for each passenger/subject, I cut the image off at the top of their hands and used proportional arithmetic to slice their body up in to zones (variable heights and all.)  I used three frames for each zone: front facing pose, back facing pose, and one side facing pose.  I also mirrored the back facing pose so that the same arithmetic could be used (the right arm, for example, was in the same place of the input image for both the front and back poses that way).  I also randomly flipped the images horizontally during training as part of the transformations.</p>\n\n<p>The final input to my net was two 240x180 images (one for each file type) composed of three 80x180 plates stacked vertically.  At the time that seemed like a good idea, because it captured enough of a view of every zone to get a good glimpse of a threat if it was present; I'm thinking now that the disadvantage to it is that as the input data passed through the net, it may have made things too sensitive to possible threats -- I'm just pondering the behavior of the final softmax layer and suspecting that it may have been a better idea to include more \"threat free\" portions of the input images to pull the output of the softmax toward the most likely option (most zones had no threats.)  I may be entirely wrong, who knows.</p>\n\n<p>I used the convolutions from two Xception nets, each devoted to one file type, and concatenated their pooled output before applying a final softmax layer for a binary classification.  I didn't clip the output, I just let the architecture say there was a 100% chance of a threat being present if it wanted to.  I don't know if that hurt my score or not (probably did by quite a bit, lol.)</p>\n\n<p>I wish I had enough experience to provide more lessons than I can, but I'm hesitant to offer advice given that I'm still learning myself.  I hope this information is useful, and I'd like to sincerely say that I've been incredibly impressed by the Kaggle community: you folks are great, and I'm privileged to have competed with you all :)</p>\n\n<p><em>Edit</em>: I just tried clipping my submission file and re-submitting it, and it didn't make much of a difference (hit 0.17 instead of 0.18).  I guess my architecture did it's job pretty well, at least.</p>",
      "rawMarkdown": "I used both the .aps and .a3daps files for a public leaderboard score of 0.09559, and a private score of 0.2 -- I tried retraining with the additional 100 images that were used as the test set for stage 1 (I had some complications that demanded my attention, so I didn't get a chance to do that before stage 1 ended), but only hit 0.189 as a late submission with them.\n\nLike JamesThornton, I segmented the input images into \"zone plates.\"  I ended up discarding most of the input data though, which in hindsight was a major mistake I think; for each passenger/subject, I cut the image off at the top of their hands and used proportional arithmetic to slice their body up in to zones (variable heights and all.)  I used three frames for each zone: front facing pose, back facing pose, and one side facing pose.  I also mirrored the back facing pose so that the same arithmetic could be used (the right arm, for example, was in the same place of the input image for both the front and back poses that way).  I also randomly flipped the images horizontally during training as part of the transformations.\n\nThe final input to my net was two 240x180 images (one for each file type) composed of three 80x180 plates stacked vertically.  At the time that seemed like a good idea, because it captured enough of a view of every zone to get a good glimpse of a threat if it was present; I'm thinking now that the disadvantage to it is that as the input data passed through the net, it may have made things too sensitive to possible threats -- I'm just pondering the behavior of the final softmax layer and suspecting that it may have been a better idea to include more \"threat free\" portions of the input images to pull the output of the softmax toward the most likely option (most zones had no threats.)  I may be entirely wrong, who knows.\n\nI used the convolutions from two Xception nets, each devoted to one file type, and concatenated their pooled output before applying a final softmax layer for a binary classification.  I didn't clip the output, I just let the architecture say there was a 100% chance of a threat being present if it wanted to.  I don't know if that hurt my score or not (probably did by quite a bit, lol.)\n\nI wish I had enough experience to provide more lessons than I can, but I'm hesitant to offer advice given that I'm still learning myself.  I hope this information is useful, and I'd like to sincerely say that I've been incredibly impressed by the Kaggle community: you folks are great, and I'm privileged to have competed with you all :)\n\n*Edit*: I just tried clipping my submission file and re-submitting it, and it didn't make much of a difference (hit 0.17 instead of 0.18).  I guess my architecture did it's job pretty well, at least.",
      "votes": null
    },
    {
      "id": "260202",
      "postDate": "12/19/2017 21:34:42",
      "content": "<p>Hi Oleg, any chance you can share how you actually accomplished this? Even just a one liner with your actual architecture - it would be very helpful to hear how you were able to build a highly effective model that could share knowledge of threat appearance between zones.</p>",
      "rawMarkdown": "Hi Oleg, any chance you can share how you actually accomplished this? Even just a one liner with your actual architecture - it would be very helpful to hear how you were able to build a highly effective model that could share knowledge of threat appearance between zones.",
      "votes": null
    },
    {
      "id": "260212",
      "postDate": "12/19/2017 21:44:11",
      "content": "<p>Hi Moejoe, thank you for sharing this solution. Dang I'll have to investigate why my MVCNN never started working (it only produced the expected probabilities).</p>\n\n<p>Anyways, here's my question - where is that LSTM bit coming from? What is the purpose and how did you know to use it? I haven't heard much about adding an RNN layer for image classification, so I'd be very curious to hear a bit more about that and why it was so applicable here.</p>",
      "rawMarkdown": "Hi Moejoe, thank you for sharing this solution. Dang I'll have to investigate why my MVCNN never started working (it only produced the expected probabilities).\n\nAnyways, here's my question - where is that LSTM bit coming from? What is the purpose and how did you know to use it? I haven't heard much about adding an RNN layer for image classification, so I'd be very curious to hear a bit more about that and why it was so applicable here.",
      "votes": null
    },
    {
      "id": "260228",
      "postDate": "12/19/2017 22:08:33",
      "content": "<p>James,</p>\n\n<p>My approach was fairly complex overall -- I don't think I could describe it briefly. Perhaps I'll do a full write-up later. </p>\n\n<p>However, the particular aspect you are asking about seems straightforward: a single NN is trained on all relevant data. When it sees threats, it predicts what zone they are in. As I argued elsewhere in this thread, having a single omniscient NN instead of 10-17 specialized experts helps with generalization. I'll be very surprised if it turns out that any of the top teams used the specialized experts approach.</p>",
      "rawMarkdown": "James,\n\nMy approach was fairly complex overall -- I don't think I could describe it briefly. Perhaps I'll do a full write-up later. \n\nHowever, the particular aspect you are asking about seems straightforward: a single NN is trained on all relevant data. When it sees threats, it predicts what zone they are in. As I argued elsewhere in this thread, having a single omniscient NN instead of 10-17 specialized experts helps with generalization. I'll be very surprised if it turns out that any of the top teams used the specialized experts approach.",
      "votes": null
    },
    {
      "id": "260233",
      "postDate": "12/19/2017 22:21:29",
      "content": "<p>File a FOIA for the code later on :)</p>",
      "rawMarkdown": "File a FOIA for the code later on :)",
      "votes": null
    },
    {
      "id": "260315",
      "postDate": "12/20/2017 02:49:01",
      "content": "<p>In MVCNN you simply take the max probabilities calculated from each frame after they've each gone through the network independently of each other.</p>\n\n<p>In LSTM the frames are no longer treated independently of each other. It is treated as a time sequence where each time step is a frame from the .aps file and there's 16 time steps. Instead of feeding the frames directly to the LSTM I first feed it through the CNN to \"compress\" it into a small high-level feature map and feed that in. The advantage of this setup over MVCNN is that the network can learn temporal dependencies. It can learn what a rotating person should look like and flag anomalies. It can learn to boost its predictions if it notices an anomaly in multiple frames as opposed to just one frame, or suppress a prediction if an anomaly was noticed in just one frame but the others didn't pick anything up. It seems to have a regularizing affect.</p>\n\n<p>I viewed the problem as a single temporal sequence rather than 16 independent problems and it seemed natural a LSTM could be applied.</p>",
      "rawMarkdown": "In MVCNN you simply take the max probabilities calculated from each frame after they've each gone through the network independently of each other.\n\nIn LSTM the frames are no longer treated independently of each other. It is treated as a time sequence where each time step is a frame from the .aps file and there's 16 time steps. Instead of feeding the frames directly to the LSTM I first feed it through the CNN to \"compress\" it into a small high-level feature map and feed that in. The advantage of this setup over MVCNN is that the network can learn temporal dependencies. It can learn what a rotating person should look like and flag anomalies. It can learn to boost its predictions if it notices an anomaly in multiple frames as opposed to just one frame, or suppress a prediction if an anomaly was noticed in just one frame but the others didn't pick anything up. It seems to have a regularizing affect.\n\nI viewed the problem as a single temporal sequence rather than 16 independent problems and it seemed natural a LSTM could be applied.",
      "votes": null
    },
    {
      "id": "260328",
      "postDate": "12/20/2017 03:06:19",
      "content": "<p><a href=\"https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\">https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution</a></p>\n\n<p>Here's my code. Model is definied in mvcnn.py.</p>\n\n<p><a href=\"https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution/blob/master/mvcnn.py\">https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution/blob/master/mvcnn.py</a></p>",
      "rawMarkdown": "https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\n\nHere's my code. Model is definied in mvcnn.py.\n\nhttps://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution/blob/master/mvcnn.py",
      "votes": null
    },
    {
      "id": "260340",
      "postDate": "12/20/2017 03:23:21",
      "content": "<p>@Moejoe -- Thank you for sharing your code. Nicely done and congratulations on your very good performance.</p>",
      "rawMarkdown": "Moejoe -- Thank you for sharing your code. Nicely done and congratulations on your very good performance.",
      "votes": null
    },
    {
      "id": "261052",
      "postDate": "12/21/2017 16:08:00",
      "content": "<p>My final submission was a little complicated, but basically it boils down to 5 variations on the strategy outlined below.  The differences were mainly in which pre-trained ImageNet model was used, what amount of augmentation was applied, size of the threat zones, and what ended up in the 3rd color channel.</p>\n\n<h2>Model Input</h2>\n\n<p>Of the various formats, only the APS was used.  Neither of the A3D formats seemed to add anything and generating an image from the AHI is still a mystery to me.</p>\n\n<p>With only ~1,200 scans and 23ish subjects, augmentation seemed critical for a model to generalize to new subjects. For me this involved rotations, translations, contrast and brightness adjustment, horizontal reflection, and gaussian blur.  In addition, ~800 more scans were generated from the training scans by using GIMP.  Mostly this involved transforming and transplanting a threat object from one location to another or from one scan to another.</p>\n\n<p>The scans are all monochrome whereas the pre-trained ImageNet models accept three channel color input.  Triplicating the the monochrome scan for each channel was one option, but it seemed like a waste.  To make use of the multiple scans per subject, I took the average and standard deviation per pixel of ten similar scans of the subject and placed them in the other two color channels.  This had the aesthetically pleasing effect of making the threats stand out in a nice red color.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/261052/8101/model_input.png\" alt=\"model_input\" title=\"\">\nOf course the solution needed to be fully automated, so selecting similar scans needed to automated as well.  The easy way is to take the average absolute difference per pixel between each of the ~1,200 scans.  Whichever ten scans have the smallest distance could be used.  I found that a VGG-style deep autoencoder worked a little better, but the idea is essentially the same. </p>\n\n<h2>Model Design</h2>\n\n<p>Initially I had started with a 17-target MVCNN similar to what Moejoe (Shayan) described in this thread, but eventually I found that splitting the image up into four overlapping zones worked a little better.  Potentially the improvement was due to using a higher resolution on each of the smaller threat zones or possibly due to eliminating irrelevant information from other parts of the scan.  In any case this split the model into four components roughly aimed at: Arms (1-4), Chest (5-7,17), Waist (8-12), and Feet (13-16).\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/261052/8102/model_zones.png\" alt=\"model_zones\" title=\"\">\nEach of the four component models had a similar design.  They took 8 of 16 equally spaced images from the whole APS scan and fed them into a pre-trained ImageNet model, resnet-50 for instance, whose weights were shared over the different inputs.  Given that we are asking the model to not just predict the threat's existence but also the location, we want to preserve spatial information so in contrast to MVCNN, which takes the maximum of all the outputs, here we just concatenate the outputs together.  This leads to a very large number of channels which need to be reduced using a 1x1 convolution.  After that we have a very standard top design with a few dense layers.  Some small amount of dropout was needed in the last few layer to prevent the model from overfitting.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/261052/8103/model_diagram2.jpg\" alt=\"model_diagram2\" title=\"\"></p>\n\n<h2>Prediction/Calibration</h2>\n\n<p>Test time augmentation was helpful here.  All the augmentation applied to the training images was applied here at three checkpoints for each of the models.  It was probably overkill, but predictions were made for each zone model on each scan around 100 times.  All the various predictions were just averaged.</p>\n\n<p>Calibrating predictions to generalize well to new subjects was also helpful.  For this a 5-fold cross-validation was performed, with scans from any particular subject all in the same fold.  Using the out-of-sample scores from cross-validation and the original labels before any corrections, a LightGBM model was built.  This LightGBM model was then applied to the stage-2 predictions.</p>",
      "rawMarkdown": "My final submission was a little complicated, but basically it boils down to 5 variations on the strategy outlined below.  The differences were mainly in which pre-trained ImageNet model was used, what amount of augmentation was applied, size of the threat zones, and what ended up in the 3rd color channel.\n\n## Model Input\nOf the various formats, only the APS was used.  Neither of the A3D formats seemed to add anything and generating an image from the AHI is still a mystery to me.\n\nWith only ~1,200 scans and 23ish subjects, augmentation seemed critical for a model to generalize to new subjects. For me this involved rotations, translations, contrast and brightness adjustment, horizontal reflection, and gaussian blur.  In addition, ~800 more scans were generated from the training scans by using GIMP.  Mostly this involved transforming and transplanting a threat object from one location to another or from one scan to another.\n\nThe scans are all monochrome whereas the pre-trained ImageNet models accept three channel color input.  Triplicating the the monochrome scan for each channel was one option, but it seemed like a waste.  To make use of the multiple scans per subject, I took the average and standard deviation per pixel of ten similar scans of the subject and placed them in the other two color channels.  This had the aesthetically pleasing effect of making the threats stand out in a nice red color.\n![model_input][1]\nOf course the solution needed to be fully automated, so selecting similar scans needed to automated as well.  The easy way is to take the average absolute difference per pixel between each of the ~1,200 scans.  Whichever ten scans have the smallest distance could be used.  I found that a VGG-style deep autoencoder worked a little better, but the idea is essentially the same. \n\n## Model Design\nInitially I had started with a 17-target MVCNN similar to what Moejoe (Shayan) described in this thread, but eventually I found that splitting the image up into four overlapping zones worked a little better.  Potentially the improvement was due to using a higher resolution on each of the smaller threat zones or possibly due to eliminating irrelevant information from other parts of the scan.  In any case this split the model into four components roughly aimed at: Arms (1-4), Chest (5-7,17), Waist (8-12), and Feet (13-16).\n![model_zones][2]\nEach of the four component models had a similar design.  They took 8 of 16 equally spaced images from the whole APS scan and fed them into a pre-trained ImageNet model, resnet-50 for instance, whose weights were shared over the different inputs.  Given that we are asking the model to not just predict the threat's existence but also the location, we want to preserve spatial information so in contrast to MVCNN, which takes the maximum of all the outputs, here we just concatenate the outputs together.  This leads to a very large number of channels which need to be reduced using a 1x1 convolution.  After that we have a very standard top design with a few dense layers.  Some small amount of dropout was needed in the last few layer to prevent the model from overfitting.\n![model_diagram2][3]\n## Prediction/Calibration\nTest time augmentation was helpful here.  All the augmentation applied to the training images was applied here at three checkpoints for each of the models.  It was probably overkill, but predictions were made for each zone model on each scan around 100 times.  All the various predictions were just averaged.\n\nCalibrating predictions to generalize well to new subjects was also helpful.  For this a 5-fold cross-validation was performed, with scans from any particular subject all in the same fold.  Using the out-of-sample scores from cross-validation and the original labels before any corrections, a LightGBM model was built.  This LightGBM model was then applied to the stage-2 predictions.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/261052/8101/model_input.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/261052/8102/model_zones.png\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/261052/8103/model_diagram2.jpg",
      "votes": null
    },
    {
      "id": "261096",
      "postDate": "12/21/2017 19:21:51",
      "content": "<p>Thank you for the great description and vizualization. It's amazing to see how the way you analyzed and augmented the data made all the difference, here. I had multiple models that were very similar to what you were doing, but never got past 0.25 LB (stage1 data). Congratulations on the first place!</p>\n\n<p>Did you ever try or consider using the .a3daps scans to augment? I tried using 8 equally spaced angles and then randomly shifting -2 to +2 steps (same for all 8 angles), but didn't see much of an effect on the model's ability to generalize.</p>",
      "rawMarkdown": "Thank you for the great description and vizualization. It's amazing to see how the way you analyzed and augmented the data made all the difference, here. I had multiple models that were very similar to what you were doing, but never got past 0.25 LB (stage1 data). Congratulations on the first place!\n\nDid you ever try or consider using the .a3daps scans to augment? I tried using 8 equally spaced angles and then randomly shifting -2 to +2 steps (same for all 8 angles), but didn't see much of an effect on the model's ability to generalize.",
      "votes": null
    },
    {
      "id": "261116",
      "postDate": "12/21/2017 20:21:02",
      "content": "<p>100 predictions!!! with Test time augmentation? WOW! we did test time augmentation too but only averaged 16 predictions per subject, and we thought that was overkill !!...  go big or go home I guess...we tried a similar approach of splitting the volume into 4 chunks for 3d convolutions at the original dimensions.  That model did not make it into our final submission because we could not reliably guarantee the zones will fall in the right chunk because the subjects had varying height. We then attempted to do registration on all the samples to make them all the same height, size e.t.c with a 12 parameter affine transformation but we were worried about the computational time and size of the state 2 data.</p>",
      "rawMarkdown": "100 predictions!!! with Test time augmentation? WOW! we did test time augmentation too but only averaged 16 predictions per subject, and we thought that was overkill !!...  go big or go home I guess...we tried a similar approach of splitting the volume into 4 chunks for 3d convolutions at the original dimensions.  That model did not make it into our final submission because we could not reliably guarantee the zones will fall in the right chunk because the subjects had varying height. We then attempted to do registration on all the samples to make them all the same height, size e.t.c with a 12 parameter affine transformation but we were worried about the computational time and size of the state 2 data.",
      "votes": null
    },
    {
      "id": "261133",
      "postDate": "12/21/2017 21:21:44",
      "content": "<p>Thanks for your write-up and congratulations on the win!</p>\n\n<blockquote>\n  <p>~800 more scans were generated from the training scans by using GIMP</p>\n</blockquote>\n\n<p>I had to Google it, thinking GIMP is some interesting algorithm. You manually drew / augmented 800 images??!! whee,...</p>",
      "rawMarkdown": "Thanks for your write-up and congratulations on the win!\n\n&gt; ~800 more scans were generated from the training scans by using GIMP\n\nI had to Google it, thinking GIMP is some interesting algorithm. You manually drew / augmented 800 images??!! whee,...",
      "votes": null
    },
    {
      "id": "261140",
      "postDate": "12/21/2017 21:53:45",
      "content": "<blockquote>\n  <p>You manually drew / augmented 800 images??!!</p>\n</blockquote>\n\n<p>800 <em>scans</em>. It must have been ~5000 photoshoppings. Congrats, idle_speculation!</p>",
      "rawMarkdown": "&gt; You manually drew / augmented 800 images??!!\n\n800 *scans*. It must have been ~5000 photoshoppings. Congrats, idle_speculation!",
      "votes": null
    },
    {
      "id": "261163",
      "postDate": "12/21/2017 22:52:59",
      "content": "<p>thanks for the great writeup and congrats. </p>\n\n<p>-- how many training steps did you run for each region (component) and what computing environment did you use?</p>\n\n<p>Should  I infer one key to Idle_Speculations's success was comparing each scan to the average (and std dev) of all scans for the subect?</p>",
      "rawMarkdown": "thanks for the great writeup and congrats. \n\n\n-- how many training steps did you run for each region (component) and what computing environment did you use?\n\nShould  I infer one key to Idle_Speculations's success was comparing each scan to the average (and std dev) of all scans for the subect?",
      "votes": null
    },
    {
      "id": "261169",
      "postDate": "12/21/2017 22:59:41",
      "content": "<p>My winning solution was an ensemble of 11 models. All models are end-to-end, i.e., they take as input 16 images per subject and spit out 17 probabilities per threats. I did not use any pretrained model or weights. All models are hand-crafted CNN based models. All models were trained on only one Titan X gpu.</p>\n\n<p><strong>Data pre-processing</strong>\n- I only used APS format.\n- For most models, we down-sampled the dataset to 256  by 330 \n- We also rotate images 90, 180 and 270 degrees and train separate models on the rotated dataset.\n- For couple of models we used the original 512 by 660 images and the rotated dataset.\n- Global zero-mean and unit-variance normalization were used. </p>\n\n<p><strong>Data Augmentation</strong>\nWe perform different image augmentations including: rotation (10-15 degrees), width and height shift (10-15%), shear, zoom (10-15 %) as well as random erasing [Ref1]. I tried to use brightness augmentation but the models were completely confused with any changes in the brightness so I dropped it from augmentation.</p>\n\n<p><strong>Architecture</strong>\nModel architecture as seen in the figure is basically a conv2D to extract features from 16 image slices one at a time, flattening features, average over 16 image slices and then generating 17 Sigmoid outputs (one per zone). I did not know about the MVCNN paper but it seems there is similarity with that method. </p>\n\n<p><strong>Insights</strong>\nRandom erasing augmentation provided some improvements. Building separate models on 90,180,270 rotated dataset was also very effective. It is also better to use the original dataset as opposed to the down-sampled data. Although, it is more time consuming to train.  </p>\n\n<p><strong>Ref1</strong>\nZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, Yi Yang “Random Erasing Data Augmentation”</p>",
      "rawMarkdown": "My winning solution was an ensemble of 11 models. All models are end-to-end, i.e., they take as input 16 images per subject and spit out 17 probabilities per threats. I did not use any pretrained model or weights. All models are hand-crafted CNN based models. All models were trained on only one Titan X gpu.\n  \n**Data pre-processing**\n- I only used APS format.\n- For most models, we down-sampled the dataset to 256  by 330 \n- We also rotate images 90, 180 and 270 degrees and train separate models on the rotated dataset.\n- For couple of models we used the original 512 by 660 images and the rotated dataset.\n- Global zero-mean and unit-variance normalization were used. \n\n**Data Augmentation**\nWe perform different image augmentations including: rotation (10-15 degrees), width and height shift (10-15%), shear, zoom (10-15 %) as well as random erasing [Ref1]. I tried to use brightness augmentation but the models were completely confused with any changes in the brightness so I dropped it from augmentation.\n\n**Architecture**\nModel architecture as seen in the figure is basically a conv2D to extract features from 16 image slices one at a time, flattening features, average over 16 image slices and then generating 17 Sigmoid outputs (one per zone). I did not know about the MVCNN paper but it seems there is similarity with that method. \n\n**Insights**\nRandom erasing augmentation provided some improvements. Building separate models on 90,180,270 rotated dataset was also very effective. It is also better to use the original dataset as opposed to the down-sampled data. Although, it is more time consuming to train.  \n\n**Ref1**\nZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, Yi Yang “Random Erasing Data Augmentation”",
      "votes": null
    },
    {
      "id": "261191",
      "postDate": "12/22/2017 00:54:09",
      "content": "<p>@Stefan</p>\n\n<p>I suspect there was something amiss in your setup if you got stuck at 0.25 on the public LB.  To answer your question, I did try using the A3DAPS in variety of ways:  Putting it another color channel, Randomly replacing an APS image with a comparable A3DAPS, Building A3DAPS specific models.  None of it really seemed to work. </p>\n\n<p>@DavidGbodiOdaibo</p>\n\n<p>I based my zones off the training set and tried to build in a little room for error.  Even so, a couple of the gentlemen in stage-2 were taller than I had anticipated.  I'm pretty sure I lost a few points in zones 5 and 17 for those guys.</p>\n\n<p>@Bastiaan/Oleg</p>\n\n<p>Yep, it took a while.  Probably there is a good way to automatically extract the threat part of an image and paste it into other scans, I wasn't able to figure it out though.  Another mind-numbing exercise I subjected myself to was drawing bounding boxes around all the threats in zones 1-4 while I explored Faster-RCNN.</p>\n\n<p>@numericLee</p>\n\n<p>Each of the zone models converged in around 16k iterations with batch size 8.  It worked out to be around 4 hours per zone on a 1080TI.  I also had some free credits for google cloud which I used to rent out Teslas for heavy lifting.  In terms of using average of similar scans, it seemed to decrease logloss by around 0.01 in cross-validation.  The benefits were perhaps larger on the test set given Mr. Dreadlocks.</p>",
      "rawMarkdown": "Stefan\n\nI suspect there was something amiss in your setup if you got stuck at 0.25 on the public LB.  To answer your question, I did try using the A3DAPS in variety of ways:  Putting it another color channel, Randomly replacing an APS image with a comparable A3DAPS, Building A3DAPS specific models.  None of it really seemed to work. \n\n@DavidGbodiOdaibo\n\nI based my zones off the training set and tried to build in a little room for error.  Even so, a couple of the gentlemen in stage-2 were taller than I had anticipated.  I'm pretty sure I lost a few points in zones 5 and 17 for those guys.\n\n@Bastiaan/Oleg\n\nYep, it took a while.  Probably there is a good way to automatically extract the threat part of an image and paste it into other scans, I wasn't able to figure it out though.  Another mind-numbing exercise I subjected myself to was drawing bounding boxes around all the threats in zones 1-4 while I explored Faster-RCNN.\n\n@numericLee\n\nEach of the zone models converged in around 16k iterations with batch size 8.  It worked out to be around 4 hours per zone on a 1080TI.  I also had some free credits for google cloud which I used to rent out Teslas for heavy lifting.  In terms of using average of similar scans, it seemed to decrease logloss by around 0.01 in cross-validation.  The benefits were perhaps larger on the test set given Mr. Dreadlocks.",
      "votes": null
    },
    {
      "id": "261235",
      "postDate": "12/22/2017 03:40:18",
      "content": "<p>I only ended up using the APS images. </p>\n\n<h1>Bounding Box</h1>\n\n<p>I spent a good amount of time annotating the zones in the front view of the images. I then used these annotations to train a bounding box model to predict where each zone was. I didn't want to take the time to do this at first, but the variance in the height/width of the subjects and the fact that some subjects held their arms at different angles led me to think it would be beneficial. I also annotated the locations of a subject's ankles, knees, elbows, shoulders, wrists on the front and side images. The reasons for this is that I wanted to align people's limbs so they they were all as vertical as possible. </p>\n\n<h1>Subject Identification</h1>\n\n<p>I realized early on that there were a limited number of subjects and that the subjects in the Phase 1 test set all appeared in the training set. I knew that this would cause overfitting, and sure enough I noticed a massive difference in models trained using proper validation and those that did not. Without proper validation I was getting &lt; 0.04 on the Public LB. Using proper validation and the same architected I was getting ~0.17 on the Public LB. I ended up with 24 different subjects. </p>\n\n<h1>Models</h1>\n\n<p>For inputs, I concatenated the APS angles into a single 13 images x 1 image (I excluded 3 angles where the zones couldn't be seen) array. I trained 2 different models: one where each image was reduced to 96x96 pixels, and another that was 128x128 pixels i.e. the final images were 1248x96 and 1664x128. As mentioned in the Bounding Box section I rotated the limbs so that they were vertical in each image to account for the fact that people had their arms/legs as different angles. My architecture ended up being the following:</p>\n\n<p>Conv2D(16, 3x3) <br>\nConv2D(16, 3x3) <br>\nMaxPooling2D(2x2)  </p>\n\n<p>Conv2D(32, 3x3) <br>\nConv2D(32, 3x3) <br>\nMaxPooling2D(2x3)  </p>\n\n<p>Conv2D(32, 3x3) <br>\nConv2D(32, 3x3) <br>\nConv2D(32, 3x3) <br>\nMaxPooling2D(2x2)  </p>\n\n<p>Conv2D(32, 3x3) <br>\nConv2D(32, 3x3) <br>\nMaxPooling2D(2x2)  </p>\n\n<p>Conv2D(32, 3x3) <br>\nConv2D(64, 3x3) <br>\nConv2D(64, 3x3) <br>\nMaxPooling2D(2x2) <br>\nDropout(0.5)  </p>\n\n<p>Conv2D(64, 3x3) <br>\nConv2D(128, 3x3) <br>\nConv2D(128, 3x3) <br>\nMaxPooling2D(3x3) <br>\nDropout(0.5)  </p>\n\n<p>Flatten()  </p>\n\n<p>Dense(512) <br>\nDropout(0.5)  </p>\n\n<p>Dense(1)  </p>\n\n<p>Anyhow, all I have time for now. I'll hopefully add a bit more detail soon.\nThis ended up being much deeper than my original model trained without using proper cross-validation.</p>",
      "rawMarkdown": "I only ended up using the APS images. \n\n# Bounding Box\nI spent a good amount of time annotating the zones in the front view of the images. I then used these annotations to train a bounding box model to predict where each zone was. I didn't want to take the time to do this at first, but the variance in the height/width of the subjects and the fact that some subjects held their arms at different angles led me to think it would be beneficial. I also annotated the locations of a subject's ankles, knees, elbows, shoulders, wrists on the front and side images. The reasons for this is that I wanted to align people's limbs so they they were all as vertical as possible. \n\n# Subject Identification\nI realized early on that there were a limited number of subjects and that the subjects in the Phase 1 test set all appeared in the training set. I knew that this would cause overfitting, and sure enough I noticed a massive difference in models trained using proper validation and those that did not. Without proper validation I was getting &lt; 0.04 on the Public LB. Using proper validation and the same architected I was getting ~0.17 on the Public LB. I ended up with 24 different subjects. \n\n# Models\nFor inputs, I concatenated the APS angles into a single 13 images x 1 image (I excluded 3 angles where the zones couldn't be seen) array. I trained 2 different models: one where each image was reduced to 96x96 pixels, and another that was 128x128 pixels i.e. the final images were 1248x96 and 1664x128. As mentioned in the Bounding Box section I rotated the limbs so that they were vertical in each image to account for the fact that people had their arms/legs as different angles. My architecture ended up being the following:\n\nConv2D(16, 3x3)  \nConv2D(16, 3x3)  \nMaxPooling2D(2x2)  \n        \nConv2D(32, 3x3)  \nConv2D(32, 3x3)  \nMaxPooling2D(2x3)  \n        \nConv2D(32, 3x3)  \nConv2D(32, 3x3)  \nConv2D(32, 3x3)  \nMaxPooling2D(2x2)  \n        \nConv2D(32, 3x3)  \nConv2D(32, 3x3)  \nMaxPooling2D(2x2)  \n        \nConv2D(32, 3x3)  \nConv2D(64, 3x3)  \nConv2D(64, 3x3)  \nMaxPooling2D(2x2)     \nDropout(0.5)  \n        \nConv2D(64, 3x3)  \nConv2D(128, 3x3)  \nConv2D(128, 3x3)  \nMaxPooling2D(3x3)      \nDropout(0.5)  \n        \nFlatten()  \n        \nDense(512)  \nDropout(0.5)  \n          \nDense(1)  \n\nAnyhow, all I have time for now. I'll hopefully add a bit more detail soon.\nThis ended up being much deeper than my original model trained without using proper cross-validation.",
      "votes": null
    },
    {
      "id": "261248",
      "postDate": "12/22/2017 04:13:32",
      "content": "<p>@idle_speculation</p>\n\n<blockquote>\n  <p>Each of the zone models converged in around 16k iterations with batch size 8.</p>\n</blockquote>\n\n<p>Did they converge in the sense that the training loss stopped changing, or did the validation loss start going up (and you used early stopping)? What were the learning rate / momentum like?</p>",
      "rawMarkdown": "idle_speculation\n\n&gt; Each of the zone models converged in around 16k iterations with batch size 8.\n\nDid they converge in the sense that the training loss stopped changing, or did the validation loss start going up (and you used early stopping)? What were the learning rate / momentum like?",
      "votes": null
    },
    {
      "id": "261365",
      "postDate": "12/22/2017 12:24:11",
      "content": "<blockquote>\n  <p>Did they converge in the sense that the training loss stopped changing, or did the validation loss start going up (and you used early stopping)? What were the learning rate / momentum like?</p>\n</blockquote>\n\n<p>They converged in the sense that the validation loss did not improve past that point.  Performance did not get any worse with further training.  It just stopped getting better.  Training loss was very close to zero near the end.</p>\n\n<p>Adagrad was used to fit all the models.  No momentum was applied.  Learning rate was gradually increased up to 0.01 in the first 1k iterations, then left at that level for the rest of the fit.</p>",
      "rawMarkdown": "&gt; Did they converge in the sense that the training loss stopped changing, or did the validation loss start going up (and you used early stopping)? What were the learning rate / momentum like?\n\nThey converged in the sense that the validation loss did not improve past that point.  Performance did not get any worse with further training.  It just stopped getting better.  Training loss was very close to zero near the end.\n\nAdagrad was used to fit all the models.  No momentum was applied.  Learning rate was gradually increased up to 0.01 in the first 1k iterations, then left at that level for the rest of the fit.",
      "votes": null
    },
    {
      "id": "261401",
      "postDate": "12/22/2017 15:04:20",
      "content": "<p>I see a potential weakness with this approach in a production environment, I am not sure if I am interpreting this correctly and will like to hear your thoughts. It seems to me that the RGB image exploits the fact that there are additional scans of the same subjects under different conditions in the training and test set and uses these additional scans to derive the “mean” and “standard deviation” green and blue channel, however, in a production/live environment the system will be seeing a subject for the first time and there will be no prior scan to derive the additional channels in the RGB image, how would you handle this? Is there a performance difference when the additional channels are not included?</p>",
      "rawMarkdown": "I see a potential weakness with this approach in a production environment, I am not sure if I am interpreting this correctly and will like to hear your thoughts. It seems to me that the RGB image exploits the fact that there are additional scans of the same subjects under different conditions in the training and test set and uses these additional scans to derive the “mean” and “standard deviation” green and blue channel, however, in a production/live environment the system will be seeing a subject for the first time and there will be no prior scan to derive the additional channels in the RGB image, how would you handle this? Is there a performance difference when the additional channels are not included?",
      "votes": null
    },
    {
      "id": "261412",
      "postDate": "12/22/2017 15:47:16",
      "content": "<p>I would ask this privately, but I can't contact users without being a \"contributor.\"  Please pardon my occasionally awkward social skills...</p>\n\n<p>I would have expected your solution to be disqualified, because it's impossible for the DHS to verify it without the additional training data.  Did they contact you to ask for that?  Please understand that I'm not trying to imply there's a reason to be suspicious, and given your standing, I'm sure your word is good enough for Kaggle (if not the sponsor as well) -- but I'm still curious how verification was handled?</p>",
      "rawMarkdown": "I would ask this privately, but I can't contact users without being a \"contributor.\"  Please pardon my occasionally awkward social skills...\n\nI would have expected your solution to be disqualified, because it's impossible for the DHS to verify it without the additional training data.  Did they contact you to ask for that?  Please understand that I'm not trying to imply there's a reason to be suspicious, and given your standing, I'm sure your word is good enough for Kaggle (if not the sponsor as well) -- but I'm still curious how verification was handled?",
      "votes": null
    },
    {
      "id": "261419",
      "postDate": "12/22/2017 16:11:11",
      "content": "<p>Murray,</p>\n\n<p>All winning solutions are verified by the sponsor and must be able to reproduce the winning leaderboard score per the rules. As the competition has closed, the model upload and verification process is underway.</p>",
      "rawMarkdown": "Murray,\n\nAll winning solutions are verified by the sponsor and must be able to reproduce the winning leaderboard score per the rules. As the competition has closed, the model upload and verification process is underway.",
      "votes": null
    },
    {
      "id": "261421",
      "postDate": "12/22/2017 16:26:11",
      "content": "<p>I'll just be perfectly frank: it doesn't bother me in the slightest, personally, and I have nothing but respect for someone who got the job done better than other teams (other teams includes me, too, lol); but in my mind, the manual augmentation of training data falls under a non-automated external source that wasn't made available to the other contestants.  But it doesn't really matter what I think, and I'm not trying to cause trouble.</p>\n\n<p>My apologies if the question was awkward.  Truthfully I'm most interested in where the boundaries of the rules are, so that I can do better myself in the future without accidentally crossing the line.  Obviously a tremendous amount of work went into idle_speculation's submission, and I think he deserves that first place slot.  Congratulations again :)</p>",
      "rawMarkdown": "I'll just be perfectly frank: it doesn't bother me in the slightest, personally, and I have nothing but respect for someone who got the job done better than other teams (other teams includes me, too, lol); but in my mind, the manual augmentation of training data falls under a non-automated external source that wasn't made available to the other contestants.  But it doesn't really matter what I think, and I'm not trying to cause trouble.\n\nMy apologies if the question was awkward.  Truthfully I'm most interested in where the boundaries of the rules are, so that I can do better myself in the future without accidentally crossing the line.  Obviously a tremendous amount of work went into idle_speculation's submission, and I think he deserves that first place slot.  Congratulations again :)",
      "votes": null
    },
    {
      "id": "261490",
      "postDate": "12/22/2017 19:37:53",
      "content": "<p>I decided to approach this competition from the viewpoint of object detection.  I annotated a subset of the training images by drawing bounding boxes around the threats in the aps images (I spent maybe 8 hrs on this).  I then trained Faster-RCNN using the annotated images, with the label of each object being its location on the body.  This way the network would learn to both detect objects and label them based on their location all in one go.</p>\n\n<p>For each aps image there are 16 viewpoints.  I ran the object detector over each viewpoint in the aps image, to generate a set of candidate object detections, with each detection having a probability over each body location.  I took the maximum of all the detections for each body location to get a final 17 dimensional confidence vector.  I then took the 17d confidence vector for each viewpoint to get a 16x17 matrix.  I used this to train an 17 gradient boosted classifiers (one for each body location) to get better calibrated probabilities.</p>\n\n<p>The only data augmentation I performed was doing a horizontal flip of the images while training the object detector.  This actually will change the label.  For example, if a threat is labeled \"right ankle\" then after flipping it will then be labeled \"left ankle\". </p>\n\n<p>I decided on using an object detector due to the limited size of the dataset and the small number of unique individuals.  Faster-RCNN is a two stage object detector.  In the first stage, it recognizes potential threats; the second stage labels the threats/decides if they are not true threats.  I thought that this built in attention mechanism would help to prevent overfitting by forcing the network to focus on the threats.</p>\n\n<p>My final model consisted of an ensemble of 5 models.  Each model using 80% of the data to train the object detector and then the final 20% to train the boosted classifier.  I then averaged the predictions over each individual model.  However, ensembling did not really seem help the performance of my model in any significant way.</p>\n\n<p>Now some things that I tried that didn't work.  I spent some time trying to use the 3d data.  I used a 3d inflated convolutional network based on the vgg16 architecture.  This just means that I turned each 2d convolution into a 3d convolution and initialized the weights of each 3d convolution by stacking the weights of the pretrained vgg model.  I then tried to use this network to perform 3d object detection.   This best I was able to do was around 0.06 on my validation set as opposed to 0.025 using my 2d model.</p>",
      "rawMarkdown": "I decided to approach this competition from the viewpoint of object detection.  I annotated a subset of the training images by drawing bounding boxes around the threats in the aps images (I spent maybe 8 hrs on this).  I then trained Faster-RCNN using the annotated images, with the label of each object being its location on the body.  This way the network would learn to both detect objects and label them based on their location all in one go.\n\nFor each aps image there are 16 viewpoints.  I ran the object detector over each viewpoint in the aps image, to generate a set of candidate object detections, with each detection having a probability over each body location.  I took the maximum of all the detections for each body location to get a final 17 dimensional confidence vector.  I then took the 17d confidence vector for each viewpoint to get a 16x17 matrix.  I used this to train an 17 gradient boosted classifiers (one for each body location) to get better calibrated probabilities.\n\nThe only data augmentation I performed was doing a horizontal flip of the images while training the object detector.  This actually will change the label.  For example, if a threat is labeled \"right ankle\" then after flipping it will then be labeled \"left ankle\". \n\nI decided on using an object detector due to the limited size of the dataset and the small number of unique individuals.  Faster-RCNN is a two stage object detector.  In the first stage, it recognizes potential threats; the second stage labels the threats/decides if they are not true threats.  I thought that this built in attention mechanism would help to prevent overfitting by forcing the network to focus on the threats.\n\nMy final model consisted of an ensemble of 5 models.  Each model using 80% of the data to train the object detector and then the final 20% to train the boosted classifier.  I then averaged the predictions over each individual model.  However, ensembling did not really seem help the performance of my model in any significant way.\n\nNow some things that I tried that didn't work.  I spent some time trying to use the 3d data.  I used a 3d inflated convolutional network based on the vgg16 architecture.  This just means that I turned each 2d convolution into a 3d convolution and initialized the weights of each 3d convolution by stacking the weights of the pretrained vgg model.  I then tried to use this network to perform 3d object detection.   This best I was able to do was around 0.06 on my validation set as opposed to 0.025 using my 2d model.",
      "votes": null
    },
    {
      "id": "261531",
      "postDate": "12/22/2017 21:02:29",
      "content": "<p>Thanks, deedrz, and congrats! Were those networks pre-trained on ImageNet?</p>",
      "rawMarkdown": "Thanks, deedrz, and congrats! Were those networks pre-trained on ImageNet?",
      "votes": null
    },
    {
      "id": "261533",
      "postDate": "12/22/2017 21:12:30",
      "content": "<p>Thanks for the writeup and congrats on the win!</p>\n\n<p>Did you consider ensembling the two models: the 3D conv and the 2D conv? I'd think the two models are fairly orthogonal?</p>",
      "rawMarkdown": "Thanks for the writeup and congrats on the win!\n\nDid you consider ensembling the two models: the 3D conv and the 2D conv? I'd think the two models are fairly orthogonal?",
      "votes": null
    },
    {
      "id": "261534",
      "postDate": "12/22/2017 21:17:13",
      "content": "<p>Thanks for the writeup!</p>\n\n<blockquote>\n  <p>I rotated the limbs so that they were vertical in each image</p>\n</blockquote>\n\n<p>How did you do that? Did you rotate the whole picture to get everything straight on average? Or did you cut up the people in separate limbs? Or did you stretch and skew the images? Or more advanced?</p>",
      "rawMarkdown": "Thanks for the writeup!\n\n&gt;I rotated the limbs so that they were vertical in each image\n\nHow did you do that? Did you rotate the whole picture to get everything straight on average? Or did you cut up the people in separate limbs? Or did you stretch and skew the images? Or more advanced?",
      "votes": null
    },
    {
      "id": "261539",
      "postDate": "12/22/2017 21:44:15",
      "content": "<p>Yep, the networks were pre-trained on ImageNet, I just replicated each image to get the 3 channel input. And Bastiaan, I think your right, combining the 2d and 3d models would've been a good thing to try</p>",
      "rawMarkdown": "Yep, the networks were pre-trained on ImageNet, I just replicated each image to get the 3 channel input. And Bastiaan, I think your right, combining the 2d and 3d models would've been a good thing to try",
      "votes": null
    },
    {
      "id": "261540",
      "postDate": "12/22/2017 22:03:47",
      "content": "<blockquote>\n  <p><strong>Murray Miron wrote</strong></p>\n  \n  <p>I would have expected your solution to be disqualified, because it's impossible for the DHS to verify it without the additional training data.  Did they contact you to ask for that?  Please understand that I'm not trying to imply there's a reason to be suspicious, and given your standing, I'm sure your word is good enough for Kaggle (if not the sponsor as well) -- but I'm still curious how verification was handled?</p>\n</blockquote>\n\n<p>Luckily I had the forethought to include the additional scans in my model upload. </p>",
      "rawMarkdown": "&gt; **Murray Miron wrote**\n&gt;\n&gt; I would have expected your solution to be disqualified, because it's impossible for the DHS to verify it without the additional training data.  Did they contact you to ask for that?  Please understand that I'm not trying to imply there's a reason to be suspicious, and given your standing, I'm sure your word is good enough for Kaggle (if not the sponsor as well) -- but I'm still curious how verification was handled?\n\nLuckily I had the forethought to include the additional scans in my model upload.",
      "votes": null
    },
    {
      "id": "261545",
      "postDate": "12/22/2017 22:22:43",
      "content": "<p>Using the right forearm as an example, I had a model which identified the point in the front and side images where the wrist and elbow were.  I then used my bounding box model to identify the where the forearm was and cropped out the rest of the image.  Then using the point locations of the wrist and elbow and using a bit of trigonometry I calculated how much I needed to rotate each image so that the forearm was vertical. </p>",
      "rawMarkdown": "Using the right forearm as an example, I had a model which identified the point in the front and side images where the wrist and elbow were.  I then used my bounding box model to identify the where the forearm was and cropped out the rest of the image.  Then using the point locations of the wrist and elbow and using a bit of trigonometry I calculated how much I needed to rotate each image so that the forearm was vertical.",
      "votes": null
    },
    {
      "id": "261546",
      "postDate": "12/22/2017 22:26:14",
      "content": "<blockquote>\n  <p><strong>DavidGbodiOdaibo wrote</strong></p>\n  \n  <p>I see a potential weakness with this approach in a production environment, I am not sure if I am interpreting this correctly and will like to hear your thoughts. It seems to me that the RGB image exploits the fact that there are additional scans of the same subjects under different conditions in the training and test set and uses these additional scans to derive the “mean” and “standard deviation” green and blue channel, however, in a production/live environment the system will be seeing a subject for the first time and there will be no prior scan to derive the additional channels in the RGB image, how would you handle this? Is there a performance difference when the additional channels are not included?</p>\n</blockquote>\n\n<p>I don't consider myself a person who travels that much.  But personally, I took three flights in the last year.  That works out to six trips though the scanners this year alone.  If you add in scans from last year or two then the approach of using ten similar scans starts sounding a lot more plausible.</p>",
      "rawMarkdown": "&gt; **DavidGbodiOdaibo wrote**\n&gt;\n&gt; I see a potential weakness with this approach in a production environment, I am not sure if I am interpreting this correctly and will like to hear your thoughts. It seems to me that the RGB image exploits the fact that there are additional scans of the same subjects under different conditions in the training and test set and uses these additional scans to derive the “mean” and “standard deviation” green and blue channel, however, in a production/live environment the system will be seeing a subject for the first time and there will be no prior scan to derive the additional channels in the RGB image, how would you handle this? Is there a performance difference when the additional channels are not included?\n\nI don't consider myself a person who travels that much.  But personally, I took three flights in the last year.  That works out to six trips though the scanners this year alone.  If you add in scans from last year or two then the approach of using ten similar scans starts sounding a lot more plausible.",
      "votes": null
    },
    {
      "id": "261583",
      "postDate": "12/23/2017 03:51:19",
      "content": "<p>I don't think the TSA is legally permitted to store the scans, actually.</p>",
      "rawMarkdown": "I don't think the TSA is legally permitted to store the scans, actually.",
      "votes": null
    },
    {
      "id": "261798",
      "postDate": "12/24/2017 01:54:59",
      "content": "<p>I posted about this in more detail <a href=\"http://suchir.io/posts/passenger-screening-algorithm-challenge-writeup.html\">here</a>, but the TL;DR is:</p>\n\n<ul>\n<li>trained a model to compute segmentations of the threats given aps/a3daps images (hand-labeled the training data)</li>\n<li>trained a model to compute segmentations of the zones given depth maps (generated synthetic data with depth maps / ground truth to train, used a3d to get real depth maps)</li>\n<li>trained a model that takes in the threat / zone segmentations from 16 angles and outputs the final predictions for each zone</li>\n</ul>",
      "rawMarkdown": "I posted about this in more detail [here][1], but the TL;DR is:\n\n- trained a model to compute segmentations of the threats given aps/a3daps images (hand-labeled the training data)\n- trained a model to compute segmentations of the zones given depth maps (generated synthetic data with depth maps / ground truth to train, used a3d to get real depth maps)\n- trained a model that takes in the threat / zone segmentations from 16 angles and outputs the final predictions for each zone\n  [1]: http://suchir.io/posts/passenger-screening-algorithm-challenge-writeup.html",
      "votes": null
    },
    {
      "id": "263940",
      "postDate": "01/01/2018 18:21:06",
      "content": "<p>The description of my cylindrical coordinates based approach is <a href=\"https://www.kaggle.com/nathanrm/full-solution-cylindrical-coordinate-method\">posted in Kernels</a>, and includes a link to the complete solution on github.</p>",
      "rawMarkdown": "The description of my cylindrical coordinates based approach is [posted in Kernels](https://www.kaggle.com/nathanrm/full-solution-cylindrical-coordinate-method), and includes a link to the complete solution on github.",
      "votes": null
    },
    {
      "id": "265258",
      "postDate": "01/05/2018 01:46:20",
      "content": "<p>@idle_speculation</p>\n\n<p>Congrats, I'm also curious since I haven't seen anybody mention about the recent published paper by Hinton in regards to Dynamic routing between capsules, even though it's relatively new but what are your thoughts on this capsule net over convolutional net? </p>\n\n<p><a href=\"https://arxiv.org/pdf/1710.09829.pdf\">https://arxiv.org/pdf/1710.09829.pdf</a></p>",
      "rawMarkdown": "idle_speculation\n\nCongrats, I'm also curious since I haven't seen anybody mention about the recent published paper by Hinton in regards to Dynamic routing between capsules, even though it's relatively new but what are your thoughts on this capsule net over convolutional net? \n\nhttps://arxiv.org/pdf/1710.09829.pdf",
      "votes": null
    },
    {
      "id": "271860",
      "postDate": "01/21/2018 17:26:10",
      "content": "<p>I participated in this competition for completing the Udacity Nanodegree capstone project requirement, you can find the full report <a href=\"https://github.com/shivarajugowda/kaggle/blob/master/TSA/report.md\">here.</a>. In summary, my solution involved these two steps. </p>\n\n<ol>\n<li>Detect head, hands, groin and threat bounding boxes using TensorFlow Object Detection API</li>\n<li>Predict Body Zone of a threat based on the bounding boxes of a threat relative to the frame number, head, hands and groin. I attempted this in two ways: A supervised learning method using XGBOOST. A custom algorithm which needed only a few ratios to be tweaked.</li>\n</ol>\n\n<p>Overall it was good learning experience, it did however take a lot of time. </p>",
      "rawMarkdown": "I participated in this competition for completing the Udacity Nanodegree capstone project requirement, you can find the full report [here.][1]. In summary, my solution involved these two steps. \n\n 1. Detect head, hands, groin and threat bounding boxes using TensorFlow Object Detection API\n 2. Predict Body Zone of a threat based on the bounding boxes of a threat relative to the frame number, head, hands and groin. I attempted this in two ways: A supervised learning method using XGBOOST. A custom algorithm which needed only a few ratios to be tweaked.\n\nOverall it was good learning experience, it did however take a lot of time. \n\n  [1]: https://github.com/shivarajugowda/kaggle/blob/master/TSA/report.md",
      "votes": null
    },
    {
      "id": "312252",
      "postDate": "04/11/2018 13:15:44",
      "content": "<p>I wonder how much time the other top-8 teams' models need for training and prediction, respectively?</p>",
      "rawMarkdown": "I wonder how much time the other top-8 teams' models need for training and prediction, respectively?",
      "votes": null
    },
    {
      "id": "457123",
      "postDate": "01/17/2019 01:28:35",
      "content": "<p>hello,I am working on this problem using your tips but changed  the net architecture,so I want to compare our result.can you share your code?</p>",
      "rawMarkdown": "hello,I am working on this problem using your tips but changed  the net architecture,so I want to compare our result.can you share your code?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 258391,
      "author_name": "moejoe",
      "author_url": "",
      "post_date": "12/16/2017 02:44:34",
      "content": "<p>Edit: <a href=\"https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\">https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution</a>\nHere's my code.</p>\n\n<p>In my first attempts I tried to use MVCNN on a3daps files and downsampled inputs to a 3rd of the original resolution. This was getting about 0.10 on the original test data. I never tried 3D convolutions as I thought it'd be too slow and memory hungry.</p>\n\n<p>I found I got better results by using a3d files and maintaining the original resolution (and lowering the batch size considerably so it would fit in memory). I tried combining a3d and a3daps by throwing them in separate channels but I found it caused overfitting.</p>\n\n<p>Later on I dropped MVCNN and went with a custom architecture where each view is fed through a pretrained ResNet-50 CNN and each feature map is pyramid pooled (to find objects/features of differing size) and fed to a LSTM with attention. My pyramid pooling doesn't actually use pooling layers but instead passes the last feature map through a 1x1, 3x3, and 5x5 convolution with stride 1, 2, and 3 respectively. I use a form of attention for CNNs that I apply to the feature map before the pyramid pooling. To make it run on my 1080 ti I increased the stride of the first couple layers from 2 to 3, which probably has a similar affect as decreasing the resolution but loses less information. I trained with SGD with a cosine annealing schedule. This was getting around 0.015 on the original test data.</p>\n\n<p>I didn't use any segmentation and had my model predict on all 17 zones at once. In hindsight maybe I should have segmented as it could allow me to train with higher resolution and focus attention on relevant zones at the cost of speed, but I was also worried that the people in the stage 2 data would be drastically different and the segmentations would be wrong. I also probably shouldn't have done experimental things like CNN attention and cosine annealing but then again maybe the risk paid off and they helped.</p>\n\n<p>Regrettably I didn't use ensembling. Just didn't have the time to train more models or think it would be too important. But looks like it could have made a difference.</p>\n\n<p>But in the end I'm still very satisfied with my score, and congratulations to everyone!</p>",
      "votes": null,
      "replies": [
        {
          "id": 258408,
          "author_name": "w9wang2",
          "author_url": "",
          "post_date": "12/16/2017 03:19:11",
          "content": "<p>Moejoe,\nThanks for sharing your experience and thoughts.  I wonder how did you decide which zone the threats locate If you did not use any segmentation. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258419,
          "author_name": "moejoe",
          "author_url": "",
          "post_date": "12/16/2017 03:43:19",
          "content": "<p>Hi Wensu, the CNN learns on its own which locations correspond to which labels without any human guidance. It outputs a size 17 vector and is trained against size 17 targets. Since the classes aren't mutually exclusive I don't use softmax and I train with binary cross entropy loss.</p>\n\n<p>The average pooling layer at the end of ResNet loses all spatial information so one would imagine this would be a problem. Instead I use a pyramid pooling scheme which pools with multiple kernel sizes and feed the output (which maintains information about location) directly to the LSTM. However, although pyramid pooling helped noticeably, I found that it wasn't absolutely necessary and even with the regular average pooling layer it still gave decent results (at least on the stage 1 data).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258581,
          "author_name": "w9wang2",
          "author_url": "",
          "post_date": "12/16/2017 14:00:10",
          "content": "<p>Hi Moejoe, Thanks for your response.  Since I spend more than a half time on segmentation, I'd like to learn how others did for this issue.  you said your model learns on its own correspond to which labels, it is great and is what I like to do, but did not figure out how.  you mentioned your model is trained against 17 targets, how did you get the 17 targets without segmentation,  could you explain further how your model did?  Thanks a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258684,
          "author_name": "moejoe",
          "author_url": "",
          "post_date": "12/16/2017 19:38:17",
          "content": "<p>The targets come from the training data. Consider the following sample from the training data:</p>\n\n<pre><code>0050492f92e22eed3474ae3a6fc907fa_Zone1,0\n0050492f92e22eed3474ae3a6fc907fa_Zone10,0\n0050492f92e22eed3474ae3a6fc907fa_Zone11,0\n0050492f92e22eed3474ae3a6fc907fa_Zone12,0\n0050492f92e22eed3474ae3a6fc907fa_Zone13,0\n0050492f92e22eed3474ae3a6fc907fa_Zone14,0\n0050492f92e22eed3474ae3a6fc907fa_Zone15,0\n0050492f92e22eed3474ae3a6fc907fa_Zone16,1\n0050492f92e22eed3474ae3a6fc907fa_Zone17,0\n0050492f92e22eed3474ae3a6fc907fa_Zone2,0\n0050492f92e22eed3474ae3a6fc907fa_Zone3,0\n0050492f92e22eed3474ae3a6fc907fa_Zone4,1\n0050492f92e22eed3474ae3a6fc907fa_Zone5,0\n0050492f92e22eed3474ae3a6fc907fa_Zone6,0\n0050492f92e22eed3474ae3a6fc907fa_Zone7,0\n0050492f92e22eed3474ae3a6fc907fa_Zone8,1\n0050492f92e22eed3474ae3a6fc907fa_Zone9,0\n</code></pre>\n\n<p>This is treated as an input target pair where 0050492f92e22eed3474ae3a6fc907fa.aps is the input and [0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0] is the target and what I train the model to output. The model is fed no other information.</p>\n\n<p>It may feel like black magic at first but it works and it works pretty well. The architecture of the model is designed so that it's more spatially aware than regular convnets to make it easier to learn.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258742,
          "author_name": "w9wang2",
          "author_url": "",
          "post_date": "12/16/2017 23:36:58",
          "content": "<p>Hi Moejoe,\nThank you very much for your explanation, it sounds abstract but really a clever approach. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258761,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/17/2017 01:06:06",
          "content": "<p>I didn't use segmentation either. I felt it wasn't possible to proper segment the image and sometimes the threat was more easily visible from a strange angle. For example the ankle strapped threat on the inside of the right leg may be best viewed from the left (with the left leg obstructing most of the view).\nI'm not too worried about \"landmarks\" as other have mentioned as the pixel position remains \"known\" throughout your convolutional layers until the resolution gets very small at which point it supposedly has figured out towards which label we're going.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258842,
          "author_name": "jamesrequa",
          "author_url": "",
          "post_date": "12/17/2017 05:37:45",
          "content": "<p>Actually segmentation wasn't very difficult to get pretty accurate if working with a3d images. While there are 16 &amp; 64 views with aps/a3daps there is really just one aerial view for a3d so all you needed to get right was the z-axis coordinates and the rest you could rely on basic symmetry of the human body. I used a simple CNN on a3daps images to predict the height range coordinates for each mirror zone pair (6/7, 2/4 etc.) and then  later any threat that my classifier found in that slice range was matched up with the corresponding zone. </p>\n\n<p>Although I definitely think Moejoe's \"automated\" solution is much more elegant and wish I had thought of it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260212,
          "author_name": "jamesthornton",
          "author_url": "",
          "post_date": "12/19/2017 21:44:11",
          "content": "<p>Hi Moejoe, thank you for sharing this solution. Dang I'll have to investigate why my MVCNN never started working (it only produced the expected probabilities).</p>\n\n<p>Anyways, here's my question - where is that LSTM bit coming from? What is the purpose and how did you know to use it? I haven't heard much about adding an RNN layer for image classification, so I'd be very curious to hear a bit more about that and why it was so applicable here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260315,
          "author_name": "moejoe",
          "author_url": "",
          "post_date": "12/20/2017 02:49:01",
          "content": "<p>In MVCNN you simply take the max probabilities calculated from each frame after they've each gone through the network independently of each other.</p>\n\n<p>In LSTM the frames are no longer treated independently of each other. It is treated as a time sequence where each time step is a frame from the .aps file and there's 16 time steps. Instead of feeding the frames directly to the LSTM I first feed it through the CNN to \"compress\" it into a small high-level feature map and feed that in. The advantage of this setup over MVCNN is that the network can learn temporal dependencies. It can learn what a rotating person should look like and flag anomalies. It can learn to boost its predictions if it notices an anomaly in multiple frames as opposed to just one frame, or suppress a prediction if an anomaly was noticed in just one frame but the others didn't pick anything up. It seems to have a regularizing affect.</p>\n\n<p>I viewed the problem as a single temporal sequence rather than 16 independent problems and it seemed natural a LSTM could be applied.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260328,
          "author_name": "moejoe",
          "author_url": "",
          "post_date": "12/20/2017 03:06:19",
          "content": "<p><a href=\"https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\">https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution</a></p>\n\n<p>Here's my code. Model is definied in mvcnn.py.</p>\n\n<p><a href=\"https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution/blob/master/mvcnn.py\">https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution/blob/master/mvcnn.py</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260340,
          "author_name": "rkrowe",
          "author_url": "",
          "post_date": "12/20/2017 03:23:21",
          "content": "<p>@Moejoe -- Thank you for sharing your code. Nicely done and congratulations on your very good performance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258430,
      "author_name": "seansoleyman",
      "author_url": "",
      "post_date": "12/16/2017 04:12:53",
      "content": "<p>Congratulations to everyone who participated in this challenging and suspenseful competition! </p>\n\n<p>I used a multi-view CNN custom ResNet that looks at the entire image (without any pre-segmentation) and outputs 17 probabilities. I used the .aps dataset downsampled to half-resolution. </p>\n\n<p>My stage1 score (around 0.025) was much better than my stage2 score (0.13589). This must be because the subjects in the stage1 training set were the same as those in the stage1 test set, but the subjects in the stage2 set were completely different people. </p>\n\n<p>During the optimization process, I used regular cross-validation. After reading comments from others, it has become clear to me that the correct way to do cross validation was to first group all of the instances depicting the same person, and then randomly assign each group of instances to a cross validation set. </p>\n\n<p>I had suspected that this would be an issue, but could not have imagined that it would make such a huge difference between stage1 and stage2 performance. I think there are two possible mechanisms that could explain this:</p>\n\n<ol>\n<li><p>Since my net does not use pre-segmentation, it probably determines the threat zones by identifying landmark features. For example, it might identify a threat between the hand and elbow, and therefore determine that the threat lies in the forearm region. However, different people have different-looking hands and elbows. This may have prevented my net from properly identifying landmarks in stage2. </p></li>\n<li><p>It is also possible that the net learns to whitelist certain objects that would otherwise be identified as threats. The Bob Marley lookalike from stage2 is a perfect example. If the stage1 training set had included someone with this type of hair, the net would most likely have learned to specifically avoid labeling dreadlocks as a threat. </p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 258760,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/17/2017 00:56:53",
          "content": "<p>I have really banged my head against walls for how other got so much better than I, with or without overfitting I don't care. Can you describe some more details how you got to 0.025? I never got past 0.11,.. See my writeup elsewhere in this thread fo rmy architecture.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258773,
          "author_name": "seansoleyman",
          "author_url": "",
          "post_date": "12/17/2017 01:43:26",
          "content": "<p>I found one paper in particular that was very helpful: \n<a href=\"https://arxiv.org/pdf/1703.07047.pdf\">High-Resolution Breast Cancer Screening with Multi-View Deep Convolutional Neural Networks</a></p>\n\n<p>My net was very similar to the one described in this paper. I added a convolutional layer with 17 output channels right before global average pooling. I optimized the layer parameters, but as described above it seems that the optimization was performed under unrealistic test conditions. </p>\n\n<p>I trained the net using SGD + Momentum + Nesterov with lr=0.4 and momentum=0.32. I settled on these unusual values after running an extensive random search, and was very surprised that the optimal momentum turned out to be so low. I also got a lot of unusual \"spikes\" in the train and test error during training. The graph below shows the log loss on one of the cross validation folds, with exponential lr decay. I will be very grateful if someone can offer insights into this unusual behavior. </p>\n\n<p><img src=\"http://seansoleyman.com/wp-content/uploads/eval.png\" alt=\"Learning Curve\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258820,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/17/2017 04:56:07",
          "content": "<p>Wohaa! You cut straight through to 0.015, you don't want to know how long I have been watching my charts slowly inching down from 0.25 to 0.18 on the test set, never getting further than that. The last bit, from 0.18 to 0.12 I did by repeating the infer at 72 different crops and resizes + mirroring. But then again, you heavily overfit making your final score not all that much better than mine ;-). I wonder if the number of feature maps in the conv layers may have to do with that. I have been very conservative with a maximum of 80 feature maps, the paper you refer to has 256. My rationale was, these standard nets (Resnet, VGG etc) are designed and tested for millions of images with thousands of classes, here we have a lot less of both so the network needs to be scaled down accordingly. Maybe I have been to aggressive with that.</p>\n\n<p>The peak in your training, is that not caused by the faulty image in the APS series? There is one that is half blacked out. Or otherwise maybe the crazy momentum?</p>\n\n<p>Did you do any tricks at infer time? Did that improve matters?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258870,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "12/17/2017 07:37:43",
          "content": "<blockquote>\n  <p><strong>Bastiaan Bergman wrote</strong></p>\n  \n  <blockquote>\n    <p>Wohaa! You cut straight through to 0.015, you don't want to know how long I have been watching my charts slowly inching down from 0.25 to 0.18 on the test set, never getting further than that. </p>\n  </blockquote>\n</blockquote>\n\n<p>A watched pot never boils. The secret is to ignore them. Do what you would normally do through the day. When you finally remember that you are training something and look at the log, then you'll find yourself thinking \"whoa, awesome results!\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259089,
          "author_name": "seansoleyman",
          "author_url": "",
          "post_date": "12/17/2017 17:11:49",
          "content": "<p>This was after I had removed the bad instance and switched the two sets of labels. This was one of the better CV folds - they averaged out to around 0.025. </p>\n\n<p>I did use a large number of feature maps (somewhere around 256 for the upper layers) because this helped speed up training and reduced the stage1 error. Looking back, I probably would have done better without such a large number of feature maps. Dropout may also have improved stage2 results by preventing the net from memorizing specific features, although it did not help much with stage1. </p>\n\n<p>Much of this is just speculation - I wouldn't read too much into these details!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259297,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/18/2017 04:48:58",
          "content": "<blockquote>\n  <p>A watched pot never boils.</p>\n</blockquote>\n\n<p>Yeah, I tried that too,.. something tells me you had more in your secret sauce ;-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258445,
      "author_name": "jamesrequa",
      "author_url": "",
      "post_date": "12/16/2017 05:23:01",
      "content": "<p>Initially, I started out by using a combination of both the aps and a3daps files. Rather than use all 16 and 64 \"views\" from each body scan format, I instead selected 8 prominent views i.e. front-center, front-left, etc. from both files and then further segmented and cropped out all 17 zones into equal 100x100 square images. I then concatenated these into a 400x400 \"mosaic\" for each passenger/zone combo. This was fed to my model that generated a binary classification. Using this strategy I was able to achieve .03 logloss on stage 1 test set. Like others have echoed on the forum threads, I too found that if I left out the aps images I would end up with lower accuracy yet I did observe that I got better accuracy when having a3daps images included in the mix rather than just using aps by itself since not all threats were visible in all cases. </p>\n\n<p>With around 2 weeks left in the competition I tried using the a3d images to see if I could improve my score further. I found that I couldn't get good results using a 3d convnet, so instead I tried a different strategy which was to treat all image slices as separate images to feed to the model in a sequence. So for each passenger/subject there was a total of 310 slices (step-size of 2 pixels and clipping everything over 620 which was mostly noise). I also segmented by zone. With this strategy I was surprised that I was able to get as low as .005 logloss on stage 1 test set (this also corresponded with my local validation loss). This was using just a single model (DenseNet169) and one 15% validation hold out. Not only was it much easier for my model to locate and classify the threats, using this strategy also really helped me with segmentation because I could also see patterns in the data like exactly how long a threat was and which image slices were most confident etc. This high score in stage 1 gave me enough confidence that I should abandon my first strategy and only use a3d images for stage 2. In hindsight, I can now see this was not such a good idea as my score ended up dropping to 0.17 on stage 2!! After reviewing my predictions, I can see that some new features of the passengers introduced in stage 2 (that didn't exist in stage 1 training set or test set) caused my model to generate a large number of false positives in a few isolated zones. Had my model at least seen examples of something similar I think it probably would have scored much higher as it would have learned to whitelist some of those new features in stage 2 (i.e. dreadlocks, suspender clips, etc.) that looked very similar to threats in stage 1 training set. </p>\n\n<p>While I am quite disappointed in my final result, but overall I think this was a really interesting challenge and was a great learning experience for me!! Hopefully someone here can also learn from my mistakes!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258447,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "12/16/2017 05:26:58",
      "content": "<p>My CV was 0.032-ish (No subject was in multiple CV folds, of course). </p>\n\n<p>The big surprise in Stage2 was the guy with massive dreadlocks, like no one had in the training dataset. They looked a lot like some of the bombs in the training  scans instead, except for the fact that they adjoined his head. My model was understandably suspicious. I had calculated that that would cost me 0.01 on the score, and it seems like it did.</p>\n\n<p>If I could have seen the Stage2 data for just 10 minutes, before finalizing the model, I would have reduced the confidence of zone 6 &amp; 7 predictions, and negated most of the 0.01 damage, but the predictions are supposed to be fully automatic (or such was my understanding).</p>",
      "votes": null,
      "replies": [
        {
          "id": 258451,
          "author_name": "jamesrequa",
          "author_url": "",
          "post_date": "12/16/2017 05:34:09",
          "content": "<p>Congrats Oleg on placing 5th in this comp, thats very impressive!</p>\n\n<p>Glad you mentioned about this guy and that I'm not the only one who had issues with it. The dreadlocks were the main culprit for pretty much all of my false positives in zones 17, 6 and 7. My model also had problems with the guy with the suspenders near the metal clips (zones 8 &amp; 10). Like you say, if this wasn't a two stage comp I would have made some adjustments to my model's sensitivity/threat thresholds to mitigate this but since I didn't anticipate needing to do this in the model upload so I couldn't make those changes. </p>\n\n<p>&gt; <strong>Oleg Trott wrote</strong>\n&gt; \n&gt; &gt; My CV was 0.032-ish (No subject was in multiple CV folds, of course). \n&gt; \n&gt; The big surprise in Stage2 was the guy with massive dreadlocks, like no one had in the training dataset. They looked a lot like some of the bombs in the training  scans instead, except for the fact that they adjoined his head. My model was understandably suspicious. I had calculated that that would cost me 0.01 on the score, and it seems like it did.\n&gt; \n&gt; If I could have seen the Stage2 data for just 10 minutes, before finalizing the model, I would have reduced the confidence of zone 6 &amp; 7 predictions, and negated most of the 0.01 damage, but the predictions are supposed to be fully automatic (or such was my understanding).\n&gt; \n&gt; \n&gt; \n&gt; </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258613,
          "author_name": "tothink",
          "author_url": "",
          "post_date": "12/16/2017 15:47:18",
          "content": "<p>oh boy, I see a future where we data scientist are to blame for people with dreadlocks going through \"special screening\" at the airport :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260202,
          "author_name": "jamesthornton",
          "author_url": "",
          "post_date": "12/19/2017 21:34:42",
          "content": "<p>Hi Oleg, any chance you can share how you actually accomplished this? Even just a one liner with your actual architecture - it would be very helpful to hear how you were able to build a highly effective model that could share knowledge of threat appearance between zones.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260228,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "12/19/2017 22:08:33",
          "content": "<p>James,</p>\n\n<p>My approach was fairly complex overall -- I don't think I could describe it briefly. Perhaps I'll do a full write-up later. </p>\n\n<p>However, the particular aspect you are asking about seems straightforward: a single NN is trained on all relevant data. When it sees threats, it predicts what zone they are in. As I argued elsewhere in this thread, having a single omniscient NN instead of 10-17 specialized experts helps with generalization. I'll be very surprised if it turns out that any of the top teams used the specialized experts approach.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260233,
          "author_name": "mmiron",
          "author_url": "",
          "post_date": "12/19/2017 22:21:29",
          "content": "<p>File a FOIA for the code later on :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258618,
      "author_name": "lwalzer",
      "author_url": "",
      "post_date": "12/16/2017 16:10:19",
      "content": "<p>we used exclusively A3DAPS (generally 7 of the 64 views), training with a 7 layer model on individual views and then combining the  views using a separate model.</p>\n\n<p>Ensemble: our logistic regression also incorporated some home-brewed algorithms (using morphology and transforms)  which added predictivity in some cases over the ML results.</p>\n\n<p>VGG16 was hard to use on my desktop, it crashed often unless we throttled back the batch size significantly\nwe had som e success training multiple zones together using left right symmetry</p>\n\n<p>feeding all views at once also crashed my desktop (16GB of RAM), I would appreciate feedback how people made that work.</p>\n\n<p>I am surprised that some contestants had success with 100 x100 crops.  we generaly needed to go 225-300 wide depending on the zone and view</p>\n\n<p>Our model seemed to work very well on 11 zones and poorly on 6 zones: arms(4), groin, and upperchest. The arm regions were noisy. I would be interested in strategies for those zones.  We tried to rotate all arms to vertical but saw no improvement. Did anyone use gender recognition to facilitate groin and upperchest? </p>\n\n<p>Did anyone use cross-zone models? In the training set there was a slight inverse correlation among zones  (ie a passenger with contraband in 2 other zones was less likely to have it in a third zone) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258673,
      "author_name": "u39kun",
      "author_url": "",
      "post_date": "12/16/2017 18:55:17",
      "content": "<p>I started out with APS images first.\nI used various CNNs (VGG, Resnet, Densenet, etc) that have been pretrained with ImageNet weights as featurizers by stripping out their classification layers.\nFor each scan, forward passes were made on the CNN to create feature maps for all 16 views, and those feature maps were flattened and concatenated into a single vector to be fed into LightGBM for classification.\nAs LightGBM does not support multi-label output, each of the 17 zones were trained/eval'd separately.\nThis got me to ~0.19 on the public LB when I used Densenet 121 as the featurizer.  I did not expect this to work too well as ImageNet and TSA images are very different.  Also, some threats were very difficult to see in APS due to occlusion and how projections seem to be computed.</p>\n\n<p>Then I started looking at A3D and thought to myself that segmenting out into various zone groups would help with classification, because:</p>\n\n<ul>\n<li>Cassifiers that solely focus on its assigned specific zone group can be created, rather than having the network figure out what to focus on</li>\n<li>More samples can be fed because mirroring can be performed independently on the body part in question</li>\n<li>By eliminating other parts, threats can be seen more easily (you get a clear shot at the zone(s) in question.)</li>\n</ul>\n\n<p>I ended up with 6 classifiers for the following zone groups: arms, legs, chest-back, torso, waist, and crotch.\nThe \"arms\" classifier will get trained on right arm images, left arm images, and their horizontal mirrors to output 2 labels (i.e., whether a threat exists in the upper arm and/or lower arm.)\nThe \"legs\" classifier will be trained in a similar fashion to output 3 labels for detecting threats on the thigh, knee, and foot zones.  Others were trained to output a single label (i.e., binary classifiers.)</p>\n\n<p>I experimented by training Conv3D-based classifiers on the segmented zone groups, but this did not work so well, and I ended up creating a variant of the Multi-View CNN (<a href=\"https://arxiv.org/pdf/1505.00880.pdf\">https://arxiv.org/pdf/1505.00880.pdf</a>) except I used a learnable convolution layer instead of a view pooling layer.  Train-time augmentation was limited to horizontal flips, and test-time augmentation was limited to 2 combinations of horizontal flips and averaging out the predictions.  No trimming of predictions were performed (boy, the models gave very confident predictions so perhaps I should have looked into this a little more to minimize the penalty on mistakes.)\nThis approach got me to ~0.019 on the public LB, and ~0.089 on the private LB.</p>\n\n<p>My final approach was preprocessing heavy due to all segmentation done based on the center of mass and assumed\nbody proportions, then rotating the zone groups in 3D to create 2D projections (I actually ended up implementing multiple methods of 2D projections, as simply doing \"max\" on the collapsed axis would sometimes hide certain threats that appear on one side only.)  Also, training was very compute intensive since there were 6 separate networks to cover all the zone groups.  And I actually trained multiple networks for each zone group for ensembling, so a total of 38 networks were trained when it was all said and done.</p>\n\n<p>Looking at other competitors posts, I feel a bit shocked and a little embarrassed to find that some did better with a much simpler and elegant pipeline of predicting all 17 labels simultaneously without any segmentation or even ensembling!  I didn't even try that because I had strong convictions in my intuitions (that I unfortunately did not bother to validate.)</p>\n\n<p>I learned a lot in this competition!</p>\n\n<ul>\n<li>Explore wider before settling down on a method (I feel that I turned down my \"learning rate\" or \"temperature\" too quickly and did not get good coverage on the solution space.)</li>\n<li>Before building a complex pipeline, try out and validate simpler solutions first.</li>\n<li>Be very thoughtful in setting up local CV.  I did not other to separate out subjects when creating folds, unlike other wise competitors did.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 258677,
          "author_name": "dmitrykovba",
          "author_url": "",
          "post_date": "12/16/2017 19:18:14",
          "content": "<p>Separating subjects when creating training/validation folds could make huge difference on the stage 2 results. I didn’t do it as well :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258716,
      "author_name": "jandjenter",
      "author_url": "",
      "post_date": "12/16/2017 21:23:39",
      "content": "<p>At first, I went through aps images and I felt that many threats seem to be hard to find, so I decided to go with a3d.\nI used APS to identify each person and make train/val set - I manually annotated a few persons at first and let NN to classify the person. But later I found that PCA can do same thing easily from discussion board :(</p>\n\n<p>When I tried to feed a3d data, the data size was too big - so I scaled down to 1/4 on each dimension and cut down mostly empty areas on front/back/side/top). I could see the threats even after 1/4 scale, so I assumed that 1/4 dowscaling doesn't reduce the accuracy much. And I adjusted vertical scale little bit for each image so that the head sits on similar position(no ML, just simple image processing). Even after scaling down, only 2-3 batches fit into 1080Ti.</p>\n\n<p>Next, I fed the data into resnet-like 3D convolution. All data is fed to 3 blocks of identify block. Then output is splitted into 17 zones(with some overlap) and connected to per-zone 3D CNN(two blocks of identify block). It is quite simple network with not much parameters, so it just takes 10 minutes per epoch and the loss stablized around 80-130 epochs. Interesting thing is that I tried to share same per-zone network for symmetric zones(like left arm and right arm), but it came out worse. I think keeping separate network for two sides made some ensemble effects. </p>\n\n<p>I used basic augmentation like rotation, zoom, shift, flip left-right.</p>\n\n<p>I tried to separate each body part and put each part into the network, but it did much worse. I guess it's because NN losses context of body and have trouble finding whether the thread is the part of body or not. </p>\n\n<p>I didn't try much other experiments due to lack of time/resources. In hindsight, I wish I put little bit more time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 258830,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/17/2017 05:14:46",
          "content": "<p>Thanks for sharing! How did you \"split the data into 17 zones with some overlap\"? You're not referring to segmentation, do you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258848,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "12/17/2017 06:09:11",
          "content": "<p>It's not segmentation, but just trick to reduce computation since 3D convolution requires a lot even at  1/4 scale.\nAfter 3 resnet block for the whole body, output dimensions are like 10(front-back) x 12(left-right) x 18(foot-hands) x # of last layer filters. I just slice this output to 17 zones - for example, I sliced it like x[2:6, 2:4, 0:1, :] for left feet, so this per-zone CNN only see small area only on center-left-bottom part.</p>\n\n<p>I think the network will eventually figure out the area of interest even if I didn't use this per-zone network, but it will require network with much more capacities and more training time. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258721,
      "author_name": "alanpo",
      "author_url": "",
      "post_date": "12/16/2017 22:02:30",
      "content": "<p>Congratulations to winners, \nthanks to all sharing their approach!</p>\n\n<p>My approach  was binary classification  with a separate set of  CNN models for each zone-view.\nClassifier models were fed with zone segments cropped from a3daps views. \nCropping coordinates were created by Unet segmentation models which I had to train. \nFor this I annotated 0,16,32 and 48-th views from 625 images.  </p>\n\n<p>To  combine models for different views of same zone segment I trained Xgboost logistic regression models.</p>\n\n<p>I used only  0,3,29,16,32,35,48 and 61 views, not all of them in each zone.\nClassifiers were Inception V3, Densenet161 and Resnet152.</p>\n\n<p>I also experimented with mosaic of 2 and 3 views for some zones. <br>\nNo ensembling of final models nor multi-fold validations.</p>\n\n<p>Stage 1 best score was 0.1074 with corresponding Stage2 of 0.1578.</p>\n\n<p>Segmentation worked pretty well on both stages. \nClassification was a real challenge.\nThe hardest zones were 2 and 4 with local validation scores of 0.156 and  0.171 respectively.\nThe easiest  zones were 5 and 17 with 0.045 and 0.012 scores.</p>\n\n<p>If I have time I would like to try approach close to Moejoe's and make some late submissions.\nI very much appreciate the ideas behind this much more elegant and effective approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 258744,
          "author_name": "w9wang2",
          "author_url": "",
          "post_date": "12/16/2017 23:52:32",
          "content": "<p>Hi Alexander, Thanks for sharing your experience.  I also used a segmentation approach with opencv which is not an effective way but works.  I'd like learn from others to improve my approach.  You used Unet segmentation models, could you explain a  little further about this model and if you can provide  any references?  Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258778,
          "author_name": "alanpo",
          "author_url": "",
          "post_date": "12/17/2017 02:14:16",
          "content": "<p>Hi Wensu,\nThis is the original paper regarding Unet:\n<a href=\"https://arxiv.org/pdf/1505.04597.pdf\">https://arxiv.org/pdf/1505.04597.pdf</a></p>\n\n<p>As for me, I adapted Unet segmentation code from Carvana Image Masking challenge posted here:\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208#latest-222649\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208#latest-222649</a></p>\n\n<p>I created very simple masks  from my annotations. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259134,
          "author_name": "w9wang2",
          "author_url": "",
          "post_date": "12/17/2017 18:57:44",
          "content": "<p>Hi Alexander, thank you very much for the references you provided.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258734,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "12/16/2017 23:10:31",
      "content": "<p>... Training separate models, one for each zone is very suboptimal. It's easy to see that with the following argument: </p>\n\n<p>If you have ankle-strapped guns in the training dataset, the model that looks at ankles will learn to recognize them. However, if forearm-strapped guns only occur in the test dataset, they will be gratuitously novel to the model that looks at forearms. In other words, you won't transfer knowledge between zones as well as you should.</p>",
      "votes": null,
      "replies": [
        {
          "id": 258768,
          "author_name": "tanderton",
          "author_url": "",
          "post_date": "12/17/2017 01:37:16",
          "content": "<p>That argument only applies if what is going on is that the network learns to detect threats and not that the network learns to detect normalcy.  It is my interpretation of my own models that they were in fact doing the later and not the former.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258776,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "12/17/2017 02:02:26",
          "content": "<p>I don't think it's an either-or type of situation. It helps to know what normal bodies look like and it helps to know what the threats look like if you aim to tell them apart. Quantitatively though, the image patches containing threats were much more scarce. Squandering them seems to me like a serious mistake.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258790,
          "author_name": "alanpo",
          "author_url": "",
          "post_date": "12/17/2017 03:05:05",
          "content": "<p>Very valuable points., Oleg. Besides, segmentation models rely on annotations and their quality. Plus complexity and extra work... Much better when single net cares about zone boundaries and threats. Still I reserve my right to make serious mistakes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258845,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/17/2017 05:50:12",
          "content": "<p>So how did you do it? Used and MVCNN? ;-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258750,
      "author_name": "bastiaanbergman",
      "author_url": "",
      "post_date": "12/17/2017 00:10:12",
      "content": "<p>My architecture:\nI worked with APS only, an MVCNN on all 16 views concurrently while also multi-task learning, in hopes of better generalization. I had 6 resnet-like layers on each view, then a average pooling to put all 16 views together and another 2 resnet-like layers then it split in two: 1) 4 layer resnet to train on the 17 threat labels and 2) a 4 layer resnet to train on the gender of the subject (which I manually labeled on the training data using PCA). I reduced the input images resolution with a factor 2 and used random crop and zoom to compensate for that. I also had the first two conv layers with a stride 2 to reduce the data.</p>\n\n<p>An interesting trick I found useful: You can <code>revolve</code> the weights from each of the view pillars. Take the weights from the 6 resnet layers of view 0 and put them in the layers of view 1, from view 1 to 2 and so on, finally put view 16 layers into view 0. If you do this every 20 epochs you'd expect you'll have to learn pretty much from start every time, but somewhat to my surprice the network doesn't lose it's ability much. Instead it prevents overfitting as  shown by the test-cross entropy which never surpassed my training cross entropy.</p>\n\n<p>Another interesting trick, I suppose most people have figured out is image mirroring for data augmentation. You have to be somewhat careful as the labels change too, e.g. a left arm threat becomes a right arm threat. Moreover the order of the APS views needs to be reversed to keep the subject as a left rotating subject. If you do that, mirror the image, reverse the views and mirror the labels you basically double the dataset. Sort of a standard trick, but with a little more to it in the case here, due to the nature of the data.</p>\n\n<p>I worked on a single GPU and couldn't get my batch size above 4, nor could I increase my network size, both due to memory constraints. Increasing the network size didn't seem to help much anyway. I'm not sure how much the gender training helped for model generalization. I did not test on proper test splits (proper as in separate subjects and threats for different folds), but my stage 1 score was 0.12 to 0.16 in stage 2.</p>\n\n<p>Things I would add next time:\n1) 5 fold training with proper separation of subjects and some boosted tree method to combine them\n2) Also separate threats, but I have no idea how\n3) Reduce image resolution more in the beginning\n4) More GPU's or different hardware to get a bigger model</p>\n\n<p>I don't think any of this will get me the winning model, wonder what the big difference is, what I'm missing. Hope to learn it here!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258775,
      "author_name": "sandshift",
      "author_url": "",
      "post_date": "12/17/2017 01:52:26",
      "content": "<p>I used two networks, one using about 15% of training data to classify threats, and the other had about 1% of training data to classify categories.  I used an IoU metric to match the location of the detection. So in total 2 resnet networks trained on a gtx 1070. Only data augmentation was horizontal flips. and only a3daps images.  I think the approach could yield a much better result with more of the training data used and more time to refine the model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 258825,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/17/2017 05:07:33",
          "content": "<p>This seems interesting, thanks for sharing! What do you mean with classify threats versus classify categories? What is the difference? And what did you do with the other 84% ot the training data?</p>\n\n<p>I got a great improvement by using random crop and zoom, which I believe has to do with the image reduction I applied. You may see a similar gain with that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 259183,
      "author_name": "jfaath",
      "author_url": "",
      "post_date": "12/17/2017 21:27:33",
      "content": "<p>I started with zone segmentation from the very beginning. The main reason for this is I've been an engineer for a long time but have only recently gotten into ML. I don't have a lot of tools in the ML toolbox but I can work through an engineering problem like segmenting an image (though a lot of that was new too!). Also I thought I might have had a puncher's chance at doing well because I figured not a lot of teams would spend time on meticulously segmenting the zones. I really didn't know how else to proceed but my solution certainly evolved as I went along.</p>\n\n<p>TL;DR - I did a reasonable job at segmenting the zones</p>\n\n<p>For my first segmentation attempt I took the max point of all slices in the dataset and tried to hand-curate the zones. I didn't like those results so then I decided if I could determine a few critical points, namely where the torso began (at the groin) and ended (at the neck) along with the torso width I could assume a universal body proportion and accurately segment the zones. I was able to determine the points pretty easily with the skimage library and it did seem like most people had similar relative proportions. Several caveats: 1) an individual could be off center on either floor axis which required an adjustment (this could be determined by looking at the center point of slices 0,8 and 4,12), 2) some subjects had excessive...um...girth so that would require an adjustment across slices and finally 3) by observing scanners at airports, I concluded that the scanner would start slow, speed up, then slow down again which required a further adjustment. To all of this, I added some padding for each zone and it seemed I had a pretty robust zone segmenter - it even looked good in the new stage 2 dataset.</p>\n\n<p>TL;DR - Used Keras' TimeDistributed layer with InceptionV3 freezing the first 172 layers and adding an LSTM and dense layer at the end. One classifier for zones 6,7,8,10-16, one for the arm zones and one each for 5, 17 and 9. Weighted threats 90/10.</p>\n\n<p>My model was relatively simple and I never experimented with anything else. I essentially viewed the input as images across time (ie. video). So I used Keras' TimeDistributed layer with a pre-trained Imagenet model (VGG16 and InceptionV3) fed into an LSTM to a dense layer and finally to a binary result. At this point it was all manual experimentation. I tried different slices in each zone, different image input sizes, different sizes for LSTM and dense layer and playing with un-freezing various layers in the pre-trained models. At first I did not fully segment the arm and leg zones (for ex. zones 11, 13, 15 were one classifier) but ultimately I achieved a nice jump on the leaderboard by segmenting the legs so that reinforced the idea, perhaps incorrectly, that segmenting was the way to go. In the end I used 150x150 size images, with 9 time slices (save for 5, 17 and 9 which had less) and then used InceptionV3 in the TimeDistributed layer with the first 172 layers frozen and the rest trainable.  I weighted threats 90/10 in gradient updates since they were so much less prevalent. I discovered the symmetry trick so I started pairing zones (like [6,7], [8,10], [11,12], etc.) by flipping horizontally. To additionally augment, I flipped the temporal axis. Then I thought why not combine more zones? So I did all the leg zones in one classifier and all the torso zones. I kept moving up in the LB so finally I just did every 9-slice zone in one giant classifier and started flipping vertically to additionally augment. The end result: one classifier for zones 6,7,8,10-16,  one for the arm zones and one each for zones 5, 17 and 9. I reached 0.108 on the public LB. I discovered all this combining at the end so I think I could've improved that if I had more time to train.</p>\n\n<p>I thought I had something with the mega-classifier. It allowed me to greatly augment the data and I figured I was relatively immune to new people being introduced in stage 2 since zone segments I would think lack landmark features and by using all the zones together and flipping them around the network could focus on the threats. One glaring problem with segmenting the zones is many of the threats seemed to span more than one zone even though they were labelled for one zone. So I'm sure that confused the network in that it maybe only identified a threat when it was in the center.</p>\n\n<p>My stage 2 results more than doubled. Bob Marley and suspender-dude didn't help. Also in my brief manual inspection, it seemed like there were new, subtle-looking threats that weren't in the stage 1 data. I also didn't account for mislabels by pulling back any 1.0 predictions to 0.999. Live and learn. This competition was a ton of fun though and looking forward to the next one!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 259200,
      "author_name": "hackerpoet",
      "author_url": "",
      "post_date": "12/17/2017 23:01:46",
      "content": "<p>As promised, I'll post my architecture.  Obviously it must have over-fit to the stage 1 participants.  In hindsight I should have split my validation by individual people.  But anyway I guess it was my first real competition and I learned a valuable lesson for next time I do a 2 stager :P</p>\n\n<h2>Normalization</h2>\n\n<p>I used only aps samples.  Height was the most variable feature, so I stretched each person to be the same height.  I chose the height based on the highest pixel above a certain threshold, it actually worked extremely well and reliably.  Once normalized, I created crop locations for each zone in each rotation, skipping occluded channels.</p>\n\n<h2>Augmentation</h2>\n\n<p>I doubled the sample size by including mirror images.  My crop areas were 160x160 pixels, but in the actual training, I randomly sampled 128x128 crops from that larger area to sample random translations.</p>\n\n<h2>Architecture</h2>\n\n<p>Each zone had a different fully convolutional network.  I couldn't use vanilla convolutions since there is no spatial locality between rotations, so I created an 'Independent Convolution' layer that partitions the rotations and applies the same convolutions independently to each rotation.  I had 5 layers of convolutions and pooling before a single dense layer that combines the partitions.</p>\n\n<p>None of my networks had more than 14,000 parameters, usually a lot less.  Keeping a low parameter count was how I avoided over-fitting (at least to the stage 1 validation I guess).  Although it was different for each zone, here's how a typical network looks:</p>\n\n<pre><code>_________________________________________________________________\nLayer (type)                 Output Shape              Param #\n=================================================================\ninput                        (None, 14, 128, 128)      0\n_________________________________________________________________\nind_conv2d_1 (IndConv2D)     (None, 168, 124, 124)     312\n_________________________________________________________________\nmax_pooling2d_1 (MaxPooling2 (None, 168, 62, 62)       0\n_________________________________________________________________\nactivation_1 (Activation)    (None, 168, 62, 62)       0\n_________________________________________________________________\nind_conv2d_2 (IndConv2D)     (None, 168, 60, 60)       1308\n_________________________________________________________________\nmax_pooling2d_2 (MaxPooling2 (None, 168, 30, 30)       0\n_________________________________________________________________\nactivation_2 (Activation)    (None, 168, 30, 30)       0\n_________________________________________________________________\nind_conv2d_3 (IndConv2D)     (None, 168, 28, 28)       1308\n_________________________________________________________________\nmax_pooling2d_3 (MaxPooling2 (None, 168, 14, 14)       0\n_________________________________________________________________\nactivation_3 (Activation)    (None, 168, 14, 14)       0\n_________________________________________________________________\nind_conv2d_4 (IndConv2D)     (None, 168, 12, 12)       1308\n_________________________________________________________________\nmax_pooling2d_4 (MaxPooling2 (None, 168, 6, 6)         0\n_________________________________________________________________\nactivation_4 (Activation)    (None, 168, 6, 6)         0\n_________________________________________________________________\nind_conv2d_5 (IndConv2D)     (None, 168, 4, 4)         1308\n_________________________________________________________________\nflatten_1 (Flatten)          (None, 2688)              0\n_________________________________________________________________\ndropout_1 (Dropout)          (None, 2688)              0\n_________________________________________________________________\ndense_1 (Dense)              (None, 1)                 2689\n_________________________________________________________________\nactivation_5 (Activation)    (None, 1)                 0\n=================================================================\nTotal params: 8,233\n</code></pre>\n\n<h2>Hyperparameter Optimization</h2>\n\n<p>The only hyper-parameters I would tune were the amount of dropout at the dense layer and the number of kernels in the convolutions, which was the same for every layer.</p>\n\n<h2>Ensembleling</h2>\n\n<p>I didn't use ensembling in the traditional sense.  Instead of using multiple models, I just also ran the mirror image through the same model (mirroring the results) as well as 9 different translations for each.  Then I averaged the results, which had to be done pre-sigmoid for statistical correctness.</p>\n\n<h2>Conclusion</h2>\n\n<p>The low parameter count and low stage 1 score gave me too much confidence that \"it couldn't possibly be over-fitting\" but obviously in hindsight I should have tested how well it generalized to an unseen body by re-splitting the data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 259294,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/18/2017 04:41:13",
          "content": "<p>Hi Kevin,\nCongratulations, a great score nevertheless! And thanks for the write up.\nI wonder what exactly your 'Independent Convolution' is about. Is it convolution on the image section of one particular threat zone? Which is different for each of the 16 views?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259299,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "12/18/2017 04:55:54",
          "content": "<p>Sort of yeah, I think its similar to a MVCNN, but instead of pooling to combine the views, I just use a dense layer.  The weights are shared for each of the 12-15 views in the zone.  I don't think any zone actually used 16 views since there was occlusion in at least some of them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259512,
          "author_name": "gaborvecsei",
          "author_url": "",
          "post_date": "12/18/2017 14:40:17",
          "content": "<p>Hi, sorry if my question is dumb, but what is this pre-sigmoid averaging? Why would this bring statistical correctness?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259574,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/18/2017 16:45:43",
          "content": "<p>I was wondering the same ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259594,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "12/18/2017 17:25:55",
          "content": "<p>So when using log-loss, the resulting probability should be interpreted as \"Confidence\".  When you have one network with 50% confidence, and one with 99.999% you definitely wouldn't want to say that it's only 75% confident!  The other network is exponentially more confident.  The difference between 99% and 99.9% is much bigger than 50% and 50.9%.  This whole argument is of course specific to log-loss scoring.  </p>\n\n<p>Because the space is stretched like this, the correct way to combine the values is to evaluate the network all the way down to right before the final sigmoid, average, and then finally apply the sigmoid once everything's been averaged together.  This transforms the weird 'sigmoidy space' to a nice linear space where averaging makes more sense.</p>\n\n<p>Overall this gave me a 30% improvement on the LB score compared to just averaging after the sigmoid.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259684,
          "author_name": "mmiron",
          "author_url": "",
          "post_date": "12/18/2017 20:56:56",
          "content": "<p>Thanks for posting this, Kevin; sorry that I was wrong about you being the obvious winner, but congratulations all the same -- hey, you beat me at least!  :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 259789,
      "author_name": "jamesthornton",
      "author_url": "",
      "post_date": "12/19/2017 01:50:46",
      "content": "<p>Overall, I concatenated threat zone crops from the aps files into \"threat plates\" and trained an ensemble of ResNet50 (Imagenet weights) using those. I normalized the data by cropping each of the 16 images around a certain threshold and resizing that to be the full image, so that threat zone locations would be more consistent. Beyond that, there were two key insights behind my solution:</p>\n\n<ol>\n<li><p>Exploit the symmetry of the human body. By switching and horizontally flipping certain images I could use the same crops for right and left forearm, right and left shin, etc. Thus I doubled the size of the dataset and only needed 9 networks for a full segmentation approach.</p></li>\n<li><p>Clipping. Based on the math of the log loss error function, on an inaccurate prediction, overconfidence will be penalized really harshly. So I clipped my predictions to the [0.015, 0.985] interval and this took me from bronze to silver.</p></li>\n</ol>\n\n<p>I really appreciate all the help I got when I asked! I think the community spirit is one of the greatest aspects of Kaggle. Here's a repo with more info and my code - <a href=\"https://github.com/jamespeterthornton/DHS\">https://github.com/jamespeterthornton/DHS</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 260002,
          "author_name": "",
          "author_url": "",
          "post_date": "12/19/2017 11:46:36",
          "content": "<p>Hello, Thanks for sharing. I don't have enough hardware to retrain the ResNet50, so can you share the re-trained ResNet50 model weights? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 260123,
      "author_name": "mmiron",
      "author_url": "",
      "post_date": "12/19/2017 16:19:56",
      "content": "<p>I used both the .aps and .a3daps files for a public leaderboard score of 0.09559, and a private score of 0.2 -- I tried retraining with the additional 100 images that were used as the test set for stage 1 (I had some complications that demanded my attention, so I didn't get a chance to do that before stage 1 ended), but only hit 0.189 as a late submission with them.</p>\n\n<p>Like JamesThornton, I segmented the input images into \"zone plates.\"  I ended up discarding most of the input data though, which in hindsight was a major mistake I think; for each passenger/subject, I cut the image off at the top of their hands and used proportional arithmetic to slice their body up in to zones (variable heights and all.)  I used three frames for each zone: front facing pose, back facing pose, and one side facing pose.  I also mirrored the back facing pose so that the same arithmetic could be used (the right arm, for example, was in the same place of the input image for both the front and back poses that way).  I also randomly flipped the images horizontally during training as part of the transformations.</p>\n\n<p>The final input to my net was two 240x180 images (one for each file type) composed of three 80x180 plates stacked vertically.  At the time that seemed like a good idea, because it captured enough of a view of every zone to get a good glimpse of a threat if it was present; I'm thinking now that the disadvantage to it is that as the input data passed through the net, it may have made things too sensitive to possible threats -- I'm just pondering the behavior of the final softmax layer and suspecting that it may have been a better idea to include more \"threat free\" portions of the input images to pull the output of the softmax toward the most likely option (most zones had no threats.)  I may be entirely wrong, who knows.</p>\n\n<p>I used the convolutions from two Xception nets, each devoted to one file type, and concatenated their pooled output before applying a final softmax layer for a binary classification.  I didn't clip the output, I just let the architecture say there was a 100% chance of a threat being present if it wanted to.  I don't know if that hurt my score or not (probably did by quite a bit, lol.)</p>\n\n<p>I wish I had enough experience to provide more lessons than I can, but I'm hesitant to offer advice given that I'm still learning myself.  I hope this information is useful, and I'd like to sincerely say that I've been incredibly impressed by the Kaggle community: you folks are great, and I'm privileged to have competed with you all :)</p>\n\n<p><em>Edit</em>: I just tried clipping my submission file and re-submitting it, and it didn't make much of a difference (hit 0.17 instead of 0.18).  I guess my architecture did it's job pretty well, at least.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 261052,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "12/21/2017 16:08:00",
      "content": "<p>My final submission was a little complicated, but basically it boils down to 5 variations on the strategy outlined below.  The differences were mainly in which pre-trained ImageNet model was used, what amount of augmentation was applied, size of the threat zones, and what ended up in the 3rd color channel.</p>\n\n<h2>Model Input</h2>\n\n<p>Of the various formats, only the APS was used.  Neither of the A3D formats seemed to add anything and generating an image from the AHI is still a mystery to me.</p>\n\n<p>With only ~1,200 scans and 23ish subjects, augmentation seemed critical for a model to generalize to new subjects. For me this involved rotations, translations, contrast and brightness adjustment, horizontal reflection, and gaussian blur.  In addition, ~800 more scans were generated from the training scans by using GIMP.  Mostly this involved transforming and transplanting a threat object from one location to another or from one scan to another.</p>\n\n<p>The scans are all monochrome whereas the pre-trained ImageNet models accept three channel color input.  Triplicating the the monochrome scan for each channel was one option, but it seemed like a waste.  To make use of the multiple scans per subject, I took the average and standard deviation per pixel of ten similar scans of the subject and placed them in the other two color channels.  This had the aesthetically pleasing effect of making the threats stand out in a nice red color.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/261052/8101/model_input.png\" alt=\"model_input\" title=\"\">\nOf course the solution needed to be fully automated, so selecting similar scans needed to automated as well.  The easy way is to take the average absolute difference per pixel between each of the ~1,200 scans.  Whichever ten scans have the smallest distance could be used.  I found that a VGG-style deep autoencoder worked a little better, but the idea is essentially the same. </p>\n\n<h2>Model Design</h2>\n\n<p>Initially I had started with a 17-target MVCNN similar to what Moejoe (Shayan) described in this thread, but eventually I found that splitting the image up into four overlapping zones worked a little better.  Potentially the improvement was due to using a higher resolution on each of the smaller threat zones or possibly due to eliminating irrelevant information from other parts of the scan.  In any case this split the model into four components roughly aimed at: Arms (1-4), Chest (5-7,17), Waist (8-12), and Feet (13-16).\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/261052/8102/model_zones.png\" alt=\"model_zones\" title=\"\">\nEach of the four component models had a similar design.  They took 8 of 16 equally spaced images from the whole APS scan and fed them into a pre-trained ImageNet model, resnet-50 for instance, whose weights were shared over the different inputs.  Given that we are asking the model to not just predict the threat's existence but also the location, we want to preserve spatial information so in contrast to MVCNN, which takes the maximum of all the outputs, here we just concatenate the outputs together.  This leads to a very large number of channels which need to be reduced using a 1x1 convolution.  After that we have a very standard top design with a few dense layers.  Some small amount of dropout was needed in the last few layer to prevent the model from overfitting.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/261052/8103/model_diagram2.jpg\" alt=\"model_diagram2\" title=\"\"></p>\n\n<h2>Prediction/Calibration</h2>\n\n<p>Test time augmentation was helpful here.  All the augmentation applied to the training images was applied here at three checkpoints for each of the models.  It was probably overkill, but predictions were made for each zone model on each scan around 100 times.  All the various predictions were just averaged.</p>\n\n<p>Calibrating predictions to generalize well to new subjects was also helpful.  For this a 5-fold cross-validation was performed, with scans from any particular subject all in the same fold.  Using the out-of-sample scores from cross-validation and the original labels before any corrections, a LightGBM model was built.  This LightGBM model was then applied to the stage-2 predictions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 261096,
          "author_name": "stefanwild",
          "author_url": "",
          "post_date": "12/21/2017 19:21:51",
          "content": "<p>Thank you for the great description and vizualization. It's amazing to see how the way you analyzed and augmented the data made all the difference, here. I had multiple models that were very similar to what you were doing, but never got past 0.25 LB (stage1 data). Congratulations on the first place!</p>\n\n<p>Did you ever try or consider using the .a3daps scans to augment? I tried using 8 equally spaced angles and then randomly shifting -2 to +2 steps (same for all 8 angles), but didn't see much of an effect on the model's ability to generalize.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261116,
          "author_name": "godaibo",
          "author_url": "",
          "post_date": "12/21/2017 20:21:02",
          "content": "<p>100 predictions!!! with Test time augmentation? WOW! we did test time augmentation too but only averaged 16 predictions per subject, and we thought that was overkill !!...  go big or go home I guess...we tried a similar approach of splitting the volume into 4 chunks for 3d convolutions at the original dimensions.  That model did not make it into our final submission because we could not reliably guarantee the zones will fall in the right chunk because the subjects had varying height. We then attempted to do registration on all the samples to make them all the same height, size e.t.c with a 12 parameter affine transformation but we were worried about the computational time and size of the state 2 data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261133,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/21/2017 21:21:44",
          "content": "<p>Thanks for your write-up and congratulations on the win!</p>\n\n<blockquote>\n  <p>~800 more scans were generated from the training scans by using GIMP</p>\n</blockquote>\n\n<p>I had to Google it, thinking GIMP is some interesting algorithm. You manually drew / augmented 800 images??!! whee,...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261140,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "12/21/2017 21:53:45",
          "content": "<blockquote>\n  <p>You manually drew / augmented 800 images??!!</p>\n</blockquote>\n\n<p>800 <em>scans</em>. It must have been ~5000 photoshoppings. Congrats, idle_speculation!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261163,
          "author_name": "lwalzer",
          "author_url": "",
          "post_date": "12/21/2017 22:52:59",
          "content": "<p>thanks for the great writeup and congrats. </p>\n\n<p>-- how many training steps did you run for each region (component) and what computing environment did you use?</p>\n\n<p>Should  I infer one key to Idle_Speculations's success was comparing each scan to the average (and std dev) of all scans for the subect?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261191,
          "author_name": "speculation",
          "author_url": "",
          "post_date": "12/22/2017 00:54:09",
          "content": "<p>@Stefan</p>\n\n<p>I suspect there was something amiss in your setup if you got stuck at 0.25 on the public LB.  To answer your question, I did try using the A3DAPS in variety of ways:  Putting it another color channel, Randomly replacing an APS image with a comparable A3DAPS, Building A3DAPS specific models.  None of it really seemed to work. </p>\n\n<p>@DavidGbodiOdaibo</p>\n\n<p>I based my zones off the training set and tried to build in a little room for error.  Even so, a couple of the gentlemen in stage-2 were taller than I had anticipated.  I'm pretty sure I lost a few points in zones 5 and 17 for those guys.</p>\n\n<p>@Bastiaan/Oleg</p>\n\n<p>Yep, it took a while.  Probably there is a good way to automatically extract the threat part of an image and paste it into other scans, I wasn't able to figure it out though.  Another mind-numbing exercise I subjected myself to was drawing bounding boxes around all the threats in zones 1-4 while I explored Faster-RCNN.</p>\n\n<p>@numericLee</p>\n\n<p>Each of the zone models converged in around 16k iterations with batch size 8.  It worked out to be around 4 hours per zone on a 1080TI.  I also had some free credits for google cloud which I used to rent out Teslas for heavy lifting.  In terms of using average of similar scans, it seemed to decrease logloss by around 0.01 in cross-validation.  The benefits were perhaps larger on the test set given Mr. Dreadlocks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261248,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "12/22/2017 04:13:32",
          "content": "<p>@idle_speculation</p>\n\n<blockquote>\n  <p>Each of the zone models converged in around 16k iterations with batch size 8.</p>\n</blockquote>\n\n<p>Did they converge in the sense that the training loss stopped changing, or did the validation loss start going up (and you used early stopping)? What were the learning rate / momentum like?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261365,
          "author_name": "speculation",
          "author_url": "",
          "post_date": "12/22/2017 12:24:11",
          "content": "<blockquote>\n  <p>Did they converge in the sense that the training loss stopped changing, or did the validation loss start going up (and you used early stopping)? What were the learning rate / momentum like?</p>\n</blockquote>\n\n<p>They converged in the sense that the validation loss did not improve past that point.  Performance did not get any worse with further training.  It just stopped getting better.  Training loss was very close to zero near the end.</p>\n\n<p>Adagrad was used to fit all the models.  No momentum was applied.  Learning rate was gradually increased up to 0.01 in the first 1k iterations, then left at that level for the rest of the fit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261401,
          "author_name": "godaibo",
          "author_url": "",
          "post_date": "12/22/2017 15:04:20",
          "content": "<p>I see a potential weakness with this approach in a production environment, I am not sure if I am interpreting this correctly and will like to hear your thoughts. It seems to me that the RGB image exploits the fact that there are additional scans of the same subjects under different conditions in the training and test set and uses these additional scans to derive the “mean” and “standard deviation” green and blue channel, however, in a production/live environment the system will be seeing a subject for the first time and there will be no prior scan to derive the additional channels in the RGB image, how would you handle this? Is there a performance difference when the additional channels are not included?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261412,
          "author_name": "mmiron",
          "author_url": "",
          "post_date": "12/22/2017 15:47:16",
          "content": "<p>I would ask this privately, but I can't contact users without being a \"contributor.\"  Please pardon my occasionally awkward social skills...</p>\n\n<p>I would have expected your solution to be disqualified, because it's impossible for the DHS to verify it without the additional training data.  Did they contact you to ask for that?  Please understand that I'm not trying to imply there's a reason to be suspicious, and given your standing, I'm sure your word is good enough for Kaggle (if not the sponsor as well) -- but I'm still curious how verification was handled?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261419,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "12/22/2017 16:11:11",
          "content": "<p>Murray,</p>\n\n<p>All winning solutions are verified by the sponsor and must be able to reproduce the winning leaderboard score per the rules. As the competition has closed, the model upload and verification process is underway.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261421,
          "author_name": "mmiron",
          "author_url": "",
          "post_date": "12/22/2017 16:26:11",
          "content": "<p>I'll just be perfectly frank: it doesn't bother me in the slightest, personally, and I have nothing but respect for someone who got the job done better than other teams (other teams includes me, too, lol); but in my mind, the manual augmentation of training data falls under a non-automated external source that wasn't made available to the other contestants.  But it doesn't really matter what I think, and I'm not trying to cause trouble.</p>\n\n<p>My apologies if the question was awkward.  Truthfully I'm most interested in where the boundaries of the rules are, so that I can do better myself in the future without accidentally crossing the line.  Obviously a tremendous amount of work went into idle_speculation's submission, and I think he deserves that first place slot.  Congratulations again :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261540,
          "author_name": "speculation",
          "author_url": "",
          "post_date": "12/22/2017 22:03:47",
          "content": "<blockquote>\n  <p><strong>Murray Miron wrote</strong></p>\n  \n  <p>I would have expected your solution to be disqualified, because it's impossible for the DHS to verify it without the additional training data.  Did they contact you to ask for that?  Please understand that I'm not trying to imply there's a reason to be suspicious, and given your standing, I'm sure your word is good enough for Kaggle (if not the sponsor as well) -- but I'm still curious how verification was handled?</p>\n</blockquote>\n\n<p>Luckily I had the forethought to include the additional scans in my model upload. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261546,
          "author_name": "speculation",
          "author_url": "",
          "post_date": "12/22/2017 22:26:14",
          "content": "<blockquote>\n  <p><strong>DavidGbodiOdaibo wrote</strong></p>\n  \n  <p>I see a potential weakness with this approach in a production environment, I am not sure if I am interpreting this correctly and will like to hear your thoughts. It seems to me that the RGB image exploits the fact that there are additional scans of the same subjects under different conditions in the training and test set and uses these additional scans to derive the “mean” and “standard deviation” green and blue channel, however, in a production/live environment the system will be seeing a subject for the first time and there will be no prior scan to derive the additional channels in the RGB image, how would you handle this? Is there a performance difference when the additional channels are not included?</p>\n</blockquote>\n\n<p>I don't consider myself a person who travels that much.  But personally, I took three flights in the last year.  That works out to six trips though the scanners this year alone.  If you add in scans from last year or two then the approach of using ten similar scans starts sounding a lot more plausible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261583,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "12/23/2017 03:51:19",
          "content": "<p>I don't think the TSA is legally permitted to store the scans, actually.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265258,
          "author_name": "byen883",
          "author_url": "",
          "post_date": "01/05/2018 01:46:20",
          "content": "<p>@idle_speculation</p>\n\n<p>Congrats, I'm also curious since I haven't seen anybody mention about the recent published paper by Hinton in regards to Dynamic routing between capsules, even though it's relatively new but what are your thoughts on this capsule net over convolutional net? </p>\n\n<p><a href=\"https://arxiv.org/pdf/1710.09829.pdf\">https://arxiv.org/pdf/1710.09829.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 261169,
      "author_name": "kaggle446",
      "author_url": "",
      "post_date": "12/21/2017 22:59:41",
      "content": "<p>My winning solution was an ensemble of 11 models. All models are end-to-end, i.e., they take as input 16 images per subject and spit out 17 probabilities per threats. I did not use any pretrained model or weights. All models are hand-crafted CNN based models. All models were trained on only one Titan X gpu.</p>\n\n<p><strong>Data pre-processing</strong>\n- I only used APS format.\n- For most models, we down-sampled the dataset to 256  by 330 \n- We also rotate images 90, 180 and 270 degrees and train separate models on the rotated dataset.\n- For couple of models we used the original 512 by 660 images and the rotated dataset.\n- Global zero-mean and unit-variance normalization were used. </p>\n\n<p><strong>Data Augmentation</strong>\nWe perform different image augmentations including: rotation (10-15 degrees), width and height shift (10-15%), shear, zoom (10-15 %) as well as random erasing [Ref1]. I tried to use brightness augmentation but the models were completely confused with any changes in the brightness so I dropped it from augmentation.</p>\n\n<p><strong>Architecture</strong>\nModel architecture as seen in the figure is basically a conv2D to extract features from 16 image slices one at a time, flattening features, average over 16 image slices and then generating 17 Sigmoid outputs (one per zone). I did not know about the MVCNN paper but it seems there is similarity with that method. </p>\n\n<p><strong>Insights</strong>\nRandom erasing augmentation provided some improvements. Building separate models on 90,180,270 rotated dataset was also very effective. It is also better to use the original dataset as opposed to the down-sampled data. Although, it is more time consuming to train.  </p>\n\n<p><strong>Ref1</strong>\nZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, Yi Yang “Random Erasing Data Augmentation”</p>",
      "votes": null,
      "replies": [
        {
          "id": 457123,
          "author_name": "yaoqinchuan",
          "author_url": "",
          "post_date": "01/17/2019 01:28:35",
          "content": "<p>hello,I am working on this problem using your tips but changed  the net architecture,so I want to compare our result.can you share your code?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 261235,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "12/22/2017 03:40:18",
      "content": "<p>I only ended up using the APS images. </p>\n\n<h1>Bounding Box</h1>\n\n<p>I spent a good amount of time annotating the zones in the front view of the images. I then used these annotations to train a bounding box model to predict where each zone was. I didn't want to take the time to do this at first, but the variance in the height/width of the subjects and the fact that some subjects held their arms at different angles led me to think it would be beneficial. I also annotated the locations of a subject's ankles, knees, elbows, shoulders, wrists on the front and side images. The reasons for this is that I wanted to align people's limbs so they they were all as vertical as possible. </p>\n\n<h1>Subject Identification</h1>\n\n<p>I realized early on that there were a limited number of subjects and that the subjects in the Phase 1 test set all appeared in the training set. I knew that this would cause overfitting, and sure enough I noticed a massive difference in models trained using proper validation and those that did not. Without proper validation I was getting &lt; 0.04 on the Public LB. Using proper validation and the same architected I was getting ~0.17 on the Public LB. I ended up with 24 different subjects. </p>\n\n<h1>Models</h1>\n\n<p>For inputs, I concatenated the APS angles into a single 13 images x 1 image (I excluded 3 angles where the zones couldn't be seen) array. I trained 2 different models: one where each image was reduced to 96x96 pixels, and another that was 128x128 pixels i.e. the final images were 1248x96 and 1664x128. As mentioned in the Bounding Box section I rotated the limbs so that they were vertical in each image to account for the fact that people had their arms/legs as different angles. My architecture ended up being the following:</p>\n\n<p>Conv2D(16, 3x3) <br>\nConv2D(16, 3x3) <br>\nMaxPooling2D(2x2)  </p>\n\n<p>Conv2D(32, 3x3) <br>\nConv2D(32, 3x3) <br>\nMaxPooling2D(2x3)  </p>\n\n<p>Conv2D(32, 3x3) <br>\nConv2D(32, 3x3) <br>\nConv2D(32, 3x3) <br>\nMaxPooling2D(2x2)  </p>\n\n<p>Conv2D(32, 3x3) <br>\nConv2D(32, 3x3) <br>\nMaxPooling2D(2x2)  </p>\n\n<p>Conv2D(32, 3x3) <br>\nConv2D(64, 3x3) <br>\nConv2D(64, 3x3) <br>\nMaxPooling2D(2x2) <br>\nDropout(0.5)  </p>\n\n<p>Conv2D(64, 3x3) <br>\nConv2D(128, 3x3) <br>\nConv2D(128, 3x3) <br>\nMaxPooling2D(3x3) <br>\nDropout(0.5)  </p>\n\n<p>Flatten()  </p>\n\n<p>Dense(512) <br>\nDropout(0.5)  </p>\n\n<p>Dense(1)  </p>\n\n<p>Anyhow, all I have time for now. I'll hopefully add a bit more detail soon.\nThis ended up being much deeper than my original model trained without using proper cross-validation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 261534,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/22/2017 21:17:13",
          "content": "<p>Thanks for the writeup!</p>\n\n<blockquote>\n  <p>I rotated the limbs so that they were vertical in each image</p>\n</blockquote>\n\n<p>How did you do that? Did you rotate the whole picture to get everything straight on average? Or did you cut up the people in separate limbs? Or did you stretch and skew the images? Or more advanced?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261545,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "12/22/2017 22:22:43",
          "content": "<p>Using the right forearm as an example, I had a model which identified the point in the front and side images where the wrist and elbow were.  I then used my bounding box model to identify the where the forearm was and cropped out the rest of the image.  Then using the point locations of the wrist and elbow and using a bit of trigonometry I calculated how much I needed to rotate each image so that the forearm was vertical. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 261490,
      "author_name": "teedrz",
      "author_url": "",
      "post_date": "12/22/2017 19:37:53",
      "content": "<p>I decided to approach this competition from the viewpoint of object detection.  I annotated a subset of the training images by drawing bounding boxes around the threats in the aps images (I spent maybe 8 hrs on this).  I then trained Faster-RCNN using the annotated images, with the label of each object being its location on the body.  This way the network would learn to both detect objects and label them based on their location all in one go.</p>\n\n<p>For each aps image there are 16 viewpoints.  I ran the object detector over each viewpoint in the aps image, to generate a set of candidate object detections, with each detection having a probability over each body location.  I took the maximum of all the detections for each body location to get a final 17 dimensional confidence vector.  I then took the 17d confidence vector for each viewpoint to get a 16x17 matrix.  I used this to train an 17 gradient boosted classifiers (one for each body location) to get better calibrated probabilities.</p>\n\n<p>The only data augmentation I performed was doing a horizontal flip of the images while training the object detector.  This actually will change the label.  For example, if a threat is labeled \"right ankle\" then after flipping it will then be labeled \"left ankle\". </p>\n\n<p>I decided on using an object detector due to the limited size of the dataset and the small number of unique individuals.  Faster-RCNN is a two stage object detector.  In the first stage, it recognizes potential threats; the second stage labels the threats/decides if they are not true threats.  I thought that this built in attention mechanism would help to prevent overfitting by forcing the network to focus on the threats.</p>\n\n<p>My final model consisted of an ensemble of 5 models.  Each model using 80% of the data to train the object detector and then the final 20% to train the boosted classifier.  I then averaged the predictions over each individual model.  However, ensembling did not really seem help the performance of my model in any significant way.</p>\n\n<p>Now some things that I tried that didn't work.  I spent some time trying to use the 3d data.  I used a 3d inflated convolutional network based on the vgg16 architecture.  This just means that I turned each 2d convolution into a 3d convolution and initialized the weights of each 3d convolution by stacking the weights of the pretrained vgg model.  I then tried to use this network to perform 3d object detection.   This best I was able to do was around 0.06 on my validation set as opposed to 0.025 using my 2d model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 261531,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "12/22/2017 21:02:29",
          "content": "<p>Thanks, deedrz, and congrats! Were those networks pre-trained on ImageNet?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261533,
          "author_name": "bastiaanbergman",
          "author_url": "",
          "post_date": "12/22/2017 21:12:30",
          "content": "<p>Thanks for the writeup and congrats on the win!</p>\n\n<p>Did you consider ensembling the two models: the 3D conv and the 2D conv? I'd think the two models are fairly orthogonal?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261539,
          "author_name": "teedrz",
          "author_url": "",
          "post_date": "12/22/2017 21:44:15",
          "content": "<p>Yep, the networks were pre-trained on ImageNet, I just replicated each image to get the 3 channel input. And Bastiaan, I think your right, combining the 2d and 3d models would've been a good thing to try</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 261798,
      "author_name": "suchir",
      "author_url": "",
      "post_date": "12/24/2017 01:54:59",
      "content": "<p>I posted about this in more detail <a href=\"http://suchir.io/posts/passenger-screening-algorithm-challenge-writeup.html\">here</a>, but the TL;DR is:</p>\n\n<ul>\n<li>trained a model to compute segmentations of the threats given aps/a3daps images (hand-labeled the training data)</li>\n<li>trained a model to compute segmentations of the zones given depth maps (generated synthetic data with depth maps / ground truth to train, used a3d to get real depth maps)</li>\n<li>trained a model that takes in the threat / zone segmentations from 16 angles and outputs the final predictions for each zone</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 263940,
      "author_name": "nathanrm",
      "author_url": "",
      "post_date": "01/01/2018 18:21:06",
      "content": "<p>The description of my cylindrical coordinates based approach is <a href=\"https://www.kaggle.com/nathanrm/full-solution-cylindrical-coordinate-method\">posted in Kernels</a>, and includes a link to the complete solution on github.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 271860,
      "author_name": "shivgowda",
      "author_url": "",
      "post_date": "01/21/2018 17:26:10",
      "content": "<p>I participated in this competition for completing the Udacity Nanodegree capstone project requirement, you can find the full report <a href=\"https://github.com/shivarajugowda/kaggle/blob/master/TSA/report.md\">here.</a>. In summary, my solution involved these two steps. </p>\n\n<ol>\n<li>Detect head, hands, groin and threat bounding boxes using TensorFlow Object Detection API</li>\n<li>Predict Body Zone of a threat based on the bounding boxes of a threat relative to the frame number, head, hands and groin. I attempted this in two ways: A supervised learning method using XGBOOST. A custom algorithm which needed only a few ratios to be tweaked.</li>\n</ol>\n\n<p>Overall it was good learning experience, it did however take a lot of time. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 312252,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "04/11/2018 13:15:44",
      "content": "<p>I wonder how much time the other top-8 teams' models need for training and prediction, respectively?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "258374": "Now that the competition is over let's have a place to discuss our architecture, findings, and hindsight?\n\nEdit: Here's my 10th place solution: https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution",
    "258391": "Edit: https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\nHere's my code.\n\nIn my first attempts I tried to use MVCNN on a3daps files and downsampled inputs to a 3rd of the original resolution. This was getting about 0.10 on the original test data. I never tried 3D convolutions as I thought it'd be too slow and memory hungry.\n\nI found I got better results by using a3d files and maintaining the original resolution (and lowering the batch size considerably so it would fit in memory). I tried combining a3d and a3daps by throwing them in separate channels but I found it caused overfitting.\n\nLater on I dropped MVCNN and went with a custom architecture where each view is fed through a pretrained ResNet-50 CNN and each feature map is pyramid pooled (to find objects/features of differing size) and fed to a LSTM with attention. My pyramid pooling doesn't actually use pooling layers but instead passes the last feature map through a 1x1, 3x3, and 5x5 convolution with stride 1, 2, and 3 respectively. I use a form of attention for CNNs that I apply to the feature map before the pyramid pooling. To make it run on my 1080 ti I increased the stride of the first couple layers from 2 to 3, which probably has a similar affect as decreasing the resolution but loses less information. I trained with SGD with a cosine annealing schedule. This was getting around 0.015 on the original test data.\n\nI didn't use any segmentation and had my model predict on all 17 zones at once. In hindsight maybe I should have segmented as it could allow me to train with higher resolution and focus attention on relevant zones at the cost of speed, but I was also worried that the people in the stage 2 data would be drastically different and the segmentations would be wrong. I also probably shouldn't have done experimental things like CNN attention and cosine annealing but then again maybe the risk paid off and they helped.\n\nRegrettably I didn't use ensembling. Just didn't have the time to train more models or think it would be too important. But looks like it could have made a difference.\n\nBut in the end I'm still very satisfied with my score, and congratulations to everyone!",
    "258408": "Moejoe,\nThanks for sharing your experience and thoughts.  I wonder how did you decide which zone the threats locate If you did not use any segmentation. Thanks.",
    "258419": "Hi Wensu, the CNN learns on its own which locations correspond to which labels without any human guidance. It outputs a size 17 vector and is trained against size 17 targets. Since the classes aren't mutually exclusive I don't use softmax and I train with binary cross entropy loss.\n\nThe average pooling layer at the end of ResNet loses all spatial information so one would imagine this would be a problem. Instead I use a pyramid pooling scheme which pools with multiple kernel sizes and feed the output (which maintains information about location) directly to the LSTM. However, although pyramid pooling helped noticeably, I found that it wasn't absolutely necessary and even with the regular average pooling layer it still gave decent results (at least on the stage 1 data).",
    "258430": "Congratulations to everyone who participated in this challenging and suspenseful competition! \n\nI used a multi-view CNN custom ResNet that looks at the entire image (without any pre-segmentation) and outputs 17 probabilities. I used the .aps dataset downsampled to half-resolution. \n\nMy stage1 score (around 0.025) was much better than my stage2 score (0.13589). This must be because the subjects in the stage1 training set were the same as those in the stage1 test set, but the subjects in the stage2 set were completely different people. \n\nDuring the optimization process, I used regular cross-validation. After reading comments from others, it has become clear to me that the correct way to do cross validation was to first group all of the instances depicting the same person, and then randomly assign each group of instances to a cross validation set. \n\nI had suspected that this would be an issue, but could not have imagined that it would make such a huge difference between stage1 and stage2 performance. I think there are two possible mechanisms that could explain this:\n\n1. Since my net does not use pre-segmentation, it probably determines the threat zones by identifying landmark features. For example, it might identify a threat between the hand and elbow, and therefore determine that the threat lies in the forearm region. However, different people have different-looking hands and elbows. This may have prevented my net from properly identifying landmarks in stage2. \n\n2. It is also possible that the net learns to whitelist certain objects that would otherwise be identified as threats. The Bob Marley lookalike from stage2 is a perfect example. If the stage1 training set had included someone with this type of hair, the net would most likely have learned to specifically avoid labeling dreadlocks as a threat.",
    "258445": "Initially, I started out by using a combination of both the aps and a3daps files. Rather than use all 16 and 64 \"views\" from each body scan format, I instead selected 8 prominent views i.e. front-center, front-left, etc. from both files and then further segmented and cropped out all 17 zones into equal 100x100 square images. I then concatenated these into a 400x400 \"mosaic\" for each passenger/zone combo. This was fed to my model that generated a binary classification. Using this strategy I was able to achieve .03 logloss on stage 1 test set. Like others have echoed on the forum threads, I too found that if I left out the aps images I would end up with lower accuracy yet I did observe that I got better accuracy when having a3daps images included in the mix rather than just using aps by itself since not all threats were visible in all cases. \n\nWith around 2 weeks left in the competition I tried using the a3d images to see if I could improve my score further. I found that I couldn't get good results using a 3d convnet, so instead I tried a different strategy which was to treat all image slices as separate images to feed to the model in a sequence. So for each passenger/subject there was a total of 310 slices (step-size of 2 pixels and clipping everything over 620 which was mostly noise). I also segmented by zone. With this strategy I was surprised that I was able to get as low as .005 logloss on stage 1 test set (this also corresponded with my local validation loss). This was using just a single model (DenseNet169) and one 15% validation hold out. Not only was it much easier for my model to locate and classify the threats, using this strategy also really helped me with segmentation because I could also see patterns in the data like exactly how long a threat was and which image slices were most confident etc. This high score in stage 1 gave me enough confidence that I should abandon my first strategy and only use a3d images for stage 2. In hindsight, I can now see this was not such a good idea as my score ended up dropping to 0.17 on stage 2!! After reviewing my predictions, I can see that some new features of the passengers introduced in stage 2 (that didn't exist in stage 1 training set or test set) caused my model to generate a large number of false positives in a few isolated zones. Had my model at least seen examples of something similar I think it probably would have scored much higher as it would have learned to whitelist some of those new features in stage 2 (i.e. dreadlocks, suspender clips, etc.) that looked very similar to threats in stage 1 training set. \n\nWhile I am quite disappointed in my final result, but overall I think this was a really interesting challenge and was a great learning experience for me!! Hopefully someone here can also learn from my mistakes!",
    "258447": "My CV was 0.032-ish (No subject was in multiple CV folds, of course). \n\nThe big surprise in Stage2 was the guy with massive dreadlocks, like no one had in the training dataset. They looked a lot like some of the bombs in the training  scans instead, except for the fact that they adjoined his head. My model was understandably suspicious. I had calculated that that would cost me 0.01 on the score, and it seems like it did.\n\nIf I could have seen the Stage2 data for just 10 minutes, before finalizing the model, I would have reduced the confidence of zone 6 &amp; 7 predictions, and negated most of the 0.01 damage, but the predictions are supposed to be fully automatic (or such was my understanding).",
    "258451": "Congrats Oleg on placing 5th in this comp, thats very impressive!\n\nGlad you mentioned about this guy and that I'm not the only one who had issues with it. The dreadlocks were the main culprit for pretty much all of my false positives in zones 17, 6 and 7. My model also had problems with the guy with the suspenders near the metal clips (zones 8 &amp; 10). Like you say, if this wasn't a two stage comp I would have made some adjustments to my model's sensitivity/threat thresholds to mitigate this but since I didn't anticipate needing to do this in the model upload so I couldn't make those changes. \n\n&gt; **Oleg Trott wrote**\n&gt; \n&gt; &gt; My CV was 0.032-ish (No subject was in multiple CV folds, of course). \n&gt; \n&gt; The big surprise in Stage2 was the guy with massive dreadlocks, like no one had in the training dataset. They looked a lot like some of the bombs in the training  scans instead, except for the fact that they adjoined his head. My model was understandably suspicious. I had calculated that that would cost me 0.01 on the score, and it seems like it did.\n&gt; \n&gt; If I could have seen the Stage2 data for just 10 minutes, before finalizing the model, I would have reduced the confidence of zone 6 &amp; 7 predictions, and negated most of the 0.01 damage, but the predictions are supposed to be fully automatic (or such was my understanding).\n&gt; \n&gt; \n&gt; \n&gt;",
    "258581": "Hi Moejoe, Thanks for your response.  Since I spend more than a half time on segmentation, I'd like to learn how others did for this issue.  you said your model learns on its own correspond to which labels, it is great and is what I like to do, but did not figure out how.  you mentioned your model is trained against 17 targets, how did you get the 17 targets without segmentation,  could you explain further how your model did?  Thanks a lot.",
    "258613": "oh boy, I see a future where we data scientist are to blame for people with dreadlocks going through \"special screening\" at the airport :)",
    "258618": "we used exclusively A3DAPS (generally 7 of the 64 views), training with a 7 layer model on individual views and then combining the  views using a separate model.\n\nEnsemble: our logistic regression also incorporated some home-brewed algorithms (using morphology and transforms)  which added predictivity in some cases over the ML results.\n\nVGG16 was hard to use on my desktop, it crashed often unless we throttled back the batch size significantly\nwe had som e success training multiple zones together using left right symmetry\n\nfeeding all views at once also crashed my desktop (16GB of RAM), I would appreciate feedback how people made that work.\n\nI am surprised that some contestants had success with 100 x100 crops.  we generaly needed to go 225-300 wide depending on the zone and view\n\nOur model seemed to work very well on 11 zones and poorly on 6 zones: arms(4), groin, and upperchest. The arm regions were noisy. I would be interested in strategies for those zones.  We tried to rotate all arms to vertical but saw no improvement. Did anyone use gender recognition to facilitate groin and upperchest? \n\nDid anyone use cross-zone models? In the training set there was a slight inverse correlation among zones  (ie a passenger with contraband in 2 other zones was less likely to have it in a third zone)",
    "258673": "I started out with APS images first.\nI used various CNNs (VGG, Resnet, Densenet, etc) that have been pretrained with ImageNet weights as featurizers by stripping out their classification layers.\nFor each scan, forward passes were made on the CNN to create feature maps for all 16 views, and those feature maps were flattened and concatenated into a single vector to be fed into LightGBM for classification.\nAs LightGBM does not support multi-label output, each of the 17 zones were trained/eval'd separately.\nThis got me to ~0.19 on the public LB when I used Densenet 121 as the featurizer.  I did not expect this to work too well as ImageNet and TSA images are very different.  Also, some threats were very difficult to see in APS due to occlusion and how projections seem to be computed.\n\nThen I started looking at A3D and thought to myself that segmenting out into various zone groups would help with classification, because:\n\n- Cassifiers that solely focus on its assigned specific zone group can be created, rather than having the network figure out what to focus on\n- More samples can be fed because mirroring can be performed independently on the body part in question\n- By eliminating other parts, threats can be seen more easily (you get a clear shot at the zone(s) in question.)\n\nI ended up with 6 classifiers for the following zone groups: arms, legs, chest-back, torso, waist, and crotch.\nThe \"arms\" classifier will get trained on right arm images, left arm images, and their horizontal mirrors to output 2 labels (i.e., whether a threat exists in the upper arm and/or lower arm.)\nThe \"legs\" classifier will be trained in a similar fashion to output 3 labels for detecting threats on the thigh, knee, and foot zones.  Others were trained to output a single label (i.e., binary classifiers.)\n\nI experimented by training Conv3D-based classifiers on the segmented zone groups, but this did not work so well, and I ended up creating a variant of the Multi-View CNN (https://arxiv.org/pdf/1505.00880.pdf) except I used a learnable convolution layer instead of a view pooling layer.  Train-time augmentation was limited to horizontal flips, and test-time augmentation was limited to 2 combinations of horizontal flips and averaging out the predictions.  No trimming of predictions were performed (boy, the models gave very confident predictions so perhaps I should have looked into this a little more to minimize the penalty on mistakes.)\nThis approach got me to ~0.019 on the public LB, and ~0.089 on the private LB.\n\nMy final approach was preprocessing heavy due to all segmentation done based on the center of mass and assumed\nbody proportions, then rotating the zone groups in 3D to create 2D projections (I actually ended up implementing multiple methods of 2D projections, as simply doing \"max\" on the collapsed axis would sometimes hide certain threats that appear on one side only.)  Also, training was very compute intensive since there were 6 separate networks to cover all the zone groups.  And I actually trained multiple networks for each zone group for ensembling, so a total of 38 networks were trained when it was all said and done.\n\nLooking at other competitors posts, I feel a bit shocked and a little embarrassed to find that some did better with a much simpler and elegant pipeline of predicting all 17 labels simultaneously without any segmentation or even ensembling!  I didn't even try that because I had strong convictions in my intuitions (that I unfortunately did not bother to validate.)\n\nI learned a lot in this competition!\n\n- Explore wider before settling down on a method (I feel that I turned down my \"learning rate\" or \"temperature\" too quickly and did not get good coverage on the solution space.)\n- Before building a complex pipeline, try out and validate simpler solutions first.\n- Be very thoughtful in setting up local CV.  I did not other to separate out subjects when creating folds, unlike other wise competitors did.",
    "258677": "Separating subjects when creating training/validation folds could make huge difference on the stage 2 results. I didn’t do it as well :(",
    "258684": "The targets come from the training data. Consider the following sample from the training data:\n\n    0050492f92e22eed3474ae3a6fc907fa_Zone1,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone10,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone11,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone12,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone13,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone14,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone15,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone16,1\n    0050492f92e22eed3474ae3a6fc907fa_Zone17,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone2,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone3,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone4,1\n    0050492f92e22eed3474ae3a6fc907fa_Zone5,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone6,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone7,0\n    0050492f92e22eed3474ae3a6fc907fa_Zone8,1\n    0050492f92e22eed3474ae3a6fc907fa_Zone9,0\n\nThis is treated as an input target pair where 0050492f92e22eed3474ae3a6fc907fa.aps is the input and [0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0] is the target and what I train the model to output. The model is fed no other information.\n\nIt may feel like black magic at first but it works and it works pretty well. The architecture of the model is designed so that it's more spatially aware than regular convnets to make it easier to learn.",
    "258716": "At first, I went through aps images and I felt that many threats seem to be hard to find, so I decided to go with a3d.\nI used APS to identify each person and make train/val set - I manually annotated a few persons at first and let NN to classify the person. But later I found that PCA can do same thing easily from discussion board :(\n\nWhen I tried to feed a3d data, the data size was too big - so I scaled down to 1/4 on each dimension and cut down mostly empty areas on front/back/side/top). I could see the threats even after 1/4 scale, so I assumed that 1/4 dowscaling doesn't reduce the accuracy much. And I adjusted vertical scale little bit for each image so that the head sits on similar position(no ML, just simple image processing). Even after scaling down, only 2-3 batches fit into 1080Ti.\n\nNext, I fed the data into resnet-like 3D convolution. All data is fed to 3 blocks of identify block. Then output is splitted into 17 zones(with some overlap) and connected to per-zone 3D CNN(two blocks of identify block). It is quite simple network with not much parameters, so it just takes 10 minutes per epoch and the loss stablized around 80-130 epochs. Interesting thing is that I tried to share same per-zone network for symmetric zones(like left arm and right arm), but it came out worse. I think keeping separate network for two sides made some ensemble effects. \n\nI used basic augmentation like rotation, zoom, shift, flip left-right.\n\nI tried to separate each body part and put each part into the network, but it did much worse. I guess it's because NN losses context of body and have trouble finding whether the thread is the part of body or not. \n\nI didn't try much other experiments due to lack of time/resources. In hindsight, I wish I put little bit more time.",
    "258721": "Congratulations to winners, \nthanks to all sharing their approach!\n\nMy approach  was binary classification  with a separate set of  CNN models for each zone-view.\nClassifier models were fed with zone segments cropped from a3daps views. \nCropping coordinates were created by Unet segmentation models which I had to train. \nFor this I annotated 0,16,32 and 48-th views from 625 images.  \n\nTo  combine models for different views of same zone segment I trained Xgboost logistic regression models.\n\nI used only  0,3,29,16,32,35,48 and 61 views, not all of them in each zone.\nClassifiers were Inception V3, Densenet161 and Resnet152.\n\nI also experimented with mosaic of 2 and 3 views for some zones.  \nNo ensembling of final models nor multi-fold validations.\n\nStage 1 best score was 0.1074 with corresponding Stage2 of 0.1578.\n\nSegmentation worked pretty well on both stages. \nClassification was a real challenge.\nThe hardest zones were 2 and 4 with local validation scores of 0.156 and  0.171 respectively.\nThe easiest  zones were 5 and 17 with 0.045 and 0.012 scores.\n\nIf I have time I would like to try approach close to Moejoe's and make some late submissions.\nI very much appreciate the ideas behind this much more elegant and effective approach.",
    "258734": "... Training separate models, one for each zone is very suboptimal. It's easy to see that with the following argument: \n\nIf you have ankle-strapped guns in the training dataset, the model that looks at ankles will learn to recognize them. However, if forearm-strapped guns only occur in the test dataset, they will be gratuitously novel to the model that looks at forearms. In other words, you won't transfer knowledge between zones as well as you should.",
    "258742": "Hi Moejoe,\nThank you very much for your explanation, it sounds abstract but really a clever approach.",
    "258744": "Hi Alexander, Thanks for sharing your experience.  I also used a segmentation approach with opencv which is not an effective way but works.  I'd like learn from others to improve my approach.  You used Unet segmentation models, could you explain a  little further about this model and if you can provide  any references?  Thank you.",
    "258750": "My architecture:\nI worked with APS only, an MVCNN on all 16 views concurrently while also multi-task learning, in hopes of better generalization. I had 6 resnet-like layers on each view, then a average pooling to put all 16 views together and another 2 resnet-like layers then it split in two: 1) 4 layer resnet to train on the 17 threat labels and 2) a 4 layer resnet to train on the gender of the subject (which I manually labeled on the training data using PCA). I reduced the input images resolution with a factor 2 and used random crop and zoom to compensate for that. I also had the first two conv layers with a stride 2 to reduce the data.\n\nAn interesting trick I found useful: You can `revolve` the weights from each of the view pillars. Take the weights from the 6 resnet layers of view 0 and put them in the layers of view 1, from view 1 to 2 and so on, finally put view 16 layers into view 0. If you do this every 20 epochs you'd expect you'll have to learn pretty much from start every time, but somewhat to my surprice the network doesn't lose it's ability much. Instead it prevents overfitting as  shown by the test-cross entropy which never surpassed my training cross entropy.\n\nAnother interesting trick, I suppose most people have figured out is image mirroring for data augmentation. You have to be somewhat careful as the labels change too, e.g. a left arm threat becomes a right arm threat. Moreover the order of the APS views needs to be reversed to keep the subject as a left rotating subject. If you do that, mirror the image, reverse the views and mirror the labels you basically double the dataset. Sort of a standard trick, but with a little more to it in the case here, due to the nature of the data.\n\nI worked on a single GPU and couldn't get my batch size above 4, nor could I increase my network size, both due to memory constraints. Increasing the network size didn't seem to help much anyway. I'm not sure how much the gender training helped for model generalization. I did not test on proper test splits (proper as in separate subjects and threats for different folds), but my stage 1 score was 0.12 to 0.16 in stage 2.\n\nThings I would add next time:\n1) 5 fold training with proper separation of subjects and some boosted tree method to combine them\n2) Also separate threats, but I have no idea how\n3) Reduce image resolution more in the beginning\n4) More GPU's or different hardware to get a bigger model\n\nI don't think any of this will get me the winning model, wonder what the big difference is, what I'm missing. Hope to learn it here!",
    "258760": "I have really banged my head against walls for how other got so much better than I, with or without overfitting I don't care. Can you describe some more details how you got to 0.025? I never got past 0.11,.. See my writeup elsewhere in this thread fo rmy architecture.",
    "258761": "I didn't use segmentation either. I felt it wasn't possible to proper segment the image and sometimes the threat was more easily visible from a strange angle. For example the ankle strapped threat on the inside of the right leg may be best viewed from the left (with the left leg obstructing most of the view).\nI'm not too worried about \"landmarks\" as other have mentioned as the pixel position remains \"known\" throughout your convolutional layers until the resolution gets very small at which point it supposedly has figured out towards which label we're going.",
    "258768": "That argument only applies if what is going on is that the network learns to detect threats and not that the network learns to detect normalcy.  It is my interpretation of my own models that they were in fact doing the later and not the former.",
    "258773": "I found one paper in particular that was very helpful: \n[High-Resolution Breast Cancer Screening with Multi-View Deep Convolutional Neural Networks][1]\n\nMy net was very similar to the one described in this paper. I added a convolutional layer with 17 output channels right before global average pooling. I optimized the layer parameters, but as described above it seems that the optimization was performed under unrealistic test conditions. \n\nI trained the net using SGD + Momentum + Nesterov with lr=0.4 and momentum=0.32. I settled on these unusual values after running an extensive random search, and was very surprised that the optimal momentum turned out to be so low. I also got a lot of unusual \"spikes\" in the train and test error during training. The graph below shows the log loss on one of the cross validation folds, with exponential lr decay. I will be very grateful if someone can offer insights into this unusual behavior. \n\n![Learning Curve][2]\n\n  [1]: https://arxiv.org/pdf/1703.07047.pdf\n  [2]: http://seansoleyman.com/wp-content/uploads/eval.png",
    "258775": "I used two networks, one using about 15% of training data to classify threats, and the other had about 1% of training data to classify categories.  I used an IoU metric to match the location of the detection. So in total 2 resnet networks trained on a gtx 1070. Only data augmentation was horizontal flips. and only a3daps images.  I think the approach could yield a much better result with more of the training data used and more time to refine the model.",
    "258776": "I don't think it's an either-or type of situation. It helps to know what normal bodies look like and it helps to know what the threats look like if you aim to tell them apart. Quantitatively though, the image patches containing threats were much more scarce. Squandering them seems to me like a serious mistake.",
    "258778": "Hi Wensu,\nThis is the original paper regarding Unet:\nhttps://arxiv.org/pdf/1505.04597.pdf\n\nAs for me, I adapted Unet segmentation code from Carvana Image Masking challenge posted here:\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208#latest-222649\n\nI created very simple masks  from my annotations.",
    "258790": "Very valuable points., Oleg. Besides, segmentation models rely on annotations and their quality. Plus complexity and extra work... Much better when single net cares about zone boundaries and threats. Still I reserve my right to make serious mistakes.",
    "258820": "Wohaa! You cut straight through to 0.015, you don't want to know how long I have been watching my charts slowly inching down from 0.25 to 0.18 on the test set, never getting further than that. The last bit, from 0.18 to 0.12 I did by repeating the infer at 72 different crops and resizes + mirroring. But then again, you heavily overfit making your final score not all that much better than mine ;-). I wonder if the number of feature maps in the conv layers may have to do with that. I have been very conservative with a maximum of 80 feature maps, the paper you refer to has 256. My rationale was, these standard nets (Resnet, VGG etc) are designed and tested for millions of images with thousands of classes, here we have a lot less of both so the network needs to be scaled down accordingly. Maybe I have been to aggressive with that.\n\nThe peak in your training, is that not caused by the faulty image in the APS series? There is one that is half blacked out. Or otherwise maybe the crazy momentum?\n\nDid you do any tricks at infer time? Did that improve matters?",
    "258825": "This seems interesting, thanks for sharing! What do you mean with classify threats versus classify categories? What is the difference? And what did you do with the other 84% ot the training data?\n\nI got a great improvement by using random crop and zoom, which I believe has to do with the image reduction I applied. You may see a similar gain with that.",
    "258830": "Thanks for sharing! How did you \"split the data into 17 zones with some overlap\"? You're not referring to segmentation, do you?",
    "258842": "Actually segmentation wasn't very difficult to get pretty accurate if working with a3d images. While there are 16 &amp; 64 views with aps/a3daps there is really just one aerial view for a3d so all you needed to get right was the z-axis coordinates and the rest you could rely on basic symmetry of the human body. I used a simple CNN on a3daps images to predict the height range coordinates for each mirror zone pair (6/7, 2/4 etc.) and then  later any threat that my classifier found in that slice range was matched up with the corresponding zone. \n\nAlthough I definitely think Moejoe's \"automated\" solution is much more elegant and wish I had thought of it :)",
    "258845": "So how did you do it? Used and MVCNN? ;-)",
    "258848": "It's not segmentation, but just trick to reduce computation since 3D convolution requires a lot even at  1/4 scale.\nAfter 3 resnet block for the whole body, output dimensions are like 10(front-back) x 12(left-right) x 18(foot-hands) x # of last layer filters. I just slice this output to 17 zones - for example, I sliced it like x[2:6, 2:4, 0:1, :] for left feet, so this per-zone CNN only see small area only on center-left-bottom part.\n\nI think the network will eventually figure out the area of interest even if I didn't use this per-zone network, but it will require network with much more capacities and more training time.",
    "258870": "&gt; **Bastiaan Bergman wrote**\n&gt; \n&gt; &gt; Wohaa! You cut straight through to 0.015, you don't want to know how long I have been watching my charts slowly inching down from 0.25 to 0.18 on the test set, never getting further than that. \n\nA watched pot never boils. The secret is to ignore them. Do what you would normally do through the day. When you finally remember that you are training something and look at the log, then you'll find yourself thinking \"whoa, awesome results!\"",
    "259089": "This was after I had removed the bad instance and switched the two sets of labels. This was one of the better CV folds - they averaged out to around 0.025. \n\nI did use a large number of feature maps (somewhere around 256 for the upper layers) because this helped speed up training and reduced the stage1 error. Looking back, I probably would have done better without such a large number of feature maps. Dropout may also have improved stage2 results by preventing the net from memorizing specific features, although it did not help much with stage1. \n\nMuch of this is just speculation - I wouldn't read too much into these details!",
    "259134": "Hi Alexander, thank you very much for the references you provided.",
    "259183": "I started with zone segmentation from the very beginning. The main reason for this is I've been an engineer for a long time but have only recently gotten into ML. I don't have a lot of tools in the ML toolbox but I can work through an engineering problem like segmenting an image (though a lot of that was new too!). Also I thought I might have had a puncher's chance at doing well because I figured not a lot of teams would spend time on meticulously segmenting the zones. I really didn't know how else to proceed but my solution certainly evolved as I went along.\n\nTL;DR - I did a reasonable job at segmenting the zones\n\nFor my first segmentation attempt I took the max point of all slices in the dataset and tried to hand-curate the zones. I didn't like those results so then I decided if I could determine a few critical points, namely where the torso began (at the groin) and ended (at the neck) along with the torso width I could assume a universal body proportion and accurately segment the zones. I was able to determine the points pretty easily with the skimage library and it did seem like most people had similar relative proportions. Several caveats: 1) an individual could be off center on either floor axis which required an adjustment (this could be determined by looking at the center point of slices 0,8 and 4,12), 2) some subjects had excessive...um...girth so that would require an adjustment across slices and finally 3) by observing scanners at airports, I concluded that the scanner would start slow, speed up, then slow down again which required a further adjustment. To all of this, I added some padding for each zone and it seemed I had a pretty robust zone segmenter - it even looked good in the new stage 2 dataset.\n\nTL;DR - Used Keras' TimeDistributed layer with InceptionV3 freezing the first 172 layers and adding an LSTM and dense layer at the end. One classifier for zones 6,7,8,10-16, one for the arm zones and one each for 5, 17 and 9. Weighted threats 90/10.\n\nMy model was relatively simple and I never experimented with anything else. I essentially viewed the input as images across time (ie. video). So I used Keras' TimeDistributed layer with a pre-trained Imagenet model (VGG16 and InceptionV3) fed into an LSTM to a dense layer and finally to a binary result. At this point it was all manual experimentation. I tried different slices in each zone, different image input sizes, different sizes for LSTM and dense layer and playing with un-freezing various layers in the pre-trained models. At first I did not fully segment the arm and leg zones (for ex. zones 11, 13, 15 were one classifier) but ultimately I achieved a nice jump on the leaderboard by segmenting the legs so that reinforced the idea, perhaps incorrectly, that segmenting was the way to go. In the end I used 150x150 size images, with 9 time slices (save for 5, 17 and 9 which had less) and then used InceptionV3 in the TimeDistributed layer with the first 172 layers frozen and the rest trainable.  I weighted threats 90/10 in gradient updates since they were so much less prevalent. I discovered the symmetry trick so I started pairing zones (like [6,7], [8,10], [11,12], etc.) by flipping horizontally. To additionally augment, I flipped the temporal axis. Then I thought why not combine more zones? So I did all the leg zones in one classifier and all the torso zones. I kept moving up in the LB so finally I just did every 9-slice zone in one giant classifier and started flipping vertically to additionally augment. The end result: one classifier for zones 6,7,8,10-16,  one for the arm zones and one each for zones 5, 17 and 9. I reached 0.108 on the public LB. I discovered all this combining at the end so I think I could've improved that if I had more time to train.\n\nI thought I had something with the mega-classifier. It allowed me to greatly augment the data and I figured I was relatively immune to new people being introduced in stage 2 since zone segments I would think lack landmark features and by using all the zones together and flipping them around the network could focus on the threats. One glaring problem with segmenting the zones is many of the threats seemed to span more than one zone even though they were labelled for one zone. So I'm sure that confused the network in that it maybe only identified a threat when it was in the center.\n\nMy stage 2 results more than doubled. Bob Marley and suspender-dude didn't help. Also in my brief manual inspection, it seemed like there were new, subtle-looking threats that weren't in the stage 1 data. I also didn't account for mislabels by pulling back any 1.0 predictions to 0.999. Live and learn. This competition was a ton of fun though and looking forward to the next one!",
    "259200": "As promised, I'll post my architecture.  Obviously it must have over-fit to the stage 1 participants.  In hindsight I should have split my validation by individual people.  But anyway I guess it was my first real competition and I learned a valuable lesson for next time I do a 2 stager :P\n\nNormalization\n-------------\nI used only aps samples.  Height was the most variable feature, so I stretched each person to be the same height.  I chose the height based on the highest pixel above a certain threshold, it actually worked extremely well and reliably.  Once normalized, I created crop locations for each zone in each rotation, skipping occluded channels.\n\nAugmentation\n-------------\nI doubled the sample size by including mirror images.  My crop areas were 160x160 pixels, but in the actual training, I randomly sampled 128x128 crops from that larger area to sample random translations.\n\nArchitecture\n-------------\nEach zone had a different fully convolutional network.  I couldn't use vanilla convolutions since there is no spatial locality between rotations, so I created an 'Independent Convolution' layer that partitions the rotations and applies the same convolutions independently to each rotation.  I had 5 layers of convolutions and pooling before a single dense layer that combines the partitions.\n\nNone of my networks had more than 14,000 parameters, usually a lot less.  Keeping a low parameter count was how I avoided over-fitting (at least to the stage 1 validation I guess).  Although it was different for each zone, here's how a typical network looks:\n\n    _________________________________________________________________\n    Layer (type)                 Output Shape              Param #\n    =================================================================\n    input                        (None, 14, 128, 128)      0\n    _________________________________________________________________\n    ind_conv2d_1 (IndConv2D)     (None, 168, 124, 124)     312\n    _________________________________________________________________\n    max_pooling2d_1 (MaxPooling2 (None, 168, 62, 62)       0\n    _________________________________________________________________\n    activation_1 (Activation)    (None, 168, 62, 62)       0\n    _________________________________________________________________\n    ind_conv2d_2 (IndConv2D)     (None, 168, 60, 60)       1308\n    _________________________________________________________________\n    max_pooling2d_2 (MaxPooling2 (None, 168, 30, 30)       0\n    _________________________________________________________________\n    activation_2 (Activation)    (None, 168, 30, 30)       0\n    _________________________________________________________________\n    ind_conv2d_3 (IndConv2D)     (None, 168, 28, 28)       1308\n    _________________________________________________________________\n    max_pooling2d_3 (MaxPooling2 (None, 168, 14, 14)       0\n    _________________________________________________________________\n    activation_3 (Activation)    (None, 168, 14, 14)       0\n    _________________________________________________________________\n    ind_conv2d_4 (IndConv2D)     (None, 168, 12, 12)       1308\n    _________________________________________________________________\n    max_pooling2d_4 (MaxPooling2 (None, 168, 6, 6)         0\n    _________________________________________________________________\n    activation_4 (Activation)    (None, 168, 6, 6)         0\n    _________________________________________________________________\n    ind_conv2d_5 (IndConv2D)     (None, 168, 4, 4)         1308\n    _________________________________________________________________\n    flatten_1 (Flatten)          (None, 2688)              0\n    _________________________________________________________________\n    dropout_1 (Dropout)          (None, 2688)              0\n    _________________________________________________________________\n    dense_1 (Dense)              (None, 1)                 2689\n    _________________________________________________________________\n    activation_5 (Activation)    (None, 1)                 0\n    =================================================================\n    Total params: 8,233\n\nHyperparameter Optimization\n-------------\nThe only hyper-parameters I would tune were the amount of dropout at the dense layer and the number of kernels in the convolutions, which was the same for every layer.\n\nEnsembleling\n-------------\nI didn't use ensembling in the traditional sense.  Instead of using multiple models, I just also ran the mirror image through the same model (mirroring the results) as well as 9 different translations for each.  Then I averaged the results, which had to be done pre-sigmoid for statistical correctness.\n\nConclusion\n-------------\nThe low parameter count and low stage 1 score gave me too much confidence that \"it couldn't possibly be over-fitting\" but obviously in hindsight I should have tested how well it generalized to an unseen body by re-splitting the data.",
    "259294": "Hi Kevin,\nCongratulations, a great score nevertheless! And thanks for the write up.\nI wonder what exactly your 'Independent Convolution' is about. Is it convolution on the image section of one particular threat zone? Which is different for each of the 16 views?",
    "259297": "&gt; A watched pot never boils.\n\nYeah, I tried that too,.. something tells me you had more in your secret sauce ;-)",
    "259299": "Sort of yeah, I think its similar to a MVCNN, but instead of pooling to combine the views, I just use a dense layer.  The weights are shared for each of the 12-15 views in the zone.  I don't think any zone actually used 16 views since there was occlusion in at least some of them.",
    "259512": "Hi, sorry if my question is dumb, but what is this pre-sigmoid averaging? Why would this bring statistical correctness?",
    "259574": "I was wondering the same ;-)",
    "259594": "So when using log-loss, the resulting probability should be interpreted as \"Confidence\".  When you have one network with 50% confidence, and one with 99.999% you definitely wouldn't want to say that it's only 75% confident!  The other network is exponentially more confident.  The difference between 99% and 99.9% is much bigger than 50% and 50.9%.  This whole argument is of course specific to log-loss scoring.  \n\nBecause the space is stretched like this, the correct way to combine the values is to evaluate the network all the way down to right before the final sigmoid, average, and then finally apply the sigmoid once everything's been averaged together.  This transforms the weird 'sigmoidy space' to a nice linear space where averaging makes more sense.\n\nOverall this gave me a 30% improvement on the LB score compared to just averaging after the sigmoid.",
    "259684": "Thanks for posting this, Kevin; sorry that I was wrong about you being the obvious winner, but congratulations all the same -- hey, you beat me at least!  :)",
    "259789": "Overall, I concatenated threat zone crops from the aps files into \"threat plates\" and trained an ensemble of ResNet50 (Imagenet weights) using those. I normalized the data by cropping each of the 16 images around a certain threshold and resizing that to be the full image, so that threat zone locations would be more consistent. Beyond that, there were two key insights behind my solution:\n\n1. Exploit the symmetry of the human body. By switching and horizontally flipping certain images I could use the same crops for right and left forearm, right and left shin, etc. Thus I doubled the size of the dataset and only needed 9 networks for a full segmentation approach.\n\n2. Clipping. Based on the math of the log loss error function, on an inaccurate prediction, overconfidence will be penalized really harshly. So I clipped my predictions to the [0.015, 0.985] interval and this took me from bronze to silver.\n\nI really appreciate all the help I got when I asked! I think the community spirit is one of the greatest aspects of Kaggle. Here's a repo with more info and my code - https://github.com/jamespeterthornton/DHS",
    "260002": "Hello, Thanks for sharing. I don't have enough hardware to retrain the ResNet50, so can you share the re-trained ResNet50 model weights?",
    "260123": "I used both the .aps and .a3daps files for a public leaderboard score of 0.09559, and a private score of 0.2 -- I tried retraining with the additional 100 images that were used as the test set for stage 1 (I had some complications that demanded my attention, so I didn't get a chance to do that before stage 1 ended), but only hit 0.189 as a late submission with them.\n\nLike JamesThornton, I segmented the input images into \"zone plates.\"  I ended up discarding most of the input data though, which in hindsight was a major mistake I think; for each passenger/subject, I cut the image off at the top of their hands and used proportional arithmetic to slice their body up in to zones (variable heights and all.)  I used three frames for each zone: front facing pose, back facing pose, and one side facing pose.  I also mirrored the back facing pose so that the same arithmetic could be used (the right arm, for example, was in the same place of the input image for both the front and back poses that way).  I also randomly flipped the images horizontally during training as part of the transformations.\n\nThe final input to my net was two 240x180 images (one for each file type) composed of three 80x180 plates stacked vertically.  At the time that seemed like a good idea, because it captured enough of a view of every zone to get a good glimpse of a threat if it was present; I'm thinking now that the disadvantage to it is that as the input data passed through the net, it may have made things too sensitive to possible threats -- I'm just pondering the behavior of the final softmax layer and suspecting that it may have been a better idea to include more \"threat free\" portions of the input images to pull the output of the softmax toward the most likely option (most zones had no threats.)  I may be entirely wrong, who knows.\n\nI used the convolutions from two Xception nets, each devoted to one file type, and concatenated their pooled output before applying a final softmax layer for a binary classification.  I didn't clip the output, I just let the architecture say there was a 100% chance of a threat being present if it wanted to.  I don't know if that hurt my score or not (probably did by quite a bit, lol.)\n\nI wish I had enough experience to provide more lessons than I can, but I'm hesitant to offer advice given that I'm still learning myself.  I hope this information is useful, and I'd like to sincerely say that I've been incredibly impressed by the Kaggle community: you folks are great, and I'm privileged to have competed with you all :)\n\n*Edit*: I just tried clipping my submission file and re-submitting it, and it didn't make much of a difference (hit 0.17 instead of 0.18).  I guess my architecture did it's job pretty well, at least.",
    "260202": "Hi Oleg, any chance you can share how you actually accomplished this? Even just a one liner with your actual architecture - it would be very helpful to hear how you were able to build a highly effective model that could share knowledge of threat appearance between zones.",
    "260212": "Hi Moejoe, thank you for sharing this solution. Dang I'll have to investigate why my MVCNN never started working (it only produced the expected probabilities).\n\nAnyways, here's my question - where is that LSTM bit coming from? What is the purpose and how did you know to use it? I haven't heard much about adding an RNN layer for image classification, so I'd be very curious to hear a bit more about that and why it was so applicable here.",
    "260228": "James,\n\nMy approach was fairly complex overall -- I don't think I could describe it briefly. Perhaps I'll do a full write-up later. \n\nHowever, the particular aspect you are asking about seems straightforward: a single NN is trained on all relevant data. When it sees threats, it predicts what zone they are in. As I argued elsewhere in this thread, having a single omniscient NN instead of 10-17 specialized experts helps with generalization. I'll be very surprised if it turns out that any of the top teams used the specialized experts approach.",
    "260233": "File a FOIA for the code later on :)",
    "260315": "In MVCNN you simply take the max probabilities calculated from each frame after they've each gone through the network independently of each other.\n\nIn LSTM the frames are no longer treated independently of each other. It is treated as a time sequence where each time step is a frame from the .aps file and there's 16 time steps. Instead of feeding the frames directly to the LSTM I first feed it through the CNN to \"compress\" it into a small high-level feature map and feed that in. The advantage of this setup over MVCNN is that the network can learn temporal dependencies. It can learn what a rotating person should look like and flag anomalies. It can learn to boost its predictions if it notices an anomaly in multiple frames as opposed to just one frame, or suppress a prediction if an anomaly was noticed in just one frame but the others didn't pick anything up. It seems to have a regularizing affect.\n\nI viewed the problem as a single temporal sequence rather than 16 independent problems and it seemed natural a LSTM could be applied.",
    "260328": "https://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution\n\nHere's my code. Model is definied in mvcnn.py.\n\nhttps://github.com/ShayanPersonal/Kaggle-Passenger-Screening-Challenge-Solution/blob/master/mvcnn.py",
    "260340": "Moejoe -- Thank you for sharing your code. Nicely done and congratulations on your very good performance.",
    "261052": "My final submission was a little complicated, but basically it boils down to 5 variations on the strategy outlined below.  The differences were mainly in which pre-trained ImageNet model was used, what amount of augmentation was applied, size of the threat zones, and what ended up in the 3rd color channel.\n\n## Model Input\nOf the various formats, only the APS was used.  Neither of the A3D formats seemed to add anything and generating an image from the AHI is still a mystery to me.\n\nWith only ~1,200 scans and 23ish subjects, augmentation seemed critical for a model to generalize to new subjects. For me this involved rotations, translations, contrast and brightness adjustment, horizontal reflection, and gaussian blur.  In addition, ~800 more scans were generated from the training scans by using GIMP.  Mostly this involved transforming and transplanting a threat object from one location to another or from one scan to another.\n\nThe scans are all monochrome whereas the pre-trained ImageNet models accept three channel color input.  Triplicating the the monochrome scan for each channel was one option, but it seemed like a waste.  To make use of the multiple scans per subject, I took the average and standard deviation per pixel of ten similar scans of the subject and placed them in the other two color channels.  This had the aesthetically pleasing effect of making the threats stand out in a nice red color.\n![model_input][1]\nOf course the solution needed to be fully automated, so selecting similar scans needed to automated as well.  The easy way is to take the average absolute difference per pixel between each of the ~1,200 scans.  Whichever ten scans have the smallest distance could be used.  I found that a VGG-style deep autoencoder worked a little better, but the idea is essentially the same. \n\n## Model Design\nInitially I had started with a 17-target MVCNN similar to what Moejoe (Shayan) described in this thread, but eventually I found that splitting the image up into four overlapping zones worked a little better.  Potentially the improvement was due to using a higher resolution on each of the smaller threat zones or possibly due to eliminating irrelevant information from other parts of the scan.  In any case this split the model into four components roughly aimed at: Arms (1-4), Chest (5-7,17), Waist (8-12), and Feet (13-16).\n![model_zones][2]\nEach of the four component models had a similar design.  They took 8 of 16 equally spaced images from the whole APS scan and fed them into a pre-trained ImageNet model, resnet-50 for instance, whose weights were shared over the different inputs.  Given that we are asking the model to not just predict the threat's existence but also the location, we want to preserve spatial information so in contrast to MVCNN, which takes the maximum of all the outputs, here we just concatenate the outputs together.  This leads to a very large number of channels which need to be reduced using a 1x1 convolution.  After that we have a very standard top design with a few dense layers.  Some small amount of dropout was needed in the last few layer to prevent the model from overfitting.\n![model_diagram2][3]\n## Prediction/Calibration\nTest time augmentation was helpful here.  All the augmentation applied to the training images was applied here at three checkpoints for each of the models.  It was probably overkill, but predictions were made for each zone model on each scan around 100 times.  All the various predictions were just averaged.\n\nCalibrating predictions to generalize well to new subjects was also helpful.  For this a 5-fold cross-validation was performed, with scans from any particular subject all in the same fold.  Using the out-of-sample scores from cross-validation and the original labels before any corrections, a LightGBM model was built.  This LightGBM model was then applied to the stage-2 predictions.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/261052/8101/model_input.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/261052/8102/model_zones.png\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/261052/8103/model_diagram2.jpg",
    "261096": "Thank you for the great description and vizualization. It's amazing to see how the way you analyzed and augmented the data made all the difference, here. I had multiple models that were very similar to what you were doing, but never got past 0.25 LB (stage1 data). Congratulations on the first place!\n\nDid you ever try or consider using the .a3daps scans to augment? I tried using 8 equally spaced angles and then randomly shifting -2 to +2 steps (same for all 8 angles), but didn't see much of an effect on the model's ability to generalize.",
    "261116": "100 predictions!!! with Test time augmentation? WOW! we did test time augmentation too but only averaged 16 predictions per subject, and we thought that was overkill !!...  go big or go home I guess...we tried a similar approach of splitting the volume into 4 chunks for 3d convolutions at the original dimensions.  That model did not make it into our final submission because we could not reliably guarantee the zones will fall in the right chunk because the subjects had varying height. We then attempted to do registration on all the samples to make them all the same height, size e.t.c with a 12 parameter affine transformation but we were worried about the computational time and size of the state 2 data.",
    "261133": "Thanks for your write-up and congratulations on the win!\n\n&gt; ~800 more scans were generated from the training scans by using GIMP\n\nI had to Google it, thinking GIMP is some interesting algorithm. You manually drew / augmented 800 images??!! whee,...",
    "261140": "&gt; You manually drew / augmented 800 images??!!\n\n800 *scans*. It must have been ~5000 photoshoppings. Congrats, idle_speculation!",
    "261163": "thanks for the great writeup and congrats. \n\n\n-- how many training steps did you run for each region (component) and what computing environment did you use?\n\nShould  I infer one key to Idle_Speculations's success was comparing each scan to the average (and std dev) of all scans for the subect?",
    "261169": "My winning solution was an ensemble of 11 models. All models are end-to-end, i.e., they take as input 16 images per subject and spit out 17 probabilities per threats. I did not use any pretrained model or weights. All models are hand-crafted CNN based models. All models were trained on only one Titan X gpu.\n  \n**Data pre-processing**\n- I only used APS format.\n- For most models, we down-sampled the dataset to 256  by 330 \n- We also rotate images 90, 180 and 270 degrees and train separate models on the rotated dataset.\n- For couple of models we used the original 512 by 660 images and the rotated dataset.\n- Global zero-mean and unit-variance normalization were used. \n\n**Data Augmentation**\nWe perform different image augmentations including: rotation (10-15 degrees), width and height shift (10-15%), shear, zoom (10-15 %) as well as random erasing [Ref1]. I tried to use brightness augmentation but the models were completely confused with any changes in the brightness so I dropped it from augmentation.\n\n**Architecture**\nModel architecture as seen in the figure is basically a conv2D to extract features from 16 image slices one at a time, flattening features, average over 16 image slices and then generating 17 Sigmoid outputs (one per zone). I did not know about the MVCNN paper but it seems there is similarity with that method. \n\n**Insights**\nRandom erasing augmentation provided some improvements. Building separate models on 90,180,270 rotated dataset was also very effective. It is also better to use the original dataset as opposed to the down-sampled data. Although, it is more time consuming to train.  \n\n**Ref1**\nZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, Yi Yang “Random Erasing Data Augmentation”",
    "261191": "Stefan\n\nI suspect there was something amiss in your setup if you got stuck at 0.25 on the public LB.  To answer your question, I did try using the A3DAPS in variety of ways:  Putting it another color channel, Randomly replacing an APS image with a comparable A3DAPS, Building A3DAPS specific models.  None of it really seemed to work. \n\n@DavidGbodiOdaibo\n\nI based my zones off the training set and tried to build in a little room for error.  Even so, a couple of the gentlemen in stage-2 were taller than I had anticipated.  I'm pretty sure I lost a few points in zones 5 and 17 for those guys.\n\n@Bastiaan/Oleg\n\nYep, it took a while.  Probably there is a good way to automatically extract the threat part of an image and paste it into other scans, I wasn't able to figure it out though.  Another mind-numbing exercise I subjected myself to was drawing bounding boxes around all the threats in zones 1-4 while I explored Faster-RCNN.\n\n@numericLee\n\nEach of the zone models converged in around 16k iterations with batch size 8.  It worked out to be around 4 hours per zone on a 1080TI.  I also had some free credits for google cloud which I used to rent out Teslas for heavy lifting.  In terms of using average of similar scans, it seemed to decrease logloss by around 0.01 in cross-validation.  The benefits were perhaps larger on the test set given Mr. Dreadlocks.",
    "261235": "I only ended up using the APS images. \n\n# Bounding Box\nI spent a good amount of time annotating the zones in the front view of the images. I then used these annotations to train a bounding box model to predict where each zone was. I didn't want to take the time to do this at first, but the variance in the height/width of the subjects and the fact that some subjects held their arms at different angles led me to think it would be beneficial. I also annotated the locations of a subject's ankles, knees, elbows, shoulders, wrists on the front and side images. The reasons for this is that I wanted to align people's limbs so they they were all as vertical as possible. \n\n# Subject Identification\nI realized early on that there were a limited number of subjects and that the subjects in the Phase 1 test set all appeared in the training set. I knew that this would cause overfitting, and sure enough I noticed a massive difference in models trained using proper validation and those that did not. Without proper validation I was getting &lt; 0.04 on the Public LB. Using proper validation and the same architected I was getting ~0.17 on the Public LB. I ended up with 24 different subjects. \n\n# Models\nFor inputs, I concatenated the APS angles into a single 13 images x 1 image (I excluded 3 angles where the zones couldn't be seen) array. I trained 2 different models: one where each image was reduced to 96x96 pixels, and another that was 128x128 pixels i.e. the final images were 1248x96 and 1664x128. As mentioned in the Bounding Box section I rotated the limbs so that they were vertical in each image to account for the fact that people had their arms/legs as different angles. My architecture ended up being the following:\n\nConv2D(16, 3x3)  \nConv2D(16, 3x3)  \nMaxPooling2D(2x2)  \n        \nConv2D(32, 3x3)  \nConv2D(32, 3x3)  \nMaxPooling2D(2x3)  \n        \nConv2D(32, 3x3)  \nConv2D(32, 3x3)  \nConv2D(32, 3x3)  \nMaxPooling2D(2x2)  \n        \nConv2D(32, 3x3)  \nConv2D(32, 3x3)  \nMaxPooling2D(2x2)  \n        \nConv2D(32, 3x3)  \nConv2D(64, 3x3)  \nConv2D(64, 3x3)  \nMaxPooling2D(2x2)     \nDropout(0.5)  \n        \nConv2D(64, 3x3)  \nConv2D(128, 3x3)  \nConv2D(128, 3x3)  \nMaxPooling2D(3x3)      \nDropout(0.5)  \n        \nFlatten()  \n        \nDense(512)  \nDropout(0.5)  \n          \nDense(1)  \n\nAnyhow, all I have time for now. I'll hopefully add a bit more detail soon.\nThis ended up being much deeper than my original model trained without using proper cross-validation.",
    "261248": "idle_speculation\n\n&gt; Each of the zone models converged in around 16k iterations with batch size 8.\n\nDid they converge in the sense that the training loss stopped changing, or did the validation loss start going up (and you used early stopping)? What were the learning rate / momentum like?",
    "261365": "&gt; Did they converge in the sense that the training loss stopped changing, or did the validation loss start going up (and you used early stopping)? What were the learning rate / momentum like?\n\nThey converged in the sense that the validation loss did not improve past that point.  Performance did not get any worse with further training.  It just stopped getting better.  Training loss was very close to zero near the end.\n\nAdagrad was used to fit all the models.  No momentum was applied.  Learning rate was gradually increased up to 0.01 in the first 1k iterations, then left at that level for the rest of the fit.",
    "261401": "I see a potential weakness with this approach in a production environment, I am not sure if I am interpreting this correctly and will like to hear your thoughts. It seems to me that the RGB image exploits the fact that there are additional scans of the same subjects under different conditions in the training and test set and uses these additional scans to derive the “mean” and “standard deviation” green and blue channel, however, in a production/live environment the system will be seeing a subject for the first time and there will be no prior scan to derive the additional channels in the RGB image, how would you handle this? Is there a performance difference when the additional channels are not included?",
    "261412": "I would ask this privately, but I can't contact users without being a \"contributor.\"  Please pardon my occasionally awkward social skills...\n\nI would have expected your solution to be disqualified, because it's impossible for the DHS to verify it without the additional training data.  Did they contact you to ask for that?  Please understand that I'm not trying to imply there's a reason to be suspicious, and given your standing, I'm sure your word is good enough for Kaggle (if not the sponsor as well) -- but I'm still curious how verification was handled?",
    "261419": "Murray,\n\nAll winning solutions are verified by the sponsor and must be able to reproduce the winning leaderboard score per the rules. As the competition has closed, the model upload and verification process is underway.",
    "261421": "I'll just be perfectly frank: it doesn't bother me in the slightest, personally, and I have nothing but respect for someone who got the job done better than other teams (other teams includes me, too, lol); but in my mind, the manual augmentation of training data falls under a non-automated external source that wasn't made available to the other contestants.  But it doesn't really matter what I think, and I'm not trying to cause trouble.\n\nMy apologies if the question was awkward.  Truthfully I'm most interested in where the boundaries of the rules are, so that I can do better myself in the future without accidentally crossing the line.  Obviously a tremendous amount of work went into idle_speculation's submission, and I think he deserves that first place slot.  Congratulations again :)",
    "261490": "I decided to approach this competition from the viewpoint of object detection.  I annotated a subset of the training images by drawing bounding boxes around the threats in the aps images (I spent maybe 8 hrs on this).  I then trained Faster-RCNN using the annotated images, with the label of each object being its location on the body.  This way the network would learn to both detect objects and label them based on their location all in one go.\n\nFor each aps image there are 16 viewpoints.  I ran the object detector over each viewpoint in the aps image, to generate a set of candidate object detections, with each detection having a probability over each body location.  I took the maximum of all the detections for each body location to get a final 17 dimensional confidence vector.  I then took the 17d confidence vector for each viewpoint to get a 16x17 matrix.  I used this to train an 17 gradient boosted classifiers (one for each body location) to get better calibrated probabilities.\n\nThe only data augmentation I performed was doing a horizontal flip of the images while training the object detector.  This actually will change the label.  For example, if a threat is labeled \"right ankle\" then after flipping it will then be labeled \"left ankle\". \n\nI decided on using an object detector due to the limited size of the dataset and the small number of unique individuals.  Faster-RCNN is a two stage object detector.  In the first stage, it recognizes potential threats; the second stage labels the threats/decides if they are not true threats.  I thought that this built in attention mechanism would help to prevent overfitting by forcing the network to focus on the threats.\n\nMy final model consisted of an ensemble of 5 models.  Each model using 80% of the data to train the object detector and then the final 20% to train the boosted classifier.  I then averaged the predictions over each individual model.  However, ensembling did not really seem help the performance of my model in any significant way.\n\nNow some things that I tried that didn't work.  I spent some time trying to use the 3d data.  I used a 3d inflated convolutional network based on the vgg16 architecture.  This just means that I turned each 2d convolution into a 3d convolution and initialized the weights of each 3d convolution by stacking the weights of the pretrained vgg model.  I then tried to use this network to perform 3d object detection.   This best I was able to do was around 0.06 on my validation set as opposed to 0.025 using my 2d model.",
    "261531": "Thanks, deedrz, and congrats! Were those networks pre-trained on ImageNet?",
    "261533": "Thanks for the writeup and congrats on the win!\n\nDid you consider ensembling the two models: the 3D conv and the 2D conv? I'd think the two models are fairly orthogonal?",
    "261534": "Thanks for the writeup!\n\n&gt;I rotated the limbs so that they were vertical in each image\n\nHow did you do that? Did you rotate the whole picture to get everything straight on average? Or did you cut up the people in separate limbs? Or did you stretch and skew the images? Or more advanced?",
    "261539": "Yep, the networks were pre-trained on ImageNet, I just replicated each image to get the 3 channel input. And Bastiaan, I think your right, combining the 2d and 3d models would've been a good thing to try",
    "261540": "&gt; **Murray Miron wrote**\n&gt;\n&gt; I would have expected your solution to be disqualified, because it's impossible for the DHS to verify it without the additional training data.  Did they contact you to ask for that?  Please understand that I'm not trying to imply there's a reason to be suspicious, and given your standing, I'm sure your word is good enough for Kaggle (if not the sponsor as well) -- but I'm still curious how verification was handled?\n\nLuckily I had the forethought to include the additional scans in my model upload.",
    "261545": "Using the right forearm as an example, I had a model which identified the point in the front and side images where the wrist and elbow were.  I then used my bounding box model to identify the where the forearm was and cropped out the rest of the image.  Then using the point locations of the wrist and elbow and using a bit of trigonometry I calculated how much I needed to rotate each image so that the forearm was vertical.",
    "261546": "&gt; **DavidGbodiOdaibo wrote**\n&gt;\n&gt; I see a potential weakness with this approach in a production environment, I am not sure if I am interpreting this correctly and will like to hear your thoughts. It seems to me that the RGB image exploits the fact that there are additional scans of the same subjects under different conditions in the training and test set and uses these additional scans to derive the “mean” and “standard deviation” green and blue channel, however, in a production/live environment the system will be seeing a subject for the first time and there will be no prior scan to derive the additional channels in the RGB image, how would you handle this? Is there a performance difference when the additional channels are not included?\n\nI don't consider myself a person who travels that much.  But personally, I took three flights in the last year.  That works out to six trips though the scanners this year alone.  If you add in scans from last year or two then the approach of using ten similar scans starts sounding a lot more plausible.",
    "261583": "I don't think the TSA is legally permitted to store the scans, actually.",
    "261798": "I posted about this in more detail [here][1], but the TL;DR is:\n\n- trained a model to compute segmentations of the threats given aps/a3daps images (hand-labeled the training data)\n- trained a model to compute segmentations of the zones given depth maps (generated synthetic data with depth maps / ground truth to train, used a3d to get real depth maps)\n- trained a model that takes in the threat / zone segmentations from 16 angles and outputs the final predictions for each zone\n  [1]: http://suchir.io/posts/passenger-screening-algorithm-challenge-writeup.html",
    "263940": "The description of my cylindrical coordinates based approach is [posted in Kernels](https://www.kaggle.com/nathanrm/full-solution-cylindrical-coordinate-method), and includes a link to the complete solution on github.",
    "265258": "idle_speculation\n\nCongrats, I'm also curious since I haven't seen anybody mention about the recent published paper by Hinton in regards to Dynamic routing between capsules, even though it's relatively new but what are your thoughts on this capsule net over convolutional net? \n\nhttps://arxiv.org/pdf/1710.09829.pdf",
    "271860": "I participated in this competition for completing the Udacity Nanodegree capstone project requirement, you can find the full report [here.][1]. In summary, my solution involved these two steps. \n\n 1. Detect head, hands, groin and threat bounding boxes using TensorFlow Object Detection API\n 2. Predict Body Zone of a threat based on the bounding boxes of a threat relative to the frame number, head, hands and groin. I attempted this in two ways: A supervised learning method using XGBOOST. A custom algorithm which needed only a few ratios to be tweaked.\n\nOverall it was good learning experience, it did however take a lot of time. \n\n  [1]: https://github.com/shivarajugowda/kaggle/blob/master/TSA/report.md",
    "312252": "I wonder how much time the other top-8 teams' models need for training and prediction, respectively?",
    "457123": "hello,I am working on this problem using your tips but changed  the net architecture,so I want to compare our result.can you share your code?"
  },
  "source": "meta"
}