{
  "id": 29829,
  "title": "0.51276 Public LB Solution",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/writeups/n0zf1tuzrb3o-0-51276-public-lb-solution",
  "author_name": "",
  "post_date": "2017-03-16T17:27:28.910Z",
  "votes": 58,
  "comment_count": 14,
  "views": 0,
  "content": "<p>To solve this problem I used: Python + Keras 1.0.8 + Theano under Windows 10. As hardware I had NVIDIA GTX 980 8 GB. I used only one GPU which worked almost 24h a day during competition. After join the team I used 2 additional GPUs (TITAN 12GB).\nIn short my solution can be described like this:</p>\n\n<ol>\n<li>My main CNN is modified UNET with input shape (20, 224, 224). It\nhas higher depth, with added batch normalization and dropout layers.\nAs loss function is used Jacquard Coefficient. I used grid search\nwith different parameters and choose the best model with highest\nvalidation score. </li>\n<li>I create separate models for each class (so 10\nindependent models) and tuned them independently. I actually think\nthat I lost some information about class interaction this way, but\nit was much easier to tune models. I partially fix it with\npostprocessing step. </li>\n<li>Each model is actually set of K different\nweights obtained with KFold. For most classes I used 5 KFold. For\nclass 7 - 2 KFold. For class 9 – 4 KFold. I split train set by image\nID (there were only 25 images). I made it once by hands\nindependently for each class. Each fold contains the same number of\nimages with existed class. So it’s actually stratified split.</li>\n</ol>\n\n<h2>Data preprocessing</h2>\n\n<p>Each “image object” (with particular img_id) had set of images made with different wave lengths. In total 20 channels. I resized all images to 3360x3360 pixels (3360 = 15*224 and 3360 is closest to panchromatic image resolution).  And join them along the axis. So at the end I had tensor of following shape: 20x3360x3360. I created the pixel masks from given polygons with same size: 3360x3360 pixels.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6037/Tensor_example_700px.png\" alt=\"Set of images example\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6029/Mask_example.png\" alt=\"Mask example\" title=\"\"></p>\n\n<p>I decided to use UNET with input shape 20x224x224. I choose 224 for 2 reasons:</p>\n\n<ol>\n<li>224 = 2*2*2*2*2*7 – I have 5 MaxPooling layers in UNET and on lowest layer has 7x7 pixel size. Which in my experience the best.</li>\n<li>The same size I could use with pretrained VGG16 or ResNet.</li>\n</ol>\n\n<p>The next step was made independently for each class:</p>\n\n<ol>\n<li>Split input images on 15*15 parts forming tensors of size\n20x224x224. </li>\n<li>Because of class imbalances we need to increase\nnumber of cases where mask exists. For some cases like class 10\n(with vehicles) I have around 99% of empty masks after first step.\nSo I added more cases using sliding window around non-zero mask\npoints.</li>\n</ol>\n\n<p>Final statistics for classes:</p>\n\n<pre><code>Number of tests for class 1: 14010. Empty files: 4431 Percent: 31.62%\nNumber of tests for class 2: 9835. Empty files: 3997 Percent: 40.64%\nNumber of tests for class 3: 20548. Empty files: 5238 Percent: 25.49%\nNumber of tests for class 4: 12411. Empty files: 2976 Percent: 23.97%\nNumber of tests for class 5: 10161. Empty files: 294 Percent: 2.89%\nNumber of tests for class 6: 10716. Empty files: 3720 Percent: 34.71%\nNumber of tests for class 7: 10328. Empty files: 5504 Percent: 53.29%\nNumber of tests for class 8: 11349. Empty files: 5469 Percent: 48.18%\nNumber of tests for class 9: 12035. Empty files: 5574 Percent: 46.31%\nNumber of tests for class 10: 16173. Empty files: 5336 Percent: 32.99%\n</code></pre>\n\n<p>UNET requires the distribution to be close to normal. Ranges for different channels was: P_3, P_P, P_M – 2048, P_A – 16384. At first I divide every channel to its maximum possible value. I’ve seen on forum that some people removed some values from the both ends of histogram but I didn’t do it.</p>\n\n<p>Then I needed to calculate mean and stdev. I calculated it for each channel independently, but at the end used the same mean and stdev for whole tensor: \nmean = 0.219613\nstdev = 0.110741</p>\n\n<h2>Creating models</h2>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6038/model_700px.png\" alt=\"Modified UNET\" title=\"\"></p>\n\n<pre><code>____________________________________________________________________________________\nLayer (type)             Output Shape   Param #     Connected to\n====================================================================================\ninput_1 (InputLayer)      (20, 224, 224)  0                                 \n____________________________________________________________________________________\nconv2d_1 (Convolution2D)  (32, 224, 224)  5792     input_1\n____________________________________________________________________________________\nbatchnorm_1 (BatchNormal  (32, 224, 224)  64       conv2d_1             \n____________________________________________________________________________________\nactivation_1 (Activation) (32, 224, 224)  0        batchnorm_1        \n____________________________________________________________________________________\nconv2d_2 (Convolution2D)  (32, 224, 224)  9248     activation_1                \n____________________________________________________________________________________\nbatchnorm_2 (BatchNormal  (32, 224, 224)  64       conv2d_2             \n____________________________________________________________________________________\nactivation_2 (Activation) (32, 224, 224)  0        batchnorm_2        \n____________________________________________________________________________________\nmaxpool2d_1 (MaxPooling2D)(32, 112, 112)  0        activation_2                \n____________________________________________________________________________________\nconv2d_3 (Convolution2D)  (64, 112, 112)  18496    maxpool2d_1              \n____________________________________________________________________________________\nbatchnorm_3 (BatchNormal  (64, 112, 112)  128      conv2d_3             \n____________________________________________________________________________________\nactivation_3 (Activation) (64, 112, 112)  0        batchnorm_3        \n____________________________________________________________________________________\nconv2d_4 (Convolution2D)  (64, 112, 112)  36928    activation_3                \n____________________________________________________________________________________\nbatchnorm_4 (BatchNormal  (64, 112, 112)  128      conv2d_4             \n____________________________________________________________________________________\nactivation_4 (Activation) (64, 112, 112)  0        batchnorm_4        \n____________________________________________________________________________________\nmaxpool2d_2 (MaxPooling2D)(64, 56, 56)    0        activation_4                \n____________________________________________________________________________________\nconv2d_5 (Convolution2D)  (128, 56, 56)   73856    maxpool2d_2              \n____________________________________________________________________________________\nbatchnorm_5 (BatchNormal  (128, 56, 56)   256      conv2d_5             \n____________________________________________________________________________________\nactivation_5 (Activation) (128, 56, 56)   0        batchnorm_5        \n____________________________________________________________________________________\nconv2d_6 (Convolution2D)  (128, 56, 56)   147584   activation_5                \n____________________________________________________________________________________\nbatchnorm_6 (BatchNormal  (128, 56, 56)   256      conv2d_6             \n____________________________________________________________________________________\nactivation_6 (Activation) (128, 56, 56)   0        batchnorm_6        \n____________________________________________________________________________________\nmaxpool2d_3 (MaxPooling2D)(128, 28, 28)   0        activation_6                \n____________________________________________________________________________________\nconv2d_7 (Convolution2D)  (256, 28, 28)   295168   maxpool2d_3              \n____________________________________________________________________________________\nbatchnorm_7 (BatchNormal  (256, 28, 28)   512      conv2d_7             \n____________________________________________________________________________________\nactivation_7 (Activation) (256, 28, 28)   0        batchnorm_7        \n____________________________________________________________________________________\nconv2d_8 (Convolution2D)  (256, 28, 28)   590080   activation_7                \n____________________________________________________________________________________\nbatchnorm_8 (BatchNormal  (256, 28, 28)   512      conv2d_8             \n____________________________________________________________________________________\nactivation_8 (Activation) (256, 28, 28)   0        batchnorm_8        \n____________________________________________________________________________________\nmaxpool2d_4 (MaxPooling2D)(256, 14, 14)   0        activation_8                \n____________________________________________________________________________________\nconv2d_9 (Convolution2D)  (512, 14, 14)   1180160  maxpool2d_4              \n____________________________________________________________________________________\nbatchnorm_9 (BatchNormal  (512, 14, 14)   1024     conv2d_9             \n____________________________________________________________________________________\nactivation_9 (Activation) (512, 14, 14)   0        batchnorm_9        \n____________________________________________________________________________________\nconv2d_10 (Convolution2D) (512, 14, 14)   2359808  activation_9                \n____________________________________________________________________________________\nbatchnorm_10 (BatchNorma  (512, 14, 14)   1024     conv2d_10            \n____________________________________________________________________________________\nactivation_10 (Activation)(512, 14, 14)   0        batchnorm_10       \n____________________________________________________________________________________\nmaxpool2d_5 (MaxPooling2D)(512, 7, 7)     0        activation_10               \n____________________________________________________________________________________\nconv2d_11 (Convolution2D) (1024, 7, 7)    4719616  maxpool2d_5              \n____________________________________________________________________________________\nbatchnorm_11 (BatchNorma  (1024, 7, 7)    2048     conv2d_11            \n____________________________________________________________________________________\nactivation_11 (Activation)(1024, 7, 7)    0        batchnorm_11       \n____________________________________________________________________________________\nconv2d_12 (Convolution2D) (1024, 7, 7)    9438208  activation_11               \n____________________________________________________________________________________\nbatchnorm_12 (BatchNorma  (1024, 7, 7)    2048     conv2d_12            \n____________________________________________________________________________________\nactivation_12 (Activation)(1024, 7, 7)    0        batchnorm_12       \n____________________________________________________________________________________\nupsamp2d_1 (UpSampling2D) (1024, 14, 14)  0        activation_12               \n____________________________________________________________________________________\nmerge_1 (Merge)           (1536, 14, 14)  0        upsamp2d_1              \n                                                      activation_10               \n____________________________________________________________________________________\nconv2d_13 (Convolution2D) (512, 14, 14)   7078400  merge_1                     \n____________________________________________________________________________________\nbatchnorm_13 (BatchNorma  (512, 14, 14)   1024     conv2d_13            \n____________________________________________________________________________________\nactivation_13 (Activation)(512, 14, 14)   0        batchnorm_13       \n____________________________________________________________________________________\nconv2d_14 (Convolution2D) (512, 14, 14)   2359808  activation_13               \n____________________________________________________________________________________\nbatchnorm_14 (BatchNorma  (512, 14, 14)   1024     conv2d_14            \n____________________________________________________________________________________\nactivation_14 (Activation)(512, 14, 14)   0        batchnorm_14       \n____________________________________________________________________________________\nupsamp2d_2 (UpSampling2D) (512, 28, 28)   0        activation_14               \n____________________________________________________________________________________\nmerge_2 (Merge)           (768, 28, 28)   0        upsamp2d_2              \n                                                                activation_8                \n____________________________________________________________________________________\nconv2d_15 (Convolution2D) (256, 28, 28)   1769728  merge_2                     \n____________________________________________________________________________________\nbatchnorm_15 (BatchNorma  (256, 28, 28)   512      conv2d_15            \n____________________________________________________________________________________\nactivation_15 (Activation)(256, 28, 28)   0        batchnorm_15       \n____________________________________________________________________________________\nconv2d_16 (Convolution2D) (256, 28, 28)   590080   activation_15               \n____________________________________________________________________________________\nbatchnorm_16 (BatchNorma  (256, 28, 28)   512      conv2d_16            \n____________________________________________________________________________________\nactivation_16 (Activation)(256, 28, 28)   0        batchnorm_16       \n____________________________________________________________________________________\nupsamp2d_3 (UpSampling2D) (256, 56, 56)   0        activation_16               \n____________________________________________________________________________________\nmerge_3 (Merge)           (384, 56, 56)   0        upsamp2d_3              \n                                                      activation_6                \n____________________________________________________________________________________\nconv2d_17 (Convolution2D) (128, 56, 56)   442496   merge_3                     \n____________________________________________________________________________________\nbatchnorm_17 (BatchNorma  (128, 56, 56)   256      conv2d_17            \n____________________________________________________________________________________\nactivation_17 (Activation)(128, 56, 56)   0        batchnorm_17       \n____________________________________________________________________________________\nconv2d_18 (Convolution2D) (128, 56, 56)   147584   activation_17               \n____________________________________________________________________________________\nbatchnorm_18 (BatchNorma  (128, 56, 56)   256      conv2d_18            \n____________________________________________________________________________________\nactivation_18 (Activation)(128, 56, 56)   0        batchnorm_18       \n____________________________________________________________________________________\nupsamp2d_4 (UpSampling2D) (128, 112, 112) 0        activation_18               \n____________________________________________________________________________________\nmerge_4 (Merge)           (192, 112, 112) 0        upsamp2d_4              \n                                                      activation_4                \n____________________________________________________________________________________\nconv2d_19 (Convolution2D) (64, 112, 112)  110656   merge_4                     \n____________________________________________________________________________________\nbatchnorm_19 (BatchNorma  (64, 112, 112)  128      conv2d_19            \n____________________________________________________________________________________\nactivation_19 (Activation)(64, 112, 112)  0        batchnorm_19       \n____________________________________________________________________________________\nconv2d_20 (Convolution2D) (64, 112, 112)  36928    activation_19               \n____________________________________________________________________________________\nbatchnorm_20 (BatchNorma  (64, 112, 112)  128      conv2d_20            \n____________________________________________________________________________________\nactivation_20 (Activation)(64, 112, 112)  0        batchnorm_20       \n____________________________________________________________________________________\nupsamp2d_5 (UpSampling2D) (64, 224, 224)  0        activation_20               \n____________________________________________________________________________________\nmerge_5 (Merge)           (96, 224, 224)  0        upsamp2d_5              \n                                                      activation_2                \n____________________________________________________________________________________\nconv2d_21 (Convolution2D) (32, 224, 224)  27680    merge_5                     \n____________________________________________________________________________________\nbatchnorm_21 (BatchNorma  (32, 224, 224)  64       conv2d_21            \n____________________________________________________________________________________\nactivation_21 (Activation)(32, 224, 224)  0        batchnorm_21       \n____________________________________________________________________________________\nconv2d_22 (Convolution2D) (32, 224, 224)  9248     activation_21               \n____________________________________________________________________________________\nbatchnorm_22 (BatchNorma  (32, 224, 224)  64       conv2d_22            \n____________________________________________________________________________________\nactivation_22 (Activation)(32, 224, 224)  0        batchnorm_23       \n____________________________________________________________________________________\nconv2d_23 (Convolution2D) (1, 224, 224)   33       activation_22               \n____________________________________________________________________________________\nbatchnorm_23 (BatchNorma  (1, 224, 224)   2        conv2d_23            \n____________________________________________________________________________________\nactivation_23 (Activation)(1, 224, 224)   0        batchnorm_23       \n====================================================================================\nTotal params: 31459619\n</code></pre>\n\n<p>LOSS Function:</p>\n\n<pre><code>def jacard_coef(y_true, y_pred):\n    y_true_f = K.flatten(y_true)\n    y_pred_f = K.flatten(y_pred)\n    intersection = K.sum(y_true_f * y_pred_f)\n    return (intersection + 1.0) / (K.sum(y_true_f) + K.sum(y_pred_f) - intersection + 1.0)\n</code></pre>\n\n<p>Typical example of training process: Class 1 (Fold 2). Parameters: lr=0.05, optim=SGD, rotation=False, dropout=enabled, UNET version=BatchNorm, patience=8, samples train: 9997, samples valid: 4013</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6031/Training.png\" alt=\"Typical example of training process\" title=\"\"></p>\n\n<h2>Tuning models and validation</h2>\n\n<p>The main problem was that training process is very unstable. Sometimes it can go to local minimum or start predicting as mask full image etc. So most of the time I spend on finding optimal parameters to maximize validation score. It was the most time consuming part. I used grid search with following parameters: </p>\n\n<ul>\n<li>Optimizer (Adam, SGD) </li>\n<li>LR SGD: (0.05, 0.01, 0.001) </li>\n<li>LR Adam: (0.01, 0.001, 0.0001) </li>\n<li>Rotation (enabled, disabled) </li>\n<li>Type of model: UNET 224x224, UNET with dropout 0.1, UNET with Batch Normalization </li>\n<li>Number of samples per epoch (Fraction from ½ up to 1)</li>\n</ul>\n\n<p>Due to limitation of computational power I checked only some of parameters. \nAt the end I use mostly Adam optimizer with BatchNormalization version of UNET with rotation enabled. Just tune learning rate.\nI use loss function value to stop training with early stopping with patience 8-15 epochs. After I obtain all 5 Folds models I had the predicted segmentation for each of train images. So I was able to predict score for full image set using the same method as it was made on Kaggle Leaderboard. Obtained score was very representative, so increasing the score locally almost always leads to increasing the score on leaderboard. Also I was able to find optimal threshold value for heatmap with this method.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6032/LB-Scores.png\" alt=\"LS vs LB Scores\" title=\"\"></p>\n\n<h2>Processing the test data</h2>\n\n<p>Test set consists of 429 images. Each test image was processed separately. Each image was resized and normalized the same way as training data to 20х3360х3360 tensor. </p>\n\n<p>In the beginning I create two zero arrays HEATMAP and COUNT with 3360х3360 shape. Then use sliding window approach to predict segmentation. All 5 folds predict on image extracted from current position and acquired probabilities sum up to HEATMAP array. COUNT array at image position added by 1. At the end HEATMAP is divided by COUNT to get real heatmap array. After this heatmap was thresholded at value acquired from validation, typically 0.5. On basis of this 2D-array I created polygons with rasterio and shapely.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6033/Sliding%20Window.png\" alt=\"Sliding window\" title=\"\"></p>\n\n<p>Additional ideas to increase the accuracy provided below. Not all of them was used in final submit because processing of test images was very long process, around 8-10 hours for whole test set.</p>\n\n<ol>\n<li><p>Decreasing sliding step from 112px to 56px and below always increased the accuracy of predictions, but increase the computational complexity as O(N^2). </p>\n\n<ul><li>For example class 4 on validation 0.385714 vs 0.403683.</li></ul></li>\n<li><p>It looks like UNET predict worse on the edges, so it’s good idea to use only central part for prediction. Example for class 5 below. </p>\n\n<ul><li>Default score: 0.507446 </li>\n<li>200x200 center part from 224x224 square + 100 px sliding window step: 0.512844 </li>\n<li>160x160 center part from 224x224 square + 80 px sliding window step: 0.514557</li></ul></li>\n<li><p>I usually used threshold 0.5 without checking optimum on validation, but this can be useful too. </p>\n\n<ul><li>Class 6 example: </li>\n<li>THR 0.1: 0.738943 </li>\n<li>THR 0.5: 0.757858 </li>\n<li>THR 0.9: 0.759850</li></ul></li>\n</ol>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6039/Masks_700px.png\" alt=\"Masks for class 5\" title=\"\"></p>\n\n<h2>Ensembles</h2>\n\n<p>The most obvious ways to do ensembles is to use heatmaps. But out team consisted of 2 people was merged at the last stage of competitions, so we actually don’t have heatmaps for all classes at this stage. So we made ensembles directly on polygons, using UNION and INTERSECTION from shapely. We have different training process and models, so for some classes we had good boost after merge of this type. </p>\n\n<pre><code>Class 6 LB: 0.08149 + LB: 0.08103 (Intersection) Score: 0.08179\nClass 8 LB: 0.03890 + LB: 0.05322 (Intersection) Score: 0.06194\nClass 9 LB: 0.02113 + LB: 0.02713 (Union) Score: 0.03254\n</code></pre>\n\n<p>Bad thing, we mostly use leaderboard to check if our ensemble gave boost, which could lead to overfitting. It would be much easier if we have 3rd independent solution to use voting mechanism.</p>\n\n<h2>Post processing</h2>\n\n<p>After analysis of class interactions in train I made the following postprocess types:\n- Remove too large polygons (from class 8, 9 and 10 for example) \n- Subtract predicted classes polygons like water from car classes</p>\n\n<h2>Class 10 with small objects</h2>\n\n<p>For class 10 default approach works not very well. The main problems with this class that it’s have very small objects, bad train segmentation and it’s total area is too small comparing to full area. So I created other UNET CNN for this class. It had input of 32x32 pixels. Data for this class was heavily augmented, rotations, different shifts etc. I also used loss function with big penalty for false positives based on <a href=\"https://en.wikipedia.org/wiki/Tversky_index\">Tversky index</a>.</p>\n\n<pre><code># https://en.wikipedia.org/wiki/Tversky_index\ndef tversky_coef(y_true, y_pred):\n    y_true_f = K.flatten(y_true)\n    y_pred_f = K.flatten(y_pred)\n    alfa = 0.1\n    false_positive = K.sum(y_pred_f * (1 - y_true_f))\n    false_negative = K.sum((1 - y_pred_f) * y_true_f)\n    true_positive = K.sum(y_true_f * y_pred_f)\n    return (true_positive + 1.0) / (false_negative*alfa + (1 - alfa) * false_positive + true_positive + 1.0)\n</code></pre>\n\n<p>It allowed me to get 0.00362 score on LB.</p>\n\n<h2>Other ideas I tried</h2>\n\n<ul>\n<li>At the early beginning I made single XGBoost model which works on squares of size 10x10. Model tries to predict to which class given square is belong. It easily got me higher than baseline. Probably creating independent XGBoost models for each class could give good results. But I switch to CNN without additional experiments with XGBoost approach.</li>\n<li>Pre-trained VGG16 for rare class localization. I tried to predict if car exists in given area with VGG16 on RGB images for later ensemble with UNET predictions.</li>\n</ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6040/vgg16_test_6060_3_2_cls_10_700px.png\" alt=\"Example of class 10 localization with VGG16\" title=\"\"></p>\n\n<ul>\n<li>Pre-trained VGG16 for segmentation. The main idea here is that we use final layer (which is now sigmoid instead of softmax) as indication in which part of 224x224 image given class exists. I used 16x16 = 256 neurons for this task. But result wasn’t very good, may be because of low resolution - 14x14 square as single pixel.</li>\n<li>There are big bunch of different indexes used for automatic segmentation without CNN:\n<a href=\"http://www.indexdatabase.de/db/i.php\">http://www.indexdatabase.de/db/i.php</a>\nIt works great for water as shown in <a href=\"https://www.kaggle.com/resolut/dstl-satellite-imagery-feature-detection/waterway-0-095-lb\">Waterway Kernel</a>. And these indexes can be added as additional planes during learning process. Unfortunately they didn’t help much in my experiments, so I abandon them. </li>\n<li>I tried to use CRF (Conditional random field) methods with <a href=\"https://pystruct.github.io/\">pystruct</a> to improve predicted masks. But it somehow won’t work for me.</li>\n</ul>\n\n<h2>Independent classes best scores</h2>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6036/Best-LB.png\" alt=\"Best Public LB\" title=\"\"></p>\n\n<h2>Segmentation example</h2>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6048/6100_1_2_mask.png\" alt=\"enter image description here\" title=\"\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6049/6100_1_2_proj_700px.jpg\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>Checkout the video:\n<strong><a href=\"https://www.youtube.com/watch?v=rpp7ZhGb1IQ\">https://www.youtube.com/watch?v=rpp7ZhGb1IQ</a></strong></p>\n\n<h2>Code</h2>\n\n<p>GitHUB link here later. I plan to release code after publication of final results.</p>",
  "messages": [
    {
      "id": "166431",
      "postDate": "03/09/2017 17:13:24",
      "content": "<p>To solve this problem I used: Python + Keras 1.0.8 + Theano under Windows 10. As hardware I had NVIDIA GTX 980 8 GB. I used only one GPU which worked almost 24h a day during competition. After join the team I used 2 additional GPUs (TITAN 12GB).\nIn short my solution can be described like this:</p>\n\n<ol>\n<li>My main CNN is modified UNET with input shape (20, 224, 224). It\nhas higher depth, with added batch normalization and dropout layers.\nAs loss function is used Jacquard Coefficient. I used grid search\nwith different parameters and choose the best model with highest\nvalidation score. </li>\n<li>I create separate models for each class (so 10\nindependent models) and tuned them independently. I actually think\nthat I lost some information about class interaction this way, but\nit was much easier to tune models. I partially fix it with\npostprocessing step. </li>\n<li>Each model is actually set of K different\nweights obtained with KFold. For most classes I used 5 KFold. For\nclass 7 - 2 KFold. For class 9 – 4 KFold. I split train set by image\nID (there were only 25 images). I made it once by hands\nindependently for each class. Each fold contains the same number of\nimages with existed class. So it’s actually stratified split.</li>\n</ol>\n\n<h2>Data preprocessing</h2>\n\n<p>Each “image object” (with particular img_id) had set of images made with different wave lengths. In total 20 channels. I resized all images to 3360x3360 pixels (3360 = 15*224 and 3360 is closest to panchromatic image resolution).  And join them along the axis. So at the end I had tensor of following shape: 20x3360x3360. I created the pixel masks from given polygons with same size: 3360x3360 pixels.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6037/Tensor_example_700px.png\" alt=\"Set of images example\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6029/Mask_example.png\" alt=\"Mask example\" title=\"\"></p>\n\n<p>I decided to use UNET with input shape 20x224x224. I choose 224 for 2 reasons:</p>\n\n<ol>\n<li>224 = 2*2*2*2*2*7 – I have 5 MaxPooling layers in UNET and on lowest layer has 7x7 pixel size. Which in my experience the best.</li>\n<li>The same size I could use with pretrained VGG16 or ResNet.</li>\n</ol>\n\n<p>The next step was made independently for each class:</p>\n\n<ol>\n<li>Split input images on 15*15 parts forming tensors of size\n20x224x224. </li>\n<li>Because of class imbalances we need to increase\nnumber of cases where mask exists. For some cases like class 10\n(with vehicles) I have around 99% of empty masks after first step.\nSo I added more cases using sliding window around non-zero mask\npoints.</li>\n</ol>\n\n<p>Final statistics for classes:</p>\n\n<pre><code>Number of tests for class 1: 14010. Empty files: 4431 Percent: 31.62%\nNumber of tests for class 2: 9835. Empty files: 3997 Percent: 40.64%\nNumber of tests for class 3: 20548. Empty files: 5238 Percent: 25.49%\nNumber of tests for class 4: 12411. Empty files: 2976 Percent: 23.97%\nNumber of tests for class 5: 10161. Empty files: 294 Percent: 2.89%\nNumber of tests for class 6: 10716. Empty files: 3720 Percent: 34.71%\nNumber of tests for class 7: 10328. Empty files: 5504 Percent: 53.29%\nNumber of tests for class 8: 11349. Empty files: 5469 Percent: 48.18%\nNumber of tests for class 9: 12035. Empty files: 5574 Percent: 46.31%\nNumber of tests for class 10: 16173. Empty files: 5336 Percent: 32.99%\n</code></pre>\n\n<p>UNET requires the distribution to be close to normal. Ranges for different channels was: P_3, P_P, P_M – 2048, P_A – 16384. At first I divide every channel to its maximum possible value. I’ve seen on forum that some people removed some values from the both ends of histogram but I didn’t do it.</p>\n\n<p>Then I needed to calculate mean and stdev. I calculated it for each channel independently, but at the end used the same mean and stdev for whole tensor: \nmean = 0.219613\nstdev = 0.110741</p>\n\n<h2>Creating models</h2>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6038/model_700px.png\" alt=\"Modified UNET\" title=\"\"></p>\n\n<pre><code>____________________________________________________________________________________\nLayer (type)             Output Shape   Param #     Connected to\n====================================================================================\ninput_1 (InputLayer)      (20, 224, 224)  0                                 \n____________________________________________________________________________________\nconv2d_1 (Convolution2D)  (32, 224, 224)  5792     input_1\n____________________________________________________________________________________\nbatchnorm_1 (BatchNormal  (32, 224, 224)  64       conv2d_1             \n____________________________________________________________________________________\nactivation_1 (Activation) (32, 224, 224)  0        batchnorm_1        \n____________________________________________________________________________________\nconv2d_2 (Convolution2D)  (32, 224, 224)  9248     activation_1                \n____________________________________________________________________________________\nbatchnorm_2 (BatchNormal  (32, 224, 224)  64       conv2d_2             \n____________________________________________________________________________________\nactivation_2 (Activation) (32, 224, 224)  0        batchnorm_2        \n____________________________________________________________________________________\nmaxpool2d_1 (MaxPooling2D)(32, 112, 112)  0        activation_2                \n____________________________________________________________________________________\nconv2d_3 (Convolution2D)  (64, 112, 112)  18496    maxpool2d_1              \n____________________________________________________________________________________\nbatchnorm_3 (BatchNormal  (64, 112, 112)  128      conv2d_3             \n____________________________________________________________________________________\nactivation_3 (Activation) (64, 112, 112)  0        batchnorm_3        \n____________________________________________________________________________________\nconv2d_4 (Convolution2D)  (64, 112, 112)  36928    activation_3                \n____________________________________________________________________________________\nbatchnorm_4 (BatchNormal  (64, 112, 112)  128      conv2d_4             \n____________________________________________________________________________________\nactivation_4 (Activation) (64, 112, 112)  0        batchnorm_4        \n____________________________________________________________________________________\nmaxpool2d_2 (MaxPooling2D)(64, 56, 56)    0        activation_4                \n____________________________________________________________________________________\nconv2d_5 (Convolution2D)  (128, 56, 56)   73856    maxpool2d_2              \n____________________________________________________________________________________\nbatchnorm_5 (BatchNormal  (128, 56, 56)   256      conv2d_5             \n____________________________________________________________________________________\nactivation_5 (Activation) (128, 56, 56)   0        batchnorm_5        \n____________________________________________________________________________________\nconv2d_6 (Convolution2D)  (128, 56, 56)   147584   activation_5                \n____________________________________________________________________________________\nbatchnorm_6 (BatchNormal  (128, 56, 56)   256      conv2d_6             \n____________________________________________________________________________________\nactivation_6 (Activation) (128, 56, 56)   0        batchnorm_6        \n____________________________________________________________________________________\nmaxpool2d_3 (MaxPooling2D)(128, 28, 28)   0        activation_6                \n____________________________________________________________________________________\nconv2d_7 (Convolution2D)  (256, 28, 28)   295168   maxpool2d_3              \n____________________________________________________________________________________\nbatchnorm_7 (BatchNormal  (256, 28, 28)   512      conv2d_7             \n____________________________________________________________________________________\nactivation_7 (Activation) (256, 28, 28)   0        batchnorm_7        \n____________________________________________________________________________________\nconv2d_8 (Convolution2D)  (256, 28, 28)   590080   activation_7                \n____________________________________________________________________________________\nbatchnorm_8 (BatchNormal  (256, 28, 28)   512      conv2d_8             \n____________________________________________________________________________________\nactivation_8 (Activation) (256, 28, 28)   0        batchnorm_8        \n____________________________________________________________________________________\nmaxpool2d_4 (MaxPooling2D)(256, 14, 14)   0        activation_8                \n____________________________________________________________________________________\nconv2d_9 (Convolution2D)  (512, 14, 14)   1180160  maxpool2d_4              \n____________________________________________________________________________________\nbatchnorm_9 (BatchNormal  (512, 14, 14)   1024     conv2d_9             \n____________________________________________________________________________________\nactivation_9 (Activation) (512, 14, 14)   0        batchnorm_9        \n____________________________________________________________________________________\nconv2d_10 (Convolution2D) (512, 14, 14)   2359808  activation_9                \n____________________________________________________________________________________\nbatchnorm_10 (BatchNorma  (512, 14, 14)   1024     conv2d_10            \n____________________________________________________________________________________\nactivation_10 (Activation)(512, 14, 14)   0        batchnorm_10       \n____________________________________________________________________________________\nmaxpool2d_5 (MaxPooling2D)(512, 7, 7)     0        activation_10               \n____________________________________________________________________________________\nconv2d_11 (Convolution2D) (1024, 7, 7)    4719616  maxpool2d_5              \n____________________________________________________________________________________\nbatchnorm_11 (BatchNorma  (1024, 7, 7)    2048     conv2d_11            \n____________________________________________________________________________________\nactivation_11 (Activation)(1024, 7, 7)    0        batchnorm_11       \n____________________________________________________________________________________\nconv2d_12 (Convolution2D) (1024, 7, 7)    9438208  activation_11               \n____________________________________________________________________________________\nbatchnorm_12 (BatchNorma  (1024, 7, 7)    2048     conv2d_12            \n____________________________________________________________________________________\nactivation_12 (Activation)(1024, 7, 7)    0        batchnorm_12       \n____________________________________________________________________________________\nupsamp2d_1 (UpSampling2D) (1024, 14, 14)  0        activation_12               \n____________________________________________________________________________________\nmerge_1 (Merge)           (1536, 14, 14)  0        upsamp2d_1              \n                                                      activation_10               \n____________________________________________________________________________________\nconv2d_13 (Convolution2D) (512, 14, 14)   7078400  merge_1                     \n____________________________________________________________________________________\nbatchnorm_13 (BatchNorma  (512, 14, 14)   1024     conv2d_13            \n____________________________________________________________________________________\nactivation_13 (Activation)(512, 14, 14)   0        batchnorm_13       \n____________________________________________________________________________________\nconv2d_14 (Convolution2D) (512, 14, 14)   2359808  activation_13               \n____________________________________________________________________________________\nbatchnorm_14 (BatchNorma  (512, 14, 14)   1024     conv2d_14            \n____________________________________________________________________________________\nactivation_14 (Activation)(512, 14, 14)   0        batchnorm_14       \n____________________________________________________________________________________\nupsamp2d_2 (UpSampling2D) (512, 28, 28)   0        activation_14               \n____________________________________________________________________________________\nmerge_2 (Merge)           (768, 28, 28)   0        upsamp2d_2              \n                                                                activation_8                \n____________________________________________________________________________________\nconv2d_15 (Convolution2D) (256, 28, 28)   1769728  merge_2                     \n____________________________________________________________________________________\nbatchnorm_15 (BatchNorma  (256, 28, 28)   512      conv2d_15            \n____________________________________________________________________________________\nactivation_15 (Activation)(256, 28, 28)   0        batchnorm_15       \n____________________________________________________________________________________\nconv2d_16 (Convolution2D) (256, 28, 28)   590080   activation_15               \n____________________________________________________________________________________\nbatchnorm_16 (BatchNorma  (256, 28, 28)   512      conv2d_16            \n____________________________________________________________________________________\nactivation_16 (Activation)(256, 28, 28)   0        batchnorm_16       \n____________________________________________________________________________________\nupsamp2d_3 (UpSampling2D) (256, 56, 56)   0        activation_16               \n____________________________________________________________________________________\nmerge_3 (Merge)           (384, 56, 56)   0        upsamp2d_3              \n                                                      activation_6                \n____________________________________________________________________________________\nconv2d_17 (Convolution2D) (128, 56, 56)   442496   merge_3                     \n____________________________________________________________________________________\nbatchnorm_17 (BatchNorma  (128, 56, 56)   256      conv2d_17            \n____________________________________________________________________________________\nactivation_17 (Activation)(128, 56, 56)   0        batchnorm_17       \n____________________________________________________________________________________\nconv2d_18 (Convolution2D) (128, 56, 56)   147584   activation_17               \n____________________________________________________________________________________\nbatchnorm_18 (BatchNorma  (128, 56, 56)   256      conv2d_18            \n____________________________________________________________________________________\nactivation_18 (Activation)(128, 56, 56)   0        batchnorm_18       \n____________________________________________________________________________________\nupsamp2d_4 (UpSampling2D) (128, 112, 112) 0        activation_18               \n____________________________________________________________________________________\nmerge_4 (Merge)           (192, 112, 112) 0        upsamp2d_4              \n                                                      activation_4                \n____________________________________________________________________________________\nconv2d_19 (Convolution2D) (64, 112, 112)  110656   merge_4                     \n____________________________________________________________________________________\nbatchnorm_19 (BatchNorma  (64, 112, 112)  128      conv2d_19            \n____________________________________________________________________________________\nactivation_19 (Activation)(64, 112, 112)  0        batchnorm_19       \n____________________________________________________________________________________\nconv2d_20 (Convolution2D) (64, 112, 112)  36928    activation_19               \n____________________________________________________________________________________\nbatchnorm_20 (BatchNorma  (64, 112, 112)  128      conv2d_20            \n____________________________________________________________________________________\nactivation_20 (Activation)(64, 112, 112)  0        batchnorm_20       \n____________________________________________________________________________________\nupsamp2d_5 (UpSampling2D) (64, 224, 224)  0        activation_20               \n____________________________________________________________________________________\nmerge_5 (Merge)           (96, 224, 224)  0        upsamp2d_5              \n                                                      activation_2                \n____________________________________________________________________________________\nconv2d_21 (Convolution2D) (32, 224, 224)  27680    merge_5                     \n____________________________________________________________________________________\nbatchnorm_21 (BatchNorma  (32, 224, 224)  64       conv2d_21            \n____________________________________________________________________________________\nactivation_21 (Activation)(32, 224, 224)  0        batchnorm_21       \n____________________________________________________________________________________\nconv2d_22 (Convolution2D) (32, 224, 224)  9248     activation_21               \n____________________________________________________________________________________\nbatchnorm_22 (BatchNorma  (32, 224, 224)  64       conv2d_22            \n____________________________________________________________________________________\nactivation_22 (Activation)(32, 224, 224)  0        batchnorm_23       \n____________________________________________________________________________________\nconv2d_23 (Convolution2D) (1, 224, 224)   33       activation_22               \n____________________________________________________________________________________\nbatchnorm_23 (BatchNorma  (1, 224, 224)   2        conv2d_23            \n____________________________________________________________________________________\nactivation_23 (Activation)(1, 224, 224)   0        batchnorm_23       \n====================================================================================\nTotal params: 31459619\n</code></pre>\n\n<p>LOSS Function:</p>\n\n<pre><code>def jacard_coef(y_true, y_pred):\n    y_true_f = K.flatten(y_true)\n    y_pred_f = K.flatten(y_pred)\n    intersection = K.sum(y_true_f * y_pred_f)\n    return (intersection + 1.0) / (K.sum(y_true_f) + K.sum(y_pred_f) - intersection + 1.0)\n</code></pre>\n\n<p>Typical example of training process: Class 1 (Fold 2). Parameters: lr=0.05, optim=SGD, rotation=False, dropout=enabled, UNET version=BatchNorm, patience=8, samples train: 9997, samples valid: 4013</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6031/Training.png\" alt=\"Typical example of training process\" title=\"\"></p>\n\n<h2>Tuning models and validation</h2>\n\n<p>The main problem was that training process is very unstable. Sometimes it can go to local minimum or start predicting as mask full image etc. So most of the time I spend on finding optimal parameters to maximize validation score. It was the most time consuming part. I used grid search with following parameters: </p>\n\n<ul>\n<li>Optimizer (Adam, SGD) </li>\n<li>LR SGD: (0.05, 0.01, 0.001) </li>\n<li>LR Adam: (0.01, 0.001, 0.0001) </li>\n<li>Rotation (enabled, disabled) </li>\n<li>Type of model: UNET 224x224, UNET with dropout 0.1, UNET with Batch Normalization </li>\n<li>Number of samples per epoch (Fraction from ½ up to 1)</li>\n</ul>\n\n<p>Due to limitation of computational power I checked only some of parameters. \nAt the end I use mostly Adam optimizer with BatchNormalization version of UNET with rotation enabled. Just tune learning rate.\nI use loss function value to stop training with early stopping with patience 8-15 epochs. After I obtain all 5 Folds models I had the predicted segmentation for each of train images. So I was able to predict score for full image set using the same method as it was made on Kaggle Leaderboard. Obtained score was very representative, so increasing the score locally almost always leads to increasing the score on leaderboard. Also I was able to find optimal threshold value for heatmap with this method.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6032/LB-Scores.png\" alt=\"LS vs LB Scores\" title=\"\"></p>\n\n<h2>Processing the test data</h2>\n\n<p>Test set consists of 429 images. Each test image was processed separately. Each image was resized and normalized the same way as training data to 20х3360х3360 tensor. </p>\n\n<p>In the beginning I create two zero arrays HEATMAP and COUNT with 3360х3360 shape. Then use sliding window approach to predict segmentation. All 5 folds predict on image extracted from current position and acquired probabilities sum up to HEATMAP array. COUNT array at image position added by 1. At the end HEATMAP is divided by COUNT to get real heatmap array. After this heatmap was thresholded at value acquired from validation, typically 0.5. On basis of this 2D-array I created polygons with rasterio and shapely.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6033/Sliding%20Window.png\" alt=\"Sliding window\" title=\"\"></p>\n\n<p>Additional ideas to increase the accuracy provided below. Not all of them was used in final submit because processing of test images was very long process, around 8-10 hours for whole test set.</p>\n\n<ol>\n<li><p>Decreasing sliding step from 112px to 56px and below always increased the accuracy of predictions, but increase the computational complexity as O(N^2). </p>\n\n<ul><li>For example class 4 on validation 0.385714 vs 0.403683.</li></ul></li>\n<li><p>It looks like UNET predict worse on the edges, so it’s good idea to use only central part for prediction. Example for class 5 below. </p>\n\n<ul><li>Default score: 0.507446 </li>\n<li>200x200 center part from 224x224 square + 100 px sliding window step: 0.512844 </li>\n<li>160x160 center part from 224x224 square + 80 px sliding window step: 0.514557</li></ul></li>\n<li><p>I usually used threshold 0.5 without checking optimum on validation, but this can be useful too. </p>\n\n<ul><li>Class 6 example: </li>\n<li>THR 0.1: 0.738943 </li>\n<li>THR 0.5: 0.757858 </li>\n<li>THR 0.9: 0.759850</li></ul></li>\n</ol>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6039/Masks_700px.png\" alt=\"Masks for class 5\" title=\"\"></p>\n\n<h2>Ensembles</h2>\n\n<p>The most obvious ways to do ensembles is to use heatmaps. But out team consisted of 2 people was merged at the last stage of competitions, so we actually don’t have heatmaps for all classes at this stage. So we made ensembles directly on polygons, using UNION and INTERSECTION from shapely. We have different training process and models, so for some classes we had good boost after merge of this type. </p>\n\n<pre><code>Class 6 LB: 0.08149 + LB: 0.08103 (Intersection) Score: 0.08179\nClass 8 LB: 0.03890 + LB: 0.05322 (Intersection) Score: 0.06194\nClass 9 LB: 0.02113 + LB: 0.02713 (Union) Score: 0.03254\n</code></pre>\n\n<p>Bad thing, we mostly use leaderboard to check if our ensemble gave boost, which could lead to overfitting. It would be much easier if we have 3rd independent solution to use voting mechanism.</p>\n\n<h2>Post processing</h2>\n\n<p>After analysis of class interactions in train I made the following postprocess types:\n- Remove too large polygons (from class 8, 9 and 10 for example) \n- Subtract predicted classes polygons like water from car classes</p>\n\n<h2>Class 10 with small objects</h2>\n\n<p>For class 10 default approach works not very well. The main problems with this class that it’s have very small objects, bad train segmentation and it’s total area is too small comparing to full area. So I created other UNET CNN for this class. It had input of 32x32 pixels. Data for this class was heavily augmented, rotations, different shifts etc. I also used loss function with big penalty for false positives based on <a href=\"https://en.wikipedia.org/wiki/Tversky_index\">Tversky index</a>.</p>\n\n<pre><code># https://en.wikipedia.org/wiki/Tversky_index\ndef tversky_coef(y_true, y_pred):\n    y_true_f = K.flatten(y_true)\n    y_pred_f = K.flatten(y_pred)\n    alfa = 0.1\n    false_positive = K.sum(y_pred_f * (1 - y_true_f))\n    false_negative = K.sum((1 - y_pred_f) * y_true_f)\n    true_positive = K.sum(y_true_f * y_pred_f)\n    return (true_positive + 1.0) / (false_negative*alfa + (1 - alfa) * false_positive + true_positive + 1.0)\n</code></pre>\n\n<p>It allowed me to get 0.00362 score on LB.</p>\n\n<h2>Other ideas I tried</h2>\n\n<ul>\n<li>At the early beginning I made single XGBoost model which works on squares of size 10x10. Model tries to predict to which class given square is belong. It easily got me higher than baseline. Probably creating independent XGBoost models for each class could give good results. But I switch to CNN without additional experiments with XGBoost approach.</li>\n<li>Pre-trained VGG16 for rare class localization. I tried to predict if car exists in given area with VGG16 on RGB images for later ensemble with UNET predictions.</li>\n</ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6040/vgg16_test_6060_3_2_cls_10_700px.png\" alt=\"Example of class 10 localization with VGG16\" title=\"\"></p>\n\n<ul>\n<li>Pre-trained VGG16 for segmentation. The main idea here is that we use final layer (which is now sigmoid instead of softmax) as indication in which part of 224x224 image given class exists. I used 16x16 = 256 neurons for this task. But result wasn’t very good, may be because of low resolution - 14x14 square as single pixel.</li>\n<li>There are big bunch of different indexes used for automatic segmentation without CNN:\n<a href=\"http://www.indexdatabase.de/db/i.php\">http://www.indexdatabase.de/db/i.php</a>\nIt works great for water as shown in <a href=\"https://www.kaggle.com/resolut/dstl-satellite-imagery-feature-detection/waterway-0-095-lb\">Waterway Kernel</a>. And these indexes can be added as additional planes during learning process. Unfortunately they didn’t help much in my experiments, so I abandon them. </li>\n<li>I tried to use CRF (Conditional random field) methods with <a href=\"https://pystruct.github.io/\">pystruct</a> to improve predicted masks. But it somehow won’t work for me.</li>\n</ul>\n\n<h2>Independent classes best scores</h2>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6036/Best-LB.png\" alt=\"Best Public LB\" title=\"\"></p>\n\n<h2>Segmentation example</h2>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6048/6100_1_2_mask.png\" alt=\"enter image description here\" title=\"\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/166431/6049/6100_1_2_proj_700px.jpg\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>Checkout the video:\n<strong><a href=\"https://www.youtube.com/watch?v=rpp7ZhGb1IQ\">https://www.youtube.com/watch?v=rpp7ZhGb1IQ</a></strong></p>\n\n<h2>Code</h2>\n\n<p>GitHUB link here later. I plan to release code after publication of final results.</p>",
      "rawMarkdown": "To solve this problem I used: Python + Keras 1.0.8 + Theano under Windows 10. As hardware I had NVIDIA GTX 980 8 GB. I used only one GPU which worked almost 24h a day during competition. After join the team I used 2 additional GPUs (TITAN 12GB).\nIn short my solution can be described like this:\n\n 1. My main CNN is modified UNET with input shape (20, 224, 224). It\n    has higher depth, with added batch normalization and dropout layers.\n    As loss function is used Jacquard Coefficient. I used grid search\n    with different parameters and choose the best model with highest\n    validation score. \n 2. I create separate models for each class (so 10\n    independent models) and tuned them independently. I actually think\n    that I lost some information about class interaction this way, but\n    it was much easier to tune models. I partially fix it with\n    postprocessing step. \n 3. Each model is actually set of K different\n    weights obtained with KFold. For most classes I used 5 KFold. For\n    class 7 - 2 KFold. For class 9 – 4 KFold. I split train set by image\n    ID (there were only 25 images). I made it once by hands\n    independently for each class. Each fold contains the same number of\n    images with existed class. So it’s actually stratified split.\n\nData preprocessing\n------------------\nEach “image object” (with particular img_id) had set of images made with different wave lengths. In total 20 channels. I resized all images to 3360x3360 pixels (3360 = 15*224 and 3360 is closest to panchromatic image resolution).  And join them along the axis. So at the end I had tensor of following shape: 20x3360x3360. I created the pixel masks from given polygons with same size: 3360x3360 pixels.\n\n![Set of images example][1]\n\n![Mask example][2]\n\nI decided to use UNET with input shape 20x224x224. I choose 224 for 2 reasons:\n\n1. 224 = 2*2*2*2*2*7 – I have 5 MaxPooling layers in UNET and on lowest layer has 7x7 pixel size. Which in my experience the best.\n2. The same size I could use with pretrained VGG16 or ResNet.\n\nThe next step was made independently for each class:\n\n 1. Split input images on 15*15 parts forming tensors of size\n    20x224x224. \n 2.\tBecause of class imbalances we need to increase\n    number of cases where mask exists. For some cases like class 10\n    (with vehicles) I have around 99% of empty masks after first step.\n    So I added more cases using sliding window around non-zero mask\n    points.\n\n\nFinal statistics for classes:\n\n    Number of tests for class 1: 14010. Empty files: 4431 Percent: 31.62%\n    Number of tests for class 2: 9835. Empty files: 3997 Percent: 40.64%\n    Number of tests for class 3: 20548. Empty files: 5238 Percent: 25.49%\n    Number of tests for class 4: 12411. Empty files: 2976 Percent: 23.97%\n    Number of tests for class 5: 10161. Empty files: 294 Percent: 2.89%\n    Number of tests for class 6: 10716. Empty files: 3720 Percent: 34.71%\n    Number of tests for class 7: 10328. Empty files: 5504 Percent: 53.29%\n    Number of tests for class 8: 11349. Empty files: 5469 Percent: 48.18%\n    Number of tests for class 9: 12035. Empty files: 5574 Percent: 46.31%\n    Number of tests for class 10: 16173. Empty files: 5336 Percent: 32.99%\n\nUNET requires the distribution to be close to normal. Ranges for different channels was: P_3, P_P, P_M – 2048, P_A – 16384. At first I divide every channel to its maximum possible value. I’ve seen on forum that some people removed some values from the both ends of histogram but I didn’t do it.\n\nThen I needed to calculate mean and stdev. I calculated it for each channel independently, but at the end used the same mean and stdev for whole tensor: \nmean = 0.219613\nstdev = 0.110741\n\nCreating models\n---------------\n![Modified UNET][3]\n\n    ____________________________________________________________________________________\n    Layer (type)             Output Shape   Param #     Connected to\n    ====================================================================================\n    input_1 (InputLayer)      (20, 224, 224)  0                                 \n    ____________________________________________________________________________________\n    conv2d_1 (Convolution2D)  (32, 224, 224)  5792     input_1\n    ____________________________________________________________________________________\n    batchnorm_1 (BatchNormal  (32, 224, 224)  64       conv2d_1             \n    ____________________________________________________________________________________\n    activation_1 (Activation) (32, 224, 224)  0        batchnorm_1        \n    ____________________________________________________________________________________\n    conv2d_2 (Convolution2D)  (32, 224, 224)  9248     activation_1                \n    ____________________________________________________________________________________\n    batchnorm_2 (BatchNormal  (32, 224, 224)  64       conv2d_2             \n    ____________________________________________________________________________________\n    activation_2 (Activation) (32, 224, 224)  0        batchnorm_2        \n    ____________________________________________________________________________________\n    maxpool2d_1 (MaxPooling2D)(32, 112, 112)  0        activation_2                \n    ____________________________________________________________________________________\n    conv2d_3 (Convolution2D)  (64, 112, 112)  18496    maxpool2d_1              \n    ____________________________________________________________________________________\n    batchnorm_3 (BatchNormal  (64, 112, 112)  128      conv2d_3             \n    ____________________________________________________________________________________\n    activation_3 (Activation) (64, 112, 112)  0        batchnorm_3        \n    ____________________________________________________________________________________\n    conv2d_4 (Convolution2D)  (64, 112, 112)  36928    activation_3                \n    ____________________________________________________________________________________\n    batchnorm_4 (BatchNormal  (64, 112, 112)  128      conv2d_4             \n    ____________________________________________________________________________________\n    activation_4 (Activation) (64, 112, 112)  0        batchnorm_4        \n    ____________________________________________________________________________________\n    maxpool2d_2 (MaxPooling2D)(64, 56, 56)    0        activation_4                \n    ____________________________________________________________________________________\n    conv2d_5 (Convolution2D)  (128, 56, 56)   73856    maxpool2d_2              \n    ____________________________________________________________________________________\n    batchnorm_5 (BatchNormal  (128, 56, 56)   256      conv2d_5             \n    ____________________________________________________________________________________\n    activation_5 (Activation) (128, 56, 56)   0        batchnorm_5        \n    ____________________________________________________________________________________\n    conv2d_6 (Convolution2D)  (128, 56, 56)   147584   activation_5                \n    ____________________________________________________________________________________\n    batchnorm_6 (BatchNormal  (128, 56, 56)   256      conv2d_6             \n    ____________________________________________________________________________________\n    activation_6 (Activation) (128, 56, 56)   0        batchnorm_6        \n    ____________________________________________________________________________________\n    maxpool2d_3 (MaxPooling2D)(128, 28, 28)   0        activation_6                \n    ____________________________________________________________________________________\n    conv2d_7 (Convolution2D)  (256, 28, 28)   295168   maxpool2d_3              \n    ____________________________________________________________________________________\n    batchnorm_7 (BatchNormal  (256, 28, 28)   512      conv2d_7             \n    ____________________________________________________________________________________\n    activation_7 (Activation) (256, 28, 28)   0        batchnorm_7        \n    ____________________________________________________________________________________\n    conv2d_8 (Convolution2D)  (256, 28, 28)   590080   activation_7                \n    ____________________________________________________________________________________\n    batchnorm_8 (BatchNormal  (256, 28, 28)   512      conv2d_8             \n    ____________________________________________________________________________________\n    activation_8 (Activation) (256, 28, 28)   0        batchnorm_8        \n    ____________________________________________________________________________________\n    maxpool2d_4 (MaxPooling2D)(256, 14, 14)   0        activation_8                \n    ____________________________________________________________________________________\n    conv2d_9 (Convolution2D)  (512, 14, 14)   1180160  maxpool2d_4              \n    ____________________________________________________________________________________\n    batchnorm_9 (BatchNormal  (512, 14, 14)   1024     conv2d_9             \n    ____________________________________________________________________________________\n    activation_9 (Activation) (512, 14, 14)   0        batchnorm_9        \n    ____________________________________________________________________________________\n    conv2d_10 (Convolution2D) (512, 14, 14)   2359808  activation_9                \n    ____________________________________________________________________________________\n    batchnorm_10 (BatchNorma  (512, 14, 14)   1024     conv2d_10            \n    ____________________________________________________________________________________\n    activation_10 (Activation)(512, 14, 14)   0        batchnorm_10       \n    ____________________________________________________________________________________\n    maxpool2d_5 (MaxPooling2D)(512, 7, 7)     0        activation_10               \n    ____________________________________________________________________________________\n    conv2d_11 (Convolution2D) (1024, 7, 7)    4719616  maxpool2d_5              \n    ____________________________________________________________________________________\n    batchnorm_11 (BatchNorma  (1024, 7, 7)    2048     conv2d_11            \n    ____________________________________________________________________________________\n    activation_11 (Activation)(1024, 7, 7)    0        batchnorm_11       \n    ____________________________________________________________________________________\n    conv2d_12 (Convolution2D) (1024, 7, 7)    9438208  activation_11               \n    ____________________________________________________________________________________\n    batchnorm_12 (BatchNorma  (1024, 7, 7)    2048     conv2d_12            \n    ____________________________________________________________________________________\n    activation_12 (Activation)(1024, 7, 7)    0        batchnorm_12       \n    ____________________________________________________________________________________\n    upsamp2d_1 (UpSampling2D) (1024, 14, 14)  0        activation_12               \n    ____________________________________________________________________________________\n    merge_1 (Merge)           (1536, 14, 14)  0        upsamp2d_1              \n                                                          activation_10               \n    ____________________________________________________________________________________\n    conv2d_13 (Convolution2D) (512, 14, 14)   7078400  merge_1                     \n    ____________________________________________________________________________________\n    batchnorm_13 (BatchNorma  (512, 14, 14)   1024     conv2d_13            \n    ____________________________________________________________________________________\n    activation_13 (Activation)(512, 14, 14)   0        batchnorm_13       \n    ____________________________________________________________________________________\n    conv2d_14 (Convolution2D) (512, 14, 14)   2359808  activation_13               \n    ____________________________________________________________________________________\n    batchnorm_14 (BatchNorma  (512, 14, 14)   1024     conv2d_14            \n    ____________________________________________________________________________________\n    activation_14 (Activation)(512, 14, 14)   0        batchnorm_14       \n    ____________________________________________________________________________________\n    upsamp2d_2 (UpSampling2D) (512, 28, 28)   0        activation_14               \n    ____________________________________________________________________________________\n    merge_2 (Merge)           (768, 28, 28)   0        upsamp2d_2              \n                                                                    activation_8                \n    ____________________________________________________________________________________\n    conv2d_15 (Convolution2D) (256, 28, 28)   1769728  merge_2                     \n    ____________________________________________________________________________________\n    batchnorm_15 (BatchNorma  (256, 28, 28)   512      conv2d_15            \n    ____________________________________________________________________________________\n    activation_15 (Activation)(256, 28, 28)   0        batchnorm_15       \n    ____________________________________________________________________________________\n    conv2d_16 (Convolution2D) (256, 28, 28)   590080   activation_15               \n    ____________________________________________________________________________________\n    batchnorm_16 (BatchNorma  (256, 28, 28)   512      conv2d_16            \n    ____________________________________________________________________________________\n    activation_16 (Activation)(256, 28, 28)   0        batchnorm_16       \n    ____________________________________________________________________________________\n    upsamp2d_3 (UpSampling2D) (256, 56, 56)   0        activation_16               \n    ____________________________________________________________________________________\n    merge_3 (Merge)           (384, 56, 56)   0        upsamp2d_3              \n                                                          activation_6                \n    ____________________________________________________________________________________\n    conv2d_17 (Convolution2D) (128, 56, 56)   442496   merge_3                     \n    ____________________________________________________________________________________\n    batchnorm_17 (BatchNorma  (128, 56, 56)   256      conv2d_17            \n    ____________________________________________________________________________________\n    activation_17 (Activation)(128, 56, 56)   0        batchnorm_17       \n    ____________________________________________________________________________________\n    conv2d_18 (Convolution2D) (128, 56, 56)   147584   activation_17               \n    ____________________________________________________________________________________\n    batchnorm_18 (BatchNorma  (128, 56, 56)   256      conv2d_18            \n    ____________________________________________________________________________________\n    activation_18 (Activation)(128, 56, 56)   0        batchnorm_18       \n    ____________________________________________________________________________________\n    upsamp2d_4 (UpSampling2D) (128, 112, 112) 0        activation_18               \n    ____________________________________________________________________________________\n    merge_4 (Merge)           (192, 112, 112) 0        upsamp2d_4              \n                                                          activation_4                \n    ____________________________________________________________________________________\n    conv2d_19 (Convolution2D) (64, 112, 112)  110656   merge_4                     \n    ____________________________________________________________________________________\n    batchnorm_19 (BatchNorma  (64, 112, 112)  128      conv2d_19            \n    ____________________________________________________________________________________\n    activation_19 (Activation)(64, 112, 112)  0        batchnorm_19       \n    ____________________________________________________________________________________\n    conv2d_20 (Convolution2D) (64, 112, 112)  36928    activation_19               \n    ____________________________________________________________________________________\n    batchnorm_20 (BatchNorma  (64, 112, 112)  128      conv2d_20            \n    ____________________________________________________________________________________\n    activation_20 (Activation)(64, 112, 112)  0        batchnorm_20       \n    ____________________________________________________________________________________\n    upsamp2d_5 (UpSampling2D) (64, 224, 224)  0        activation_20               \n    ____________________________________________________________________________________\n    merge_5 (Merge)           (96, 224, 224)  0        upsamp2d_5              \n                                                          activation_2                \n    ____________________________________________________________________________________\n    conv2d_21 (Convolution2D) (32, 224, 224)  27680    merge_5                     \n    ____________________________________________________________________________________\n    batchnorm_21 (BatchNorma  (32, 224, 224)  64       conv2d_21            \n    ____________________________________________________________________________________\n    activation_21 (Activation)(32, 224, 224)  0        batchnorm_21       \n    ____________________________________________________________________________________\n    conv2d_22 (Convolution2D) (32, 224, 224)  9248     activation_21               \n    ____________________________________________________________________________________\n    batchnorm_22 (BatchNorma  (32, 224, 224)  64       conv2d_22            \n    ____________________________________________________________________________________\n    activation_22 (Activation)(32, 224, 224)  0        batchnorm_23       \n    ____________________________________________________________________________________\n    conv2d_23 (Convolution2D) (1, 224, 224)   33       activation_22               \n    ____________________________________________________________________________________\n    batchnorm_23 (BatchNorma  (1, 224, 224)   2        conv2d_23            \n    ____________________________________________________________________________________\n    activation_23 (Activation)(1, 224, 224)   0        batchnorm_23       \n    ====================================================================================\n    Total params: 31459619\n\n\nLOSS Function:\n\n    def jacard_coef(y_true, y_pred):\n        y_true_f = K.flatten(y_true)\n        y_pred_f = K.flatten(y_pred)\n        intersection = K.sum(y_true_f * y_pred_f)\n        return (intersection + 1.0) / (K.sum(y_true_f) + K.sum(y_pred_f) - intersection + 1.0)\n\nTypical example of training process: Class 1 (Fold 2). Parameters: lr=0.05, optim=SGD, rotation=False, dropout=enabled, UNET version=BatchNorm, patience=8, samples train: 9997, samples valid: 4013\n\n![Typical example of training process][4]\n\nTuning models and validation\n----------------------------\nThe main problem was that training process is very unstable. Sometimes it can go to local minimum or start predicting as mask full image etc. So most of the time I spend on finding optimal parameters to maximize validation score. It was the most time consuming part. I used grid search with following parameters: \n\n - Optimizer (Adam, SGD) \n - LR SGD: (0.05, 0.01, 0.001) \n - LR Adam: (0.01, 0.001, 0.0001) \n - Rotation (enabled, disabled) \n - Type of model: UNET 224x224, UNET with dropout 0.1, UNET with Batch Normalization \n - Number of samples per epoch (Fraction from ½ up to 1)\n\nDue to limitation of computational power I checked only some of parameters. \nAt the end I use mostly Adam optimizer with BatchNormalization version of UNET with rotation enabled. Just tune learning rate.\nI use loss function value to stop training with early stopping with patience 8-15 epochs. After I obtain all 5 Folds models I had the predicted segmentation for each of train images. So I was able to predict score for full image set using the same method as it was made on Kaggle Leaderboard. Obtained score was very representative, so increasing the score locally almost always leads to increasing the score on leaderboard. Also I was able to find optimal threshold value for heatmap with this method.\n\n![LS vs LB Scores][5]\n\nProcessing the test data\n------------------------\nTest set consists of 429 images. Each test image was processed separately. Each image was resized and normalized the same way as training data to 20х3360х3360 tensor. \n\nIn the beginning I create two zero arrays HEATMAP and COUNT with 3360х3360 shape. Then use sliding window approach to predict segmentation. All 5 folds predict on image extracted from current position and acquired probabilities sum up to HEATMAP array. COUNT array at image position added by 1. At the end HEATMAP is divided by COUNT to get real heatmap array. After this heatmap was thresholded at value acquired from validation, typically 0.5. On basis of this 2D-array I created polygons with rasterio and shapely.\n\n![Sliding window][6]\n\nAdditional ideas to increase the accuracy provided below. Not all of them was used in final submit because processing of test images was very long process, around 8-10 hours for whole test set.\n\n\n1.\tDecreasing sliding step from 112px to 56px and below always increased the accuracy of predictions, but increase the computational complexity as O(N^2). \n- For example class 4 on validation 0.385714 vs 0.403683.\n\n2.\tIt looks like UNET predict worse on the edges, so it’s good idea to use only central part for prediction. Example for class 5 below. \n- Default score: 0.507446 \n- 200x200 center part from 224x224 square + 100 px sliding window step: 0.512844 \n- 160x160 center part from 224x224 square + 80 px sliding window step: 0.514557\n\n3.\tI usually used threshold 0.5 without checking optimum on validation, but this can be useful too. \n- Class 6 example: \n- THR 0.1: 0.738943 \n- THR 0.5: 0.757858 \n- THR 0.9: 0.759850\n\n![Masks for class 5][7]\n\nEnsembles\n---------\nThe most obvious ways to do ensembles is to use heatmaps. But out team consisted of 2 people was merged at the last stage of competitions, so we actually don’t have heatmaps for all classes at this stage. So we made ensembles directly on polygons, using UNION and INTERSECTION from shapely. We have different training process and models, so for some classes we had good boost after merge of this type. \n\n    Class 6 LB: 0.08149 + LB: 0.08103 (Intersection) Score: 0.08179\n    Class 8 LB: 0.03890 + LB: 0.05322 (Intersection) Score: 0.06194\n    Class 9 LB: 0.02113 + LB: 0.02713 (Union) Score: 0.03254\n\nBad thing, we mostly use leaderboard to check if our ensemble gave boost, which could lead to overfitting. It would be much easier if we have 3rd independent solution to use voting mechanism.\n\nPost processing\n---------------\nAfter analysis of class interactions in train I made the following postprocess types:\n- Remove too large polygons (from class 8, 9 and 10 for example) \n- Subtract predicted classes polygons like water from car classes\n\nClass 10 with small objects\n---------------------------\n\nFor class 10 default approach works not very well. The main problems with this class that it’s have very small objects, bad train segmentation and it’s total area is too small comparing to full area. So I created other UNET CNN for this class. It had input of 32x32 pixels. Data for this class was heavily augmented, rotations, different shifts etc. I also used loss function with big penalty for false positives based on [Tversky index][8].\n\n    # https://en.wikipedia.org/wiki/Tversky_index\n    def tversky_coef(y_true, y_pred):\n        y_true_f = K.flatten(y_true)\n        y_pred_f = K.flatten(y_pred)\n        alfa = 0.1\n        false_positive = K.sum(y_pred_f * (1 - y_true_f))\n        false_negative = K.sum((1 - y_pred_f) * y_true_f)\n        true_positive = K.sum(y_true_f * y_pred_f)\n        return (true_positive + 1.0) / (false_negative*alfa + (1 - alfa) * false_positive + true_positive + 1.0)\n\nIt allowed me to get 0.00362 score on LB.\n\nOther ideas I tried\n-------------------\n- At the early beginning I made single XGBoost model which works on squares of size 10x10. Model tries to predict to which class given square is belong. It easily got me higher than baseline. Probably creating independent XGBoost models for each class could give good results. But I switch to CNN without additional experiments with XGBoost approach.\n- Pre-trained VGG16 for rare class localization. I tried to predict if car exists in given area with VGG16 on RGB images for later ensemble with UNET predictions.\n\n![Example of class 10 localization with VGG16][9]\n\n- Pre-trained VGG16 for segmentation. The main idea here is that we use final layer (which is now sigmoid instead of softmax) as indication in which part of 224x224 image given class exists. I used 16x16 = 256 neurons for this task. But result wasn’t very good, may be because of low resolution - 14x14 square as single pixel.\n- There are big bunch of different indexes used for automatic segmentation without CNN:\nhttp://www.indexdatabase.de/db/i.php\nIt works great for water as shown in [Waterway Kernel][10]. And these indexes can be added as additional planes during learning process. Unfortunately they didn’t help much in my experiments, so I abandon them. \n- I tried to use CRF (Conditional random field) methods with [pystruct][11] to improve predicted masks. But it somehow won’t work for me.\n\nIndependent classes best scores\n-------------------------------\n![Best Public LB][12]\n\nSegmentation example\n--------------------\n![enter image description here][13]\n![enter image description here][14]\n\nCheckout the video:\n**[https://www.youtube.com/watch?v=rpp7ZhGb1IQ][15]**\n\nCode\n----\n\nGitHUB link here later. I plan to release code after publication of final results.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6037/Tensor_example_700px.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6029/Mask_example.png\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6038/model_700px.png\n  [4]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6031/Training.png\n  [5]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6032/LB-Scores.png\n  [6]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6033/Sliding%20Window.png\n  [7]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6039/Masks_700px.png\n  [8]: https://en.wikipedia.org/wiki/Tversky_index\n  [9]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6040/vgg16_test_6060_3_2_cls_10_700px.png\n  [10]: https://www.kaggle.com/resolut/dstl-satellite-imagery-feature-detection/waterway-0-095-lb\n  [11]: https://pystruct.github.io/\n  [12]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6036/Best-LB.png\n  [13]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6048/6100_1_2_mask.png\n  [14]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6049/6100_1_2_proj_700px.jpg\n  [15]: https://www.youtube.com/watch?v=rpp7ZhGb1IQ",
      "votes": null
    },
    {
      "id": "166458",
      "postDate": "03/09/2017 18:55:07",
      "content": "<p>Great write up and solution! For crf I tried using <a href=\"https://github.com/lucasb-eyer/pydensecrf\">https://github.com/lucasb-eyer/pydensecrf</a> which is a partial wrapper for the crf used in deep lab, it showed improvement for validation scores, but would have taken a little too long for all the data at panchromatic resolution on my hardware. Here is dense crf paper <a href=\"https://arxiv.org/abs/1210.5644\">https://arxiv.org/abs/1210.5644</a>. </p>\n\n<p>With respect to edges. I used valid border mode instead of same, and it removes the problem with borders. The original unet paper says they used valid borders as well.  </p>",
      "rawMarkdown": "Great write up and solution! For crf I tried using https://github.com/lucasb-eyer/pydensecrf which is a partial wrapper for the crf used in deep lab, it showed improvement for validation scores, but would have taken a little too long for all the data at panchromatic resolution on my hardware. Here is dense crf paper https://arxiv.org/abs/1210.5644. \n\nWith respect to edges. I used valid border mode instead of same, and it removes the problem with borders. The original unet paper says they used valid borders as well.",
      "votes": null
    },
    {
      "id": "166460",
      "postDate": "03/09/2017 19:06:35",
      "content": "<p>Actually I had the same issue - very long run time for full bunch of my train images. And CRF run on small subset lead to bad result. May be the link you provided is better for this task.</p>",
      "rawMarkdown": "Actually I had the same issue - very long run time for full bunch of my train images. And CRF run on small subset lead to bad result. May be the link you provided is better for this task.",
      "votes": null
    },
    {
      "id": "166542",
      "postDate": "03/10/2017 02:56:23",
      "content": "<p>Why do you just split the image to 15*15 but not randomly sample from the image ?</p>",
      "rawMarkdown": "Why do you just split the image to 15*15 but not randomly sample from the image ?",
      "votes": null
    },
    {
      "id": "166592",
      "postDate": "03/10/2017 08:38:42",
      "content": "<p>Good question. If made straigtforward random sample from image, then it could lead to bad result due to class imbalances. Most of batches will contain small ammount of predicted class. What is worse in some cases batch won't contain predicted class. Probably using pregenerated positive cases in union with random sample from images will work better than my solution. But I think my solution is ok, due to large variance of samples.\nThe other good thing about my approach: </p>\n\n<p>1) I generated samples one time with balanced classes and can use them \"as is\" just with random() function.</p>\n\n<p>2) Reading all train images in memory (to random sample from them) requires ~32 GB of RAM. My current solution don't require big ammount of memory. But require SSD because of heavy reading from disk.</p>",
      "rawMarkdown": "Good question. If made straigtforward random sample from image, then it could lead to bad result due to class imbalances. Most of batches will contain small ammount of predicted class. What is worse in some cases batch won't contain predicted class. Probably using pregenerated positive cases in union with random sample from images will work better than my solution. But I think my solution is ok, due to large variance of samples.\nThe other good thing about my approach: \n\n1) I generated samples one time with balanced classes and can use them \"as is\" just with random() function.\n\n2) Reading all train images in memory (to random sample from them) requires ~32 GB of RAM. My current solution don't require big ammount of memory. But require SSD because of heavy reading from disk.",
      "votes": null
    },
    {
      "id": "166684",
      "postDate": "03/10/2017 16:34:11",
      "content": "<ol>\n<li>I got better results when I started using batch with 128 samples, I would believe exactly due to the reasons that you mentioned.</li>\n<li>One may save train data in float16 format =&gt; ~13Gb for the whole train set.</li>\n</ol>",
      "rawMarkdown": "1. I got better results when I started using batch with 128 samples, I would believe exactly due to the reasons that you mentioned.\n 2. One may save train data in float16 format => ~13Gb for the whole train set.",
      "votes": null
    },
    {
      "id": "166686",
      "postDate": "03/10/2017 16:38:13",
      "content": "<p>Hi, <br>\nDid you try to use dropout? I tried in some of the classes but surprisingly the cross-validation score was lower...</p>",
      "rawMarkdown": "Hi,   \nDid you try to use dropout? I tried in some of the classes but surprisingly the cross-validation score was lower...",
      "votes": null
    },
    {
      "id": "166704",
      "postDate": "03/10/2017 17:16:55",
      "content": "<p>I couldn't use big batches because I choose to big image as input. My batch was 22 - higher batches didn't fit in memory. Probably if I'd start from scratch I'd use 112x112 input.</p>\n\n<p>I afraid of float16 )</p>",
      "rawMarkdown": "I couldn't use big batches because I choose to big image as input. My batch was 22 - higher batches didn't fit in memory. Probably if I'd start from scratch I'd use 112x112 input.\n\nI afraid of float16 )",
      "votes": null
    },
    {
      "id": "166706",
      "postDate": "03/10/2017 17:18:46",
      "content": "<p>Yes I used it (Dropout(0.1)) in my latest models. I can't say for sure if it was good or not, but my score became better.</p>",
      "rawMarkdown": "Yes I used it (Dropout(0.1)) in my latest models. I can't say for sure if it was good or not, but my score became better.",
      "votes": null
    },
    {
      "id": "168194",
      "postDate": "03/16/2017 14:59:28",
      "content": "<p>@ZFTurbo, Would you please publish final code for our learning ?</p>",
      "rawMarkdown": "ZFTurbo, Would you please publish final code for our learning ?",
      "votes": null
    },
    {
      "id": "168238",
      "postDate": "03/16/2017 17:24:58",
      "content": "<p>I made video with segmentation examples:</p>\n\n<p><strong><a href=\"https://www.youtube.com/watch?v=rpp7ZhGb1IQ\">https://www.youtube.com/watch?v=rpp7ZhGb1IQ</a></strong></p>",
      "rawMarkdown": "I made video with segmentation examples:\n\n**[https://www.youtube.com/watch?v=rpp7ZhGb1IQ][1]**\n\n\n  [1]: https://www.youtube.com/watch?v=rpp7ZhGb1IQ",
      "votes": null
    },
    {
      "id": "168240",
      "postDate": "03/16/2017 17:26:32",
      "content": "<p>Currently preparing the code.</p>",
      "rawMarkdown": "Currently preparing the code.",
      "votes": null
    },
    {
      "id": "170186",
      "postDate": "03/24/2017 14:05:37",
      "content": "<p>@ZFTurbo Would you pls share the final code ?</p>",
      "rawMarkdown": "ZFTurbo Would you pls share the final code ?",
      "votes": null
    },
    {
      "id": "170192",
      "postDate": "03/24/2017 14:21:07",
      "content": "<p>I asked kaggle support about sharing the code and they said as a prize winner I can't publish it. Sorry. (</p>",
      "rawMarkdown": "I asked kaggle support about sharing the code and they said as a prize winner I can't publish it. Sorry. (",
      "votes": null
    },
    {
      "id": "170214",
      "postDate": "03/24/2017 16:18:04",
      "content": "<p>@ZFTurbo ,\n really? i have seen the prize winners sharing the code like by 1st prize winners deepsense.io in Right Whale Recognition competition. Strange. That's not good by the kaggle support.</p>",
      "rawMarkdown": "ZFTurbo ,\n really? i have seen the prize winners sharing the code like by 1st prize winners deepsense.io in Right Whale Recognition competition. Strange. That's not good by the kaggle support.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 166458,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "03/09/2017 18:55:07",
      "content": "<p>Great write up and solution! For crf I tried using <a href=\"https://github.com/lucasb-eyer/pydensecrf\">https://github.com/lucasb-eyer/pydensecrf</a> which is a partial wrapper for the crf used in deep lab, it showed improvement for validation scores, but would have taken a little too long for all the data at panchromatic resolution on my hardware. Here is dense crf paper <a href=\"https://arxiv.org/abs/1210.5644\">https://arxiv.org/abs/1210.5644</a>. </p>\n\n<p>With respect to edges. I used valid border mode instead of same, and it removes the problem with borders. The original unet paper says they used valid borders as well.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 166460,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "03/09/2017 19:06:35",
          "content": "<p>Actually I had the same issue - very long run time for full bunch of my train images. And CRF run on small subset lead to bad result. May be the link you provided is better for this task.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 166542,
      "author_name": "zeliek",
      "author_url": "",
      "post_date": "03/10/2017 02:56:23",
      "content": "<p>Why do you just split the image to 15*15 but not randomly sample from the image ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 166592,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "03/10/2017 08:38:42",
          "content": "<p>Good question. If made straigtforward random sample from image, then it could lead to bad result due to class imbalances. Most of batches will contain small ammount of predicted class. What is worse in some cases batch won't contain predicted class. Probably using pregenerated positive cases in union with random sample from images will work better than my solution. But I think my solution is ok, due to large variance of samples.\nThe other good thing about my approach: </p>\n\n<p>1) I generated samples one time with balanced classes and can use them \"as is\" just with random() function.</p>\n\n<p>2) Reading all train images in memory (to random sample from them) requires ~32 GB of RAM. My current solution don't require big ammount of memory. But require SSD because of heavy reading from disk.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 166684,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "03/10/2017 16:34:11",
          "content": "<ol>\n<li>I got better results when I started using batch with 128 samples, I would believe exactly due to the reasons that you mentioned.</li>\n<li>One may save train data in float16 format =&gt; ~13Gb for the whole train set.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 166704,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "03/10/2017 17:16:55",
          "content": "<p>I couldn't use big batches because I choose to big image as input. My batch was 22 - higher batches didn't fit in memory. Probably if I'd start from scratch I'd use 112x112 input.</p>\n\n<p>I afraid of float16 )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 166686,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "03/10/2017 16:38:13",
      "content": "<p>Hi, <br>\nDid you try to use dropout? I tried in some of the classes but surprisingly the cross-validation score was lower...</p>",
      "votes": null,
      "replies": [
        {
          "id": 166706,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "03/10/2017 17:18:46",
          "content": "<p>Yes I used it (Dropout(0.1)) in my latest models. I can't say for sure if it was good or not, but my score became better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 168194,
      "author_name": "samihaq",
      "author_url": "",
      "post_date": "03/16/2017 14:59:28",
      "content": "<p>@ZFTurbo, Would you please publish final code for our learning ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 168240,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "03/16/2017 17:26:32",
          "content": "<p>Currently preparing the code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 170186,
          "author_name": "samihaq",
          "author_url": "",
          "post_date": "03/24/2017 14:05:37",
          "content": "<p>@ZFTurbo Would you pls share the final code ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 170192,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "03/24/2017 14:21:07",
          "content": "<p>I asked kaggle support about sharing the code and they said as a prize winner I can't publish it. Sorry. (</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 170214,
          "author_name": "samihaq",
          "author_url": "",
          "post_date": "03/24/2017 16:18:04",
          "content": "<p>@ZFTurbo ,\n really? i have seen the prize winners sharing the code like by 1st prize winners deepsense.io in Right Whale Recognition competition. Strange. That's not good by the kaggle support.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 168238,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "03/16/2017 17:24:58",
      "content": "<p>I made video with segmentation examples:</p>\n\n<p><strong><a href=\"https://www.youtube.com/watch?v=rpp7ZhGb1IQ\">https://www.youtube.com/watch?v=rpp7ZhGb1IQ</a></strong></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "166431": "To solve this problem I used: Python + Keras 1.0.8 + Theano under Windows 10. As hardware I had NVIDIA GTX 980 8 GB. I used only one GPU which worked almost 24h a day during competition. After join the team I used 2 additional GPUs (TITAN 12GB).\nIn short my solution can be described like this:\n\n 1. My main CNN is modified UNET with input shape (20, 224, 224). It\n    has higher depth, with added batch normalization and dropout layers.\n    As loss function is used Jacquard Coefficient. I used grid search\n    with different parameters and choose the best model with highest\n    validation score. \n 2. I create separate models for each class (so 10\n    independent models) and tuned them independently. I actually think\n    that I lost some information about class interaction this way, but\n    it was much easier to tune models. I partially fix it with\n    postprocessing step. \n 3. Each model is actually set of K different\n    weights obtained with KFold. For most classes I used 5 KFold. For\n    class 7 - 2 KFold. For class 9 – 4 KFold. I split train set by image\n    ID (there were only 25 images). I made it once by hands\n    independently for each class. Each fold contains the same number of\n    images with existed class. So it’s actually stratified split.\n\nData preprocessing\n------------------\nEach “image object” (with particular img_id) had set of images made with different wave lengths. In total 20 channels. I resized all images to 3360x3360 pixels (3360 = 15*224 and 3360 is closest to panchromatic image resolution).  And join them along the axis. So at the end I had tensor of following shape: 20x3360x3360. I created the pixel masks from given polygons with same size: 3360x3360 pixels.\n\n![Set of images example][1]\n\n![Mask example][2]\n\nI decided to use UNET with input shape 20x224x224. I choose 224 for 2 reasons:\n\n1. 224 = 2*2*2*2*2*7 – I have 5 MaxPooling layers in UNET and on lowest layer has 7x7 pixel size. Which in my experience the best.\n2. The same size I could use with pretrained VGG16 or ResNet.\n\nThe next step was made independently for each class:\n\n 1. Split input images on 15*15 parts forming tensors of size\n    20x224x224. \n 2.\tBecause of class imbalances we need to increase\n    number of cases where mask exists. For some cases like class 10\n    (with vehicles) I have around 99% of empty masks after first step.\n    So I added more cases using sliding window around non-zero mask\n    points.\n\n\nFinal statistics for classes:\n\n    Number of tests for class 1: 14010. Empty files: 4431 Percent: 31.62%\n    Number of tests for class 2: 9835. Empty files: 3997 Percent: 40.64%\n    Number of tests for class 3: 20548. Empty files: 5238 Percent: 25.49%\n    Number of tests for class 4: 12411. Empty files: 2976 Percent: 23.97%\n    Number of tests for class 5: 10161. Empty files: 294 Percent: 2.89%\n    Number of tests for class 6: 10716. Empty files: 3720 Percent: 34.71%\n    Number of tests for class 7: 10328. Empty files: 5504 Percent: 53.29%\n    Number of tests for class 8: 11349. Empty files: 5469 Percent: 48.18%\n    Number of tests for class 9: 12035. Empty files: 5574 Percent: 46.31%\n    Number of tests for class 10: 16173. Empty files: 5336 Percent: 32.99%\n\nUNET requires the distribution to be close to normal. Ranges for different channels was: P_3, P_P, P_M – 2048, P_A – 16384. At first I divide every channel to its maximum possible value. I’ve seen on forum that some people removed some values from the both ends of histogram but I didn’t do it.\n\nThen I needed to calculate mean and stdev. I calculated it for each channel independently, but at the end used the same mean and stdev for whole tensor: \nmean = 0.219613\nstdev = 0.110741\n\nCreating models\n---------------\n![Modified UNET][3]\n\n    ____________________________________________________________________________________\n    Layer (type)             Output Shape   Param #     Connected to\n    ====================================================================================\n    input_1 (InputLayer)      (20, 224, 224)  0                                 \n    ____________________________________________________________________________________\n    conv2d_1 (Convolution2D)  (32, 224, 224)  5792     input_1\n    ____________________________________________________________________________________\n    batchnorm_1 (BatchNormal  (32, 224, 224)  64       conv2d_1             \n    ____________________________________________________________________________________\n    activation_1 (Activation) (32, 224, 224)  0        batchnorm_1        \n    ____________________________________________________________________________________\n    conv2d_2 (Convolution2D)  (32, 224, 224)  9248     activation_1                \n    ____________________________________________________________________________________\n    batchnorm_2 (BatchNormal  (32, 224, 224)  64       conv2d_2             \n    ____________________________________________________________________________________\n    activation_2 (Activation) (32, 224, 224)  0        batchnorm_2        \n    ____________________________________________________________________________________\n    maxpool2d_1 (MaxPooling2D)(32, 112, 112)  0        activation_2                \n    ____________________________________________________________________________________\n    conv2d_3 (Convolution2D)  (64, 112, 112)  18496    maxpool2d_1              \n    ____________________________________________________________________________________\n    batchnorm_3 (BatchNormal  (64, 112, 112)  128      conv2d_3             \n    ____________________________________________________________________________________\n    activation_3 (Activation) (64, 112, 112)  0        batchnorm_3        \n    ____________________________________________________________________________________\n    conv2d_4 (Convolution2D)  (64, 112, 112)  36928    activation_3                \n    ____________________________________________________________________________________\n    batchnorm_4 (BatchNormal  (64, 112, 112)  128      conv2d_4             \n    ____________________________________________________________________________________\n    activation_4 (Activation) (64, 112, 112)  0        batchnorm_4        \n    ____________________________________________________________________________________\n    maxpool2d_2 (MaxPooling2D)(64, 56, 56)    0        activation_4                \n    ____________________________________________________________________________________\n    conv2d_5 (Convolution2D)  (128, 56, 56)   73856    maxpool2d_2              \n    ____________________________________________________________________________________\n    batchnorm_5 (BatchNormal  (128, 56, 56)   256      conv2d_5             \n    ____________________________________________________________________________________\n    activation_5 (Activation) (128, 56, 56)   0        batchnorm_5        \n    ____________________________________________________________________________________\n    conv2d_6 (Convolution2D)  (128, 56, 56)   147584   activation_5                \n    ____________________________________________________________________________________\n    batchnorm_6 (BatchNormal  (128, 56, 56)   256      conv2d_6             \n    ____________________________________________________________________________________\n    activation_6 (Activation) (128, 56, 56)   0        batchnorm_6        \n    ____________________________________________________________________________________\n    maxpool2d_3 (MaxPooling2D)(128, 28, 28)   0        activation_6                \n    ____________________________________________________________________________________\n    conv2d_7 (Convolution2D)  (256, 28, 28)   295168   maxpool2d_3              \n    ____________________________________________________________________________________\n    batchnorm_7 (BatchNormal  (256, 28, 28)   512      conv2d_7             \n    ____________________________________________________________________________________\n    activation_7 (Activation) (256, 28, 28)   0        batchnorm_7        \n    ____________________________________________________________________________________\n    conv2d_8 (Convolution2D)  (256, 28, 28)   590080   activation_7                \n    ____________________________________________________________________________________\n    batchnorm_8 (BatchNormal  (256, 28, 28)   512      conv2d_8             \n    ____________________________________________________________________________________\n    activation_8 (Activation) (256, 28, 28)   0        batchnorm_8        \n    ____________________________________________________________________________________\n    maxpool2d_4 (MaxPooling2D)(256, 14, 14)   0        activation_8                \n    ____________________________________________________________________________________\n    conv2d_9 (Convolution2D)  (512, 14, 14)   1180160  maxpool2d_4              \n    ____________________________________________________________________________________\n    batchnorm_9 (BatchNormal  (512, 14, 14)   1024     conv2d_9             \n    ____________________________________________________________________________________\n    activation_9 (Activation) (512, 14, 14)   0        batchnorm_9        \n    ____________________________________________________________________________________\n    conv2d_10 (Convolution2D) (512, 14, 14)   2359808  activation_9                \n    ____________________________________________________________________________________\n    batchnorm_10 (BatchNorma  (512, 14, 14)   1024     conv2d_10            \n    ____________________________________________________________________________________\n    activation_10 (Activation)(512, 14, 14)   0        batchnorm_10       \n    ____________________________________________________________________________________\n    maxpool2d_5 (MaxPooling2D)(512, 7, 7)     0        activation_10               \n    ____________________________________________________________________________________\n    conv2d_11 (Convolution2D) (1024, 7, 7)    4719616  maxpool2d_5              \n    ____________________________________________________________________________________\n    batchnorm_11 (BatchNorma  (1024, 7, 7)    2048     conv2d_11            \n    ____________________________________________________________________________________\n    activation_11 (Activation)(1024, 7, 7)    0        batchnorm_11       \n    ____________________________________________________________________________________\n    conv2d_12 (Convolution2D) (1024, 7, 7)    9438208  activation_11               \n    ____________________________________________________________________________________\n    batchnorm_12 (BatchNorma  (1024, 7, 7)    2048     conv2d_12            \n    ____________________________________________________________________________________\n    activation_12 (Activation)(1024, 7, 7)    0        batchnorm_12       \n    ____________________________________________________________________________________\n    upsamp2d_1 (UpSampling2D) (1024, 14, 14)  0        activation_12               \n    ____________________________________________________________________________________\n    merge_1 (Merge)           (1536, 14, 14)  0        upsamp2d_1              \n                                                          activation_10               \n    ____________________________________________________________________________________\n    conv2d_13 (Convolution2D) (512, 14, 14)   7078400  merge_1                     \n    ____________________________________________________________________________________\n    batchnorm_13 (BatchNorma  (512, 14, 14)   1024     conv2d_13            \n    ____________________________________________________________________________________\n    activation_13 (Activation)(512, 14, 14)   0        batchnorm_13       \n    ____________________________________________________________________________________\n    conv2d_14 (Convolution2D) (512, 14, 14)   2359808  activation_13               \n    ____________________________________________________________________________________\n    batchnorm_14 (BatchNorma  (512, 14, 14)   1024     conv2d_14            \n    ____________________________________________________________________________________\n    activation_14 (Activation)(512, 14, 14)   0        batchnorm_14       \n    ____________________________________________________________________________________\n    upsamp2d_2 (UpSampling2D) (512, 28, 28)   0        activation_14               \n    ____________________________________________________________________________________\n    merge_2 (Merge)           (768, 28, 28)   0        upsamp2d_2              \n                                                                    activation_8                \n    ____________________________________________________________________________________\n    conv2d_15 (Convolution2D) (256, 28, 28)   1769728  merge_2                     \n    ____________________________________________________________________________________\n    batchnorm_15 (BatchNorma  (256, 28, 28)   512      conv2d_15            \n    ____________________________________________________________________________________\n    activation_15 (Activation)(256, 28, 28)   0        batchnorm_15       \n    ____________________________________________________________________________________\n    conv2d_16 (Convolution2D) (256, 28, 28)   590080   activation_15               \n    ____________________________________________________________________________________\n    batchnorm_16 (BatchNorma  (256, 28, 28)   512      conv2d_16            \n    ____________________________________________________________________________________\n    activation_16 (Activation)(256, 28, 28)   0        batchnorm_16       \n    ____________________________________________________________________________________\n    upsamp2d_3 (UpSampling2D) (256, 56, 56)   0        activation_16               \n    ____________________________________________________________________________________\n    merge_3 (Merge)           (384, 56, 56)   0        upsamp2d_3              \n                                                          activation_6                \n    ____________________________________________________________________________________\n    conv2d_17 (Convolution2D) (128, 56, 56)   442496   merge_3                     \n    ____________________________________________________________________________________\n    batchnorm_17 (BatchNorma  (128, 56, 56)   256      conv2d_17            \n    ____________________________________________________________________________________\n    activation_17 (Activation)(128, 56, 56)   0        batchnorm_17       \n    ____________________________________________________________________________________\n    conv2d_18 (Convolution2D) (128, 56, 56)   147584   activation_17               \n    ____________________________________________________________________________________\n    batchnorm_18 (BatchNorma  (128, 56, 56)   256      conv2d_18            \n    ____________________________________________________________________________________\n    activation_18 (Activation)(128, 56, 56)   0        batchnorm_18       \n    ____________________________________________________________________________________\n    upsamp2d_4 (UpSampling2D) (128, 112, 112) 0        activation_18               \n    ____________________________________________________________________________________\n    merge_4 (Merge)           (192, 112, 112) 0        upsamp2d_4              \n                                                          activation_4                \n    ____________________________________________________________________________________\n    conv2d_19 (Convolution2D) (64, 112, 112)  110656   merge_4                     \n    ____________________________________________________________________________________\n    batchnorm_19 (BatchNorma  (64, 112, 112)  128      conv2d_19            \n    ____________________________________________________________________________________\n    activation_19 (Activation)(64, 112, 112)  0        batchnorm_19       \n    ____________________________________________________________________________________\n    conv2d_20 (Convolution2D) (64, 112, 112)  36928    activation_19               \n    ____________________________________________________________________________________\n    batchnorm_20 (BatchNorma  (64, 112, 112)  128      conv2d_20            \n    ____________________________________________________________________________________\n    activation_20 (Activation)(64, 112, 112)  0        batchnorm_20       \n    ____________________________________________________________________________________\n    upsamp2d_5 (UpSampling2D) (64, 224, 224)  0        activation_20               \n    ____________________________________________________________________________________\n    merge_5 (Merge)           (96, 224, 224)  0        upsamp2d_5              \n                                                          activation_2                \n    ____________________________________________________________________________________\n    conv2d_21 (Convolution2D) (32, 224, 224)  27680    merge_5                     \n    ____________________________________________________________________________________\n    batchnorm_21 (BatchNorma  (32, 224, 224)  64       conv2d_21            \n    ____________________________________________________________________________________\n    activation_21 (Activation)(32, 224, 224)  0        batchnorm_21       \n    ____________________________________________________________________________________\n    conv2d_22 (Convolution2D) (32, 224, 224)  9248     activation_21               \n    ____________________________________________________________________________________\n    batchnorm_22 (BatchNorma  (32, 224, 224)  64       conv2d_22            \n    ____________________________________________________________________________________\n    activation_22 (Activation)(32, 224, 224)  0        batchnorm_23       \n    ____________________________________________________________________________________\n    conv2d_23 (Convolution2D) (1, 224, 224)   33       activation_22               \n    ____________________________________________________________________________________\n    batchnorm_23 (BatchNorma  (1, 224, 224)   2        conv2d_23            \n    ____________________________________________________________________________________\n    activation_23 (Activation)(1, 224, 224)   0        batchnorm_23       \n    ====================================================================================\n    Total params: 31459619\n\n\nLOSS Function:\n\n    def jacard_coef(y_true, y_pred):\n        y_true_f = K.flatten(y_true)\n        y_pred_f = K.flatten(y_pred)\n        intersection = K.sum(y_true_f * y_pred_f)\n        return (intersection + 1.0) / (K.sum(y_true_f) + K.sum(y_pred_f) - intersection + 1.0)\n\nTypical example of training process: Class 1 (Fold 2). Parameters: lr=0.05, optim=SGD, rotation=False, dropout=enabled, UNET version=BatchNorm, patience=8, samples train: 9997, samples valid: 4013\n\n![Typical example of training process][4]\n\nTuning models and validation\n----------------------------\nThe main problem was that training process is very unstable. Sometimes it can go to local minimum or start predicting as mask full image etc. So most of the time I spend on finding optimal parameters to maximize validation score. It was the most time consuming part. I used grid search with following parameters: \n\n - Optimizer (Adam, SGD) \n - LR SGD: (0.05, 0.01, 0.001) \n - LR Adam: (0.01, 0.001, 0.0001) \n - Rotation (enabled, disabled) \n - Type of model: UNET 224x224, UNET with dropout 0.1, UNET with Batch Normalization \n - Number of samples per epoch (Fraction from ½ up to 1)\n\nDue to limitation of computational power I checked only some of parameters. \nAt the end I use mostly Adam optimizer with BatchNormalization version of UNET with rotation enabled. Just tune learning rate.\nI use loss function value to stop training with early stopping with patience 8-15 epochs. After I obtain all 5 Folds models I had the predicted segmentation for each of train images. So I was able to predict score for full image set using the same method as it was made on Kaggle Leaderboard. Obtained score was very representative, so increasing the score locally almost always leads to increasing the score on leaderboard. Also I was able to find optimal threshold value for heatmap with this method.\n\n![LS vs LB Scores][5]\n\nProcessing the test data\n------------------------\nTest set consists of 429 images. Each test image was processed separately. Each image was resized and normalized the same way as training data to 20х3360х3360 tensor. \n\nIn the beginning I create two zero arrays HEATMAP and COUNT with 3360х3360 shape. Then use sliding window approach to predict segmentation. All 5 folds predict on image extracted from current position and acquired probabilities sum up to HEATMAP array. COUNT array at image position added by 1. At the end HEATMAP is divided by COUNT to get real heatmap array. After this heatmap was thresholded at value acquired from validation, typically 0.5. On basis of this 2D-array I created polygons with rasterio and shapely.\n\n![Sliding window][6]\n\nAdditional ideas to increase the accuracy provided below. Not all of them was used in final submit because processing of test images was very long process, around 8-10 hours for whole test set.\n\n\n1.\tDecreasing sliding step from 112px to 56px and below always increased the accuracy of predictions, but increase the computational complexity as O(N^2). \n- For example class 4 on validation 0.385714 vs 0.403683.\n\n2.\tIt looks like UNET predict worse on the edges, so it’s good idea to use only central part for prediction. Example for class 5 below. \n- Default score: 0.507446 \n- 200x200 center part from 224x224 square + 100 px sliding window step: 0.512844 \n- 160x160 center part from 224x224 square + 80 px sliding window step: 0.514557\n\n3.\tI usually used threshold 0.5 without checking optimum on validation, but this can be useful too. \n- Class 6 example: \n- THR 0.1: 0.738943 \n- THR 0.5: 0.757858 \n- THR 0.9: 0.759850\n\n![Masks for class 5][7]\n\nEnsembles\n---------\nThe most obvious ways to do ensembles is to use heatmaps. But out team consisted of 2 people was merged at the last stage of competitions, so we actually don’t have heatmaps for all classes at this stage. So we made ensembles directly on polygons, using UNION and INTERSECTION from shapely. We have different training process and models, so for some classes we had good boost after merge of this type. \n\n    Class 6 LB: 0.08149 + LB: 0.08103 (Intersection) Score: 0.08179\n    Class 8 LB: 0.03890 + LB: 0.05322 (Intersection) Score: 0.06194\n    Class 9 LB: 0.02113 + LB: 0.02713 (Union) Score: 0.03254\n\nBad thing, we mostly use leaderboard to check if our ensemble gave boost, which could lead to overfitting. It would be much easier if we have 3rd independent solution to use voting mechanism.\n\nPost processing\n---------------\nAfter analysis of class interactions in train I made the following postprocess types:\n- Remove too large polygons (from class 8, 9 and 10 for example) \n- Subtract predicted classes polygons like water from car classes\n\nClass 10 with small objects\n---------------------------\n\nFor class 10 default approach works not very well. The main problems with this class that it’s have very small objects, bad train segmentation and it’s total area is too small comparing to full area. So I created other UNET CNN for this class. It had input of 32x32 pixels. Data for this class was heavily augmented, rotations, different shifts etc. I also used loss function with big penalty for false positives based on [Tversky index][8].\n\n    # https://en.wikipedia.org/wiki/Tversky_index\n    def tversky_coef(y_true, y_pred):\n        y_true_f = K.flatten(y_true)\n        y_pred_f = K.flatten(y_pred)\n        alfa = 0.1\n        false_positive = K.sum(y_pred_f * (1 - y_true_f))\n        false_negative = K.sum((1 - y_pred_f) * y_true_f)\n        true_positive = K.sum(y_true_f * y_pred_f)\n        return (true_positive + 1.0) / (false_negative*alfa + (1 - alfa) * false_positive + true_positive + 1.0)\n\nIt allowed me to get 0.00362 score on LB.\n\nOther ideas I tried\n-------------------\n- At the early beginning I made single XGBoost model which works on squares of size 10x10. Model tries to predict to which class given square is belong. It easily got me higher than baseline. Probably creating independent XGBoost models for each class could give good results. But I switch to CNN without additional experiments with XGBoost approach.\n- Pre-trained VGG16 for rare class localization. I tried to predict if car exists in given area with VGG16 on RGB images for later ensemble with UNET predictions.\n\n![Example of class 10 localization with VGG16][9]\n\n- Pre-trained VGG16 for segmentation. The main idea here is that we use final layer (which is now sigmoid instead of softmax) as indication in which part of 224x224 image given class exists. I used 16x16 = 256 neurons for this task. But result wasn’t very good, may be because of low resolution - 14x14 square as single pixel.\n- There are big bunch of different indexes used for automatic segmentation without CNN:\nhttp://www.indexdatabase.de/db/i.php\nIt works great for water as shown in [Waterway Kernel][10]. And these indexes can be added as additional planes during learning process. Unfortunately they didn’t help much in my experiments, so I abandon them. \n- I tried to use CRF (Conditional random field) methods with [pystruct][11] to improve predicted masks. But it somehow won’t work for me.\n\nIndependent classes best scores\n-------------------------------\n![Best Public LB][12]\n\nSegmentation example\n--------------------\n![enter image description here][13]\n![enter image description here][14]\n\nCheckout the video:\n**[https://www.youtube.com/watch?v=rpp7ZhGb1IQ][15]**\n\nCode\n----\n\nGitHUB link here later. I plan to release code after publication of final results.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6037/Tensor_example_700px.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6029/Mask_example.png\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6038/model_700px.png\n  [4]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6031/Training.png\n  [5]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6032/LB-Scores.png\n  [6]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6033/Sliding%20Window.png\n  [7]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6039/Masks_700px.png\n  [8]: https://en.wikipedia.org/wiki/Tversky_index\n  [9]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6040/vgg16_test_6060_3_2_cls_10_700px.png\n  [10]: https://www.kaggle.com/resolut/dstl-satellite-imagery-feature-detection/waterway-0-095-lb\n  [11]: https://pystruct.github.io/\n  [12]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6036/Best-LB.png\n  [13]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6048/6100_1_2_mask.png\n  [14]: https://kaggle2.blob.core.windows.net/forum-message-attachments/166431/6049/6100_1_2_proj_700px.jpg\n  [15]: https://www.youtube.com/watch?v=rpp7ZhGb1IQ",
    "166458": "Great write up and solution! For crf I tried using https://github.com/lucasb-eyer/pydensecrf which is a partial wrapper for the crf used in deep lab, it showed improvement for validation scores, but would have taken a little too long for all the data at panchromatic resolution on my hardware. Here is dense crf paper https://arxiv.org/abs/1210.5644. \n\nWith respect to edges. I used valid border mode instead of same, and it removes the problem with borders. The original unet paper says they used valid borders as well.",
    "166460": "Actually I had the same issue - very long run time for full bunch of my train images. And CRF run on small subset lead to bad result. May be the link you provided is better for this task.",
    "166542": "Why do you just split the image to 15*15 but not randomly sample from the image ?",
    "166592": "Good question. If made straigtforward random sample from image, then it could lead to bad result due to class imbalances. Most of batches will contain small ammount of predicted class. What is worse in some cases batch won't contain predicted class. Probably using pregenerated positive cases in union with random sample from images will work better than my solution. But I think my solution is ok, due to large variance of samples.\nThe other good thing about my approach: \n\n1) I generated samples one time with balanced classes and can use them \"as is\" just with random() function.\n\n2) Reading all train images in memory (to random sample from them) requires ~32 GB of RAM. My current solution don't require big ammount of memory. But require SSD because of heavy reading from disk.",
    "166684": "1. I got better results when I started using batch with 128 samples, I would believe exactly due to the reasons that you mentioned.\n 2. One may save train data in float16 format => ~13Gb for the whole train set.",
    "166686": "Hi,   \nDid you try to use dropout? I tried in some of the classes but surprisingly the cross-validation score was lower...",
    "166704": "I couldn't use big batches because I choose to big image as input. My batch was 22 - higher batches didn't fit in memory. Probably if I'd start from scratch I'd use 112x112 input.\n\nI afraid of float16 )",
    "166706": "Yes I used it (Dropout(0.1)) in my latest models. I can't say for sure if it was good or not, but my score became better.",
    "168194": "ZFTurbo, Would you please publish final code for our learning ?",
    "168238": "I made video with segmentation examples:\n\n**[https://www.youtube.com/watch?v=rpp7ZhGb1IQ][1]**\n\n\n  [1]: https://www.youtube.com/watch?v=rpp7ZhGb1IQ",
    "168240": "Currently preparing the code.",
    "170186": "ZFTurbo Would you pls share the final code ?",
    "170192": "I asked kaggle support about sharing the code and they said as a prize winner I can't publish it. Sorry. (",
    "170214": "ZFTurbo ,\n really? i have seen the prize winners sharing the code like by 1st prize winners deepsense.io in Right Whale Recognition competition. Strange. That's not good by the kaggle support."
  },
  "source": "meta"
}