{
  "id": 29790,
  "title": "0.53 Public LB Solution",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/29790",
  "author_name": "Guillermo Barbadillo",
  "post_date": "2017-03-08T20:40:51.799000",
  "votes": 43,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>Summary of the DSTL Challenge</h1>\n\n<p>In this document I will review all my work in this challenge</p>\n\n<h2>1. First steps</h2>\n\n<h3>Preprocessing the images</h3>\n\n<p>When I looked at the rgb images I noticed that the white balance was no good. So I applied white balance like in the link below. <br>\n<a href=\"https://docs.gimp.org/en/gimp-layer-white-balance.html\">https://docs.gimp.org/en/gimp-layer-white-balance.html</a> <br>\nInstead of applying white balance to single images I made collages(for example all the images that start by 6010) and I applied white balance to the whole collage. I made this for all the satellite channels.\nI also corrected the missalignement between the different channels using simple traslations.</p>\n\n<h3>First trainings</h3>\n\n<p>At the beginning I tried to train a single model to predict all the classes. I found it very difficult and I think it was because the great imbalance between the classes. <br>\n<img src=\"http://imgur.com/exlz22p.png\" alt=\"enter image description here\" title=\"\"> <br>\nI decided that I was going to train a different model for each class, I tried that and it worked better so I continue with that aproach until the end.</p>\n\n<h3>Sampling</h3>\n\n<p>The low frequency of the classes make hard to train them. So I decided that instead of training each epoch with all the samples I was going to take some samples with positive instances and some samples without. That way I could change the frequency of the classes and help the model to learn.</p>\n\n<h3>Creating the submission</h3>\n\n<p>This was probably the most painfull part of the challenge. The slow evaluation method of kaggle forced me to simplify the predictions. I could not make a succesfull submission for trees without simplifying. <br>\nI used shapely for making the simplifications, and at the start I was simplifying without preserving the topology so the scores were badly hurt. I discovered this when I made a tool for visualizing submissions, I was just visualizing predictions so this shows how important is to visualize each step of the pipeline. <br>\nMy last submission was made with tolerance 8e-6 ~ 4 pixels.</p>\n\n<h2>2. Tools</h2>\n\n<p>I have used python and the main libraries used are:\n* Keras\n* Opencv\n* Shapely  </p>\n\n<p>I started the challenge using a laptop with a 980m gpu and 8GB of RAM. At the half of the challenge I bought a pc with a gpu 1080 and 32GB of RAM, and in the last week of the challenge I bought another gpu 1080 and another 32GB of RAM :)</p>\n\n<h2>3. Models</h2>\n\n<p>I think this competition was pleasing because it was really nice to watch how the model was learning to paint a map just by using the images. Below you can find some visualizations of images from the test set.   </p>\n\n<p><img src=\"http://imgur.com/1vdZ5em.png\" alt=\"enter image description here\" title=\"\">\n<img src=\"http://imgur.com/8jPC2UM.png\" alt=\"enter image description here\" title=\"\">\n<img src=\"http://imgur.com/R5QQiJG.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>Main characteristics of the models   </p>\n\n<ul>\n<li>I used unet architecture on all my models.   </li>\n<li>I trained using cross-validation with 3 and 5 folds.  </li>\n<li>For each class I tuned the depth of the model, the window size and the resize of the images.  </li>\n<li>I used as input all satellite bands except for the models with finer detail. In the case of structures and vehicles I only used rgb images and p band.  </li>\n</ul>\n\n<p><img src=\"http://imgur.com/7LHvaOr.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<h3>Water</h3>\n\n<p>This is my most worked model. At the beginning I tried to train a model for rivers and another for lakes. But then I realized that it was really hard for a model to segment the image pixel by pixel and also make an abstraction and decide if something is a lake or a river.  </p>\n\n<p>So I decided to train a model just for predicting water, and after that add a high level classifier to make the division between lakes and rivers.</p>\n\n<p>Another idea was to use images with different scale for training the model. I realized that unlike the other classes a lake or a river can have all sort of shapes and sizes (all cars will have a similar size). So I trained the model with images with scale 1:1, 1/2, 1/4 and 1/8. That way my training set was much bigger.</p>\n\n<p>Before feeding the images to the high level classifier I made collages with them because sometimes in a single image we don't have enought information to decide to which category they belong.   </p>\n\n<p><img src=\"http://imgur.com/9K6XUdg.png\" alt=\"enter image description here\" title=\"\">\nThe high level classifier had only two steps: in the first I use a threshold on the contour size to search for big rivers. In the second I use the distance to the big rivers to decide if a contour was a lake or a river. </p>\n\n<p>Having a good water model was important, because later I used this model to clean the output of some of the others.</p>\n\n<h3>Building, road, tree, crops, track</h3>\n\n<p>I don't think this models have something special. I just tried with different parameters such as: resize, window size, depth, sampling until I find a good configuration.</p>\n\n<h3>Structure</h3>\n\n<p>This model was hard to train. I had to use extracost of 2 or 4 on white pixels to force the model to learn. Otherwise it painted all black.</p>\n\n<h3>Small and big vehicles</h3>\n\n<p>I spend the last days of the contest trying to improve this two categories. <br>\nMy final submission consisted on averaging the prediction of a lot of models (up to 15 different models). For trainign different models I used cross-validation with different random seeds. <br>\nI cleaned the prediction using the prediction of the other models. I used the water, building and tree model for cleaning.</p>\n\n<h2>4. Scores</h2>\n\n<p><img src=\"http://imgur.com/9EJggch.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>Please let me now if something is not clear or if I have to elaborate more  on some topic. <br>\nI will be adding more info if necessary to the first post.   </p>\n\n<p>Thanks to the organizers and the participants. <br>\nironbar</p>",
  "messages": [
    {
      "id": 166230,
      "postDate": "2017-03-08T20:40:51.800Z",
      "content": "<h1>Summary of the DSTL Challenge</h1>\n\n<p>In this document I will review all my work in this challenge</p>\n\n<h2>1. First steps</h2>\n\n<h3>Preprocessing the images</h3>\n\n<p>When I looked at the rgb images I noticed that the white balance was no good. So I applied white balance like in the link below. <br>\n<a href=\"https://docs.gimp.org/en/gimp-layer-white-balance.html\">https://docs.gimp.org/en/gimp-layer-white-balance.html</a> <br>\nInstead of applying white balance to single images I made collages(for example all the images that start by 6010) and I applied white balance to the whole collage. I made this for all the satellite channels.\nI also corrected the missalignement between the different channels using simple traslations.</p>\n\n<h3>First trainings</h3>\n\n<p>At the beginning I tried to train a single model to predict all the classes. I found it very difficult and I think it was because the great imbalance between the classes. <br>\n<img src=\"http://imgur.com/exlz22p.png\" alt=\"enter image description here\" title=\"\"> <br>\nI decided that I was going to train a different model for each class, I tried that and it worked better so I continue with that aproach until the end.</p>\n\n<h3>Sampling</h3>\n\n<p>The low frequency of the classes make hard to train them. So I decided that instead of training each epoch with all the samples I was going to take some samples with positive instances and some samples without. That way I could change the frequency of the classes and help the model to learn.</p>\n\n<h3>Creating the submission</h3>\n\n<p>This was probably the most painfull part of the challenge. The slow evaluation method of kaggle forced me to simplify the predictions. I could not make a succesfull submission for trees without simplifying. <br>\nI used shapely for making the simplifications, and at the start I was simplifying without preserving the topology so the scores were badly hurt. I discovered this when I made a tool for visualizing submissions, I was just visualizing predictions so this shows how important is to visualize each step of the pipeline. <br>\nMy last submission was made with tolerance 8e-6 ~ 4 pixels.</p>\n\n<h2>2. Tools</h2>\n\n<p>I have used python and the main libraries used are:\n* Keras\n* Opencv\n* Shapely  </p>\n\n<p>I started the challenge using a laptop with a 980m gpu and 8GB of RAM. At the half of the challenge I bought a pc with a gpu 1080 and 32GB of RAM, and in the last week of the challenge I bought another gpu 1080 and another 32GB of RAM :)</p>\n\n<h2>3. Models</h2>\n\n<p>I think this competition was pleasing because it was really nice to watch how the model was learning to paint a map just by using the images. Below you can find some visualizations of images from the test set.   </p>\n\n<p><img src=\"http://imgur.com/1vdZ5em.png\" alt=\"enter image description here\" title=\"\">\n<img src=\"http://imgur.com/8jPC2UM.png\" alt=\"enter image description here\" title=\"\">\n<img src=\"http://imgur.com/R5QQiJG.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>Main characteristics of the models   </p>\n\n<ul>\n<li>I used unet architecture on all my models.   </li>\n<li>I trained using cross-validation with 3 and 5 folds.  </li>\n<li>For each class I tuned the depth of the model, the window size and the resize of the images.  </li>\n<li>I used as input all satellite bands except for the models with finer detail. In the case of structures and vehicles I only used rgb images and p band.  </li>\n</ul>\n\n<p><img src=\"http://imgur.com/7LHvaOr.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<h3>Water</h3>\n\n<p>This is my most worked model. At the beginning I tried to train a model for rivers and another for lakes. But then I realized that it was really hard for a model to segment the image pixel by pixel and also make an abstraction and decide if something is a lake or a river.  </p>\n\n<p>So I decided to train a model just for predicting water, and after that add a high level classifier to make the division between lakes and rivers.</p>\n\n<p>Another idea was to use images with different scale for training the model. I realized that unlike the other classes a lake or a river can have all sort of shapes and sizes (all cars will have a similar size). So I trained the model with images with scale 1:1, 1/2, 1/4 and 1/8. That way my training set was much bigger.</p>\n\n<p>Before feeding the images to the high level classifier I made collages with them because sometimes in a single image we don't have enought information to decide to which category they belong.   </p>\n\n<p><img src=\"http://imgur.com/9K6XUdg.png\" alt=\"enter image description here\" title=\"\">\nThe high level classifier had only two steps: in the first I use a threshold on the contour size to search for big rivers. In the second I use the distance to the big rivers to decide if a contour was a lake or a river. </p>\n\n<p>Having a good water model was important, because later I used this model to clean the output of some of the others.</p>\n\n<h3>Building, road, tree, crops, track</h3>\n\n<p>I don't think this models have something special. I just tried with different parameters such as: resize, window size, depth, sampling until I find a good configuration.</p>\n\n<h3>Structure</h3>\n\n<p>This model was hard to train. I had to use extracost of 2 or 4 on white pixels to force the model to learn. Otherwise it painted all black.</p>\n\n<h3>Small and big vehicles</h3>\n\n<p>I spend the last days of the contest trying to improve this two categories. <br>\nMy final submission consisted on averaging the prediction of a lot of models (up to 15 different models). For trainign different models I used cross-validation with different random seeds. <br>\nI cleaned the prediction using the prediction of the other models. I used the water, building and tree model for cleaning.</p>\n\n<h2>4. Scores</h2>\n\n<p><img src=\"http://imgur.com/9EJggch.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>Please let me now if something is not clear or if I have to elaborate more  on some topic. <br>\nI will be adding more info if necessary to the first post.   </p>\n\n<p>Thanks to the organizers and the participants. <br>\nironbar</p>",
      "rawMarkdown": "# Summary of the DSTL Challenge\nIn this document I will review all my work in this challenge\n\n## 1. First steps\n### Preprocessing the images\nWhen I looked at the rgb images I noticed that the white balance was no good. So I applied white balance like in the link below.  \nhttps://docs.gimp.org/en/gimp-layer-white-balance.html  \nInstead of applying white balance to single images I made collages(for example all the images that start by 6010) and I applied white balance to the whole collage. I made this for all the satellite channels.\nI also corrected the missalignement between the different channels using simple traslations.\n### First trainings\nAt the beginning I tried to train a single model to predict all the classes. I found it very difficult and I think it was because the great imbalance between the classes.  \n![enter image description here][1]   \nI decided that I was going to train a different model for each class, I tried that and it worked better so I continue with that aproach until the end.\n### Sampling \nThe low frequency of the classes make hard to train them. So I decided that instead of training each epoch with all the samples I was going to take some samples with positive instances and some samples without. That way I could change the frequency of the classes and help the model to learn.\n### Creating the submission\nThis was probably the most painfull part of the challenge. The slow evaluation method of kaggle forced me to simplify the predictions. I could not make a succesfull submission for trees without simplifying.   \nI used shapely for making the simplifications, and at the start I was simplifying without preserving the topology so the scores were badly hurt. I discovered this when I made a tool for visualizing submissions, I was just visualizing predictions so this shows how important is to visualize each step of the pipeline.  \nMy last submission was made with tolerance 8e-6 ~ 4 pixels.\n\n## 2. Tools\nI have used python and the main libraries used are:\n* Keras\n* Opencv\n* Shapely  \n\nI started the challenge using a laptop with a 980m gpu and 8GB of RAM. At the half of the challenge I bought a pc with a gpu 1080 and 32GB of RAM, and in the last week of the challenge I bought another gpu 1080 and another 32GB of RAM :)\n\n## 3. Models\nI think this competition was pleasing because it was really nice to watch how the model was learning to paint a map just by using the images. Below you can find some visualizations of images from the test set.   \n \n![enter image description here][2]\n![enter image description here][3]\n![enter image description here][4]\n\nMain characteristics of the models   \n\n* I used unet architecture on all my models.   \n* I trained using cross-validation with 3 and 5 folds.  \n* For each class I tuned the depth of the model, the window size and the resize of the images.  \n* I used as input all satellite bands except for the models with finer detail. In the case of structures and vehicles I only used rgb images and p band.  \n\n![enter image description here][5]\n\n### Water\nThis is my most worked model. At the beginning I tried to train a model for rivers and another for lakes. But then I realized that it was really hard for a model to segment the image pixel by pixel and also make an abstraction and decide if something is a lake or a river.  \n\nSo I decided to train a model just for predicting water, and after that add a high level classifier to make the division between lakes and rivers.\n\nAnother idea was to use images with different scale for training the model. I realized that unlike the other classes a lake or a river can have all sort of shapes and sizes (all cars will have a similar size). So I trained the model with images with scale 1:1, 1/2, 1/4 and 1/8. That way my training set was much bigger.\n\nBefore feeding the images to the high level classifier I made collages with them because sometimes in a single image we don't have enought information to decide to which category they belong.   \n\n![enter image description here][6]\nThe high level classifier had only two steps: in the first I use a threshold on the contour size to search for big rivers. In the second I use the distance to the big rivers to decide if a contour was a lake or a river. \n\nHaving a good water model was important, because later I used this model to clean the output of some of the others.\n\n\n\n### Building, road, tree, crops, track\nI don't think this models have something special. I just tried with different parameters such as: resize, window size, depth, sampling until I find a good configuration.\n\n### Structure\nThis model was hard to train. I had to use extracost of 2 or 4 on white pixels to force the model to learn. Otherwise it painted all black.\n\n### Small and big vehicles\nI spend the last days of the contest trying to improve this two categories.  \nMy final submission consisted on averaging the prediction of a lot of models (up to 15 different models). For trainign different models I used cross-validation with different random seeds.  \nI cleaned the prediction using the prediction of the other models. I used the water, building and tree model for cleaning.\n\n\n## 4. Scores\n![enter image description here][7]\n\n\nPlease let me now if something is not clear or if I have to elaborate more  on some topic.  \nI will be adding more info if necessary to the first post.   \n\nThanks to the organizers and the participants.  \nironbar\n\n\n  [1]: http://imgur.com/exlz22p.png\n  [2]: http://imgur.com/1vdZ5em.png\n  [3]: http://imgur.com/8jPC2UM.png\n  [4]: http://imgur.com/R5QQiJG.png\n  [5]: http://imgur.com/7LHvaOr.png\n  [6]: http://imgur.com/9K6XUdg.png\n  [7]: http://imgur.com/9EJggch.png",
      "votes": 43
    },
    {
      "id": 166336,
      "postDate": "2017-03-09T08:31:37.173Z",
      "content": "<p>Pls share code for us to learn. Thank you very much.</p>",
      "rawMarkdown": "Pls share code for us to learn. Thank you very much.",
      "votes": 1
    },
    {
      "id": 166289,
      "postDate": "2017-03-09T03:45:56.247Z",
      "content": "<p>Tks for sharing, do you use only M bands of all the bands?</p>",
      "rawMarkdown": "Tks for sharing, do you use only M bands of all the bands?",
      "votes": 1,
      "replies": [
        {
          "id": 166321,
          "postDate": "2017-03-09T06:29:27.773Z",
          "content": "<p>I used as input all satellite bands except for the models with finer detail. In the case of structures and vehicles I only used rgb images and p band.  </p>",
          "rawMarkdown": "I used as input all satellite bands except for the models with finer detail. In the case of structures and vehicles I only used rgb images and p band.  "
        }
      ]
    },
    {
      "id": 183594,
      "postDate": "2017-05-18T17:21:01.397Z",
      "content": "<p>grt work!\nbut one silly question... how can we get lb score for one class?</p>",
      "rawMarkdown": "grt work!\nbut one silly question... how can we get lb score for one class?\n",
      "replies": [
        {
          "id": 183623,
          "postDate": "2017-05-18T19:06:26.220Z",
          "content": "<p>Just submit this class alone and multiply score by ten.</p>",
          "rawMarkdown": "Just submit this class alone and multiply score by ten."
        },
        {
          "id": 183624,
          "postDate": "2017-05-18T19:07:53.680Z",
          "content": "<p>Very easy, make a submission leaving all the other classes empty. The score is the average of all the classes, so when you make a submission of only one class the score is 1/10 of the score of the class.</p>",
          "rawMarkdown": "Very easy, make a submission leaving all the other classes empty. The score is the average of all the classes, so when you make a submission of only one class the score is 1/10 of the score of the class."
        }
      ]
    },
    {
      "id": 166949,
      "postDate": "2017-03-12T05:23:49.380Z",
      "content": "<p>Thank a lot for sharing.</p>\n\n<p>I have a question:</p>\n\n<p>\"Instead of applying white balance to single images I made collages(for example all the images that start by 6010) and I applied white balance to the whole collage. I made this for all the satellite channels. \"</p>\n\n<ol>\n<li><p>It seems you are using channel stretching, what is the upper and lower percentage thresholds you used? </p></li>\n<li><p>Somehow, you applied stretching to whole area, for example, in training set, it is 5x5, so for each channel, you read all 25 images and then do channel stretching and save the images?</p></li>\n</ol>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Thank a lot for sharing.\n\nI have a question:\n\n\"Instead of applying white balance to single images I made collages(for example all the images that start by 6010) and I applied white balance to the whole collage. I made this for all the satellite channels. \"\n\n1. It seems you are using channel stretching, what is the upper and lower percentage thresholds you used? \n\n2. Somehow, you applied stretching to whole area, for example, in training set, it is 5x5, so for each channel, you read all 25 images and then do channel stretching and save the images?\n\nThanks!\n\n\n\n\n",
      "replies": [
        {
          "id": 166979,
          "postDate": "2017-03-12T09:09:03.627Z",
          "content": "<p>Hi Raymond, <br>\nI did just like the gimp link. <br>\n<a href=\"https://docs.gimp.org/en/gimp-layer-white-balance.html\">https://docs.gimp.org/en/gimp-layer-white-balance.html</a>  </p>\n\n<blockquote>\n  <p>To do this, it discards pixel colors at each end of the Red, Green and\n  Blue histograms which are used by only 0.05% of the pixels in the\n  image and stretches the remaining range as much as possible.</p>\n</blockquote>\n\n<p>Yes you are right on the second question. I did this because there are some images that for example only have trees. I thought that by combining the 25 images I will have a better white balance. <br>\nI also tried combining all the images of the dataset but the results seemed to be worse. </p>",
          "rawMarkdown": "Hi Raymond,   \nI did just like the gimp link.  \nhttps://docs.gimp.org/en/gimp-layer-white-balance.html  \n\n> To do this, it discards pixel colors at each end of the Red, Green and\n> Blue histograms which are used by only 0.05% of the pixels in the\n> image and stretches the remaining range as much as possible.\n\nYes you are right on the second question. I did this because there are some images that for example only have trees. I thought that by combining the 25 images I will have a better white balance.  \nI also tried combining all the images of the dataset but the results seemed to be worse. "
        }
      ]
    },
    {
      "id": 166340,
      "postDate": "2017-03-09T08:45:02.153Z",
      "content": "<p>Thanks for sharing this! Could you elaborate on the method you used to integrate different bands?</p>",
      "rawMarkdown": "Thanks for sharing this! Could you elaborate on the method you used to integrate different bands?",
      "replies": [
        {
          "id": 166464,
          "postDate": "2017-03-09T19:28:13.097Z",
          "content": "<p>I upsize the different bands to match the size of RGB channels and simply stack all the layers together.</p>",
          "rawMarkdown": "I upsize the different bands to match the size of RGB channels and simply stack all the layers together."
        }
      ]
    },
    {
      "id": 234261,
      "postDate": "2017-10-22T19:39:19.330Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 166336,
      "author_name": "quest",
      "author_url": "",
      "post_date": "2017-03-09T08:31:37.173000",
      "content": "<p>Pls share code for us to learn. Thank you very much.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 166289,
      "author_name": "Argmen",
      "author_url": "",
      "post_date": "2017-03-09T03:45:56.247000",
      "content": "<p>Tks for sharing, do you use only M bands of all the bands?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 166321,
          "author_name": "Guillermo Barbadillo",
          "author_url": "",
          "post_date": "2017-03-09T06:29:27.773000",
          "content": "<p>I used as input all satellite bands except for the models with finer detail. In the case of structures and vehicles I only used rgb images and p band.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 183594,
      "author_name": "HimanshuPareek",
      "author_url": "",
      "post_date": "2017-05-18T17:21:01.397000",
      "content": "<p>grt work!\nbut one silly question... how can we get lb score for one class?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 183623,
          "author_name": "Cut Onion",
          "author_url": "",
          "post_date": "2017-05-18T19:06:26.220000",
          "content": "<p>Just submit this class alone and multiply score by ten.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 183624,
          "author_name": "Guillermo Barbadillo",
          "author_url": "",
          "post_date": "2017-05-18T19:07:53.680000",
          "content": "<p>Very easy, make a submission leaving all the other classes empty. The score is the average of all the classes, so when you make a submission of only one class the score is 1/10 of the score of the class.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 166949,
      "author_name": "CRaymond",
      "author_url": "",
      "post_date": "2017-03-12T05:23:49.380000",
      "content": "<p>Thank a lot for sharing.</p>\n\n<p>I have a question:</p>\n\n<p>\"Instead of applying white balance to single images I made collages(for example all the images that start by 6010) and I applied white balance to the whole collage. I made this for all the satellite channels. \"</p>\n\n<ol>\n<li><p>It seems you are using channel stretching, what is the upper and lower percentage thresholds you used? </p></li>\n<li><p>Somehow, you applied stretching to whole area, for example, in training set, it is 5x5, so for each channel, you read all 25 images and then do channel stretching and save the images?</p></li>\n</ol>\n\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 166979,
          "author_name": "Guillermo Barbadillo",
          "author_url": "",
          "post_date": "2017-03-12T09:09:03.627000",
          "content": "<p>Hi Raymond, <br>\nI did just like the gimp link. <br>\n<a href=\"https://docs.gimp.org/en/gimp-layer-white-balance.html\">https://docs.gimp.org/en/gimp-layer-white-balance.html</a>  </p>\n\n<blockquote>\n  <p>To do this, it discards pixel colors at each end of the Red, Green and\n  Blue histograms which are used by only 0.05% of the pixels in the\n  image and stretches the remaining range as much as possible.</p>\n</blockquote>\n\n<p>Yes you are right on the second question. I did this because there are some images that for example only have trees. I thought that by combining the 25 images I will have a better white balance. <br>\nI also tried combining all the images of the dataset but the results seemed to be worse. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 166340,
      "author_name": "Egor Panfilov",
      "author_url": "",
      "post_date": "2017-03-09T08:45:02.153000",
      "content": "<p>Thanks for sharing this! Could you elaborate on the method you used to integrate different bands?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 166464,
          "author_name": "Guillermo Barbadillo",
          "author_url": "",
          "post_date": "2017-03-09T19:28:13.097000",
          "content": "<p>I upsize the different bands to match the size of RGB channels and simply stack all the layers together.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 234261,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-10-22T19:39:19.330000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "166230": "# Summary of the DSTL Challenge\nIn this document I will review all my work in this challenge\n\n## 1. First steps\n### Preprocessing the images\nWhen I looked at the rgb images I noticed that the white balance was no good. So I applied white balance like in the link below.  \nhttps://docs.gimp.org/en/gimp-layer-white-balance.html  \nInstead of applying white balance to single images I made collages(for example all the images that start by 6010) and I applied white balance to the whole collage. I made this for all the satellite channels.\nI also corrected the missalignement between the different channels using simple traslations.\n### First trainings\nAt the beginning I tried to train a single model to predict all the classes. I found it very difficult and I think it was because the great imbalance between the classes.  \n![enter image description here][1]   \nI decided that I was going to train a different model for each class, I tried that and it worked better so I continue with that aproach until the end.\n### Sampling \nThe low frequency of the classes make hard to train them. So I decided that instead of training each epoch with all the samples I was going to take some samples with positive instances and some samples without. That way I could change the frequency of the classes and help the model to learn.\n### Creating the submission\nThis was probably the most painfull part of the challenge. The slow evaluation method of kaggle forced me to simplify the predictions. I could not make a succesfull submission for trees without simplifying.   \nI used shapely for making the simplifications, and at the start I was simplifying without preserving the topology so the scores were badly hurt. I discovered this when I made a tool for visualizing submissions, I was just visualizing predictions so this shows how important is to visualize each step of the pipeline.  \nMy last submission was made with tolerance 8e-6 ~ 4 pixels.\n\n## 2. Tools\nI have used python and the main libraries used are:\n* Keras\n* Opencv\n* Shapely  \n\nI started the challenge using a laptop with a 980m gpu and 8GB of RAM. At the half of the challenge I bought a pc with a gpu 1080 and 32GB of RAM, and in the last week of the challenge I bought another gpu 1080 and another 32GB of RAM :)\n\n## 3. Models\nI think this competition was pleasing because it was really nice to watch how the model was learning to paint a map just by using the images. Below you can find some visualizations of images from the test set.   \n \n![enter image description here][2]\n![enter image description here][3]\n![enter image description here][4]\n\nMain characteristics of the models   \n\n* I used unet architecture on all my models.   \n* I trained using cross-validation with 3 and 5 folds.  \n* For each class I tuned the depth of the model, the window size and the resize of the images.  \n* I used as input all satellite bands except for the models with finer detail. In the case of structures and vehicles I only used rgb images and p band.  \n\n![enter image description here][5]\n\n### Water\nThis is my most worked model. At the beginning I tried to train a model for rivers and another for lakes. But then I realized that it was really hard for a model to segment the image pixel by pixel and also make an abstraction and decide if something is a lake or a river.  \n\nSo I decided to train a model just for predicting water, and after that add a high level classifier to make the division between lakes and rivers.\n\nAnother idea was to use images with different scale for training the model. I realized that unlike the other classes a lake or a river can have all sort of shapes and sizes (all cars will have a similar size). So I trained the model with images with scale 1:1, 1/2, 1/4 and 1/8. That way my training set was much bigger.\n\nBefore feeding the images to the high level classifier I made collages with them because sometimes in a single image we don't have enought information to decide to which category they belong.   \n\n![enter image description here][6]\nThe high level classifier had only two steps: in the first I use a threshold on the contour size to search for big rivers. In the second I use the distance to the big rivers to decide if a contour was a lake or a river. \n\nHaving a good water model was important, because later I used this model to clean the output of some of the others.\n\n\n\n### Building, road, tree, crops, track\nI don't think this models have something special. I just tried with different parameters such as: resize, window size, depth, sampling until I find a good configuration.\n\n### Structure\nThis model was hard to train. I had to use extracost of 2 or 4 on white pixels to force the model to learn. Otherwise it painted all black.\n\n### Small and big vehicles\nI spend the last days of the contest trying to improve this two categories.  \nMy final submission consisted on averaging the prediction of a lot of models (up to 15 different models). For trainign different models I used cross-validation with different random seeds.  \nI cleaned the prediction using the prediction of the other models. I used the water, building and tree model for cleaning.\n\n\n## 4. Scores\n![enter image description here][7]\n\n\nPlease let me now if something is not clear or if I have to elaborate more  on some topic.  \nI will be adding more info if necessary to the first post.   \n\nThanks to the organizers and the participants.  \nironbar\n\n\n  [1]: http://imgur.com/exlz22p.png\n  [2]: http://imgur.com/1vdZ5em.png\n  [3]: http://imgur.com/8jPC2UM.png\n  [4]: http://imgur.com/R5QQiJG.png\n  [5]: http://imgur.com/7LHvaOr.png\n  [6]: http://imgur.com/9K6XUdg.png\n  [7]: http://imgur.com/9EJggch.png",
    "166336": "Pls share code for us to learn. Thank you very much.",
    "166289": "Tks for sharing, do you use only M bands of all the bands?",
    "183594": "grt work!\nbut one silly question... how can we get lb score for one class?\n",
    "166949": "Thank a lot for sharing.\n\nI have a question:\n\n\"Instead of applying white balance to single images I made collages(for example all the images that start by 6010) and I applied white balance to the whole collage. I made this for all the satellite channels. \"\n\n1. It seems you are using channel stretching, what is the upper and lower percentage thresholds you used? \n\n2. Somehow, you applied stretching to whole area, for example, in training set, it is 5x5, so for each channel, you read all 25 images and then do channel stretching and save the images?\n\nThanks!\n\n\n\n\n",
    "166340": "Thanks for sharing this! Could you elaborate on the method you used to integrate different bands?",
    "234261": ""
  }
}