{
  "id": 32076,
  "title": "133rd place solution with code;)",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/32076",
  "author_name": "Devin Anzelmo",
  "post_date": "2017-04-25T18:28:02.542000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Here is code derived from my 133rd place solution trained on M images \n<a href=\"https://github.com/danzelmo/dstl-competition\">https://github.com/danzelmo/dstl-competition</a> (0.39 public, 0.35 private)</p>\n\n<p><strong>Edit</strong> in description of what this model is:\ncnn used here is based on dilated convolutions, it does not use max pooling. No padding is used in order to have good border predictions. Dilation is used early in network to reduce size. Above score is from a very simple version that is not tuned, higher scores achievable if resolution is increased. This is something I am experimenting with.  </p>\n\n<p>My actual solution was made up of 9 of these models one for each class trained on P image resolution data. </p>\n\n<p>Thanks to everyone who posted solutions, and helpful starter code. I learned so much about image processing and cnn from reading the forums. </p>\n\n<p>This competition reminded me of various important ideas.  </p>\n\n<p><strong>Look at the full data before anything else.</strong>\nspent a lot of time looking at images, didn't put the grids together and take a look until half way through, and realizing this perspective simplified thinking about the problem immensely. </p>\n\n<p><strong>Good software design, and engineering habits are really important</strong>\nWrote a bit of brittle code that required re-engineering whenever I changed my mind on how to solve things. Wasted time and got frustrated.\nAlso unit testing would have helped a lot initially, spent a lot of time using pdb which is a plus. </p>\n\n<p><strong>Reduce the problem size, get the simplest full solution done immediately, then iterate.</strong>\nMost of the complications I encountered were because I neglected to follow this very good advice. And I knew I wasn't the whole time, so doubly silly of me. I didn't downsample the images. I assumed that the polygon labels would correspond best to the rgb resolution, and went about resizing and then aligning the M, and A images. Now my dataset is 200+ gigs and I have no ssd drive:)</p>\n\n<p><strong>Wandering around, and trying weird solutions can be fun and informative</strong> \nBecause I  chose to neglect some of the the above advice, I read many interesting papers, learned several things that I would not have if I took the practical route. Some times following whatever aspect is interesting can be very rewarding. </p>\n\n<h3>Somewhat chronological notes on things I tried.</h3>\n\n<ul>\n<li><p>I made the decision to base my first solution off of logistic regression and lightgbm. </p>\n\n<ul><li>But with 4 billion pixels to predict this actually takes a while with lightgbm. </li>\n<li>With subsampling was able to do hold one image out cross validation using logistic regression and lightgbm</li></ul></li>\n<li><p>Found this paper: Features, color spaces, and boosting: New insights on semantic classification of remote sensing images <a href=\"https://www.researchgate.net/profile/Jan_Wegner2/publication/264564101_Features_Color_Spaces_and_Boosting_New_Insights_on_Semantic_Classification_of_Remote_Sensing_Images/links/553e2b7e0cf294deef6fc6bf.pdf\">link to paper</a></p>\n\n<ul><li>Seemed easy enough to implement, and would be faster with lightgbm</li>\n<li>Learned a tiny bit of cython and what integral images are so that I could implement a lot of this feature processing code.</li>\n<li>Realized I could do it much faster with opencv</li>\n<li>Saving anything intermediate to disk was not feasible(up to 14k features per pixel location)</li>\n<li>Implemented this so all features computed on-line</li>\n<li>Still way to cumbersome so abandoned idea. </li></ul></li>\n<li>Wrote various sampling code for training lightgbm. </li>\n<li>Wrote too much code before having a good idea what my solution would be like.\n<ul><li>Alternate the above with coming up with new approaches + too much data equals a ton of frustration with my code base.</li>\n<li>Several failed approaches</li>\n<li>Ended up just using jupyter notebooks, and made sure each one was somewhat self contained and left something on disk when finished.</li></ul></li>\n<li><p>Learned a bit about <a href=\"https://github.com/spotify/luigi\">luigi</a> and used it to organize feature generation batch jobs.</p></li>\n<li><p>I struggled with ethical concerns regarding looking at the test data, as this seemed like a really good way to validate my methods. There was no explicit hand labeling, but a quick glance at the test data could give a lot of information.</p></li>\n<li>Used kmeans to cluster individual images\n<ul><li>much easier to do this on full grids.</li></ul></li>\n<li>Figured I could reduce, computation time if I just never predicted on images that had low probability of having certain classes.  </li>\n<li><p>Generated 128x128 patches of the rgb resolution data</p>\n\n<ul><li>generate descriptive features for patches, train lightgbm for each class, and multiplied by four because I split the data into forest, desert, mountain, and deforested groups. Thinking like a decision tree.\n<ul><li>Lead to engineering headaches.</li></ul></li>\n<li>At test time only predict on patches that are classified to contain the given class.\n<ul><li>Implies I also need individual classifiers for each class, and for the four partitions of the data, plus keep everything organized\n<ul><li>too much overhead, eventually abandon to work on neural nets.</li>\n<li>Even if cnn take more effort to learn to use they are very efficient time wise for segmentation.</li></ul></li></ul></li></ul></li>\n<li><p>got into using <a href=\"http://toolz.readthedocs.io/en/latest/\">cytoolz</a>, and learned about currying. This was really fun. </p>\n\n<ul><li>Learned that some api's(like sklearn) work really nicely with currying and functional style.</li>\n<li>Others do not pandas was a pain to try to wrap. Things didn't translate well.</li>\n<li>Wrote a lot of silly bad functions, this was fun:) </li>\n<li>Abandoned this because anything I tried I could do faster and easier using default api for numpy/pandas</li></ul></li>\n<li><p>I didn't have any experience using convolutional neural nets, and wanted to learn.</p></li>\n<li>Saw this <a href=\"http://juliandewit.github.io/kaggle-ndsb/\">blog post</a> by Julian de Wit, using unet. Translated his mxnet code into keras.</li>\n<li>Unet posted on kaggle, I found I had bugs in my unet code. This helped get it running.</li>\n<li><p>Decided to try batchnormalization, forgot to add activation function</p>\n\n<ul><li>cnn can learn something even under bad circumstances, makes debugging a pain</li>\n<li>I found several bugs, that should have made learning very difficult, but still somehow it learns a little.</li>\n<li>Unit testing critical since the net won't tell you when its broken and only half learning.</li></ul></li>\n<li><p>Read this <a href=\"https://arxiv.org/abs/1606.00915\">deeplab crf</a> paper posted by Vladimir Iglovikov</p></li>\n<li><p>Found <a href=\"https://github.com/lucasb-eyer/pydensecrf\">pydensecrf</a> a partial wrapper for the crf code used in the deeplab paper</p>\n\n<ul><li>This gave quite a bit of improvement(validation on separate held out images), for some classes, not much for others</li>\n<li>Also took quite a while to run on full resolution images. </li></ul></li>\n<li><p>Read <a href=\"https://arxiv.org/abs/1511.07122\">Multi-Scale Context Aggregation by Dilated Convolutions</a></p></li>\n<li><p>Read many github comments, and academic papers. </p>\n\n<ul><li>Noticed a trend of people being dissatisfied with using max pooling for segmentation networks.</li>\n<li>Noticed that all current segmentation nets are based on architectures designed explicitly for classification</li>\n<li>Noticed classification nets use a lot of padding which is not good for edge predictions.</li>\n<li>Decided to tinker around with network architectures that could bypass these issues. </li></ul></li>\n<li><p>network architecture doesn't use padding and has cropping combined with dilated convolution early in the network to cut down on image size. </p>\n\n<ul><li>Doesn't have coarse grained view of data(no max pooling)</li>\n<li>Nice details but miss larger relationships</li>\n<li>Tried some average pooling for decimation.  </li>\n<li>Endless variation to try.  </li></ul></li>\n<li><p>Designing architecture for cnn is time consuming. 1.5 months, I just tinkered and tried things, was a ton of fun</p>\n\n<ul><li>Eventually competition is ending and I realized I still have not got to probably the most important part for having a good score in the competition: conversion between raster and polygon.</li>\n<li>Many ways to get creative with post processing(subtracting water from other class, remove small object, etc etc)</li>\n<li>Make one submission consisting of 9 of my nets, still experimenting with architecture so not are all the same, </li>\n<li>Ran out of time several of the nets were not fully trained</li></ul></li>\n<li><p>Get a very nice score on public leaderboard;) </p>\n\n<ul><li>Since I haven't been paying attention to practicalities of kaggle don't realize public dataset is tiny(evidence was there from some of the stats people were posting)</li>\n<li>Blend my solution with public benchmark for safety. </li>\n<li>Drop 113 places</li></ul></li>\n<li><p>Post competition, simplified everything, then simplified it some more. </p></li>\n<li>Net on all classes, on only M images gets 0.39 public, 0.35 private(I wouldn't have selected such a model with bad understanding of leaderboard.)  </li>\n</ul>\n\n<p>Overall this competition was really fun. Managing goals, whether caring about learning various things verses just getting done what works was a little frustrating. Lots of useful code posted on the forums helped me bypass steps and improve naive solutions. It would be nice if the competition was basically a collaborative effort, still learning how to balance information sharing with competition.</p>\n\n<h3>A couple of unanswered questions,</h3>\n\n<ul>\n<li>low resolution labels generate bad solutions on the leaderboard and in validation\n<ul><li>I didn't notice this so much when working with full resolution, very noticeable at M resolution</li>\n<li>Just upsampling the final low resolution predictions before conversion to polygons gave a large increases in score.</li>\n<li>It seems like there should be better ways to handle label conversion and creation to prevent mismatch. Did anyone have a great solution to this?</li></ul></li>\n<li><p>Did anyone else split the images into groups based on terrain type(I had four groups: desert, mountain, forest, and deforested). </p>\n\n<ul><li>Selecting thresholds by group instead of over all the images seemed to improve validation score.</li>\n<li>It seems like splitting like this would simplify things the learning algorithm has to learn, or would it not learn as general of a solution, thus degrading performance on new types of terrain in test set?</li></ul></li>\n<li><p>Are there any other segmentation architectures out there that don't use max pooling?</p></li>\n<li>Any ideas on creating the equivalent to pretrained weights similar to imagenet but with segmentation nets?. Has this been done anywhere?</li>\n</ul>",
  "messages": [
    {
      "id": 177741,
      "postDate": "2017-04-25T18:28:02.543Z",
      "content": "<p>Here is code derived from my 133rd place solution trained on M images \n<a href=\"https://github.com/danzelmo/dstl-competition\">https://github.com/danzelmo/dstl-competition</a> (0.39 public, 0.35 private)</p>\n\n<p><strong>Edit</strong> in description of what this model is:\ncnn used here is based on dilated convolutions, it does not use max pooling. No padding is used in order to have good border predictions. Dilation is used early in network to reduce size. Above score is from a very simple version that is not tuned, higher scores achievable if resolution is increased. This is something I am experimenting with.  </p>\n\n<p>My actual solution was made up of 9 of these models one for each class trained on P image resolution data. </p>\n\n<p>Thanks to everyone who posted solutions, and helpful starter code. I learned so much about image processing and cnn from reading the forums. </p>\n\n<p>This competition reminded me of various important ideas.  </p>\n\n<p><strong>Look at the full data before anything else.</strong>\nspent a lot of time looking at images, didn't put the grids together and take a look until half way through, and realizing this perspective simplified thinking about the problem immensely. </p>\n\n<p><strong>Good software design, and engineering habits are really important</strong>\nWrote a bit of brittle code that required re-engineering whenever I changed my mind on how to solve things. Wasted time and got frustrated.\nAlso unit testing would have helped a lot initially, spent a lot of time using pdb which is a plus. </p>\n\n<p><strong>Reduce the problem size, get the simplest full solution done immediately, then iterate.</strong>\nMost of the complications I encountered were because I neglected to follow this very good advice. And I knew I wasn't the whole time, so doubly silly of me. I didn't downsample the images. I assumed that the polygon labels would correspond best to the rgb resolution, and went about resizing and then aligning the M, and A images. Now my dataset is 200+ gigs and I have no ssd drive:)</p>\n\n<p><strong>Wandering around, and trying weird solutions can be fun and informative</strong> \nBecause I  chose to neglect some of the the above advice, I read many interesting papers, learned several things that I would not have if I took the practical route. Some times following whatever aspect is interesting can be very rewarding. </p>\n\n<h3>Somewhat chronological notes on things I tried.</h3>\n\n<ul>\n<li><p>I made the decision to base my first solution off of logistic regression and lightgbm. </p>\n\n<ul><li>But with 4 billion pixels to predict this actually takes a while with lightgbm. </li>\n<li>With subsampling was able to do hold one image out cross validation using logistic regression and lightgbm</li></ul></li>\n<li><p>Found this paper: Features, color spaces, and boosting: New insights on semantic classification of remote sensing images <a href=\"https://www.researchgate.net/profile/Jan_Wegner2/publication/264564101_Features_Color_Spaces_and_Boosting_New_Insights_on_Semantic_Classification_of_Remote_Sensing_Images/links/553e2b7e0cf294deef6fc6bf.pdf\">link to paper</a></p>\n\n<ul><li>Seemed easy enough to implement, and would be faster with lightgbm</li>\n<li>Learned a tiny bit of cython and what integral images are so that I could implement a lot of this feature processing code.</li>\n<li>Realized I could do it much faster with opencv</li>\n<li>Saving anything intermediate to disk was not feasible(up to 14k features per pixel location)</li>\n<li>Implemented this so all features computed on-line</li>\n<li>Still way to cumbersome so abandoned idea. </li></ul></li>\n<li>Wrote various sampling code for training lightgbm. </li>\n<li>Wrote too much code before having a good idea what my solution would be like.\n<ul><li>Alternate the above with coming up with new approaches + too much data equals a ton of frustration with my code base.</li>\n<li>Several failed approaches</li>\n<li>Ended up just using jupyter notebooks, and made sure each one was somewhat self contained and left something on disk when finished.</li></ul></li>\n<li><p>Learned a bit about <a href=\"https://github.com/spotify/luigi\">luigi</a> and used it to organize feature generation batch jobs.</p></li>\n<li><p>I struggled with ethical concerns regarding looking at the test data, as this seemed like a really good way to validate my methods. There was no explicit hand labeling, but a quick glance at the test data could give a lot of information.</p></li>\n<li>Used kmeans to cluster individual images\n<ul><li>much easier to do this on full grids.</li></ul></li>\n<li>Figured I could reduce, computation time if I just never predicted on images that had low probability of having certain classes.  </li>\n<li><p>Generated 128x128 patches of the rgb resolution data</p>\n\n<ul><li>generate descriptive features for patches, train lightgbm for each class, and multiplied by four because I split the data into forest, desert, mountain, and deforested groups. Thinking like a decision tree.\n<ul><li>Lead to engineering headaches.</li></ul></li>\n<li>At test time only predict on patches that are classified to contain the given class.\n<ul><li>Implies I also need individual classifiers for each class, and for the four partitions of the data, plus keep everything organized\n<ul><li>too much overhead, eventually abandon to work on neural nets.</li>\n<li>Even if cnn take more effort to learn to use they are very efficient time wise for segmentation.</li></ul></li></ul></li></ul></li>\n<li><p>got into using <a href=\"http://toolz.readthedocs.io/en/latest/\">cytoolz</a>, and learned about currying. This was really fun. </p>\n\n<ul><li>Learned that some api's(like sklearn) work really nicely with currying and functional style.</li>\n<li>Others do not pandas was a pain to try to wrap. Things didn't translate well.</li>\n<li>Wrote a lot of silly bad functions, this was fun:) </li>\n<li>Abandoned this because anything I tried I could do faster and easier using default api for numpy/pandas</li></ul></li>\n<li><p>I didn't have any experience using convolutional neural nets, and wanted to learn.</p></li>\n<li>Saw this <a href=\"http://juliandewit.github.io/kaggle-ndsb/\">blog post</a> by Julian de Wit, using unet. Translated his mxnet code into keras.</li>\n<li>Unet posted on kaggle, I found I had bugs in my unet code. This helped get it running.</li>\n<li><p>Decided to try batchnormalization, forgot to add activation function</p>\n\n<ul><li>cnn can learn something even under bad circumstances, makes debugging a pain</li>\n<li>I found several bugs, that should have made learning very difficult, but still somehow it learns a little.</li>\n<li>Unit testing critical since the net won't tell you when its broken and only half learning.</li></ul></li>\n<li><p>Read this <a href=\"https://arxiv.org/abs/1606.00915\">deeplab crf</a> paper posted by Vladimir Iglovikov</p></li>\n<li><p>Found <a href=\"https://github.com/lucasb-eyer/pydensecrf\">pydensecrf</a> a partial wrapper for the crf code used in the deeplab paper</p>\n\n<ul><li>This gave quite a bit of improvement(validation on separate held out images), for some classes, not much for others</li>\n<li>Also took quite a while to run on full resolution images. </li></ul></li>\n<li><p>Read <a href=\"https://arxiv.org/abs/1511.07122\">Multi-Scale Context Aggregation by Dilated Convolutions</a></p></li>\n<li><p>Read many github comments, and academic papers. </p>\n\n<ul><li>Noticed a trend of people being dissatisfied with using max pooling for segmentation networks.</li>\n<li>Noticed that all current segmentation nets are based on architectures designed explicitly for classification</li>\n<li>Noticed classification nets use a lot of padding which is not good for edge predictions.</li>\n<li>Decided to tinker around with network architectures that could bypass these issues. </li></ul></li>\n<li><p>network architecture doesn't use padding and has cropping combined with dilated convolution early in the network to cut down on image size. </p>\n\n<ul><li>Doesn't have coarse grained view of data(no max pooling)</li>\n<li>Nice details but miss larger relationships</li>\n<li>Tried some average pooling for decimation.  </li>\n<li>Endless variation to try.  </li></ul></li>\n<li><p>Designing architecture for cnn is time consuming. 1.5 months, I just tinkered and tried things, was a ton of fun</p>\n\n<ul><li>Eventually competition is ending and I realized I still have not got to probably the most important part for having a good score in the competition: conversion between raster and polygon.</li>\n<li>Many ways to get creative with post processing(subtracting water from other class, remove small object, etc etc)</li>\n<li>Make one submission consisting of 9 of my nets, still experimenting with architecture so not are all the same, </li>\n<li>Ran out of time several of the nets were not fully trained</li></ul></li>\n<li><p>Get a very nice score on public leaderboard;) </p>\n\n<ul><li>Since I haven't been paying attention to practicalities of kaggle don't realize public dataset is tiny(evidence was there from some of the stats people were posting)</li>\n<li>Blend my solution with public benchmark for safety. </li>\n<li>Drop 113 places</li></ul></li>\n<li><p>Post competition, simplified everything, then simplified it some more. </p></li>\n<li>Net on all classes, on only M images gets 0.39 public, 0.35 private(I wouldn't have selected such a model with bad understanding of leaderboard.)  </li>\n</ul>\n\n<p>Overall this competition was really fun. Managing goals, whether caring about learning various things verses just getting done what works was a little frustrating. Lots of useful code posted on the forums helped me bypass steps and improve naive solutions. It would be nice if the competition was basically a collaborative effort, still learning how to balance information sharing with competition.</p>\n\n<h3>A couple of unanswered questions,</h3>\n\n<ul>\n<li>low resolution labels generate bad solutions on the leaderboard and in validation\n<ul><li>I didn't notice this so much when working with full resolution, very noticeable at M resolution</li>\n<li>Just upsampling the final low resolution predictions before conversion to polygons gave a large increases in score.</li>\n<li>It seems like there should be better ways to handle label conversion and creation to prevent mismatch. Did anyone have a great solution to this?</li></ul></li>\n<li><p>Did anyone else split the images into groups based on terrain type(I had four groups: desert, mountain, forest, and deforested). </p>\n\n<ul><li>Selecting thresholds by group instead of over all the images seemed to improve validation score.</li>\n<li>It seems like splitting like this would simplify things the learning algorithm has to learn, or would it not learn as general of a solution, thus degrading performance on new types of terrain in test set?</li></ul></li>\n<li><p>Are there any other segmentation architectures out there that don't use max pooling?</p></li>\n<li>Any ideas on creating the equivalent to pretrained weights similar to imagenet but with segmentation nets?. Has this been done anywhere?</li>\n</ul>",
      "rawMarkdown": "Here is code derived from my 133rd place solution trained on M images \nhttps://github.com/danzelmo/dstl-competition (0.39 public, 0.35 private)\n\n**Edit** in description of what this model is:\ncnn used here is based on dilated convolutions, it does not use max pooling. No padding is used in order to have good border predictions. Dilation is used early in network to reduce size. Above score is from a very simple version that is not tuned, higher scores achievable if resolution is increased. This is something I am experimenting with.  \n\nMy actual solution was made up of 9 of these models one for each class trained on P image resolution data. \n\nThanks to everyone who posted solutions, and helpful starter code. I learned so much about image processing and cnn from reading the forums. \n\nThis competition reminded me of various important ideas.  \n \n\n**Look at the full data before anything else.**\nspent a lot of time looking at images, didn't put the grids together and take a look until half way through, and realizing this perspective simplified thinking about the problem immensely. \n\n**Good software design, and engineering habits are really important**\nWrote a bit of brittle code that required re-engineering whenever I changed my mind on how to solve things. Wasted time and got frustrated.\nAlso unit testing would have helped a lot initially, spent a lot of time using pdb which is a plus. \n\n**Reduce the problem size, get the simplest full solution done immediately, then iterate.**\nMost of the complications I encountered were because I neglected to follow this very good advice. And I knew I wasn't the whole time, so doubly silly of me. I didn't downsample the images. I assumed that the polygon labels would correspond best to the rgb resolution, and went about resizing and then aligning the M, and A images. Now my dataset is 200+ gigs and I have no ssd drive:)\n\n**Wandering around, and trying weird solutions can be fun and informative** \nBecause I  chose to neglect some of the the above advice, I read many interesting papers, learned several things that I would not have if I took the practical route. Some times following whatever aspect is interesting can be very rewarding. \n\n### Somewhat chronological notes on things I tried.  \n\n* I made the decision to base my first solution off of logistic regression and lightgbm. \n    * But with 4 billion pixels to predict this actually takes a while with lightgbm. \n    * With subsampling was able to do hold one image out cross validation using logistic regression and lightgbm\n \n* Found this paper: Features, color spaces, and boosting: New insights on semantic classification of remote sensing images [link to paper](https://www.researchgate.net/profile/Jan_Wegner2/publication/264564101_Features_Color_Spaces_and_Boosting_New_Insights_on_Semantic_Classification_of_Remote_Sensing_Images/links/553e2b7e0cf294deef6fc6bf.pdf)\n    * Seemed easy enough to implement, and would be faster with lightgbm\n    * Learned a tiny bit of cython and what integral images are so that I could implement a lot of this feature processing code.\n    * Realized I could do it much faster with opencv\n    * Saving anything intermediate to disk was not feasible(up to 14k features per pixel location)\n    * Implemented this so all features computed on-line\n    * Still way to cumbersome so abandoned idea. \n* Wrote various sampling code for training lightgbm. \n* Wrote too much code before having a good idea what my solution would be like.\n    * Alternate the above with coming up with new approaches + too much data equals a ton of frustration with my code base.\n    * Several failed approaches\n    * Ended up just using jupyter notebooks, and made sure each one was somewhat self contained and left something on disk when finished.\n* Learned a bit about [luigi](https://github.com/spotify/luigi) and used it to organize feature generation batch jobs.\n\n* I struggled with ethical concerns regarding looking at the test data, as this seemed like a really good way to validate my methods. There was no explicit hand labeling, but a quick glance at the test data could give a lot of information.\n* Used kmeans to cluster individual images\n    * much easier to do this on full grids.\n* Figured I could reduce, computation time if I just never predicted on images that had low probability of having certain classes.  \n* Generated 128x128 patches of the rgb resolution data\n    * generate descriptive features for patches, train lightgbm for each class, and multiplied by four because I split the data into forest, desert, mountain, and deforested groups. Thinking like a decision tree.\n        * Lead to engineering headaches.\n    * At test time only predict on patches that are classified to contain the given class.\n        * Implies I also need individual classifiers for each class, and for the four partitions of the data, plus keep everything organized\n            * too much overhead, eventually abandon to work on neural nets.\n            * Even if cnn take more effort to learn to use they are very efficient time wise for segmentation.\n\n* got into using [cytoolz](http://toolz.readthedocs.io/en/latest/), and learned about currying. This was really fun. \n    * Learned that some api's(like sklearn) work really nicely with currying and functional style.\n    * Others do not pandas was a pain to try to wrap. Things didn't translate well.\n    * Wrote a lot of silly bad functions, this was fun:) \n    * Abandoned this because anything I tried I could do faster and easier using default api for numpy/pandas\n\n* I didn't have any experience using convolutional neural nets, and wanted to learn.\n* Saw this [blog post](http://juliandewit.github.io/kaggle-ndsb/) by Julian de Wit, using unet. Translated his mxnet code into keras.\n* Unet posted on kaggle, I found I had bugs in my unet code. This helped get it running.\n* Decided to try batchnormalization, forgot to add activation function\n    * cnn can learn something even under bad circumstances, makes debugging a pain\n    * I found several bugs, that should have made learning very difficult, but still somehow it learns a little.\n    * Unit testing critical since the net won't tell you when its broken and only half learning.\n\n* Read this [deeplab crf](https://arxiv.org/abs/1606.00915) paper posted by Vladimir Iglovikov\n* Found [pydensecrf](https://github.com/lucasb-eyer/pydensecrf) a partial wrapper for the crf code used in the deeplab paper\n    * This gave quite a bit of improvement(validation on separate held out images), for some classes, not much for others\n    * Also took quite a while to run on full resolution images. \n\n* Read [Multi-Scale Context Aggregation by Dilated Convolutions](https://arxiv.org/abs/1511.07122)\n* Read many github comments, and academic papers. \n    * Noticed a trend of people being dissatisfied with using max pooling for segmentation networks.\n    * Noticed that all current segmentation nets are based on architectures designed explicitly for classification\n    * Noticed classification nets use a lot of padding which is not good for edge predictions.\n    * Decided to tinker around with network architectures that could bypass these issues. \n\n* network architecture doesn't use padding and has cropping combined with dilated convolution early in the network to cut down on image size. \n    * Doesn't have coarse grained view of data(no max pooling)\n    * Nice details but miss larger relationships\n    * Tried some average pooling for decimation.  \n    * Endless variation to try.  \n\n* Designing architecture for cnn is time consuming. 1.5 months, I just tinkered and tried things, was a ton of fun\n    * Eventually competition is ending and I realized I still have not got to probably the most important part for having a good score in the competition: conversion between raster and polygon.\n    * Many ways to get creative with post processing(subtracting water from other class, remove small object, etc etc)\n    * Make one submission consisting of 9 of my nets, still experimenting with architecture so not are all the same, \n    * Ran out of time several of the nets were not fully trained\n\n* Get a very nice score on public leaderboard;) \n    * Since I haven't been paying attention to practicalities of kaggle don't realize public dataset is tiny(evidence was there from some of the stats people were posting)\n    * Blend my solution with public benchmark for safety. \n    * Drop 113 places\n\n* Post competition, simplified everything, then simplified it some more. \n* Net on all classes, on only M images gets 0.39 public, 0.35 private(I wouldn't have selected such a model with bad understanding of leaderboard.)  \n\nOverall this competition was really fun. Managing goals, whether caring about learning various things verses just getting done what works was a little frustrating. Lots of useful code posted on the forums helped me bypass steps and improve naive solutions. It would be nice if the competition was basically a collaborative effort, still learning how to balance information sharing with competition.\n\n### A couple of unanswered questions, \n* low resolution labels generate bad solutions on the leaderboard and in validation\n    *  I didn't notice this so much when working with full resolution, very noticeable at M resolution\n    *  Just upsampling the final low resolution predictions before conversion to polygons gave a large increases in score.\n    *  It seems like there should be better ways to handle label conversion and creation to prevent mismatch. Did anyone have a great solution to this?\n* Did anyone else split the images into groups based on terrain type(I had four groups: desert, mountain, forest, and deforested). \n    * Selecting thresholds by group instead of over all the images seemed to improve validation score.\n    * It seems like splitting like this would simplify things the learning algorithm has to learn, or would it not learn as general of a solution, thus degrading performance on new types of terrain in test set?\n\n* Are there any other segmentation architectures out there that don't use max pooling?\n* Any ideas on creating the equivalent to pretrained weights similar to imagenet but with segmentation nets?. Has this been done anywhere?",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "177741": "Here is code derived from my 133rd place solution trained on M images \nhttps://github.com/danzelmo/dstl-competition (0.39 public, 0.35 private)\n\n**Edit** in description of what this model is:\ncnn used here is based on dilated convolutions, it does not use max pooling. No padding is used in order to have good border predictions. Dilation is used early in network to reduce size. Above score is from a very simple version that is not tuned, higher scores achievable if resolution is increased. This is something I am experimenting with.  \n\nMy actual solution was made up of 9 of these models one for each class trained on P image resolution data. \n\nThanks to everyone who posted solutions, and helpful starter code. I learned so much about image processing and cnn from reading the forums. \n\nThis competition reminded me of various important ideas.  \n \n\n**Look at the full data before anything else.**\nspent a lot of time looking at images, didn't put the grids together and take a look until half way through, and realizing this perspective simplified thinking about the problem immensely. \n\n**Good software design, and engineering habits are really important**\nWrote a bit of brittle code that required re-engineering whenever I changed my mind on how to solve things. Wasted time and got frustrated.\nAlso unit testing would have helped a lot initially, spent a lot of time using pdb which is a plus. \n\n**Reduce the problem size, get the simplest full solution done immediately, then iterate.**\nMost of the complications I encountered were because I neglected to follow this very good advice. And I knew I wasn't the whole time, so doubly silly of me. I didn't downsample the images. I assumed that the polygon labels would correspond best to the rgb resolution, and went about resizing and then aligning the M, and A images. Now my dataset is 200+ gigs and I have no ssd drive:)\n\n**Wandering around, and trying weird solutions can be fun and informative** \nBecause I  chose to neglect some of the the above advice, I read many interesting papers, learned several things that I would not have if I took the practical route. Some times following whatever aspect is interesting can be very rewarding. \n\n### Somewhat chronological notes on things I tried.  \n\n* I made the decision to base my first solution off of logistic regression and lightgbm. \n    * But with 4 billion pixels to predict this actually takes a while with lightgbm. \n    * With subsampling was able to do hold one image out cross validation using logistic regression and lightgbm\n \n* Found this paper: Features, color spaces, and boosting: New insights on semantic classification of remote sensing images [link to paper](https://www.researchgate.net/profile/Jan_Wegner2/publication/264564101_Features_Color_Spaces_and_Boosting_New_Insights_on_Semantic_Classification_of_Remote_Sensing_Images/links/553e2b7e0cf294deef6fc6bf.pdf)\n    * Seemed easy enough to implement, and would be faster with lightgbm\n    * Learned a tiny bit of cython and what integral images are so that I could implement a lot of this feature processing code.\n    * Realized I could do it much faster with opencv\n    * Saving anything intermediate to disk was not feasible(up to 14k features per pixel location)\n    * Implemented this so all features computed on-line\n    * Still way to cumbersome so abandoned idea. \n* Wrote various sampling code for training lightgbm. \n* Wrote too much code before having a good idea what my solution would be like.\n    * Alternate the above with coming up with new approaches + too much data equals a ton of frustration with my code base.\n    * Several failed approaches\n    * Ended up just using jupyter notebooks, and made sure each one was somewhat self contained and left something on disk when finished.\n* Learned a bit about [luigi](https://github.com/spotify/luigi) and used it to organize feature generation batch jobs.\n\n* I struggled with ethical concerns regarding looking at the test data, as this seemed like a really good way to validate my methods. There was no explicit hand labeling, but a quick glance at the test data could give a lot of information.\n* Used kmeans to cluster individual images\n    * much easier to do this on full grids.\n* Figured I could reduce, computation time if I just never predicted on images that had low probability of having certain classes.  \n* Generated 128x128 patches of the rgb resolution data\n    * generate descriptive features for patches, train lightgbm for each class, and multiplied by four because I split the data into forest, desert, mountain, and deforested groups. Thinking like a decision tree.\n        * Lead to engineering headaches.\n    * At test time only predict on patches that are classified to contain the given class.\n        * Implies I also need individual classifiers for each class, and for the four partitions of the data, plus keep everything organized\n            * too much overhead, eventually abandon to work on neural nets.\n            * Even if cnn take more effort to learn to use they are very efficient time wise for segmentation.\n\n* got into using [cytoolz](http://toolz.readthedocs.io/en/latest/), and learned about currying. This was really fun. \n    * Learned that some api's(like sklearn) work really nicely with currying and functional style.\n    * Others do not pandas was a pain to try to wrap. Things didn't translate well.\n    * Wrote a lot of silly bad functions, this was fun:) \n    * Abandoned this because anything I tried I could do faster and easier using default api for numpy/pandas\n\n* I didn't have any experience using convolutional neural nets, and wanted to learn.\n* Saw this [blog post](http://juliandewit.github.io/kaggle-ndsb/) by Julian de Wit, using unet. Translated his mxnet code into keras.\n* Unet posted on kaggle, I found I had bugs in my unet code. This helped get it running.\n* Decided to try batchnormalization, forgot to add activation function\n    * cnn can learn something even under bad circumstances, makes debugging a pain\n    * I found several bugs, that should have made learning very difficult, but still somehow it learns a little.\n    * Unit testing critical since the net won't tell you when its broken and only half learning.\n\n* Read this [deeplab crf](https://arxiv.org/abs/1606.00915) paper posted by Vladimir Iglovikov\n* Found [pydensecrf](https://github.com/lucasb-eyer/pydensecrf) a partial wrapper for the crf code used in the deeplab paper\n    * This gave quite a bit of improvement(validation on separate held out images), for some classes, not much for others\n    * Also took quite a while to run on full resolution images. \n\n* Read [Multi-Scale Context Aggregation by Dilated Convolutions](https://arxiv.org/abs/1511.07122)\n* Read many github comments, and academic papers. \n    * Noticed a trend of people being dissatisfied with using max pooling for segmentation networks.\n    * Noticed that all current segmentation nets are based on architectures designed explicitly for classification\n    * Noticed classification nets use a lot of padding which is not good for edge predictions.\n    * Decided to tinker around with network architectures that could bypass these issues. \n\n* network architecture doesn't use padding and has cropping combined with dilated convolution early in the network to cut down on image size. \n    * Doesn't have coarse grained view of data(no max pooling)\n    * Nice details but miss larger relationships\n    * Tried some average pooling for decimation.  \n    * Endless variation to try.  \n\n* Designing architecture for cnn is time consuming. 1.5 months, I just tinkered and tried things, was a ton of fun\n    * Eventually competition is ending and I realized I still have not got to probably the most important part for having a good score in the competition: conversion between raster and polygon.\n    * Many ways to get creative with post processing(subtracting water from other class, remove small object, etc etc)\n    * Make one submission consisting of 9 of my nets, still experimenting with architecture so not are all the same, \n    * Ran out of time several of the nets were not fully trained\n\n* Get a very nice score on public leaderboard;) \n    * Since I haven't been paying attention to practicalities of kaggle don't realize public dataset is tiny(evidence was there from some of the stats people were posting)\n    * Blend my solution with public benchmark for safety. \n    * Drop 113 places\n\n* Post competition, simplified everything, then simplified it some more. \n* Net on all classes, on only M images gets 0.39 public, 0.35 private(I wouldn't have selected such a model with bad understanding of leaderboard.)  \n\nOverall this competition was really fun. Managing goals, whether caring about learning various things verses just getting done what works was a little frustrating. Lots of useful code posted on the forums helped me bypass steps and improve naive solutions. It would be nice if the competition was basically a collaborative effort, still learning how to balance information sharing with competition.\n\n### A couple of unanswered questions, \n* low resolution labels generate bad solutions on the leaderboard and in validation\n    *  I didn't notice this so much when working with full resolution, very noticeable at M resolution\n    *  Just upsampling the final low resolution predictions before conversion to polygons gave a large increases in score.\n    *  It seems like there should be better ways to handle label conversion and creation to prevent mismatch. Did anyone have a great solution to this?\n* Did anyone else split the images into groups based on terrain type(I had four groups: desert, mountain, forest, and deforested). \n    * Selecting thresholds by group instead of over all the images seemed to improve validation score.\n    * It seems like splitting like this would simplify things the learning algorithm has to learn, or would it not learn as general of a solution, thus degrading performance on new types of terrain in test set?\n\n* Are there any other segmentation architectures out there that don't use max pooling?\n* Any ideas on creating the equivalent to pretrained weights similar to imagenet but with segmentation nets?. Has this been done anywhere?"
  }
}