{
  "id": 26512,
  "title": " A little redundant",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/26512",
  "author_name": "",
  "post_date": "2016-12-15T04:50:32.573Z",
  "votes": -9,
  "comment_count": 8,
  "views": 1549,
  "content": "<p>Are  we  going to  predict the ,, multy polygons  here''?</p>",
  "messages": [
    {
      "id": "150421",
      "postDate": "12/15/2016 04:50:32",
      "content": "<p>Are  we  going to  predict the ,, multy polygons  here''?</p>",
      "rawMarkdown": "Are  we  going to  predict the ,, multy polygons  here''?",
      "votes": null
    },
    {
      "id": "150466",
      "postDate": "12/15/2016 09:59:52",
      "content": "<p>I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .</p>",
      "rawMarkdown": "I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .",
      "votes": null
    },
    {
      "id": "150483",
      "postDate": "12/15/2016 11:56:04",
      "content": "<p>[quote=Daia Alexandru;150466]</p>\n\n<p>I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .</p>\n\n<p>[/quote]</p>\n\n<p>Polygons are common in mapping and a multipolygon is just a list of polygons.</p>\n\n<p>For instance GoogleMaps makes available the use of polylines and polygons for use with their API (or when you want to select areas / draw lines, that's polygons / polylines). This is not counterintuitive as it allows you to define \"mathematically\" what you want in a dimensional space (if you were to store all the labels for each pixel, it would explode the file size).</p>\n\n<p>[quote=Daia Alexandru;150421]</p>\n\n<p>Are  we  going to  predict the ,, multy polygons  here''?</p>\n\n<p>[/quote]</p>\n\n<p>yes, you are going to predict the multipolygons.</p>\n\n<p>If you want to brute force things, you can raster the multipolygons to a full sized polygon, then convert to label per pixel. Afterwards, throw into a ML model to predict each pixel by using as features the pixel + its surroundings + use a convolution, but I don't think it is a good way to start working on this (although it could work). If you work on a 3,000x3,000 pixel file, then you have 9,000,000 observations, urgh!</p>",
      "rawMarkdown": "[quote=Daia Alexandru;150466]\r\n\r\nI am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .\r\n\r\n[/quote]\r\n\r\nPolygons are common in mapping and a multipolygon is just a list of polygons.\r\n\r\nFor instance GoogleMaps makes available the use of polylines and polygons for use with their API (or when you want to select areas / draw lines, that's polygons / polylines). This is not counterintuitive as it allows you to define \"mathematically\" what you want in a dimensional space (if you were to store all the labels for each pixel, it would explode the file size).\r\n\r\n[quote=Daia Alexandru;150421]\r\n\r\nAre  we  going to  predict the ,, multy polygons  here''?\r\n\r\n[/quote]\r\n\r\nyes, you are going to predict the multipolygons.\r\n\r\nIf you want to brute force things, you can raster the multipolygons to a full sized polygon, then convert to label per pixel. Afterwards, throw into a ML model to predict each pixel by using as features the pixel + its surroundings + use a convolution, but I don't think it is a good way to start working on this (although it could work). If you work on a 3,000x3,000 pixel file, then you have 9,000,000 observations, urgh!",
      "votes": null
    },
    {
      "id": "150510",
      "postDate": "12/15/2016 14:22:02",
      "content": "<p>[quote=Daia Alexandru;150466]</p>\n\n<p>I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .</p>\n\n<p>[/quote]</p>\n\n<p>What, in particular, is wrong with unclever?</p>",
      "rawMarkdown": "[quote=Daia Alexandru;150466]\r\n\r\nI am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .\r\n\r\n[/quote]\r\n\r\nWhat, in particular, is wrong with unclever?",
      "votes": null
    },
    {
      "id": "150561",
      "postDate": "12/15/2016 19:49:30",
      "content": "<p>[quote=inversion;150510]</p>\n\n<p>[quote=Daia Alexandru;150466]</p>\n\n<p>I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .</p>\n\n<p>[/quote]</p>\n\n<p>What, in particular, is wrong with unclever?</p>\n\n<p>Usualy  unclever is    equivelent  to   wrong.\nThe  manner in which  the  competition  details are posted   are   really  unclever  and   wrong.\nBy the  way   Mr.  Inversion ,   do you  know  what  is  supposed  to   be predicted  here?\nI am  feeling  funny abouyt   the   kaggle  hoax    :)))  I  have  more sumed downvotes then  number  of  competitors until now  :)) Ha Ha\n[/quote]</p>",
      "rawMarkdown": "[quote=inversion;150510]\r\n\r\n[quote=Daia Alexandru;150466]\r\n\r\nI am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .\r\n\r\n[/quote]\r\n\r\nWhat, in particular, is wrong with unclever?\r\n\r\nUsualy  unclever is    equivelent  to   wrong.\r\nThe  manner in which  the  competition  details are posted   are   really  unclever  and   wrong.\r\nBy the  way   Mr.  Inversion ,   do you  know  what  is  supposed  to   be predicted  here?\r\nI am  feeling  funny abouyt   the   kaggle  hoax    :)))  I  have  more sumed downvotes then  number  of  competitors until now  :)) Ha Ha\r\n[/quote]",
      "votes": null
    },
    {
      "id": "150585",
      "postDate": "12/15/2016 21:48:08",
      "content": "<p>Please don't let it get to you, Daia. <a href=\"https://www.kaggle.com/c/home-depot-product-search-relevance/forums/t/20248/any-way-around-new-submission-deadline-utc-is-misleading\">The Home Depot Forum</a>  had waaayyy more downvotes for Tom!</p>",
      "rawMarkdown": "Please don't let it get to you, Daia. [The Home Depot Forum](https://www.kaggle.com/c/home-depot-product-search-relevance/forums/t/20248/any-way-around-new-submission-deadline-utc-is-misleading)  had waaayyy more downvotes for Tom!",
      "votes": null
    },
    {
      "id": "150640",
      "postDate": "12/16/2016 05:23:25",
      "content": "<p>[quote=JohnM;150585]</p>\n\n<p>Please don't let it get to you, Daia. <a href=\"https://www.kaggle.com/c/home-depot-product-search-relevance/forums/t/20248/any-way-around-new-submission-deadline-utc-is-misleading\">The Home Depot Forum</a>  had waaayyy more downvotes for Tom!</p>\n\n<p>[/quote]\nHa , yeah , I notticed ,   it is something  specific   to  kaggle .</p>",
      "rawMarkdown": "[quote=JohnM;150585]\r\n\r\nPlease don't let it get to you, Daia. [The Home Depot Forum](https://www.kaggle.com/c/home-depot-product-search-relevance/forums/t/20248/any-way-around-new-submission-deadline-utc-is-misleading)  had waaayyy more downvotes for Tom!\r\n\r\n[/quote]\r\nHa , yeah , I notticed ,   it is something  specific   to  kaggle .",
      "votes": null
    },
    {
      "id": "150958",
      "postDate": "12/17/2016 21:12:59",
      "content": "<p>[quote=Laurae;150483]\nIf you want to brute force things, you can raster the multipolygons to a full sized polygon, then convert to label per pixel. Afterwards, throw into a ML model to predict each pixel by using as features the pixel + its surroundings + use a convolution, but I don't think it is a good way to start working on this (although it could work). If you work on a 3,000x3,000 pixel file, then you have 9,000,000 observations, urgh!\n[/quote]</p>\n\n<p>You described the exact process that came to my mind after looking at this dataset, but it would be very annoying/inefficient to actually implement. What is a 'good' way to approach this problem in your opinion?</p>",
      "rawMarkdown": "[quote=Laurae;150483]\r\nIf you want to brute force things, you can raster the multipolygons to a full sized polygon, then convert to label per pixel. Afterwards, throw into a ML model to predict each pixel by using as features the pixel + its surroundings + use a convolution, but I don't think it is a good way to start working on this (although it could work). If you work on a 3,000x3,000 pixel file, then you have 9,000,000 observations, urgh!\r\n[/quote]\r\n\r\nYou described the exact process that came to my mind after looking at this dataset, but it would be very annoying/inefficient to actually implement. What is a 'good' way to approach this problem in your opinion?",
      "votes": null
    },
    {
      "id": "150966",
      "postDate": "12/17/2016 22:58:09",
      "content": "<p>[quote=Jeff Delaney;150958]</p>\n\n<p>You described the exact process that came to my mind after looking at this dataset, but it would be very annoying/inefficient to actually implement. What is a 'good' way to approach this problem in your opinion?</p>\n\n<p>[/quote]</p>\n\n<p>The data is too much unbalanced for a 10-class classification, and we should reduce first the data otherwise it will explode (9M x 22). It might be worth trying a small batched neural network to see its performance (brute force method). Actually, it might be the best solution if all other solutions we might find are failing.</p>\n\n<p>I don't think there's a \"good\" solution until we find one working with an adequate (and better than brute force) performance, but some ideas coming to my mind which should be combined:</p>\n\n<ul>\n<li>Tailoring the process per label and/or using filters (like <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26508/manual-steps-allowed\">my example for dealing with roads using filtering methods</a>) - could be also a pre-processing method for feature extraction</li>\n<li>Semi-supervised approach (unsupervised+supervised nearest neighbors) to match convolutions/non-convolutions, then taking only parts to segregate learning</li>\n<li>Clustering on images (but for some labels it might be a disaster)</li>\n<li>Clustering on pixel features to get rid of redundant data using similarity (this method should get rid of 95%+ of predictions if not more - remember if you delete something, it should not skew the local validation and therefore you must weight back correctly each prediction)</li>\n<li>Hierarchical approach (find the difference between labeled / non-labeled) for data reduction, but to use this without supervised machine learning, one must find the correct sequence of filtering methods to separate labels and non-labels</li>\n<li>Any other method for reducing data (who is going to train on 9M * 22 pictures, and predict on the \"big bunch\" of test pictures?)</li>\n<li>Similarity-based approach combined with a regression-type model, this is the equivalent of turning the 10-class classification model to a regression model and vice-versa (from which you can threshold at your own will to optimize the performance metric backwards) - this should speed up learning by ~10x at the expense of a lower performance (then, you need to use a continuous optimizer to get the best thresholds for each label)</li>\n</ul>\n\n<p>When I saw <a href=\"https://www.kaggle.com/torrinos/dstl-satellite-imagery-feature-detection/exploration-and-plotting/run/553107\">this kernel</a> and I looked at it made me remember that not everything is covered, along with the <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/evaluation\">performance metric</a>, TP &gt;&gt; FP performance reward as long as TP doesn't come close to infinite, so a very simple machine learning model (like a linear model) tuned to optimize TP = FP &gt; ~1 should be able to get a better score than the benchmark (and fast). Therefore, rasterizing the multipolygon to pixels, and throwing it to a machine learning, and picking nearly pixels may make more sense currently (I was thinking at first that nearly everything was covered by a label, which is not the case). The only issue is the pre-processing the data and holding everything in memory, and adding the 11th class \"empty\" which is going to kill models' performance.</p>\n\n<p>Going \"incremental\" is probably the key to succeed in this competition (not learning all the stuff but learning tailored parts).</p>\n\n<p>I recently worked on a face scoring model, and time is clearly better spent reducing the data (too much is not always useful), crafting features, and using appropriate filters that seems obvious to the data scientist for the task than throwing a bazooka and get the results to start (\"garbage in garbage out\"). Data reduction might help a lot also (not predicting on all data but taking only a non-similar / appropriate subset, like what was done on Santander). But in the end I think we will all have to go through that bazooka (with \"good quality data\", not the \"garbage\") for the best performance.</p>\n\n<p>As the data is very large (when you take pixel per pixel), one might be interested into creating his/her own Private-Pubilc LB by using only a part (for public, or all for semi-private) of the test data to predict to avoid overfitting severely (\"hello skewed trees\", only ~22 big training pictures). That could be separating images, taking only the 50% first vertical of images, the 50% first horizontal...</p>\n\n<p><img src=\"https://www.kaggle.io/svf/553107/98eeb79dcf041472e247ff53602b7473/__results___files/__results___8_0.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "[quote=Jeff Delaney;150958]\r\n\r\nYou described the exact process that came to my mind after looking at this dataset, but it would be very annoying/inefficient to actually implement. What is a 'good' way to approach this problem in your opinion?\r\n\r\n[/quote]\r\n\r\nThe data is too much unbalanced for a 10-class classification, and we should reduce first the data otherwise it will explode (9M x 22). It might be worth trying a small batched neural network to see its performance (brute force method). Actually, it might be the best solution if all other solutions we might find are failing.\r\n\r\nI don't think there's a \"good\" solution until we find one working with an adequate (and better than brute force) performance, but some ideas coming to my mind which should be combined:\r\n\r\n* Tailoring the process per label and/or using filters (like [my example for dealing with roads using filtering methods][1]) - could be also a pre-processing method for feature extraction\r\n* Semi-supervised approach (unsupervised+supervised nearest neighbors) to match convolutions/non-convolutions, then taking only parts to segregate learning\r\n* Clustering on images (but for some labels it might be a disaster)\r\n* Clustering on pixel features to get rid of redundant data using similarity (this method should get rid of 95%+ of predictions if not more - remember if you delete something, it should not skew the local validation and therefore you must weight back correctly each prediction)\r\n* Hierarchical approach (find the difference between labeled / non-labeled) for data reduction, but to use this without supervised machine learning, one must find the correct sequence of filtering methods to separate labels and non-labels\r\n* Any other method for reducing data (who is going to train on 9M * 22 pictures, and predict on the \"big bunch\" of test pictures?)\r\n* Similarity-based approach combined with a regression-type model, this is the equivalent of turning the 10-class classification model to a regression model and vice-versa (from which you can threshold at your own will to optimize the performance metric backwards) - this should speed up learning by ~10x at the expense of a lower performance (then, you need to use a continuous optimizer to get the best thresholds for each label)\r\n\r\nWhen I saw [this kernel][2] and I looked at it made me remember that not everything is covered, along with the [performance metric][3], TP >> FP performance reward as long as TP doesn't come close to infinite, so a very simple machine learning model (like a linear model) tuned to optimize TP = FP > ~1 should be able to get a better score than the benchmark (and fast). Therefore, rasterizing the multipolygon to pixels, and throwing it to a machine learning, and picking nearly pixels may make more sense currently (I was thinking at first that nearly everything was covered by a label, which is not the case). The only issue is the pre-processing the data and holding everything in memory, and adding the 11th class \"empty\" which is going to kill models' performance.\r\n\r\nGoing \"incremental\" is probably the key to succeed in this competition (not learning all the stuff but learning tailored parts).\r\n\r\nI recently worked on a face scoring model, and time is clearly better spent reducing the data (too much is not always useful), crafting features, and using appropriate filters that seems obvious to the data scientist for the task than throwing a bazooka and get the results to start (\"garbage in garbage out\"). Data reduction might help a lot also (not predicting on all data but taking only a non-similar / appropriate subset, like what was done on Santander). But in the end I think we will all have to go through that bazooka (with \"good quality data\", not the \"garbage\") for the best performance.\r\n\r\nAs the data is very large (when you take pixel per pixel), one might be interested into creating his/her own Private-Pubilc LB by using only a part (for public, or all for semi-private) of the test data to predict to avoid overfitting severely (\"hello skewed trees\", only ~22 big training pictures). That could be separating images, taking only the 50% first vertical of images, the 50% first horizontal...\r\n\r\n![enter image description here][4]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26508/manual-steps-allowed\r\n  [2]: https://www.kaggle.com/torrinos/dstl-satellite-imagery-feature-detection/exploration-and-plotting/run/553107\r\n  [3]: https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/evaluation\r\n  [4]: https://www.kaggle.io/svf/553107/98eeb79dcf041472e247ff53602b7473/__results___files/__results___8_0.png",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 150466,
      "author_name": "",
      "author_url": "",
      "post_date": "12/15/2016 09:59:52",
      "content": "<p>I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .</p>",
      "votes": null,
      "replies": [
        {
          "id": 150483,
          "author_name": "laurae2",
          "author_url": "",
          "post_date": "12/15/2016 11:56:04",
          "content": "<p>[quote=Daia Alexandru;150466]</p>\n\n<p>I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .</p>\n\n<p>[/quote]</p>\n\n<p>Polygons are common in mapping and a multipolygon is just a list of polygons.</p>\n\n<p>For instance GoogleMaps makes available the use of polylines and polygons for use with their API (or when you want to select areas / draw lines, that's polygons / polylines). This is not counterintuitive as it allows you to define \"mathematically\" what you want in a dimensional space (if you were to store all the labels for each pixel, it would explode the file size).</p>\n\n<p>[quote=Daia Alexandru;150421]</p>\n\n<p>Are  we  going to  predict the ,, multy polygons  here''?</p>\n\n<p>[/quote]</p>\n\n<p>yes, you are going to predict the multipolygons.</p>\n\n<p>If you want to brute force things, you can raster the multipolygons to a full sized polygon, then convert to label per pixel. Afterwards, throw into a ML model to predict each pixel by using as features the pixel + its surroundings + use a convolution, but I don't think it is a good way to start working on this (although it could work). If you work on a 3,000x3,000 pixel file, then you have 9,000,000 observations, urgh!</p>",
          "votes": null,
          "replies": [
            {
              "id": 150958,
              "author_name": "jeffd23",
              "author_url": "",
              "post_date": "12/17/2016 21:12:59",
              "content": "<p>[quote=Laurae;150483]\nIf you want to brute force things, you can raster the multipolygons to a full sized polygon, then convert to label per pixel. Afterwards, throw into a ML model to predict each pixel by using as features the pixel + its surroundings + use a convolution, but I don't think it is a good way to start working on this (although it could work). If you work on a 3,000x3,000 pixel file, then you have 9,000,000 observations, urgh!\n[/quote]</p>\n\n<p>You described the exact process that came to my mind after looking at this dataset, but it would be very annoying/inefficient to actually implement. What is a 'good' way to approach this problem in your opinion?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 150966,
                  "author_name": "laurae2",
                  "author_url": "",
                  "post_date": "12/17/2016 22:58:09",
                  "content": "<p>[quote=Jeff Delaney;150958]</p>\n\n<p>You described the exact process that came to my mind after looking at this dataset, but it would be very annoying/inefficient to actually implement. What is a 'good' way to approach this problem in your opinion?</p>\n\n<p>[/quote]</p>\n\n<p>The data is too much unbalanced for a 10-class classification, and we should reduce first the data otherwise it will explode (9M x 22). It might be worth trying a small batched neural network to see its performance (brute force method). Actually, it might be the best solution if all other solutions we might find are failing.</p>\n\n<p>I don't think there's a \"good\" solution until we find one working with an adequate (and better than brute force) performance, but some ideas coming to my mind which should be combined:</p>\n\n<ul>\n<li>Tailoring the process per label and/or using filters (like <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26508/manual-steps-allowed\">my example for dealing with roads using filtering methods</a>) - could be also a pre-processing method for feature extraction</li>\n<li>Semi-supervised approach (unsupervised+supervised nearest neighbors) to match convolutions/non-convolutions, then taking only parts to segregate learning</li>\n<li>Clustering on images (but for some labels it might be a disaster)</li>\n<li>Clustering on pixel features to get rid of redundant data using similarity (this method should get rid of 95%+ of predictions if not more - remember if you delete something, it should not skew the local validation and therefore you must weight back correctly each prediction)</li>\n<li>Hierarchical approach (find the difference between labeled / non-labeled) for data reduction, but to use this without supervised machine learning, one must find the correct sequence of filtering methods to separate labels and non-labels</li>\n<li>Any other method for reducing data (who is going to train on 9M * 22 pictures, and predict on the \"big bunch\" of test pictures?)</li>\n<li>Similarity-based approach combined with a regression-type model, this is the equivalent of turning the 10-class classification model to a regression model and vice-versa (from which you can threshold at your own will to optimize the performance metric backwards) - this should speed up learning by ~10x at the expense of a lower performance (then, you need to use a continuous optimizer to get the best thresholds for each label)</li>\n</ul>\n\n<p>When I saw <a href=\"https://www.kaggle.com/torrinos/dstl-satellite-imagery-feature-detection/exploration-and-plotting/run/553107\">this kernel</a> and I looked at it made me remember that not everything is covered, along with the <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/evaluation\">performance metric</a>, TP &gt;&gt; FP performance reward as long as TP doesn't come close to infinite, so a very simple machine learning model (like a linear model) tuned to optimize TP = FP &gt; ~1 should be able to get a better score than the benchmark (and fast). Therefore, rasterizing the multipolygon to pixels, and throwing it to a machine learning, and picking nearly pixels may make more sense currently (I was thinking at first that nearly everything was covered by a label, which is not the case). The only issue is the pre-processing the data and holding everything in memory, and adding the 11th class \"empty\" which is going to kill models' performance.</p>\n\n<p>Going \"incremental\" is probably the key to succeed in this competition (not learning all the stuff but learning tailored parts).</p>\n\n<p>I recently worked on a face scoring model, and time is clearly better spent reducing the data (too much is not always useful), crafting features, and using appropriate filters that seems obvious to the data scientist for the task than throwing a bazooka and get the results to start (\"garbage in garbage out\"). Data reduction might help a lot also (not predicting on all data but taking only a non-similar / appropriate subset, like what was done on Santander). But in the end I think we will all have to go through that bazooka (with \"good quality data\", not the \"garbage\") for the best performance.</p>\n\n<p>As the data is very large (when you take pixel per pixel), one might be interested into creating his/her own Private-Pubilc LB by using only a part (for public, or all for semi-private) of the test data to predict to avoid overfitting severely (\"hello skewed trees\", only ~22 big training pictures). That could be separating images, taking only the 50% first vertical of images, the 50% first horizontal...</p>\n\n<p><img src=\"https://www.kaggle.io/svf/553107/98eeb79dcf041472e247ff53602b7473/__results___files/__results___8_0.png\" alt=\"enter image description here\" title=\"\"></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 150510,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "12/15/2016 14:22:02",
          "content": "<p>[quote=Daia Alexandru;150466]</p>\n\n<p>I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .</p>\n\n<p>[/quote]</p>\n\n<p>What, in particular, is wrong with unclever?</p>",
          "votes": null,
          "replies": [
            {
              "id": 150561,
              "author_name": "",
              "author_url": "",
              "post_date": "12/15/2016 19:49:30",
              "content": "<p>[quote=inversion;150510]</p>\n\n<p>[quote=Daia Alexandru;150466]</p>\n\n<p>I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .</p>\n\n<p>[/quote]</p>\n\n<p>What, in particular, is wrong with unclever?</p>\n\n<p>Usualy  unclever is    equivelent  to   wrong.\nThe  manner in which  the  competition  details are posted   are   really  unclever  and   wrong.\nBy the  way   Mr.  Inversion ,   do you  know  what  is  supposed  to   be predicted  here?\nI am  feeling  funny abouyt   the   kaggle  hoax    :)))  I  have  more sumed downvotes then  number  of  competitors until now  :)) Ha Ha\n[/quote]</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 150585,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "12/15/2016 21:48:08",
      "content": "<p>Please don't let it get to you, Daia. <a href=\"https://www.kaggle.com/c/home-depot-product-search-relevance/forums/t/20248/any-way-around-new-submission-deadline-utc-is-misleading\">The Home Depot Forum</a>  had waaayyy more downvotes for Tom!</p>",
      "votes": null,
      "replies": [
        {
          "id": 150640,
          "author_name": "",
          "author_url": "",
          "post_date": "12/16/2016 05:23:25",
          "content": "<p>[quote=JohnM;150585]</p>\n\n<p>Please don't let it get to you, Daia. <a href=\"https://www.kaggle.com/c/home-depot-product-search-relevance/forums/t/20248/any-way-around-new-submission-deadline-utc-is-misleading\">The Home Depot Forum</a>  had waaayyy more downvotes for Tom!</p>\n\n<p>[/quote]\nHa , yeah , I notticed ,   it is something  specific   to  kaggle .</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "150421": "Are  we  going to  predict the ,, multy polygons  here''?",
    "150466": "I am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .",
    "150483": "[quote=Daia Alexandru;150466]\r\n\r\nI am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .\r\n\r\n[/quote]\r\n\r\nPolygons are common in mapping and a multipolygon is just a list of polygons.\r\n\r\nFor instance GoogleMaps makes available the use of polylines and polygons for use with their API (or when you want to select areas / draw lines, that's polygons / polylines). This is not counterintuitive as it allows you to define \"mathematically\" what you want in a dimensional space (if you were to store all the labels for each pixel, it would explode the file size).\r\n\r\n[quote=Daia Alexandru;150421]\r\n\r\nAre  we  going to  predict the ,, multy polygons  here''?\r\n\r\n[/quote]\r\n\r\nyes, you are going to predict the multipolygons.\r\n\r\nIf you want to brute force things, you can raster the multipolygons to a full sized polygon, then convert to label per pixel. Afterwards, throw into a ML model to predict each pixel by using as features the pixel + its surroundings + use a convolution, but I don't think it is a good way to start working on this (although it could work). If you work on a 3,000x3,000 pixel file, then you have 9,000,000 observations, urgh!",
    "150510": "[quote=Daia Alexandru;150466]\r\n\r\nI am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .\r\n\r\n[/quote]\r\n\r\nWhat, in particular, is wrong with unclever?",
    "150561": "[quote=inversion;150510]\r\n\r\n[quote=Daia Alexandru;150466]\r\n\r\nI am  asking  because  of   so ,  this is  the  most   unclever thing  I  header   until now .\r\n\r\n[/quote]\r\n\r\nWhat, in particular, is wrong with unclever?\r\n\r\nUsualy  unclever is    equivelent  to   wrong.\r\nThe  manner in which  the  competition  details are posted   are   really  unclever  and   wrong.\r\nBy the  way   Mr.  Inversion ,   do you  know  what  is  supposed  to   be predicted  here?\r\nI am  feeling  funny abouyt   the   kaggle  hoax    :)))  I  have  more sumed downvotes then  number  of  competitors until now  :)) Ha Ha\r\n[/quote]",
    "150585": "Please don't let it get to you, Daia. [The Home Depot Forum](https://www.kaggle.com/c/home-depot-product-search-relevance/forums/t/20248/any-way-around-new-submission-deadline-utc-is-misleading)  had waaayyy more downvotes for Tom!",
    "150640": "[quote=JohnM;150585]\r\n\r\nPlease don't let it get to you, Daia. [The Home Depot Forum](https://www.kaggle.com/c/home-depot-product-search-relevance/forums/t/20248/any-way-around-new-submission-deadline-utc-is-misleading)  had waaayyy more downvotes for Tom!\r\n\r\n[/quote]\r\nHa , yeah , I notticed ,   it is something  specific   to  kaggle .",
    "150958": "[quote=Laurae;150483]\r\nIf you want to brute force things, you can raster the multipolygons to a full sized polygon, then convert to label per pixel. Afterwards, throw into a ML model to predict each pixel by using as features the pixel + its surroundings + use a convolution, but I don't think it is a good way to start working on this (although it could work). If you work on a 3,000x3,000 pixel file, then you have 9,000,000 observations, urgh!\r\n[/quote]\r\n\r\nYou described the exact process that came to my mind after looking at this dataset, but it would be very annoying/inefficient to actually implement. What is a 'good' way to approach this problem in your opinion?",
    "150966": "[quote=Jeff Delaney;150958]\r\n\r\nYou described the exact process that came to my mind after looking at this dataset, but it would be very annoying/inefficient to actually implement. What is a 'good' way to approach this problem in your opinion?\r\n\r\n[/quote]\r\n\r\nThe data is too much unbalanced for a 10-class classification, and we should reduce first the data otherwise it will explode (9M x 22). It might be worth trying a small batched neural network to see its performance (brute force method). Actually, it might be the best solution if all other solutions we might find are failing.\r\n\r\nI don't think there's a \"good\" solution until we find one working with an adequate (and better than brute force) performance, but some ideas coming to my mind which should be combined:\r\n\r\n* Tailoring the process per label and/or using filters (like [my example for dealing with roads using filtering methods][1]) - could be also a pre-processing method for feature extraction\r\n* Semi-supervised approach (unsupervised+supervised nearest neighbors) to match convolutions/non-convolutions, then taking only parts to segregate learning\r\n* Clustering on images (but for some labels it might be a disaster)\r\n* Clustering on pixel features to get rid of redundant data using similarity (this method should get rid of 95%+ of predictions if not more - remember if you delete something, it should not skew the local validation and therefore you must weight back correctly each prediction)\r\n* Hierarchical approach (find the difference between labeled / non-labeled) for data reduction, but to use this without supervised machine learning, one must find the correct sequence of filtering methods to separate labels and non-labels\r\n* Any other method for reducing data (who is going to train on 9M * 22 pictures, and predict on the \"big bunch\" of test pictures?)\r\n* Similarity-based approach combined with a regression-type model, this is the equivalent of turning the 10-class classification model to a regression model and vice-versa (from which you can threshold at your own will to optimize the performance metric backwards) - this should speed up learning by ~10x at the expense of a lower performance (then, you need to use a continuous optimizer to get the best thresholds for each label)\r\n\r\nWhen I saw [this kernel][2] and I looked at it made me remember that not everything is covered, along with the [performance metric][3], TP >> FP performance reward as long as TP doesn't come close to infinite, so a very simple machine learning model (like a linear model) tuned to optimize TP = FP > ~1 should be able to get a better score than the benchmark (and fast). Therefore, rasterizing the multipolygon to pixels, and throwing it to a machine learning, and picking nearly pixels may make more sense currently (I was thinking at first that nearly everything was covered by a label, which is not the case). The only issue is the pre-processing the data and holding everything in memory, and adding the 11th class \"empty\" which is going to kill models' performance.\r\n\r\nGoing \"incremental\" is probably the key to succeed in this competition (not learning all the stuff but learning tailored parts).\r\n\r\nI recently worked on a face scoring model, and time is clearly better spent reducing the data (too much is not always useful), crafting features, and using appropriate filters that seems obvious to the data scientist for the task than throwing a bazooka and get the results to start (\"garbage in garbage out\"). Data reduction might help a lot also (not predicting on all data but taking only a non-similar / appropriate subset, like what was done on Santander). But in the end I think we will all have to go through that bazooka (with \"good quality data\", not the \"garbage\") for the best performance.\r\n\r\nAs the data is very large (when you take pixel per pixel), one might be interested into creating his/her own Private-Pubilc LB by using only a part (for public, or all for semi-private) of the test data to predict to avoid overfitting severely (\"hello skewed trees\", only ~22 big training pictures). That could be separating images, taking only the 50% first vertical of images, the 50% first horizontal...\r\n\r\n![enter image description here][4]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26508/manual-steps-allowed\r\n  [2]: https://www.kaggle.com/torrinos/dstl-satellite-imagery-feature-detection/exploration-and-plotting/run/553107\r\n  [3]: https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/evaluation\r\n  [4]: https://www.kaggle.io/svf/553107/98eeb79dcf041472e247ff53602b7473/__results___files/__results___8_0.png"
  },
  "source": "meta"
}