{
  "id": 27120,
  "title": "Modelling strategies",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/27120",
  "author_name": "raddar",
  "post_date": "2016-12-30T14:27:16.820000",
  "votes": 4,
  "comment_count": 25,
  "views": 1701,
  "content": "<p>I have spent few days on these images and still have not figured out what would be the best way to start off. Im thinking of building single pixel/pooled pixels model using colour band values of a pixel and its neighbouring pixels.</p>\n\n<p>The challange for me is how to reconstruct polygons from pixel probabilities. Anyone has any tips/examples/papers on that?</p>\n\n<p>Anyone going this way?</p>",
  "messages": [
    {
      "id": 153240,
      "postDate": "2016-12-30T15:11:21.080Z",
      "content": "<p>For reconstructing polygons, convert the pixels probabilities to a class prediction using some probability cutoff. Then convert that image to polygons using rasterio.features.shapes() function. Other people have mentioned using opencv to do it as well. </p>\n\n<p>I do it one class at a time. So for each class prediction I have a mask image of 0, and 1.  Here <code>mask</code> is one of those class masks, and all_polygons will become a shapely multipolygon object of the prediction. The checks at the end deal with some errors that may happen in the polygon creation.</p>\n\n<pre><code>def mask_to_polygons(mask):\n    all_polygons=[]\n    for shape, value in features.shapes(mask.astype(np.int16),\n                                mask = (mask==1),\n                                transform = rasterio.Affine(1.0, 0, 0, 0, 1.0, 0)):\n\n        all_polygons.append(shapely.geometry.shape(shape))\n\n    all_polygons = shapely.geometry.MultiPolygon(all_polygons)\n    if not all_polygons.is_valid:\n        all_polygons = all_polygons.buffer(0)\n        #Sometimes buffer() converts a simple Multipolygon to just a Polygon,\n        #need to keep it a Multi throughout\n        if all_polygons.type == 'Polygon':\n            all_polygons = shapely.geometry.MultiPolygon([all_polygons])\n    return all_polygons\n</code></pre>",
      "rawMarkdown": "For reconstructing polygons, convert the pixels probabilities to a class prediction using some probability cutoff. Then convert that image to polygons using rasterio.features.shapes() function. Other people have mentioned using opencv to do it as well. \r\n\r\nI do it one class at a time. So for each class prediction I have a mask image of 0, and 1.  Here `mask` is one of those class masks, and all_polygons will become a shapely multipolygon object of the prediction. The checks at the end deal with some errors that may happen in the polygon creation.\r\n\r\n\r\n    def mask_to_polygons(mask):\r\n        all_polygons=[]\r\n        for shape, value in features.shapes(mask.astype(np.int16),\r\n                                    mask = (mask==1),\r\n                                    transform = rasterio.Affine(1.0, 0, 0, 0, 1.0, 0)):\r\n    \r\n            all_polygons.append(shapely.geometry.shape(shape))\r\n    \r\n        all_polygons = shapely.geometry.MultiPolygon(all_polygons)\r\n        if not all_polygons.is_valid:\r\n            all_polygons = all_polygons.buffer(0)\r\n            #Sometimes buffer() converts a simple Multipolygon to just a Polygon,\r\n            #need to keep it a Multi throughout\r\n            if all_polygons.type == 'Polygon':\r\n                all_polygons = shapely.geometry.MultiPolygon([all_polygons])\r\n        return all_polygons\r\n\r\n",
      "votes": 14
    },
    {
      "id": 153441,
      "postDate": "2017-01-01T13:42:35.767Z",
      "content": "<p>Here <a href=\"https://github.com/trailbehind/DeepOSM\">DeepOSM</a>, in reading section there are several interesting approaches. Might be good lecture for obtaining a strategy.</p>",
      "rawMarkdown": "Here [DeepOSM][1], in reading section there are several interesting approaches. Might be good lecture for obtaining a strategy.\r\n\r\n\r\n  [1]: https://github.com/trailbehind/DeepOSM",
      "votes": 3
    },
    {
      "id": 153223,
      "postDate": "2016-12-30T14:27:16.820Z",
      "content": "<p>I have spent few days on these images and still have not figured out what would be the best way to start off. Im thinking of building single pixel/pooled pixels model using colour band values of a pixel and its neighbouring pixels.</p>\n\n<p>The challange for me is how to reconstruct polygons from pixel probabilities. Anyone has any tips/examples/papers on that?</p>\n\n<p>Anyone going this way?</p>",
      "rawMarkdown": "I have spent few days on these images and still have not figured out what would be the best way to start off. Im thinking of building single pixel/pooled pixels model using colour band values of a pixel and its neighbouring pixels.\r\n\r\nThe challange for me is how to reconstruct polygons from pixel probabilities. Anyone has any tips/examples/papers on that?\r\n\r\nAnyone going this way?",
      "votes": 4
    },
    {
      "id": 153793,
      "postDate": "2017-01-03T11:39:00.093Z",
      "content": "<p>Because we're talking about <strong>area</strong> and something empty does not have an area?!</p>",
      "rawMarkdown": "Because we're talking about **area** and something empty does not have an area?!",
      "votes": 1
    },
    {
      "id": 153790,
      "postDate": "2017-01-03T11:29:20.627Z",
      "content": "<p>Sorry, TP=0 and we can't say anything about TN, FP, FN but TP being 0 means that image will not add anything to the class score.</p>",
      "rawMarkdown": "Sorry, TP=0 and we can't say anything about TN, FP, FN but TP being 0 means that image will not add anything to the class score.",
      "votes": 1
    },
    {
      "id": 153789,
      "postDate": "2017-01-03T11:28:06.720Z",
      "content": "<p>I'm pretty sure you're wrong about  this :-))) The definition of Jaccard index for sets does not apply here. If you use Empty, TP=0, TN=0, FP=0, FN=0, etc. so that image does not contribute to your score.</p>\n\n<p>p.s.\nwhat's up with image 22? </p>",
      "rawMarkdown": "I'm pretty sure you're wrong about  this :-))) The definition of Jaccard index for sets does not apply here. If you use Empty, TP=0, TN=0, FP=0, FN=0, etc. so that image does not contribute to your score.\r\n\r\np.s.\r\nwhat's up with image 22? ",
      "votes": 1
    },
    {
      "id": 153764,
      "postDate": "2017-01-03T09:30:21.473Z",
      "content": "<p>Btw, 1 is the maximum possible score. There are 10 classes and each has a maximum score of 0.1.</p>",
      "rawMarkdown": "Btw, 1 is the maximum possible score. There are 10 classes and each has a maximum score of 0.1.\r\n",
      "votes": 1
    },
    {
      "id": 153761,
      "postDate": "2017-01-03T09:21:14.770Z",
      "content": "<p>I think empty polygon always result in tp=0 and Jaccard is calculated over the whole class, not for a single image. </p>",
      "rawMarkdown": "I think empty polygon always result in tp=0 and Jaccard is calculated over the whole class, not for a single image. ",
      "votes": 1
    },
    {
      "id": 153315,
      "postDate": "2016-12-31T02:15:20.700Z",
      "content": "<p>Vladimir, <code>features.shapes</code> is from the rasterio package.  So in the header of this file I have <br>\n<code>from rasterio import features</code></p>\n\n<p>See my script here about transforming the polygons using shapely to match the images using xmax and ymin. The output from the above script would get the inverse transformation to go back to the original polygon space. </p>\n\n<p><a href=\"https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/comments\">https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/comments</a></p>",
      "rawMarkdown": "Vladimir, `features.shapes` is from the rasterio package.  So in the header of this file I have   \r\n`from rasterio import features`\r\n\r\nSee my script here about transforming the polygons using shapely to match the images using xmax and ymin. The output from the above script would get the inverse transformation to go back to the original polygon space. \r\n\r\nhttps://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/comments",
      "votes": 2
    },
    {
      "id": 153229,
      "postDate": "2016-12-30T14:53:30.260Z",
      "content": "<p>[quote=raddar;153223]</p>\n\n<p>The challange for me is how to reconstruct polygons from pixel probabilities. Anyone has any tips/examples/papers on that?</p>\n\n<p>[/quote]</p>\n\n<p>@raddar I think the challenge is having the probabilities for each pixel :-) If you have that you can multiply the probability by 255 to have a grayscale image and use cv2.findContours for example (before running findContours you need to smooth out using a filter and threshold).</p>",
      "rawMarkdown": "[quote=raddar;153223]\r\n\r\nThe challange for me is how to reconstruct polygons from pixel probabilities. Anyone has any tips/examples/papers on that?\r\n\r\n[/quote]\r\n\r\n@raddar I think the challenge is having the probabilities for each pixel :-) If you have that you can multiply the probability by 255 to have a grayscale image and use cv2.findContours for example (before running findContours you need to smooth out using a filter and threshold).\r\n",
      "votes": 2
    },
    {
      "id": 153796,
      "postDate": "2017-01-03T11:45:22.650Z",
      "content": "<p>Sorry, my bad, i overlooked that TN was not used in a formula :) glad to have this sorted out. Nice discussion tho!</p>",
      "rawMarkdown": "Sorry, my bad, i overlooked that TN was not used in a formula :) glad to have this sorted out. Nice discussion tho!"
    },
    {
      "id": 153794,
      "postDate": "2017-01-03T11:41:05.213Z",
      "content": "<p>TN is not part of the formula, I made that up, sorry :-)</p>",
      "rawMarkdown": "TN is not part of the formula, I made that up, sorry :-)"
    },
    {
      "id": 153792,
      "postDate": "2017-01-03T11:36:01.933Z",
      "content": "<p>why TN=0? empty polygon would mean i am 100% accurate with detecting TN on empty-class image.</p>\n\n<p>regarding 22, many smallish-size FP polygons:) my overall table was before applying filters. reasonable size filters deal with this pretty well</p>",
      "rawMarkdown": "why TN=0? empty polygon would mean i am 100% accurate with detecting TN on empty-class image.\r\n\r\nregarding 22, many smallish-size FP polygons:) my overall table was before applying filters. reasonable size filters deal with this pretty well"
    },
    {
      "id": 153788,
      "postDate": "2017-01-03T11:25:50.450Z",
      "content": "<p>So as I brought this topic up, I think i was wrong to see it as a problem in first place. I think organizers set up the Jaccard index in the way to prevent such marginal cases with empty polygons, as small predicted polygons in an empty class image would have only minor impact to the overall score. summing TP's ,FP's and FN's over all images for each class seems to be a smart move from organizer's perspective.</p>",
      "rawMarkdown": "So as I brought this topic up, I think i was wrong to see it as a problem in first place. I think organizers set up the Jaccard index in the way to prevent such marginal cases with empty polygons, as small predicted polygons in an empty class image would have only minor impact to the overall score. summing TP's ,FP's and FN's over all images for each class seems to be a smart move from organizer's perspective."
    },
    {
      "id": 153782,
      "postDate": "2017-01-03T10:41:46.177Z",
      "content": "<p>Not sure if this is too obvious or if you missed, but when you replaced  SP&lt;50000 with Empty, you removed only FP's, so of course your score would improve. In the table above all rows with SP&lt;50000 have exactly 0 == TP.</p>\n\n<p>(my first comment about empty polygons was wrong, I edited it slightly. UPDATE: edited again).</p>",
      "rawMarkdown": "Not sure if this is too obvious or if you missed, but when you replaced  SP<50000 with Empty, you removed only FP's, so of course your score would improve. In the table above all rows with SP<50000 have exactly 0 == TP.\r\n\r\n(my first comment about empty polygons was wrong, I edited it slightly. UPDATE: edited again)."
    },
    {
      "id": 153778,
      "postDate": "2017-01-03T10:19:58.137Z",
      "content": "<p>In the table above 0.224 would be the Jaccard index for class 1? It doesn't seem right to me.</p>",
      "rawMarkdown": "In the table above 0.224 would be the Jaccard index for class 1? It doesn't seem right to me.",
      "replies": [
        {
          "id": 153787,
          "postDate": "2017-01-03T11:18:39.427Z",
          "content": "<p>[quote=amaia;153778]</p>\n\n<p>In the table above 0.224 would be the Jaccard index for class 1? It doesn't seem right to me.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, its average value of every images' jaccard index (in the way I calculated, not necessarily true to kaggle's evaluation function)</p>\n\n<p>[quote=amaia;153782]</p>\n\n<p>Not sure if this is too obvious or if you missed, but when you replaced  SP&lt;50000 with Empty, you removed only FP's, so of course your score would improve. In the table above all rows with SP&lt;50000 have exactly 0 == TP.</p>\n\n<p>(my first comment about empty polygons was wrong, I edited it slightly. UPDATE: edited again).</p>\n\n<p>[/quote]</p>\n\n<p>I removed both FP and TP's, therefore image 8 in my table would also be scored 0, but all other images where ST=0 would be scored as 1 :)</p>",
          "rawMarkdown": "[quote=amaia;153778]\r\n\r\nIn the table above 0.224 would be the Jaccard index for class 1? It doesn't seem right to me.\r\n\r\n[/quote]\r\n\r\nYes, its average value of every images' jaccard index (in the way I calculated, not necessarily true to kaggle's evaluation function)\r\n\r\n\r\n[quote=amaia;153782]\r\n\r\nNot sure if this is too obvious or if you missed, but when you replaced  SP<50000 with Empty, you removed only FP's, so of course your score would improve. In the table above all rows with SP<50000 have exactly 0 == TP.\r\n\r\n(my first comment about empty polygons was wrong, I edited it slightly. UPDATE: edited again).\r\n\r\n[/quote]\r\n\r\nI removed both FP and TP's, therefore image 8 in my table would also be scored 0, but all other images where ST=0 would be scored as 1 :)\r\n\r\n"
        }
      ]
    },
    {
      "id": 153775,
      "postDate": "2017-01-03T10:09:12.100Z",
      "content": "<p>my personal model for class1 : (ST number of TRUE pixels, SP: number of Predicted pixels)</p>\n\n<pre><code>   imageId       jac      ST      SP\n1        1 0.0000000       0      20\n2        2 0.0000000       0      27\n3        3 0.0000000       0     632\n4        4 0.0000000       0     329\n5        5 0.0000000       0    4004\n6        6 0.0000000       0    3812\n7        7 0.0000000       0      92\n8        8 0.1378115   19619   24676\n9        9 0.2889307  190811  369357\n10      10 0.0000000       0    1621\n11      11 0.4796804 1179770  923720\n12      12 0.5414345  422030  284226\n13      13 0.5687705 1094665  894707\n14      14 0.5251849 1700869 1169866\n15      15 0.5929104  590729  650599\n16      16 0.6435233  318399  307550\n17      17 0.3048369  246119   96432\n18      18 0.5211798 2660299 1962565\n19      19 0.5113194 1675009 1544591\n20      20 0.4946393  670967  949648\n21      21 0.0000000       0    8974\n22      22 0.0000000       0   31104\n23      23 0.0000000       0    2283\n24      24 0.0000000       0     720\n25      25 0.0000000       0    3460\n[1] 0.2244089\n</code></pre>\n\n<p>now, if i made ad-hoc rule, that i predict empy polygon when SP&lt;50000, i get jaccard index of 0.7388964 for this specific class on training set</p>\n\n<p>Edit: table above has incorrect Jaccard index calculations (and mean value as well), but will keep it for the reference in the context of discussion</p>",
      "rawMarkdown": "my personal model for class1 : (ST number of TRUE pixels, SP: number of Predicted pixels)\r\n\r\n       imageId       jac      ST      SP\r\n    1        1 0.0000000       0      20\r\n    2        2 0.0000000       0      27\r\n    3        3 0.0000000       0     632\r\n    4        4 0.0000000       0     329\r\n    5        5 0.0000000       0    4004\r\n    6        6 0.0000000       0    3812\r\n    7        7 0.0000000       0      92\r\n    8        8 0.1378115   19619   24676\r\n    9        9 0.2889307  190811  369357\r\n    10      10 0.0000000       0    1621\r\n    11      11 0.4796804 1179770  923720\r\n    12      12 0.5414345  422030  284226\r\n    13      13 0.5687705 1094665  894707\r\n    14      14 0.5251849 1700869 1169866\r\n    15      15 0.5929104  590729  650599\r\n    16      16 0.6435233  318399  307550\r\n    17      17 0.3048369  246119   96432\r\n    18      18 0.5211798 2660299 1962565\r\n    19      19 0.5113194 1675009 1544591\r\n    20      20 0.4946393  670967  949648\r\n    21      21 0.0000000       0    8974\r\n    22      22 0.0000000       0   31104\r\n    23      23 0.0000000       0    2283\r\n    24      24 0.0000000       0     720\r\n    25      25 0.0000000       0    3460\r\n    [1] 0.2244089\r\n\r\nnow, if i made ad-hoc rule, that i predict empy polygon when SP<50000, i get jaccard index of 0.7388964 for this specific class on training set\r\n\r\n\r\nEdit: table above has incorrect Jaccard index calculations (and mean value as well), but will keep it for the reference in the context of discussion\r\n"
    },
    {
      "id": 153773,
      "postDate": "2017-01-03T10:06:13.503Z",
      "content": "<p>I think you are wrong in both your comments:</p>\n\n<p>For <strong>each object class</strong>, of <strong>each image</strong>, we calculate the TP, FP, and FN areas. We then sum the total TP, total FP, and total FN across all the images, then the Jaccard is calculated for that class using total TP, total FP, and total FN. Then, we average all the Jaccard Indexes for all the 10 classes. </p>\n\n<p>Thats why I think some clarification on evaluation is needed:) Looking at wiki, it is explicitly said:\n\"If A and B are both empty, we define J(A,B) = 1\"</p>\n\n<p>Edit: On the other hand, reading the evaluation metric description, I might be wrong as well :)</p>",
      "rawMarkdown": "I think you are wrong in both your comments:\r\n\r\nFor **each object class**, of **each image**, we calculate the TP, FP, and FN areas. We then sum the total TP, total FP, and total FN across all the images, then the Jaccard is calculated for that class using total TP, total FP, and total FN. Then, we average all the Jaccard Indexes for all the 10 classes. \r\n\r\n\r\nThats why I think some clarification on evaluation is needed:) Looking at wiki, it is explicitly said:\r\n\"If A and B are both empty, we define J(A,B) = 1\"\r\n\r\nEdit: On the other hand, reading the evaluation metric description, I might be wrong as well :)\r\n"
    },
    {
      "id": 153758,
      "postDate": "2017-01-03T09:00:06.540Z",
      "content": "<p>If I understand correctly, empty polygon prediction should give Jaccard=1, when the image has no class pixels; On the other hand, you will get Jaccard = 0, if you predict at least one polygon on such image no matter what size it is. I think there will be some fine-tuning whether to submit empty polygon or not in those extremely-low frequency predictions in some of the images.</p>\n\n<p>Admins, could you confirm, that in Jaccard computation empty polygons on empty class images gives score of 1?</p>",
      "rawMarkdown": "If I understand correctly, empty polygon prediction should give Jaccard=1, when the image has no class pixels; On the other hand, you will get Jaccard = 0, if you predict at least one polygon on such image no matter what size it is. I think there will be some fine-tuning whether to submit empty polygon or not in those extremely-low frequency predictions in some of the images.\r\n\r\nAdmins, could you confirm, that in Jaccard computation empty polygons on empty class images gives score of 1?",
      "replies": [
        {
          "id": 153894,
          "postDate": "2017-01-03T23:17:09.237Z",
          "content": "<p>@raddar, </p>\n\n<p>A little late to the game and seems like you already sorted out. But I'd like to clarify that @amaia is correct: The Jaccard index is not per-image, it's calculating a total TP/FP/FN over a bunch of images for a given class. So if they are both empty, then TP=FP=FN=0, so totalTP/totalFP/totalFN don't change for this image.</p>\n\n<p>Strategy-wise, you are correct that smaller regions contribute less to the formula. All the TP/FP/FN are areas, so we're measuring your performance by the area, not each polygon or each image. </p>",
          "rawMarkdown": "@raddar, \n\nA little late to the game and seems like you already sorted out. But I'd like to clarify that @amaia is correct: The Jaccard index is not per-image, it's calculating a total TP/FP/FN over a bunch of images for a given class. So if they are both empty, then TP=FP=FN=0, so totalTP/totalFP/totalFN don't change for this image.\n\nStrategy-wise, you are correct that smaller regions contribute less to the formula. All the TP/FP/FN are areas, so we're measuring your performance by the area, not each polygon or each image. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 153312,
      "postDate": "2016-12-31T01:56:52.480Z",
      "content": "<p>@shawn, Thank you</p>\n\n<p>I have couple questions:</p>\n\n<ol>\n<li>Third line: <code>features.shapes</code>, what is this variable <code>features</code>?</li>\n<li>At what moment <code>x_max, y_min</code> come into play? Do we just do coordinate transformation on the polygon coordinates that are returned by your function?</li>\n</ol>",
      "rawMarkdown": "@shawn, Thank you\r\n\r\nI have couple questions:\r\n\r\n 1. Third line: `features.shapes`, what is this variable `features`?\r\n 2. At what moment `x_max, y_min` come into play? Do we just do coordinate transformation on the polygon coordinates that are returned by your function?"
    },
    {
      "id": 153230,
      "postDate": "2016-12-30T14:53:57.740Z",
      "content": "<p>I work with hyper spectral data using a RAMAN spectrograph (<a href=\"http://www.headwallphotonics.com/spectral-imaging/raman\">http://www.headwallphotonics.com/spectral-imaging/raman</a> ), but I have had time to solve the problem of reading the images, drawing the polygons, extracting the pixels for later link them to a class. The problem is basically \"classification\"</p>",
      "rawMarkdown": "\r\nI work with hyper spectral data using a RAMAN spectrograph (http://www.headwallphotonics.com/spectral-imaging/raman ), but I have had time to solve the problem of reading the images, drawing the polygons, extracting the pixels for later link them to a class. The problem is basically \"classification\""
    },
    {
      "id": 153225,
      "postDate": "2016-12-30T14:35:00.520Z",
      "content": "<p>Me too</p>",
      "rawMarkdown": "Me too"
    },
    {
      "id": 153224,
      "postDate": "2016-12-30T14:31:27.720Z",
      "content": "<p>nope....Stuck with images... :/</p>",
      "rawMarkdown": "nope....Stuck with images... :/"
    },
    {
      "id": 153336,
      "postDate": "2016-12-31T08:44:45.413Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 153240,
      "author_name": "shawn",
      "author_url": "",
      "post_date": "2016-12-30T15:11:21.080000",
      "content": "<p>For reconstructing polygons, convert the pixels probabilities to a class prediction using some probability cutoff. Then convert that image to polygons using rasterio.features.shapes() function. Other people have mentioned using opencv to do it as well. </p>\n\n<p>I do it one class at a time. So for each class prediction I have a mask image of 0, and 1.  Here <code>mask</code> is one of those class masks, and all_polygons will become a shapely multipolygon object of the prediction. The checks at the end deal with some errors that may happen in the polygon creation.</p>\n\n<pre><code>def mask_to_polygons(mask):\n    all_polygons=[]\n    for shape, value in features.shapes(mask.astype(np.int16),\n                                mask = (mask==1),\n                                transform = rasterio.Affine(1.0, 0, 0, 0, 1.0, 0)):\n\n        all_polygons.append(shapely.geometry.shape(shape))\n\n    all_polygons = shapely.geometry.MultiPolygon(all_polygons)\n    if not all_polygons.is_valid:\n        all_polygons = all_polygons.buffer(0)\n        #Sometimes buffer() converts a simple Multipolygon to just a Polygon,\n        #need to keep it a Multi throughout\n        if all_polygons.type == 'Polygon':\n            all_polygons = shapely.geometry.MultiPolygon([all_polygons])\n    return all_polygons\n</code></pre>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 153441,
      "author_name": "RafaLeo",
      "author_url": "",
      "post_date": "2017-01-01T13:42:35.767000",
      "content": "<p>Here <a href=\"https://github.com/trailbehind/DeepOSM\">DeepOSM</a>, in reading section there are several interesting approaches. Might be good lecture for obtaining a strategy.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 153793,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-03T11:39:00.093000",
      "content": "<p>Because we're talking about <strong>area</strong> and something empty does not have an area?!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 153790,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-03T11:29:20.627000",
      "content": "<p>Sorry, TP=0 and we can't say anything about TN, FP, FN but TP being 0 means that image will not add anything to the class score.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 153789,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-03T11:28:06.720000",
      "content": "<p>I'm pretty sure you're wrong about  this :-))) The definition of Jaccard index for sets does not apply here. If you use Empty, TP=0, TN=0, FP=0, FN=0, etc. so that image does not contribute to your score.</p>\n\n<p>p.s.\nwhat's up with image 22? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 153764,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-03T09:30:21.473000",
      "content": "<p>Btw, 1 is the maximum possible score. There are 10 classes and each has a maximum score of 0.1.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 153761,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-03T09:21:14.770000",
      "content": "<p>I think empty polygon always result in tp=0 and Jaccard is calculated over the whole class, not for a single image. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 153315,
      "author_name": "shawn",
      "author_url": "",
      "post_date": "2016-12-31T02:15:20.700000",
      "content": "<p>Vladimir, <code>features.shapes</code> is from the rasterio package.  So in the header of this file I have <br>\n<code>from rasterio import features</code></p>\n\n<p>See my script here about transforming the polygons using shapely to match the images using xmax and ymin. The output from the above script would get the inverse transformation to go back to the original polygon space. </p>\n\n<p><a href=\"https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/comments\">https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/comments</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 153229,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2016-12-30T14:53:30.260000",
      "content": "<p>[quote=raddar;153223]</p>\n\n<p>The challange for me is how to reconstruct polygons from pixel probabilities. Anyone has any tips/examples/papers on that?</p>\n\n<p>[/quote]</p>\n\n<p>@raddar I think the challenge is having the probabilities for each pixel :-) If you have that you can multiply the probability by 255 to have a grayscale image and use cv2.findContours for example (before running findContours you need to smooth out using a filter and threshold).</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 153796,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2017-01-03T11:45:22.650000",
      "content": "<p>Sorry, my bad, i overlooked that TN was not used in a formula :) glad to have this sorted out. Nice discussion tho!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153794,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-03T11:41:05.213000",
      "content": "<p>TN is not part of the formula, I made that up, sorry :-)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153792,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2017-01-03T11:36:01.933000",
      "content": "<p>why TN=0? empty polygon would mean i am 100% accurate with detecting TN on empty-class image.</p>\n\n<p>regarding 22, many smallish-size FP polygons:) my overall table was before applying filters. reasonable size filters deal with this pretty well</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153788,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2017-01-03T11:25:50.450000",
      "content": "<p>So as I brought this topic up, I think i was wrong to see it as a problem in first place. I think organizers set up the Jaccard index in the way to prevent such marginal cases with empty polygons, as small predicted polygons in an empty class image would have only minor impact to the overall score. summing TP's ,FP's and FN's over all images for each class seems to be a smart move from organizer's perspective.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153782,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-03T10:41:46.177000",
      "content": "<p>Not sure if this is too obvious or if you missed, but when you replaced  SP&lt;50000 with Empty, you removed only FP's, so of course your score would improve. In the table above all rows with SP&lt;50000 have exactly 0 == TP.</p>\n\n<p>(my first comment about empty polygons was wrong, I edited it slightly. UPDATE: edited again).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153778,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-03T10:19:58.137000",
      "content": "<p>In the table above 0.224 would be the Jaccard index for class 1? It doesn't seem right to me.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 153787,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2017-01-03T11:18:39.427000",
          "content": "<p>[quote=amaia;153778]</p>\n\n<p>In the table above 0.224 would be the Jaccard index for class 1? It doesn't seem right to me.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, its average value of every images' jaccard index (in the way I calculated, not necessarily true to kaggle's evaluation function)</p>\n\n<p>[quote=amaia;153782]</p>\n\n<p>Not sure if this is too obvious or if you missed, but when you replaced  SP&lt;50000 with Empty, you removed only FP's, so of course your score would improve. In the table above all rows with SP&lt;50000 have exactly 0 == TP.</p>\n\n<p>(my first comment about empty polygons was wrong, I edited it slightly. UPDATE: edited again).</p>\n\n<p>[/quote]</p>\n\n<p>I removed both FP and TP's, therefore image 8 in my table would also be scored 0, but all other images where ST=0 would be scored as 1 :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 153775,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2017-01-03T10:09:12.100000",
      "content": "<p>my personal model for class1 : (ST number of TRUE pixels, SP: number of Predicted pixels)</p>\n\n<pre><code>   imageId       jac      ST      SP\n1        1 0.0000000       0      20\n2        2 0.0000000       0      27\n3        3 0.0000000       0     632\n4        4 0.0000000       0     329\n5        5 0.0000000       0    4004\n6        6 0.0000000       0    3812\n7        7 0.0000000       0      92\n8        8 0.1378115   19619   24676\n9        9 0.2889307  190811  369357\n10      10 0.0000000       0    1621\n11      11 0.4796804 1179770  923720\n12      12 0.5414345  422030  284226\n13      13 0.5687705 1094665  894707\n14      14 0.5251849 1700869 1169866\n15      15 0.5929104  590729  650599\n16      16 0.6435233  318399  307550\n17      17 0.3048369  246119   96432\n18      18 0.5211798 2660299 1962565\n19      19 0.5113194 1675009 1544591\n20      20 0.4946393  670967  949648\n21      21 0.0000000       0    8974\n22      22 0.0000000       0   31104\n23      23 0.0000000       0    2283\n24      24 0.0000000       0     720\n25      25 0.0000000       0    3460\n[1] 0.2244089\n</code></pre>\n\n<p>now, if i made ad-hoc rule, that i predict empy polygon when SP&lt;50000, i get jaccard index of 0.7388964 for this specific class on training set</p>\n\n<p>Edit: table above has incorrect Jaccard index calculations (and mean value as well), but will keep it for the reference in the context of discussion</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153773,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2017-01-03T10:06:13.503000",
      "content": "<p>I think you are wrong in both your comments:</p>\n\n<p>For <strong>each object class</strong>, of <strong>each image</strong>, we calculate the TP, FP, and FN areas. We then sum the total TP, total FP, and total FN across all the images, then the Jaccard is calculated for that class using total TP, total FP, and total FN. Then, we average all the Jaccard Indexes for all the 10 classes. </p>\n\n<p>Thats why I think some clarification on evaluation is needed:) Looking at wiki, it is explicitly said:\n\"If A and B are both empty, we define J(A,B) = 1\"</p>\n\n<p>Edit: On the other hand, reading the evaluation metric description, I might be wrong as well :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153758,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2017-01-03T09:00:06.540000",
      "content": "<p>If I understand correctly, empty polygon prediction should give Jaccard=1, when the image has no class pixels; On the other hand, you will get Jaccard = 0, if you predict at least one polygon on such image no matter what size it is. I think there will be some fine-tuning whether to submit empty polygon or not in those extremely-low frequency predictions in some of the images.</p>\n\n<p>Admins, could you confirm, that in Jaccard computation empty polygons on empty class images gives score of 1?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 153894,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-03T23:17:09.237000",
          "content": "<p>@raddar, </p>\n\n<p>A little late to the game and seems like you already sorted out. But I'd like to clarify that @amaia is correct: The Jaccard index is not per-image, it's calculating a total TP/FP/FN over a bunch of images for a given class. So if they are both empty, then TP=FP=FN=0, so totalTP/totalFP/totalFN don't change for this image.</p>\n\n<p>Strategy-wise, you are correct that smaller regions contribute less to the formula. All the TP/FP/FN are areas, so we're measuring your performance by the area, not each polygon or each image. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 153312,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2016-12-31T01:56:52.480000",
      "content": "<p>@shawn, Thank you</p>\n\n<p>I have couple questions:</p>\n\n<ol>\n<li>Third line: <code>features.shapes</code>, what is this variable <code>features</code>?</li>\n<li>At what moment <code>x_max, y_min</code> come into play? Do we just do coordinate transformation on the polygon coordinates that are returned by your function?</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153230,
      "author_name": "Rodrigo Salas",
      "author_url": "",
      "post_date": "2016-12-30T14:53:57.740000",
      "content": "<p>I work with hyper spectral data using a RAMAN spectrograph (<a href=\"http://www.headwallphotonics.com/spectral-imaging/raman\">http://www.headwallphotonics.com/spectral-imaging/raman</a> ), but I have had time to solve the problem of reading the images, drawing the polygons, extracting the pixels for later link them to a class. The problem is basically \"classification\"</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153225,
      "author_name": "Santiago Mota",
      "author_url": "",
      "post_date": "2016-12-30T14:35:00.520000",
      "content": "<p>Me too</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153224,
      "author_name": "Nashit",
      "author_url": "",
      "post_date": "2016-12-30T14:31:27.720000",
      "content": "<p>nope....Stuck with images... :/</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153336,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-31T08:44:45.413000",
      "content": "",
      "votes": 4,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "153240": "For reconstructing polygons, convert the pixels probabilities to a class prediction using some probability cutoff. Then convert that image to polygons using rasterio.features.shapes() function. Other people have mentioned using opencv to do it as well. \r\n\r\nI do it one class at a time. So for each class prediction I have a mask image of 0, and 1.  Here `mask` is one of those class masks, and all_polygons will become a shapely multipolygon object of the prediction. The checks at the end deal with some errors that may happen in the polygon creation.\r\n\r\n\r\n    def mask_to_polygons(mask):\r\n        all_polygons=[]\r\n        for shape, value in features.shapes(mask.astype(np.int16),\r\n                                    mask = (mask==1),\r\n                                    transform = rasterio.Affine(1.0, 0, 0, 0, 1.0, 0)):\r\n    \r\n            all_polygons.append(shapely.geometry.shape(shape))\r\n    \r\n        all_polygons = shapely.geometry.MultiPolygon(all_polygons)\r\n        if not all_polygons.is_valid:\r\n            all_polygons = all_polygons.buffer(0)\r\n            #Sometimes buffer() converts a simple Multipolygon to just a Polygon,\r\n            #need to keep it a Multi throughout\r\n            if all_polygons.type == 'Polygon':\r\n                all_polygons = shapely.geometry.MultiPolygon([all_polygons])\r\n        return all_polygons\r\n\r\n",
    "153441": "Here [DeepOSM][1], in reading section there are several interesting approaches. Might be good lecture for obtaining a strategy.\r\n\r\n\r\n  [1]: https://github.com/trailbehind/DeepOSM",
    "153223": "I have spent few days on these images and still have not figured out what would be the best way to start off. Im thinking of building single pixel/pooled pixels model using colour band values of a pixel and its neighbouring pixels.\r\n\r\nThe challange for me is how to reconstruct polygons from pixel probabilities. Anyone has any tips/examples/papers on that?\r\n\r\nAnyone going this way?",
    "153793": "Because we're talking about **area** and something empty does not have an area?!",
    "153790": "Sorry, TP=0 and we can't say anything about TN, FP, FN but TP being 0 means that image will not add anything to the class score.",
    "153789": "I'm pretty sure you're wrong about  this :-))) The definition of Jaccard index for sets does not apply here. If you use Empty, TP=0, TN=0, FP=0, FN=0, etc. so that image does not contribute to your score.\r\n\r\np.s.\r\nwhat's up with image 22? ",
    "153764": "Btw, 1 is the maximum possible score. There are 10 classes and each has a maximum score of 0.1.\r\n",
    "153761": "I think empty polygon always result in tp=0 and Jaccard is calculated over the whole class, not for a single image. ",
    "153315": "Vladimir, `features.shapes` is from the rasterio package.  So in the header of this file I have   \r\n`from rasterio import features`\r\n\r\nSee my script here about transforming the polygons using shapely to match the images using xmax and ymin. The output from the above script would get the inverse transformation to go back to the original polygon space. \r\n\r\nhttps://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/comments",
    "153229": "[quote=raddar;153223]\r\n\r\nThe challange for me is how to reconstruct polygons from pixel probabilities. Anyone has any tips/examples/papers on that?\r\n\r\n[/quote]\r\n\r\n@raddar I think the challenge is having the probabilities for each pixel :-) If you have that you can multiply the probability by 255 to have a grayscale image and use cv2.findContours for example (before running findContours you need to smooth out using a filter and threshold).\r\n",
    "153796": "Sorry, my bad, i overlooked that TN was not used in a formula :) glad to have this sorted out. Nice discussion tho!",
    "153794": "TN is not part of the formula, I made that up, sorry :-)",
    "153792": "why TN=0? empty polygon would mean i am 100% accurate with detecting TN on empty-class image.\r\n\r\nregarding 22, many smallish-size FP polygons:) my overall table was before applying filters. reasonable size filters deal with this pretty well",
    "153788": "So as I brought this topic up, I think i was wrong to see it as a problem in first place. I think organizers set up the Jaccard index in the way to prevent such marginal cases with empty polygons, as small predicted polygons in an empty class image would have only minor impact to the overall score. summing TP's ,FP's and FN's over all images for each class seems to be a smart move from organizer's perspective.",
    "153782": "Not sure if this is too obvious or if you missed, but when you replaced  SP<50000 with Empty, you removed only FP's, so of course your score would improve. In the table above all rows with SP<50000 have exactly 0 == TP.\r\n\r\n(my first comment about empty polygons was wrong, I edited it slightly. UPDATE: edited again).",
    "153778": "In the table above 0.224 would be the Jaccard index for class 1? It doesn't seem right to me.",
    "153775": "my personal model for class1 : (ST number of TRUE pixels, SP: number of Predicted pixels)\r\n\r\n       imageId       jac      ST      SP\r\n    1        1 0.0000000       0      20\r\n    2        2 0.0000000       0      27\r\n    3        3 0.0000000       0     632\r\n    4        4 0.0000000       0     329\r\n    5        5 0.0000000       0    4004\r\n    6        6 0.0000000       0    3812\r\n    7        7 0.0000000       0      92\r\n    8        8 0.1378115   19619   24676\r\n    9        9 0.2889307  190811  369357\r\n    10      10 0.0000000       0    1621\r\n    11      11 0.4796804 1179770  923720\r\n    12      12 0.5414345  422030  284226\r\n    13      13 0.5687705 1094665  894707\r\n    14      14 0.5251849 1700869 1169866\r\n    15      15 0.5929104  590729  650599\r\n    16      16 0.6435233  318399  307550\r\n    17      17 0.3048369  246119   96432\r\n    18      18 0.5211798 2660299 1962565\r\n    19      19 0.5113194 1675009 1544591\r\n    20      20 0.4946393  670967  949648\r\n    21      21 0.0000000       0    8974\r\n    22      22 0.0000000       0   31104\r\n    23      23 0.0000000       0    2283\r\n    24      24 0.0000000       0     720\r\n    25      25 0.0000000       0    3460\r\n    [1] 0.2244089\r\n\r\nnow, if i made ad-hoc rule, that i predict empy polygon when SP<50000, i get jaccard index of 0.7388964 for this specific class on training set\r\n\r\n\r\nEdit: table above has incorrect Jaccard index calculations (and mean value as well), but will keep it for the reference in the context of discussion\r\n",
    "153773": "I think you are wrong in both your comments:\r\n\r\nFor **each object class**, of **each image**, we calculate the TP, FP, and FN areas. We then sum the total TP, total FP, and total FN across all the images, then the Jaccard is calculated for that class using total TP, total FP, and total FN. Then, we average all the Jaccard Indexes for all the 10 classes. \r\n\r\n\r\nThats why I think some clarification on evaluation is needed:) Looking at wiki, it is explicitly said:\r\n\"If A and B are both empty, we define J(A,B) = 1\"\r\n\r\nEdit: On the other hand, reading the evaluation metric description, I might be wrong as well :)\r\n",
    "153758": "If I understand correctly, empty polygon prediction should give Jaccard=1, when the image has no class pixels; On the other hand, you will get Jaccard = 0, if you predict at least one polygon on such image no matter what size it is. I think there will be some fine-tuning whether to submit empty polygon or not in those extremely-low frequency predictions in some of the images.\r\n\r\nAdmins, could you confirm, that in Jaccard computation empty polygons on empty class images gives score of 1?",
    "153312": "@shawn, Thank you\r\n\r\nI have couple questions:\r\n\r\n 1. Third line: `features.shapes`, what is this variable `features`?\r\n 2. At what moment `x_max, y_min` come into play? Do we just do coordinate transformation on the polygon coordinates that are returned by your function?",
    "153230": "\r\nI work with hyper spectral data using a RAMAN spectrograph (http://www.headwallphotonics.com/spectral-imaging/raman ), but I have had time to solve the problem of reading the images, drawing the polygons, extracting the pixels for later link them to a class. The problem is basically \"classification\"",
    "153225": "Me too",
    "153224": "nope....Stuck with images... :/",
    "153336": ""
  }
}