{
  "id": 26807,
  "title": "Class overlap table (pixel based)",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/26807",
  "author_name": "",
  "post_date": "2016-12-22T06:22:49.510Z",
  "votes": 6,
  "comment_count": 6,
  "views": 367,
  "content": "<p>I'd like to share my class overlap table.  It's calculated using a pixel by pixel comparison over the large image size. More specifically, my class rasters have been created with the size of the large images(3348 by 3403).</p>\n\n<p>For each of the \"is\" classes (rows), I gathered all pixels for each of the \"is also\" (cols) and combined them all.  Hopefully the calculations are correct.</p>\n\n<p>Examples of how to read the table in English:</p>\n\n<p>\"Of all pixels classified as a Tree, 0.35% of them are also Buildings.\"</p>\n\n<p>\"22.3% of 'Structure' pixels are also 'Crop' pixels and/but only 0.57% of 'Crop' pixels are also 'Structure' pixels.\"  </p>\n\n<p>One can see that 'Waterway' is the most pure class.  If there was a 'Boat' class then they'd be more overlap with it. I suppose the small amount of 'waterway' that is also 'road' are bridges (although riverside trees figures a wee bit more).</p>\n\n<p><img src=\"http://funvert.com/kaggle/is_isalso.png\" alt=\"percentage comparison\" title=\"\"></p>",
  "messages": [
    {
      "id": "151807",
      "postDate": "12/22/2016 06:22:49",
      "content": "<p>I'd like to share my class overlap table.  It's calculated using a pixel by pixel comparison over the large image size. More specifically, my class rasters have been created with the size of the large images(3348 by 3403).</p>\n\n<p>For each of the \"is\" classes (rows), I gathered all pixels for each of the \"is also\" (cols) and combined them all.  Hopefully the calculations are correct.</p>\n\n<p>Examples of how to read the table in English:</p>\n\n<p>\"Of all pixels classified as a Tree, 0.35% of them are also Buildings.\"</p>\n\n<p>\"22.3% of 'Structure' pixels are also 'Crop' pixels and/but only 0.57% of 'Crop' pixels are also 'Structure' pixels.\"  </p>\n\n<p>One can see that 'Waterway' is the most pure class.  If there was a 'Boat' class then they'd be more overlap with it. I suppose the small amount of 'waterway' that is also 'road' are bridges (although riverside trees figures a wee bit more).</p>\n\n<p><img src=\"http://funvert.com/kaggle/is_isalso.png\" alt=\"percentage comparison\" title=\"\"></p>",
      "rawMarkdown": "I'd like to share my class overlap table.  It's calculated using a pixel by pixel comparison over the large image size. More specifically, my class rasters have been created with the size of the large images(3348 by 3403).\r\n\r\nFor each of the \"is\" classes (rows), I gathered all pixels for each of the \"is also\" (cols) and combined them all.  Hopefully the calculations are correct.\r\n\r\nExamples of how to read the table in English:\r\n\r\n\"Of all pixels classified as a Tree, 0.35% of them are also Buildings.\"\r\n\r\n\"22.3% of 'Structure' pixels are also 'Crop' pixels and/but only 0.57% of 'Crop' pixels are also 'Structure' pixels.\"  \r\n\r\nOne can see that 'Waterway' is the most pure class.  If there was a 'Boat' class then they'd be more overlap with it. I suppose the small amount of 'waterway' that is also 'road' are bridges (although riverside trees figures a wee bit more).\r\n\r\n![percentage comparison][1]\r\n  [1]: http://funvert.com/kaggle/is_isalso.png",
      "votes": null
    },
    {
      "id": "152128",
      "postDate": "12/23/2016 15:01:03",
      "content": "<p>Could you share the code? ....\nThanks</p>",
      "rawMarkdown": "Could you share the code? ....\r\nThanks",
      "votes": null
    },
    {
      "id": "152130",
      "postDate": "12/23/2016 15:04:07",
      "content": "<p>Nice! Very informative when you choose the train/test splits!</p>\n\n<p>Can you confirm that these results are with v4 wkt or with geojson? (aka the \"corrected' versions of labeled data)</p>\n\n<p>Thank you!</p>",
      "rawMarkdown": "Nice! Very informative when you choose the train/test splits!\r\n\r\nCan you confirm that these results are with v4 wkt or with geojson? (aka the \"corrected' versions of labeled data)\r\n\r\nThank you!",
      "votes": null
    },
    {
      "id": "152163",
      "postDate": "12/23/2016 19:02:06",
      "content": "<p>[quote=visoft;152130]</p>\n\n<p>Can you confirm that these results are with v4 wkt or with geojson? (aka the \"corrected' versions of labeled data)</p>\n\n<p>[/quote]</p>\n\n<p>yes, version 4 wkt.</p>",
      "rawMarkdown": "[quote=visoft;152130]\r\n\r\nCan you confirm that these results are with v4 wkt or with geojson? (aka the \"corrected' versions of labeled data)\r\n\r\n[/quote]\r\n\r\nyes, version 4 wkt.",
      "votes": null
    },
    {
      "id": "152165",
      "postDate": "12/23/2016 19:09:33",
      "content": "<p>When choosing which class that a training pixel should be from overlapping classes, I think it should be the one that is seen from the sky. For example  Tree -&gt; Crop, Car -&gt; Road, Building -&gt; Crop, Structure -&gt; Waterway, etc. Any thoughts?</p>",
      "rawMarkdown": "When choosing which class that a training pixel should be from overlapping classes, I think it should be the one that is seen from the sky. For example  Tree -> Crop, Car -> Road, Building -> Crop, Structure -> Waterway, etc. Any thoughts?",
      "votes": null
    },
    {
      "id": "152170",
      "postDate": "12/23/2016 19:44:21",
      "content": "<p>[quote=shawn;152165]</p>\n\n<p>When choosing which class that a training pixel should be from overlapping classes, I think it should be the one that is seen from the sky. For example  Tree -&gt; Crop, Car -&gt; Road, Building -&gt; Crop, Structure -&gt; Waterway, etc. Any thoughts?</p>\n\n<p>[/quote]</p>\n\n<p>I think I'll try using all the labels and model at the same time or have specialist models...depends how it goes.  When they evaluate, they'll evaluate each class so you'll need to have <em>both</em> trees and crops even if they have overlap.  I'm guessing the <em>crop model</em> (crops as in planting stuff) will learn to ignore trees and predict crop over large areas.  A tree model on the other hand (or car model) will need higher resolution and will label smaller groups of pixels.  Just rambling off the top of my head...net result is I think we should use all labels...create 10 truth masks per image.</p>",
      "rawMarkdown": "[quote=shawn;152165]\r\n\r\nWhen choosing which class that a training pixel should be from overlapping classes, I think it should be the one that is seen from the sky. For example  Tree -> Crop, Car -> Road, Building -> Crop, Structure -> Waterway, etc. Any thoughts?\r\n\r\n[/quote]\r\n\r\nI think I'll try using all the labels and model at the same time or have specialist models...depends how it goes.  When they evaluate, they'll evaluate each class so you'll need to have *both* trees and crops even if they have overlap.  I'm guessing the *crop model* (crops as in planting stuff) will learn to ignore trees and predict crop over large areas.  A tree model on the other hand (or car model) will need higher resolution and will label smaller groups of pixels.  Just rambling off the top of my head...net result is I think we should use all labels...create 10 truth masks per image.",
      "votes": null
    },
    {
      "id": "152172",
      "postDate": "12/23/2016 19:46:31",
      "content": "<p>[quote=Rodrigo Salas;152128]</p>\n\n<p>Could you share the code? ....\nThanks</p>\n\n<p>[/quote]</p>\n\n<p>maybe after the xmas madness.  I've asked for gdal to be included in the kernels...that would help since I used gdal and rasterio for some processing.\n( <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26876/gdal-in-kernels\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26876/gdal-in-kernels</a> )</p>",
      "rawMarkdown": "[quote=Rodrigo Salas;152128]\r\n\r\n\r\nCould you share the code? ....\r\nThanks\r\n\r\n[/quote]\r\n\r\nmaybe after the xmas madness.  I've asked for gdal to be included in the kernels...that would help since I used gdal and rasterio for some processing.\r\n( https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26876/gdal-in-kernels )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 152128,
      "author_name": "rsalaschile",
      "author_url": "",
      "post_date": "12/23/2016 15:01:03",
      "content": "<p>Could you share the code? ....\nThanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 152172,
          "author_name": "zerozero",
          "author_url": "",
          "post_date": "12/23/2016 19:46:31",
          "content": "<p>[quote=Rodrigo Salas;152128]</p>\n\n<p>Could you share the code? ....\nThanks</p>\n\n<p>[/quote]</p>\n\n<p>maybe after the xmas madness.  I've asked for gdal to be included in the kernels...that would help since I used gdal and rasterio for some processing.\n( <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26876/gdal-in-kernels\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26876/gdal-in-kernels</a> )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 152130,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "12/23/2016 15:04:07",
      "content": "<p>Nice! Very informative when you choose the train/test splits!</p>\n\n<p>Can you confirm that these results are with v4 wkt or with geojson? (aka the \"corrected' versions of labeled data)</p>\n\n<p>Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 152163,
          "author_name": "zerozero",
          "author_url": "",
          "post_date": "12/23/2016 19:02:06",
          "content": "<p>[quote=visoft;152130]</p>\n\n<p>Can you confirm that these results are with v4 wkt or with geojson? (aka the \"corrected' versions of labeled data)</p>\n\n<p>[/quote]</p>\n\n<p>yes, version 4 wkt.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 152165,
      "author_name": "shawn775",
      "author_url": "",
      "post_date": "12/23/2016 19:09:33",
      "content": "<p>When choosing which class that a training pixel should be from overlapping classes, I think it should be the one that is seen from the sky. For example  Tree -&gt; Crop, Car -&gt; Road, Building -&gt; Crop, Structure -&gt; Waterway, etc. Any thoughts?</p>",
      "votes": null,
      "replies": [
        {
          "id": 152170,
          "author_name": "zerozero",
          "author_url": "",
          "post_date": "12/23/2016 19:44:21",
          "content": "<p>[quote=shawn;152165]</p>\n\n<p>When choosing which class that a training pixel should be from overlapping classes, I think it should be the one that is seen from the sky. For example  Tree -&gt; Crop, Car -&gt; Road, Building -&gt; Crop, Structure -&gt; Waterway, etc. Any thoughts?</p>\n\n<p>[/quote]</p>\n\n<p>I think I'll try using all the labels and model at the same time or have specialist models...depends how it goes.  When they evaluate, they'll evaluate each class so you'll need to have <em>both</em> trees and crops even if they have overlap.  I'm guessing the <em>crop model</em> (crops as in planting stuff) will learn to ignore trees and predict crop over large areas.  A tree model on the other hand (or car model) will need higher resolution and will label smaller groups of pixels.  Just rambling off the top of my head...net result is I think we should use all labels...create 10 truth masks per image.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "151807": "I'd like to share my class overlap table.  It's calculated using a pixel by pixel comparison over the large image size. More specifically, my class rasters have been created with the size of the large images(3348 by 3403).\r\n\r\nFor each of the \"is\" classes (rows), I gathered all pixels for each of the \"is also\" (cols) and combined them all.  Hopefully the calculations are correct.\r\n\r\nExamples of how to read the table in English:\r\n\r\n\"Of all pixels classified as a Tree, 0.35% of them are also Buildings.\"\r\n\r\n\"22.3% of 'Structure' pixels are also 'Crop' pixels and/but only 0.57% of 'Crop' pixels are also 'Structure' pixels.\"  \r\n\r\nOne can see that 'Waterway' is the most pure class.  If there was a 'Boat' class then they'd be more overlap with it. I suppose the small amount of 'waterway' that is also 'road' are bridges (although riverside trees figures a wee bit more).\r\n\r\n![percentage comparison][1]\r\n  [1]: http://funvert.com/kaggle/is_isalso.png",
    "152128": "Could you share the code? ....\r\nThanks",
    "152130": "Nice! Very informative when you choose the train/test splits!\r\n\r\nCan you confirm that these results are with v4 wkt or with geojson? (aka the \"corrected' versions of labeled data)\r\n\r\nThank you!",
    "152163": "[quote=visoft;152130]\r\n\r\nCan you confirm that these results are with v4 wkt or with geojson? (aka the \"corrected' versions of labeled data)\r\n\r\n[/quote]\r\n\r\nyes, version 4 wkt.",
    "152165": "When choosing which class that a training pixel should be from overlapping classes, I think it should be the one that is seen from the sky. For example  Tree -> Crop, Car -> Road, Building -> Crop, Structure -> Waterway, etc. Any thoughts?",
    "152170": "[quote=shawn;152165]\r\n\r\nWhen choosing which class that a training pixel should be from overlapping classes, I think it should be the one that is seen from the sky. For example  Tree -> Crop, Car -> Road, Building -> Crop, Structure -> Waterway, etc. Any thoughts?\r\n\r\n[/quote]\r\n\r\nI think I'll try using all the labels and model at the same time or have specialist models...depends how it goes.  When they evaluate, they'll evaluate each class so you'll need to have *both* trees and crops even if they have overlap.  I'm guessing the *crop model* (crops as in planting stuff) will learn to ignore trees and predict crop over large areas.  A tree model on the other hand (or car model) will need higher resolution and will label smaller groups of pixels.  Just rambling off the top of my head...net result is I think we should use all labels...create 10 truth masks per image.",
    "152172": "[quote=Rodrigo Salas;152128]\r\n\r\n\r\nCould you share the code? ....\r\nThanks\r\n\r\n[/quote]\r\n\r\nmaybe after the xmas madness.  I've asked for gdal to be included in the kernels...that would help since I used gdal and rasterio for some processing.\r\n( https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26876/gdal-in-kernels )"
  },
  "source": "meta"
}