{
  "id": 26624,
  "title": "Meaning of Xmax and Ymin",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/26624",
  "author_name": "",
  "post_date": "2016-12-17T22:29:20.280Z",
  "votes": 6,
  "comment_count": 8,
  "views": 465,
  "content": "<blockquote>\n  <p>To utilize these images, we provide the grid coordinates of each image\n  so you know how to scale them and align them with the images in\n  pixels. In grid_sizes.csv, you are given the Xmax and Ymin values for\n  each imageId.\n  For each image, you should be able to get the width (W) and height (H)\n  from the image raster. For a 3-band image that is 3391 x 3349 x 3, W\n  is 3349, and H is 3391.</p>\n</blockquote>\n\n<p>Just trying to check my understanding here.  It seems that the Xmax and Ymin values for a given ImageId are to be used to calculate a consistent physical area per pixel (even if we don't know exactly what that physical area is) and have nothing to do with the location of the image on the surface of Earth or relative to another ImageId.   Is that right?  </p>\n\n<p>For example, here are a few rows from grid_sizes.csv ... </p>\n\n<pre><code>  ImageId      Xmax      Ymin  \n 6070_2_3  0.009013 -0.009045   \n 6040_2_2  0.009158 -0.009043   \n 6170_2_4  0.009150 -0.009042  \n 6100_2_3  0.009156 -0.009042   \n 6040_1_0  0.009156 -0.009042  \n</code></pre>\n\n<p>Should we interpret this data as meaning the W=3349 pixels in the 3-band image with ImageId 6040_1_0 cover a physical distance of 0.009156 and that the H=3391 pixels cover a physical distance of 0.009042?  </p>\n\n<p>I'm having trouble understanding the utility of the W^{\\prime}=Wp and H^{\\prime}=Hp variables in the tutorial as well.  For example, using the formula from the tutorial we get Wp = W * W / (W + 1) = 3348.00029851.  In the next equation we have x/xmax.  What is x supposed to be here? If x ranges from 0 to xmax, then x_prime ranges from 0 to Wp.  Is this useful in some way that I'm not seeing?  </p>",
  "messages": [
    {
      "id": "150964",
      "postDate": "12/17/2016 22:29:20",
      "content": "<blockquote>\n  <p>To utilize these images, we provide the grid coordinates of each image\n  so you know how to scale them and align them with the images in\n  pixels. In grid_sizes.csv, you are given the Xmax and Ymin values for\n  each imageId.\n  For each image, you should be able to get the width (W) and height (H)\n  from the image raster. For a 3-band image that is 3391 x 3349 x 3, W\n  is 3349, and H is 3391.</p>\n</blockquote>\n\n<p>Just trying to check my understanding here.  It seems that the Xmax and Ymin values for a given ImageId are to be used to calculate a consistent physical area per pixel (even if we don't know exactly what that physical area is) and have nothing to do with the location of the image on the surface of Earth or relative to another ImageId.   Is that right?  </p>\n\n<p>For example, here are a few rows from grid_sizes.csv ... </p>\n\n<pre><code>  ImageId      Xmax      Ymin  \n 6070_2_3  0.009013 -0.009045   \n 6040_2_2  0.009158 -0.009043   \n 6170_2_4  0.009150 -0.009042  \n 6100_2_3  0.009156 -0.009042   \n 6040_1_0  0.009156 -0.009042  \n</code></pre>\n\n<p>Should we interpret this data as meaning the W=3349 pixels in the 3-band image with ImageId 6040_1_0 cover a physical distance of 0.009156 and that the H=3391 pixels cover a physical distance of 0.009042?  </p>\n\n<p>I'm having trouble understanding the utility of the W^{\\prime}=Wp and H^{\\prime}=Hp variables in the tutorial as well.  For example, using the formula from the tutorial we get Wp = W * W / (W + 1) = 3348.00029851.  In the next equation we have x/xmax.  What is x supposed to be here? If x ranges from 0 to xmax, then x_prime ranges from 0 to Wp.  Is this useful in some way that I'm not seeing?  </p>",
      "rawMarkdown": "> To utilize these images, we provide the grid coordinates of each image\r\n> so you know how to scale them and align them with the images in\r\n> pixels. In grid_sizes.csv, you are given the Xmax and Ymin values for\r\n> each imageId.\r\n> For each image, you should be able to get the width (W) and height (H)\r\n> from the image raster. For a 3-band image that is 3391 x 3349 x 3, W\r\n> is 3349, and H is 3391.\r\n\r\nJust trying to check my understanding here.  It seems that the Xmax and Ymin values for a given ImageId are to be used to calculate a consistent physical area per pixel (even if we don't know exactly what that physical area is) and have nothing to do with the location of the image on the surface of Earth or relative to another ImageId.   Is that right?  \r\n\r\nFor example, here are a few rows from grid_sizes.csv ... \r\n\r\n      ImageId      Xmax      Ymin  \r\n     6070_2_3  0.009013 -0.009045   \r\n     6040_2_2  0.009158 -0.009043   \r\n     6170_2_4  0.009150 -0.009042  \r\n     6100_2_3  0.009156 -0.009042   \r\n     6040_1_0  0.009156 -0.009042  \r\n\r\nShould we interpret this data as meaning the W=3349 pixels in the 3-band image with ImageId 6040_1_0 cover a physical distance of 0.009156 and that the H=3391 pixels cover a physical distance of 0.009042?  \r\n\r\nI'm having trouble understanding the utility of the W^{\\prime}=Wp and H^{\\prime}=Hp variables in the tutorial as well.  For example, using the formula from the tutorial we get Wp = W * W / (W + 1) = 3348.00029851.  In the next equation we have x/xmax.  What is x supposed to be here? If x ranges from 0 to xmax, then x_prime ranges from 0 to Wp.  Is this useful in some way that I'm not seeing?",
      "votes": null
    },
    {
      "id": "150965",
      "postDate": "12/17/2016 22:47:31",
      "content": "<p>I have the same issue concerning  these coordinates. As @Gabriel mentioned, according to the conversion formulas in the tutorial, this could imply the existence of nested images (as in the attached png). However, this would imply different physical sizes of images, which is not coherent with what is stated on the Data page:</p>\n\n<p>\" In this competition, Dstl provides you with <strong>1km x 1km</strong> satellite images in both 3-band and 16-band formats. Your goal is to detect and classify the types of objects found in these regions. \"</p>\n\n<p>I should miss something too...</p>",
      "rawMarkdown": "I have the same issue concerning  these coordinates. As @Gabriel mentioned, according to the conversion formulas in the tutorial, this could imply the existence of nested images (as in the attached png). However, this would imply different physical sizes of images, which is not coherent with what is stated on the Data page:\r\n\r\n\" In this competition, Dstl provides you with **1km x 1km** satellite images in both 3-band and 16-band formats. Your goal is to detect and classify the types of objects found in these regions. \"\r\n\r\nI should miss something too...",
      "votes": null
    },
    {
      "id": "151010",
      "postDate": "12/18/2016 08:57:16",
      "content": "<p>Can you run a \"group by\" on the Xmax Xmin? To see how many variants are?</p>\n\n<p>Thanx</p>",
      "rawMarkdown": "Can you run a \"group by\" on the Xmax Xmin? To see how many variants are?\r\n\r\nThanx",
      "votes": null
    },
    {
      "id": "151014",
      "postDate": "12/18/2016 09:13:05",
      "content": "<p>you can have a look at the following kernel:</p>\n\n<p><a href=\"https://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/number-of-xmax-and-ymin-variants/code\">https://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/number-of-xmax-and-ymin-variants/code</a></p>\n\n<p>and this notebook:</p>\n\n<p><a href=\"https://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/exploration-of-nearby-images\">https://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/exploration-of-nearby-images</a></p>",
      "rawMarkdown": "you can have a look at the following kernel:\r\n\r\nhttps://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/number-of-xmax-and-ymin-variants/code\r\n\r\nand this notebook:\r\n\r\nhttps://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/exploration-of-nearby-images",
      "votes": null
    },
    {
      "id": "151052",
      "postDate": "12/18/2016 15:48:13",
      "content": "<p>I think I got it. The transformations have to be done on the polygons to match an image, so for each of the training images, a different transformation is done for each image type. So x and y in the tutorial equations are the points for the polygons, and x' and y' are the new transformed points.</p>\n\n<p>This can be done using  <code>shapely.affinity.transform</code> . I made an example here <a href=\"https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/\">https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/</a></p>",
      "rawMarkdown": "I think I got it. The transformations have to be done on the polygons to match an image, so for each of the training images, a different transformation is done for each image type. So x and y in the tutorial equations are the points for the polygons, and x' and y' are the new transformed points.\r\n\r\nThis can be done using  `shapely.affinity.transform` . I made an example here https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/",
      "votes": null
    },
    {
      "id": "151054",
      "postDate": "12/18/2016 16:01:50",
      "content": "<p>thank you! So If I get it right, each image has its own xMax and xMin (which are not related to a global coordinates) and the (x',y') coordinates correspond to pixels indices in each image, whereas (x,y) values are only relevant for submission (in the polygons), am I right?</p>",
      "rawMarkdown": "thank you! So If I get it right, each image has its own xMax and xMin (which are not related to a global coordinates) and the (x',y') coordinates correspond to pixels indices in each image, whereas (x,y) values are only relevant for submission (in the polygons), am I right?",
      "votes": null
    },
    {
      "id": "151056",
      "postDate": "12/18/2016 16:22:10",
      "content": "<p>@LorisMichel, I think we're on the same track, but let me try and clarify things.</p>\n\n<p>Lets call each image a scene instead . A scene includes polygons of 10 classes, and 4 types of images. Each image has a different number of bands, and a different number of x and y pixels.</p>\n\n<p>Each scene has a unique Xmax and Ymin. Within each scene, each image will have its own transformation for the class polygons (using Xmax, yMin, and the images x and y pixel count). The polygons must be transformed 4 times, once for each image type within a scene. When making classifications, the transformations have to be done in reverse to get them back to the original scale for submission. </p>",
      "rawMarkdown": "LorisMichel, I think we're on the same track, but let me try and clarify things.\r\n\r\nLets call each image a scene instead . A scene includes polygons of 10 classes, and 4 types of images. Each image has a different number of bands, and a different number of x and y pixels.\r\n\r\nEach scene has a unique Xmax and Ymin. Within each scene, each image will have its own transformation for the class polygons (using Xmax, yMin, and the images x and y pixel count). The polygons must be transformed 4 times, once for each image type within a scene. When making classifications, the transformations have to be done in reverse to get them back to the original scale for submission.",
      "votes": null
    },
    {
      "id": "151159",
      "postDate": "12/19/2016 05:06:14",
      "content": "<p>thanks LorisMichel and shawn for the discussion and notebook.  I think we can also say that the \"scenes\" can be composed into 5x5 grids using the single integers in the file names (e.g. 6120_2_4) is in the 3rd row and 5th column (0-based indexing) of a 5x5 grid of scenes. see ..</p>\n\n<p>in Python - <a href=\"https://www.kaggle.com/gabrielaltay/dstl-satellite-imagery-feature-detection/polygons-over-images-and-5x5-mosaics\">https://www.kaggle.com/gabrielaltay/dstl-satellite-imagery-feature-detection/polygons-over-images-and-5x5-mosaics</a></p>\n\n<p>in R - <a href=\"https://www.kaggle.com/jeffhebert/dstl-satellite-imagery-feature-detection/stitch-a-16-channel-image-together/discussion\">https://www.kaggle.com/jeffhebert/dstl-satellite-imagery-feature-detection/stitch-a-16-channel-image-together/discussion</a> </p>",
      "rawMarkdown": "thanks LorisMichel and shawn for the discussion and notebook.  I think we can also say that the \"scenes\" can be composed into 5x5 grids using the single integers in the file names (e.g. 6120_2_4) is in the 3rd row and 5th column (0-based indexing) of a 5x5 grid of scenes. see ..\r\n\r\nin Python - https://www.kaggle.com/gabrielaltay/dstl-satellite-imagery-feature-detection/polygons-over-images-and-5x5-mosaics\r\n\r\nin R - https://www.kaggle.com/jeffhebert/dstl-satellite-imagery-feature-detection/stitch-a-16-channel-image-together/discussion",
      "votes": null
    },
    {
      "id": "151241",
      "postDate": "12/19/2016 16:24:50",
      "content": "<p>Seems like the understanding of the coordinates is under control.  One other thing you can do to see how big the image pixels are is to look at the quoted resolution of the particular data.  ~0.3m resolution would mean ~3333 pixels across an image for a 1km distance.  Coarser resolution, say 10m would mean an image with only 100 pixels across.  This explains the size difference of the various images coming from different sensors.</p>",
      "rawMarkdown": "Seems like the understanding of the coordinates is under control.  One other thing you can do to see how big the image pixels are is to look at the quoted resolution of the particular data.  ~0.3m resolution would mean ~3333 pixels across an image for a 1km distance.  Coarser resolution, say 10m would mean an image with only 100 pixels across.  This explains the size difference of the various images coming from different sensors.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 150965,
      "author_name": "lorismichel",
      "author_url": "",
      "post_date": "12/17/2016 22:47:31",
      "content": "<p>I have the same issue concerning  these coordinates. As @Gabriel mentioned, according to the conversion formulas in the tutorial, this could imply the existence of nested images (as in the attached png). However, this would imply different physical sizes of images, which is not coherent with what is stated on the Data page:</p>\n\n<p>\" In this competition, Dstl provides you with <strong>1km x 1km</strong> satellite images in both 3-band and 16-band formats. Your goal is to detect and classify the types of objects found in these regions. \"</p>\n\n<p>I should miss something too...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 151010,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "12/18/2016 08:57:16",
      "content": "<p>Can you run a \"group by\" on the Xmax Xmin? To see how many variants are?</p>\n\n<p>Thanx</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 151014,
      "author_name": "lorismichel",
      "author_url": "",
      "post_date": "12/18/2016 09:13:05",
      "content": "<p>you can have a look at the following kernel:</p>\n\n<p><a href=\"https://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/number-of-xmax-and-ymin-variants/code\">https://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/number-of-xmax-and-ymin-variants/code</a></p>\n\n<p>and this notebook:</p>\n\n<p><a href=\"https://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/exploration-of-nearby-images\">https://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/exploration-of-nearby-images</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 151052,
      "author_name": "shawn775",
      "author_url": "",
      "post_date": "12/18/2016 15:48:13",
      "content": "<p>I think I got it. The transformations have to be done on the polygons to match an image, so for each of the training images, a different transformation is done for each image type. So x and y in the tutorial equations are the points for the polygons, and x' and y' are the new transformed points.</p>\n\n<p>This can be done using  <code>shapely.affinity.transform</code> . I made an example here <a href=\"https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/\">https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 151054,
      "author_name": "lorismichel",
      "author_url": "",
      "post_date": "12/18/2016 16:01:50",
      "content": "<p>thank you! So If I get it right, each image has its own xMax and xMin (which are not related to a global coordinates) and the (x',y') coordinates correspond to pixels indices in each image, whereas (x,y) values are only relevant for submission (in the polygons), am I right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 151056,
      "author_name": "shawn775",
      "author_url": "",
      "post_date": "12/18/2016 16:22:10",
      "content": "<p>@LorisMichel, I think we're on the same track, but let me try and clarify things.</p>\n\n<p>Lets call each image a scene instead . A scene includes polygons of 10 classes, and 4 types of images. Each image has a different number of bands, and a different number of x and y pixels.</p>\n\n<p>Each scene has a unique Xmax and Ymin. Within each scene, each image will have its own transformation for the class polygons (using Xmax, yMin, and the images x and y pixel count). The polygons must be transformed 4 times, once for each image type within a scene. When making classifications, the transformations have to be done in reverse to get them back to the original scale for submission. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 151159,
      "author_name": "gabrielaltay",
      "author_url": "",
      "post_date": "12/19/2016 05:06:14",
      "content": "<p>thanks LorisMichel and shawn for the discussion and notebook.  I think we can also say that the \"scenes\" can be composed into 5x5 grids using the single integers in the file names (e.g. 6120_2_4) is in the 3rd row and 5th column (0-based indexing) of a 5x5 grid of scenes. see ..</p>\n\n<p>in Python - <a href=\"https://www.kaggle.com/gabrielaltay/dstl-satellite-imagery-feature-detection/polygons-over-images-and-5x5-mosaics\">https://www.kaggle.com/gabrielaltay/dstl-satellite-imagery-feature-detection/polygons-over-images-and-5x5-mosaics</a></p>\n\n<p>in R - <a href=\"https://www.kaggle.com/jeffhebert/dstl-satellite-imagery-feature-detection/stitch-a-16-channel-image-together/discussion\">https://www.kaggle.com/jeffhebert/dstl-satellite-imagery-feature-detection/stitch-a-16-channel-image-together/discussion</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 151241,
      "author_name": "zerozero",
      "author_url": "",
      "post_date": "12/19/2016 16:24:50",
      "content": "<p>Seems like the understanding of the coordinates is under control.  One other thing you can do to see how big the image pixels are is to look at the quoted resolution of the particular data.  ~0.3m resolution would mean ~3333 pixels across an image for a 1km distance.  Coarser resolution, say 10m would mean an image with only 100 pixels across.  This explains the size difference of the various images coming from different sensors.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "150964": "> To utilize these images, we provide the grid coordinates of each image\r\n> so you know how to scale them and align them with the images in\r\n> pixels. In grid_sizes.csv, you are given the Xmax and Ymin values for\r\n> each imageId.\r\n> For each image, you should be able to get the width (W) and height (H)\r\n> from the image raster. For a 3-band image that is 3391 x 3349 x 3, W\r\n> is 3349, and H is 3391.\r\n\r\nJust trying to check my understanding here.  It seems that the Xmax and Ymin values for a given ImageId are to be used to calculate a consistent physical area per pixel (even if we don't know exactly what that physical area is) and have nothing to do with the location of the image on the surface of Earth or relative to another ImageId.   Is that right?  \r\n\r\nFor example, here are a few rows from grid_sizes.csv ... \r\n\r\n      ImageId      Xmax      Ymin  \r\n     6070_2_3  0.009013 -0.009045   \r\n     6040_2_2  0.009158 -0.009043   \r\n     6170_2_4  0.009150 -0.009042  \r\n     6100_2_3  0.009156 -0.009042   \r\n     6040_1_0  0.009156 -0.009042  \r\n\r\nShould we interpret this data as meaning the W=3349 pixels in the 3-band image with ImageId 6040_1_0 cover a physical distance of 0.009156 and that the H=3391 pixels cover a physical distance of 0.009042?  \r\n\r\nI'm having trouble understanding the utility of the W^{\\prime}=Wp and H^{\\prime}=Hp variables in the tutorial as well.  For example, using the formula from the tutorial we get Wp = W * W / (W + 1) = 3348.00029851.  In the next equation we have x/xmax.  What is x supposed to be here? If x ranges from 0 to xmax, then x_prime ranges from 0 to Wp.  Is this useful in some way that I'm not seeing?",
    "150965": "I have the same issue concerning  these coordinates. As @Gabriel mentioned, according to the conversion formulas in the tutorial, this could imply the existence of nested images (as in the attached png). However, this would imply different physical sizes of images, which is not coherent with what is stated on the Data page:\r\n\r\n\" In this competition, Dstl provides you with **1km x 1km** satellite images in both 3-band and 16-band formats. Your goal is to detect and classify the types of objects found in these regions. \"\r\n\r\nI should miss something too...",
    "151010": "Can you run a \"group by\" on the Xmax Xmin? To see how many variants are?\r\n\r\nThanx",
    "151014": "you can have a look at the following kernel:\r\n\r\nhttps://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/number-of-xmax-and-ymin-variants/code\r\n\r\nand this notebook:\r\n\r\nhttps://www.kaggle.com/lorismichel/dstl-satellite-imagery-feature-detection/exploration-of-nearby-images",
    "151052": "I think I got it. The transformations have to be done on the polygons to match an image, so for each of the training images, a different transformation is done for each image type. So x and y in the tutorial equations are the points for the polygons, and x' and y' are the new transformed points.\r\n\r\nThis can be done using  `shapely.affinity.transform` . I made an example here https://www.kaggle.com/shawn775/dstl-satellite-imagery-feature-detection/polygon-transformation-to-match-image/",
    "151054": "thank you! So If I get it right, each image has its own xMax and xMin (which are not related to a global coordinates) and the (x',y') coordinates correspond to pixels indices in each image, whereas (x,y) values are only relevant for submission (in the polygons), am I right?",
    "151056": "LorisMichel, I think we're on the same track, but let me try and clarify things.\r\n\r\nLets call each image a scene instead . A scene includes polygons of 10 classes, and 4 types of images. Each image has a different number of bands, and a different number of x and y pixels.\r\n\r\nEach scene has a unique Xmax and Ymin. Within each scene, each image will have its own transformation for the class polygons (using Xmax, yMin, and the images x and y pixel count). The polygons must be transformed 4 times, once for each image type within a scene. When making classifications, the transformations have to be done in reverse to get them back to the original scale for submission.",
    "151159": "thanks LorisMichel and shawn for the discussion and notebook.  I think we can also say that the \"scenes\" can be composed into 5x5 grids using the single integers in the file names (e.g. 6120_2_4) is in the 3rd row and 5th column (0-based indexing) of a 5x5 grid of scenes. see ..\r\n\r\nin Python - https://www.kaggle.com/gabrielaltay/dstl-satellite-imagery-feature-detection/polygons-over-images-and-5x5-mosaics\r\n\r\nin R - https://www.kaggle.com/jeffhebert/dstl-satellite-imagery-feature-detection/stitch-a-16-channel-image-together/discussion",
    "151241": "Seems like the understanding of the coordinates is under control.  One other thing you can do to see how big the image pixels are is to look at the quoted resolution of the particular data.  ~0.3m resolution would mean ~3333 pixels across an image for a 1km distance.  Coarser resolution, say 10m would mean an image with only 100 pixels across.  This explains the size difference of the various images coming from different sensors."
  },
  "source": "meta"
}