{
  "id": 27319,
  "title": "identical images?",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/27319",
  "author_name": "",
  "post_date": "2017-01-04T22:08:59.550Z",
  "votes": null,
  "comment_count": 5,
  "views": 229,
  "content": "<p>when I look at these 17 tiffs in the 3-band files (17 of the 25 training samples), they look identical to each other.</p>\n\n<p>I imagine identical images would have been found on day 1 by this group, so I must be doing something really dumb, but the rest of the training set looks unique.  </p>\n\n<p>Any thoughts on where I've gone wrong?</p>\n\n<pre><code>import tifffile as tiff\n\nfs = map(lambda x: 'data/three_band/' + x + '.tif', \n         ['6070_2_3', '6090_2_0', '6100_1_3', '6100_2_2', \n          '6100_2_3', '6110_1_2', '6110_3_1', '6110_4_0', \n          '6120_2_0', '6120_2_2', '6140_1_2', '6140_3_1', \n          '6150_2_3', '6160_2_1', '6170_0_4', '6170_2_4', '6170_4_1'])\n\nfor f in fs:\n  P = tiff.imread(f) \n  tiff.imshow(P)\n</code></pre>",
  "messages": [
    {
      "id": "154132",
      "postDate": "01/04/2017 22:08:59",
      "content": "<p>when I look at these 17 tiffs in the 3-band files (17 of the 25 training samples), they look identical to each other.</p>\n\n<p>I imagine identical images would have been found on day 1 by this group, so I must be doing something really dumb, but the rest of the training set looks unique.  </p>\n\n<p>Any thoughts on where I've gone wrong?</p>\n\n<pre><code>import tifffile as tiff\n\nfs = map(lambda x: 'data/three_band/' + x + '.tif', \n         ['6070_2_3', '6090_2_0', '6100_1_3', '6100_2_2', \n          '6100_2_3', '6110_1_2', '6110_3_1', '6110_4_0', \n          '6120_2_0', '6120_2_2', '6140_1_2', '6140_3_1', \n          '6150_2_3', '6160_2_1', '6170_0_4', '6170_2_4', '6170_4_1'])\n\nfor f in fs:\n  P = tiff.imread(f) \n  tiff.imshow(P)\n</code></pre>",
      "rawMarkdown": "when I look at these 17 tiffs in the 3-band files (17 of the 25 training samples), they look identical to each other.\r\n\r\nI imagine identical images would have been found on day 1 by this group, so I must be doing something really dumb, but the rest of the training set looks unique.  \r\n\r\nAny thoughts on where I've gone wrong?\r\n\r\n    import tifffile as tiff\r\n    \r\n    fs = map(lambda x: 'data/three_band/' + x + '.tif', \r\n             ['6070_2_3', '6090_2_0', '6100_1_3', '6100_2_2', \r\n              '6100_2_3', '6110_1_2', '6110_3_1', '6110_4_0', \r\n              '6120_2_0', '6120_2_2', '6140_1_2', '6140_3_1', \r\n              '6150_2_3', '6160_2_1', '6170_0_4', '6170_2_4', '6170_4_1'])\r\n    \r\n    for f in fs:\r\n      P = tiff.imread(f) \r\n      tiff.imshow(P)",
      "votes": null
    },
    {
      "id": "154135",
      "postDate": "01/04/2017 22:20:14",
      "content": "<p>Hi, 6110 and 6140 (all 25 images) correspond to the exact same location but on different days (you can tell that by clouds and other features such as the color of crops). But I suspect that's not what you're talking about. The images above are different :-)</p>",
      "rawMarkdown": "Hi, 6110 and 6140 (all 25 images) correspond to the exact same location but on different days (you can tell that by clouds and other features such as the color of crops). But I suspect that's not what you're talking about. The images above are different :-)",
      "votes": null
    },
    {
      "id": "154147",
      "postDate": "01/04/2017 23:17:06",
      "content": "<p>hi Amaia...</p>\n\n<p>thanks for the info.  I don't think that's my issue...  but I also didn't know that...  this is a really small dataset then, eh?  Just 18 unique locations?  do I understand what you're saying correctly?</p>\n\n<p>this elementwise comparison of the 17 files in question shows every one to be exactly identical.  this can't be right, but I can't imagine how I misreading these.</p>\n\n<pre><code>import itertools\nimport numpy as np \nimport tifffile as tiff\n\ndata_loc = 'data/three_band/'\nfs = map(lambda x: data_loc + x + '.tif', \n         ['6070_2_3', '6090_2_0', '6100_1_3', '6100_2_2', \n          '6100_2_3', '6110_1_2', '6110_3_1', '6110_4_0', \n          '6120_2_0', '6120_2_2', '6140_1_2', '6140_3_1', \n          '6150_2_3', '6160_2_1', '6170_0_4', '6170_2_4', '6170_4_1'])\n\n# returns True if elementwise comparison of 2 files show them to be identical\ndef compare(tup):\n  f1, f2 = tup \n  arr1 = tiff.imread(f1)  # tiff as 3d np.array\n  arr2 = tiff.imread(f2)  # tiff as 3d np.array\n  return (arr1==arr2).all() # elementwise comparison shows them to be identical\n\n\ncross = list(itertools.product(fs, fs))      # len - 289 = 17*17\ntest = map(lambda tup: compare(tup), cross)  # compare every combination\nnp.array(test, dtype=int).sum()              # result: 289 - all identical\n</code></pre>",
      "rawMarkdown": "hi Amaia...\r\n\r\nthanks for the info.  I don't think that's my issue...  but I also didn't know that...  this is a really small dataset then, eh?  Just 18 unique locations?  do I understand what you're saying correctly?\r\n\r\nthis elementwise comparison of the 17 files in question shows every one to be exactly identical.  this can't be right, but I can't imagine how I misreading these.\r\n\r\n    import itertools\r\n    import numpy as np \r\n    import tifffile as tiff\r\n    \r\n    data_loc = 'data/three_band/'\r\n    fs = map(lambda x: data_loc + x + '.tif', \r\n             ['6070_2_3', '6090_2_0', '6100_1_3', '6100_2_2', \r\n              '6100_2_3', '6110_1_2', '6110_3_1', '6110_4_0', \r\n              '6120_2_0', '6120_2_2', '6140_1_2', '6140_3_1', \r\n              '6150_2_3', '6160_2_1', '6170_0_4', '6170_2_4', '6170_4_1'])\r\n    \r\n    # returns True if elementwise comparison of 2 files show them to be identical\r\n    def compare(tup):\r\n      f1, f2 = tup \r\n      arr1 = tiff.imread(f1)  # tiff as 3d np.array\r\n      arr2 = tiff.imread(f2)  # tiff as 3d np.array\r\n      return (arr1==arr2).all() # elementwise comparison shows them to be identical\r\n    \r\n    \r\n    cross = list(itertools.product(fs, fs))      # len - 289 = 17*17\r\n    test = map(lambda tup: compare(tup), cross)  # compare every combination\r\n    np.array(test, dtype=int).sum()              # result: 289 - all identical",
      "votes": null
    },
    {
      "id": "154151",
      "postDate": "01/04/2017 23:42:10",
      "content": "<p>Your first snippet correctly plot different images. I copy&amp;pasted it into a kernel and saw different images visually. The second snippet contains a bug, \"cross\" is an empty list for me, and the last line in not valid in Python 3. </p>\n\n<blockquote>\n  <p>thanks for the info. I don't think that's my issue... but I also didn't know that... this is a really small dataset then, eh? Just 18 unique locations? do I understand what you're saying correctly?</p>\n</blockquote>\n\n<p>You mean 18 x 5 x 5? I'm just saying 6110_x_y is the same location as 6140_x_y (25 images in total, the same location but different images).</p>",
      "rawMarkdown": "Your first snippet correctly plot different images. I copy&pasted it into a kernel and saw different images visually. The second snippet contains a bug, \"cross\" is an empty list for me, and the last line in not valid in Python 3. \r\n\r\n> thanks for the info. I don't think that's my issue... but I also didn't know that... this is a really small dataset then, eh? Just 18 unique locations? do I understand what you're saying correctly?\r\n\r\nYou mean 18 x 5 x 5? I'm just saying 6110_x_y is the same location as 6140_x_y (25 images in total, the same location but different images).",
      "votes": null
    },
    {
      "id": "154156",
      "postDate": "01/05/2017 00:28:32",
      "content": "<p>Ah...  got it...  now I understand the naming terminology.</p>\n\n<p>Thanks for giving my sample a run.  I still use python 2.7, but if you visually saw different images, I'm deleting everything and downloading again.  I viewed those files with 4 different libraries and had the same result with each.</p>\n\n<p>thanks again for engaging on the boards with me on this!</p>",
      "rawMarkdown": "Ah...  got it...  now I understand the naming terminology.\r\n\r\nThanks for giving my sample a run.  I still use python 2.7, but if you visually saw different images, I'm deleting everything and downloading again.  I viewed those files with 4 different libraries and had the same result with each.\r\n\r\nthanks again for engaging on the boards with me on this!",
      "votes": null
    },
    {
      "id": "154157",
      "postDate": "01/05/2017 00:30:01",
      "content": "<p>Before deleting, try using QGIS. It's really helpful to have something you know works for sure as comparison.</p>",
      "rawMarkdown": "Before deleting, try using QGIS. It's really helpful to have something you know works for sure as comparison.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 154135,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "01/04/2017 22:20:14",
      "content": "<p>Hi, 6110 and 6140 (all 25 images) correspond to the exact same location but on different days (you can tell that by clouds and other features such as the color of crops). But I suspect that's not what you're talking about. The images above are different :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154147,
      "author_name": "kevinhinson",
      "author_url": "",
      "post_date": "01/04/2017 23:17:06",
      "content": "<p>hi Amaia...</p>\n\n<p>thanks for the info.  I don't think that's my issue...  but I also didn't know that...  this is a really small dataset then, eh?  Just 18 unique locations?  do I understand what you're saying correctly?</p>\n\n<p>this elementwise comparison of the 17 files in question shows every one to be exactly identical.  this can't be right, but I can't imagine how I misreading these.</p>\n\n<pre><code>import itertools\nimport numpy as np \nimport tifffile as tiff\n\ndata_loc = 'data/three_band/'\nfs = map(lambda x: data_loc + x + '.tif', \n         ['6070_2_3', '6090_2_0', '6100_1_3', '6100_2_2', \n          '6100_2_3', '6110_1_2', '6110_3_1', '6110_4_0', \n          '6120_2_0', '6120_2_2', '6140_1_2', '6140_3_1', \n          '6150_2_3', '6160_2_1', '6170_0_4', '6170_2_4', '6170_4_1'])\n\n# returns True if elementwise comparison of 2 files show them to be identical\ndef compare(tup):\n  f1, f2 = tup \n  arr1 = tiff.imread(f1)  # tiff as 3d np.array\n  arr2 = tiff.imread(f2)  # tiff as 3d np.array\n  return (arr1==arr2).all() # elementwise comparison shows them to be identical\n\n\ncross = list(itertools.product(fs, fs))      # len - 289 = 17*17\ntest = map(lambda tup: compare(tup), cross)  # compare every combination\nnp.array(test, dtype=int).sum()              # result: 289 - all identical\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154151,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "01/04/2017 23:42:10",
      "content": "<p>Your first snippet correctly plot different images. I copy&amp;pasted it into a kernel and saw different images visually. The second snippet contains a bug, \"cross\" is an empty list for me, and the last line in not valid in Python 3. </p>\n\n<blockquote>\n  <p>thanks for the info. I don't think that's my issue... but I also didn't know that... this is a really small dataset then, eh? Just 18 unique locations? do I understand what you're saying correctly?</p>\n</blockquote>\n\n<p>You mean 18 x 5 x 5? I'm just saying 6110_x_y is the same location as 6140_x_y (25 images in total, the same location but different images).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154156,
      "author_name": "kevinhinson",
      "author_url": "",
      "post_date": "01/05/2017 00:28:32",
      "content": "<p>Ah...  got it...  now I understand the naming terminology.</p>\n\n<p>Thanks for giving my sample a run.  I still use python 2.7, but if you visually saw different images, I'm deleting everything and downloading again.  I viewed those files with 4 different libraries and had the same result with each.</p>\n\n<p>thanks again for engaging on the boards with me on this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154157,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "01/05/2017 00:30:01",
      "content": "<p>Before deleting, try using QGIS. It's really helpful to have something you know works for sure as comparison.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "154132": "when I look at these 17 tiffs in the 3-band files (17 of the 25 training samples), they look identical to each other.\r\n\r\nI imagine identical images would have been found on day 1 by this group, so I must be doing something really dumb, but the rest of the training set looks unique.  \r\n\r\nAny thoughts on where I've gone wrong?\r\n\r\n    import tifffile as tiff\r\n    \r\n    fs = map(lambda x: 'data/three_band/' + x + '.tif', \r\n             ['6070_2_3', '6090_2_0', '6100_1_3', '6100_2_2', \r\n              '6100_2_3', '6110_1_2', '6110_3_1', '6110_4_0', \r\n              '6120_2_0', '6120_2_2', '6140_1_2', '6140_3_1', \r\n              '6150_2_3', '6160_2_1', '6170_0_4', '6170_2_4', '6170_4_1'])\r\n    \r\n    for f in fs:\r\n      P = tiff.imread(f) \r\n      tiff.imshow(P)",
    "154135": "Hi, 6110 and 6140 (all 25 images) correspond to the exact same location but on different days (you can tell that by clouds and other features such as the color of crops). But I suspect that's not what you're talking about. The images above are different :-)",
    "154147": "hi Amaia...\r\n\r\nthanks for the info.  I don't think that's my issue...  but I also didn't know that...  this is a really small dataset then, eh?  Just 18 unique locations?  do I understand what you're saying correctly?\r\n\r\nthis elementwise comparison of the 17 files in question shows every one to be exactly identical.  this can't be right, but I can't imagine how I misreading these.\r\n\r\n    import itertools\r\n    import numpy as np \r\n    import tifffile as tiff\r\n    \r\n    data_loc = 'data/three_band/'\r\n    fs = map(lambda x: data_loc + x + '.tif', \r\n             ['6070_2_3', '6090_2_0', '6100_1_3', '6100_2_2', \r\n              '6100_2_3', '6110_1_2', '6110_3_1', '6110_4_0', \r\n              '6120_2_0', '6120_2_2', '6140_1_2', '6140_3_1', \r\n              '6150_2_3', '6160_2_1', '6170_0_4', '6170_2_4', '6170_4_1'])\r\n    \r\n    # returns True if elementwise comparison of 2 files show them to be identical\r\n    def compare(tup):\r\n      f1, f2 = tup \r\n      arr1 = tiff.imread(f1)  # tiff as 3d np.array\r\n      arr2 = tiff.imread(f2)  # tiff as 3d np.array\r\n      return (arr1==arr2).all() # elementwise comparison shows them to be identical\r\n    \r\n    \r\n    cross = list(itertools.product(fs, fs))      # len - 289 = 17*17\r\n    test = map(lambda tup: compare(tup), cross)  # compare every combination\r\n    np.array(test, dtype=int).sum()              # result: 289 - all identical",
    "154151": "Your first snippet correctly plot different images. I copy&pasted it into a kernel and saw different images visually. The second snippet contains a bug, \"cross\" is an empty list for me, and the last line in not valid in Python 3. \r\n\r\n> thanks for the info. I don't think that's my issue... but I also didn't know that... this is a really small dataset then, eh? Just 18 unique locations? do I understand what you're saying correctly?\r\n\r\nYou mean 18 x 5 x 5? I'm just saying 6110_x_y is the same location as 6140_x_y (25 images in total, the same location but different images).",
    "154156": "Ah...  got it...  now I understand the naming terminology.\r\n\r\nThanks for giving my sample a run.  I still use python 2.7, but if you visually saw different images, I'm deleting everything and downloading again.  I viewed those files with 4 different libraries and had the same result with each.\r\n\r\nthanks again for engaging on the boards with me on this!",
    "154157": "Before deleting, try using QGIS. It's really helpful to have something you know works for sure as comparison."
  },
  "source": "meta"
}