{
  "id": 224883,
  "title": "Data still have problem?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/224883",
  "author_name": "zihuanqiu",
  "post_date": "2021-03-10T04:33:21.212000",
  "votes": 12,
  "comment_count": 19,
  "views": 0,
  "content": "<p>It seems d488c759a.tiff in test set only have 1 channel, whats wrong?</p>",
  "messages": [
    {
      "id": 1233186,
      "postDate": "2021-03-10T08:15:07.840Z",
      "content": "<p>I think new data is fine, 1 channel artifact was also in previous data. It depends how you read data, do you use RasterIO or GDAL package? Most of us are using these packages that are memory friendly to read TIFF files. New kaggle image comes with GDAL 3.1.4 which requires to use <code>subdatasets</code> to load some images. You have a working sample at:<br>\n<a href=\"https://www.kaggle.com/mpware/masks-quick-eda-updated-data\" target=\"_blank\">https://www.kaggle.com/mpware/masks-quick-eda-updated-data</a></p>\n<pre><code>with rasterio.open(TRAIN_HOME + image_id + \".tiff\") as file:\n    if file.count == 3:\n        image = file.read([1,2,3]).transpose(1,2,0).copy()\n    else:\n        h, w = (file.height, file.width)\n        subdatasets = file.subdatasets\n        if len(subdatasets) &gt; 0:\n            image = np.zeros((h, w, len(subdatasets)), dtype=np.uint8)\n            for i, subdataset in enumerate(subdatasets, 0):\n                with rasterio.open(subdataset) as layer:\n                    image[:,:,i] = layer.read(1)\n</code></pre>",
      "rawMarkdown": "I think new data is fine, 1 channel artifact was also in previous data. It depends how you read data, do you use RasterIO or GDAL package? Most of us are using these packages that are memory friendly to read TIFF files. New kaggle image comes with GDAL 3.1.4 which requires to use `subdatasets` to load some images. You have a working sample at:\nhttps://www.kaggle.com/mpware/masks-quick-eda-updated-data\n\n```\nwith rasterio.open(TRAIN_HOME + image_id + \".tiff\") as file:\n\tif file.count == 3:\n\t\timage = file.read([1,2,3]).transpose(1,2,0).copy()\n\telse:\n\t\th, w = (file.height, file.width)\n\t\tsubdatasets = file.subdatasets\n\t\tif len(subdatasets) > 0:\n\t\t\timage = np.zeros((h, w, len(subdatasets)), dtype=np.uint8)\n\t\t\tfor i, subdataset in enumerate(subdatasets, 0):\n\t\t\t\twith rasterio.open(subdataset) as layer:\n\t\t\t\t\timage[:,:,i] = layer.read(1)\n```",
      "votes": 22,
      "replies": [
        {
          "id": 1233203,
          "postDate": "2021-03-10T08:36:02.683Z",
          "content": "<p>good job!     </p>",
          "rawMarkdown": "good job!     ",
          "votes": 2
        },
        {
          "id": 1233217,
          "postDate": "2021-03-10T08:43:17.447Z",
          "content": "<p>Subdataset path looks like:<br>\n<code>'aa05346ff': ['GTIFF_DIR:1:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff', 'GTIFF_DIR:2:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff', 'GTIFF_DIR:3:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff']</code></p>\n<p>However, if, like us, you're using <strong>latest Kaggle image in your preferences</strong> then all submissions will either fail or return zero score due to this GDAL upgrade.</p>",
          "rawMarkdown": "Subdataset path looks like:\n`'aa05346ff': ['GTIFF_DIR:1:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff', 'GTIFF_DIR:2:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff', 'GTIFF_DIR:3:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff']`\n\nHowever, if, like us, you're using **latest Kaggle image in your preferences** then all submissions will either fail or return zero score due to this GDAL upgrade.",
          "votes": 1
        },
        {
          "id": 1233296,
          "postDate": "2021-03-10T09:34:18.520Z",
          "content": "<p>yeah, i was thinking about the same thing.. how all the subs that use rasterio will fail and i won`t be able use my previous subs to estimate how good my previous models would handle the new testset. the organizers have been doing a really great job! not</p>",
          "rawMarkdown": "yeah, i was thinking about the same thing.. how all the subs that use rasterio will fail and i won`t be able use my previous subs to estimate how good my previous models would handle the new testset. the organizers have been doing a really great job! not",
          "votes": 1
        },
        {
          "id": 1233932,
          "postDate": "2021-03-10T19:40:05.733Z",
          "content": "<p>Amazing job, thanks for sharing this!<br>\nI've noticed that this line:<br>\n<code>image = np.zeros((h, w, len(subdatasets)), dtype=np.uint8)</code><br>\nis RAM consuming a lot, thus it seems to be much safe to read using rasterio like this:</p>\n<pre><code>...\n        if dataset.count != 3:\n            print('Image file with subdatasets as channels')\n            layers = [rasterio.open(subd) for subd in dataset.subdatasets]\n\n        for (x1,x2,y1,y2) in slices:\n            if dataset.count == 3: # normal\n                image = dataset.read([1,2,3],\n                            window=Window.from_slices((x1,x2),(y1,y2)))\n                image = np.moveaxis(image, 0, -1)\n            else: # with subdatasets/layers\n                image = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n                for fl in range(3):\n                    image[:,:,fl] = layers[fl].read(window=Window.from_slices((x1,x2),(y1,y2)))\n...\n</code></pre>\n<p>I've tested this in <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-subm\" target=\"_blank\">notebook</a> and it works well without issues - both for public and private test data. Once again thank you!</p>",
          "rawMarkdown": "Amazing job, thanks for sharing this!\nI've noticed that this line:\n`image = np.zeros((h, w, len(subdatasets)), dtype=np.uint8)`\nis RAM consuming a lot, thus it seems to be much safe to read using rasterio like this:\n```\n...\n        if dataset.count != 3:\n            print('Image file with subdatasets as channels')\n            layers = [rasterio.open(subd) for subd in dataset.subdatasets]\n            \n        for (x1,x2,y1,y2) in slices:\n            if dataset.count == 3: # normal\n                image = dataset.read([1,2,3],\n                            window=Window.from_slices((x1,x2),(y1,y2)))\n                image = np.moveaxis(image, 0, -1)\n            else: # with subdatasets/layers\n                image = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n                for fl in range(3):\n                    image[:,:,fl] = layers[fl].read(window=Window.from_slices((x1,x2),(y1,y2)))\n...\n```\nI've tested this in [notebook](https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-subm) and it works well without issues - both for public and private test data. Once again thank you!",
          "votes": 13
        },
        {
          "id": 1235496,
          "postDate": "2021-03-12T07:35:37.233Z",
          "content": "<p>Thank you very much for the excellent code snippet. Indeed the memory requirements for your approach are so much better. </p>",
          "rawMarkdown": "Thank you very much for the excellent code snippet. Indeed the memory requirements for your approach are so much better. ",
          "votes": 1
        },
        {
          "id": 1236469,
          "postDate": "2021-03-13T07:03:21.353Z",
          "content": "<p>Thanks, great info!</p>",
          "rawMarkdown": "Thanks, great info!"
        }
      ]
    },
    {
      "id": 1232988,
      "postDate": "2021-03-10T04:33:21.213Z",
      "content": "<p>It seems d488c759a.tiff in test set only have 1 channel, whats wrong?</p>",
      "rawMarkdown": "It seems d488c759a.tiff in test set only have 1 channel, whats wrong?",
      "votes": 10
    },
    {
      "id": 1233007,
      "postDate": "2021-03-10T04:56:31.020Z",
      "content": "<p>and also <br>\n57512b7f1 aa05346ff in test set<br>\n095bf7a1f 1e2425f28 26dc41664 4ef6695ce c68fe75ea in train set have this problem</p>",
      "rawMarkdown": "and also \n57512b7f1 aa05346ff in test set\n095bf7a1f 1e2425f28 26dc41664 4ef6695ce c68fe75ea in train set have this problem",
      "votes": 1,
      "replies": [
        {
          "id": 1233146,
          "postDate": "2021-03-10T07:21:59.453Z",
          "content": "<p>you need to use transpose method, I have met that before. But I dont know how to solve it by using rasterio. <strong><a href=\"https://github.com/mapbox/rasterio/blob/master/rasterio/__init__.py\" target=\"_blank\">https://github.com/mapbox/rasterio/blob/master/rasterio/__init__.py</a></strong> . I was considering making new test set by myself.</p>",
          "rawMarkdown": "you need to use transpose method, I have met that before. But I dont know how to solve it by using rasterio. **https://github.com/mapbox/rasterio/blob/master/rasterio/__init__.py** . I was considering making new test set by myself."
        },
        {
          "id": 1233156,
          "postDate": "2021-03-10T07:35:26.913Z",
          "content": "<p>I understand.</p>",
          "rawMarkdown": "I understand."
        },
        {
          "id": 1233165,
          "postDate": "2021-03-10T07:48:01.450Z",
          "content": "<p>As <a href=\"https://www.kaggle.com/frankx7\" target=\"_blank\">@frankx7</a> said:</p>\n<p>Testing data </p>\n<blockquote>\n  <p><strong>(3, 30720, 47340) aa05346ff</strong><br>\n  (23990, 47723, 3) 2ec3f1bb9<br>\n  <strong>(3, 33240, 43160) 57512b7f1</strong><br>\n  (29433, 22165, 3) 3589adb90<br>\n  <strong>(3, 46660, 29020) d488c759a</strong></p>\n</blockquote>\n<p>I've tested and obtained the same result. Not means one channel as I thought. maybe you just need to transpose them.</p>",
          "rawMarkdown": "As @frankx7 said:\n\nTesting data \n> **(3, 30720, 47340) aa05346ff**\n> (23990, 47723, 3) 2ec3f1bb9\n> **(3, 33240, 43160) 57512b7f1**\n> (29433, 22165, 3) 3589adb90\n> **(3, 46660, 29020) d488c759a**\n\nI've tested and obtained the same result. Not means one channel as I thought. maybe you just need to transpose them."
        },
        {
          "id": 1233248,
          "postDate": "2021-03-10T08:56:56.683Z",
          "content": "<p>With that code, It seems to work fine in every case: </p>\n<pre><code>def read_tiff(image_file):\n    image = tiff.imread(image_file)\n    print(f\"reads {image_file}; shape={image.shape}\")\n    if image.shape[0] == 3 and image.ndim == 3:\n        image = image.transpose(1, 2, 0)\n    elif image.shape[2] == 3 and image.ndim == 5 :\n        image = np.transpose(image.squeeze(), (1, 2, 0))\n    image = np.ascontiguousarray(image)\n    return image\n</code></pre>",
          "rawMarkdown": "With that code, It seems to work fine in every case: \n\n```\n\ndef read_tiff(image_file):\n    image = tiff.imread(image_file)\n    print(f\"reads {image_file}; shape={image.shape}\")\n    if image.shape[0] == 3 and image.ndim == 3:\n        image = image.transpose(1, 2, 0)\n    elif image.shape[2] == 3 and image.ndim == 5 :\n        image = np.transpose(image.squeeze(), (1, 2, 0))\n    image = np.ascontiguousarray(image)\n    return image\n```",
          "votes": 6
        }
      ]
    },
    {
      "id": 1233077,
      "postDate": "2021-03-10T05:53:13.483Z",
      "content": "<p>Yeah, I have met the same problem.</p>",
      "rawMarkdown": "Yeah, I have met the same problem.",
      "replies": [
        {
          "id": 1235970,
          "postDate": "2021-03-12T16:24:07.430Z",
          "content": "<p>Try running <a href=\"https://www.kaggle.com/markalavin/list-image-shapes;\" target=\"_blank\">https://www.kaggle.com/markalavin/list-image-shapes;</a> it uses <code>tifffile</code>; sorry it's so slow, I'm reading all of each Train and Test images.</p>",
          "rawMarkdown": "Try running https://www.kaggle.com/markalavin/list-image-shapes; it uses ```tifffile```; sorry it's so slow, I'm reading all of each Train and Test images."
        }
      ]
    },
    {
      "id": 1233042,
      "postDate": "2021-03-10T05:23:59.467Z",
      "content": "<p>thanks Wiseman.X<br>\nuse tiff.imread can return the correct size<br>\nimg = tiff.imread('../input/hubmap-kidney-segmentation/train/c68fe75ea.tiff')<br>\nimg.shape</p>\n<blockquote>\n  <p>(3, 26840, 49780)</p>\n</blockquote>",
      "rawMarkdown": "thanks Wiseman.X\nuse tiff.imread can return the correct size\nimg = tiff.imread('../input/hubmap-kidney-segmentation/train/c68fe75ea.tiff')\nimg.shape\n> (3, 26840, 49780)"
    },
    {
      "id": 1233030,
      "postDate": "2021-03-10T05:10:21.040Z",
      "content": "<p>I use this code to load tiff<br>\ndata = rasterio.open('hubmap-kidney-segmentation/train/c68fe75ea.tiff')<br>\ndata.read().shape</p>\n<blockquote>\n  <p>(1, 26840, 49780)</p>\n</blockquote>",
      "rawMarkdown": "I use this code to load tiff\ndata = rasterio.open('hubmap-kidney-segmentation/train/c68fe75ea.tiff')\ndata.read().shape\n> (1, 26840, 49780)"
    },
    {
      "id": 1233911,
      "postDate": "2021-03-10T19:13:29.670Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1233055,
      "postDate": "2021-03-10T05:32:48.557Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1233001,
      "postDate": "2021-03-10T04:51:19.790Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1233186,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-03-10T08:15:07.840000",
      "content": "<p>I think new data is fine, 1 channel artifact was also in previous data. It depends how you read data, do you use RasterIO or GDAL package? Most of us are using these packages that are memory friendly to read TIFF files. New kaggle image comes with GDAL 3.1.4 which requires to use <code>subdatasets</code> to load some images. You have a working sample at:<br>\n<a href=\"https://www.kaggle.com/mpware/masks-quick-eda-updated-data\" target=\"_blank\">https://www.kaggle.com/mpware/masks-quick-eda-updated-data</a></p>\n<pre><code>with rasterio.open(TRAIN_HOME + image_id + \".tiff\") as file:\n    if file.count == 3:\n        image = file.read([1,2,3]).transpose(1,2,0).copy()\n    else:\n        h, w = (file.height, file.width)\n        subdatasets = file.subdatasets\n        if len(subdatasets) &gt; 0:\n            image = np.zeros((h, w, len(subdatasets)), dtype=np.uint8)\n            for i, subdataset in enumerate(subdatasets, 0):\n                with rasterio.open(subdataset) as layer:\n                    image[:,:,i] = layer.read(1)\n</code></pre>",
      "votes": 22,
      "replies": [
        {
          "id": 1233203,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-10T08:36:02.683000",
          "content": "<p>good job!     </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1233217,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-03-10T08:43:17.447000",
          "content": "<p>Subdataset path looks like:<br>\n<code>'aa05346ff': ['GTIFF_DIR:1:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff', 'GTIFF_DIR:2:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff', 'GTIFF_DIR:3:../input/hubmap-kidney-segmentation/test/aa05346ff.tiff']</code></p>\n<p>However, if, like us, you're using <strong>latest Kaggle image in your preferences</strong> then all submissions will either fail or return zero score due to this GDAL upgrade.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1233296,
          "author_name": "rosuluc",
          "author_url": "",
          "post_date": "2021-03-10T09:34:18.520000",
          "content": "<p>yeah, i was thinking about the same thing.. how all the subs that use rasterio will fail and i won`t be able use my previous subs to estimate how good my previous models would handle the new testset. the organizers have been doing a really great job! not</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1233932,
          "author_name": "Wojtek Rosa",
          "author_url": "",
          "post_date": "2021-03-10T19:40:05.733000",
          "content": "<p>Amazing job, thanks for sharing this!<br>\nI've noticed that this line:<br>\n<code>image = np.zeros((h, w, len(subdatasets)), dtype=np.uint8)</code><br>\nis RAM consuming a lot, thus it seems to be much safe to read using rasterio like this:</p>\n<pre><code>...\n        if dataset.count != 3:\n            print('Image file with subdatasets as channels')\n            layers = [rasterio.open(subd) for subd in dataset.subdatasets]\n\n        for (x1,x2,y1,y2) in slices:\n            if dataset.count == 3: # normal\n                image = dataset.read([1,2,3],\n                            window=Window.from_slices((x1,x2),(y1,y2)))\n                image = np.moveaxis(image, 0, -1)\n            else: # with subdatasets/layers\n                image = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n                for fl in range(3):\n                    image[:,:,fl] = layers[fl].read(window=Window.from_slices((x1,x2),(y1,y2)))\n...\n</code></pre>\n<p>I've tested this in <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-subm\" target=\"_blank\">notebook</a> and it works well without issues - both for public and private test data. Once again thank you!</p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 1235496,
          "author_name": "gil fernandes",
          "author_url": "",
          "post_date": "2021-03-12T07:35:37.233000",
          "content": "<p>Thank you very much for the excellent code snippet. Indeed the memory requirements for your approach are so much better. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1236469,
          "author_name": "Geir Drange",
          "author_url": "",
          "post_date": "2021-03-13T07:03:21.353000",
          "content": "<p>Thanks, great info!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1233007,
      "author_name": "zihuanqiu",
      "author_url": "",
      "post_date": "2021-03-10T04:56:31.020000",
      "content": "<p>and also <br>\n57512b7f1 aa05346ff in test set<br>\n095bf7a1f 1e2425f28 26dc41664 4ef6695ce c68fe75ea in train set have this problem</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1233146,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-10T07:21:59.453000",
          "content": "<p>you need to use transpose method, I have met that before. But I dont know how to solve it by using rasterio. <strong><a href=\"https://github.com/mapbox/rasterio/blob/master/rasterio/__init__.py\" target=\"_blank\">https://github.com/mapbox/rasterio/blob/master/rasterio/__init__.py</a></strong> . I was considering making new test set by myself.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1233156,
          "author_name": "shinewine",
          "author_url": "",
          "post_date": "2021-03-10T07:35:26.913000",
          "content": "<p>I understand.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1233165,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-10T07:48:01.450000",
          "content": "<p>As <a href=\"https://www.kaggle.com/frankx7\" target=\"_blank\">@frankx7</a> said:</p>\n<p>Testing data </p>\n<blockquote>\n  <p><strong>(3, 30720, 47340) aa05346ff</strong><br>\n  (23990, 47723, 3) 2ec3f1bb9<br>\n  <strong>(3, 33240, 43160) 57512b7f1</strong><br>\n  (29433, 22165, 3) 3589adb90<br>\n  <strong>(3, 46660, 29020) d488c759a</strong></p>\n</blockquote>\n<p>I've tested and obtained the same result. Not means one channel as I thought. maybe you just need to transpose them.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1233248,
          "author_name": "FabienDaniel",
          "author_url": "",
          "post_date": "2021-03-10T08:56:56.683000",
          "content": "<p>With that code, It seems to work fine in every case: </p>\n<pre><code>def read_tiff(image_file):\n    image = tiff.imread(image_file)\n    print(f\"reads {image_file}; shape={image.shape}\")\n    if image.shape[0] == 3 and image.ndim == 3:\n        image = image.transpose(1, 2, 0)\n    elif image.shape[2] == 3 and image.ndim == 5 :\n        image = np.transpose(image.squeeze(), (1, 2, 0))\n    image = np.ascontiguousarray(image)\n    return image\n</code></pre>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1233077,
      "author_name": "shinewine",
      "author_url": "",
      "post_date": "2021-03-10T05:53:13.483000",
      "content": "<p>Yeah, I have met the same problem.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1235970,
          "author_name": "Mark A Lavin",
          "author_url": "",
          "post_date": "2021-03-12T16:24:07.430000",
          "content": "<p>Try running <a href=\"https://www.kaggle.com/markalavin/list-image-shapes;\" target=\"_blank\">https://www.kaggle.com/markalavin/list-image-shapes;</a> it uses <code>tifffile</code>; sorry it's so slow, I'm reading all of each Train and Test images.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1233042,
      "author_name": "zihuanqiu",
      "author_url": "",
      "post_date": "2021-03-10T05:23:59.467000",
      "content": "<p>thanks Wiseman.X<br>\nuse tiff.imread can return the correct size<br>\nimg = tiff.imread('../input/hubmap-kidney-segmentation/train/c68fe75ea.tiff')<br>\nimg.shape</p>\n<blockquote>\n  <p>(3, 26840, 49780)</p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1233030,
      "author_name": "zihuanqiu",
      "author_url": "",
      "post_date": "2021-03-10T05:10:21.040000",
      "content": "<p>I use this code to load tiff<br>\ndata = rasterio.open('hubmap-kidney-segmentation/train/c68fe75ea.tiff')<br>\ndata.read().shape</p>\n<blockquote>\n  <p>(1, 26840, 49780)</p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1233911,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-10T19:13:29.670000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1233055,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-10T05:32:48.557000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1233001,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-10T04:51:19.790000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1233186": "I think new data is fine, 1 channel artifact was also in previous data. It depends how you read data, do you use RasterIO or GDAL package? Most of us are using these packages that are memory friendly to read TIFF files. New kaggle image comes with GDAL 3.1.4 which requires to use `subdatasets` to load some images. You have a working sample at:\nhttps://www.kaggle.com/mpware/masks-quick-eda-updated-data\n\n```\nwith rasterio.open(TRAIN_HOME + image_id + \".tiff\") as file:\n\tif file.count == 3:\n\t\timage = file.read([1,2,3]).transpose(1,2,0).copy()\n\telse:\n\t\th, w = (file.height, file.width)\n\t\tsubdatasets = file.subdatasets\n\t\tif len(subdatasets) > 0:\n\t\t\timage = np.zeros((h, w, len(subdatasets)), dtype=np.uint8)\n\t\t\tfor i, subdataset in enumerate(subdatasets, 0):\n\t\t\t\twith rasterio.open(subdataset) as layer:\n\t\t\t\t\timage[:,:,i] = layer.read(1)\n```",
    "1232988": "It seems d488c759a.tiff in test set only have 1 channel, whats wrong?",
    "1233007": "and also \n57512b7f1 aa05346ff in test set\n095bf7a1f 1e2425f28 26dc41664 4ef6695ce c68fe75ea in train set have this problem",
    "1233077": "Yeah, I have met the same problem.",
    "1233042": "thanks Wiseman.X\nuse tiff.imread can return the correct size\nimg = tiff.imread('../input/hubmap-kidney-segmentation/train/c68fe75ea.tiff')\nimg.shape\n> (3, 26840, 49780)",
    "1233030": "I use this code to load tiff\ndata = rasterio.open('hubmap-kidney-segmentation/train/c68fe75ea.tiff')\ndata.read().shape\n> (1, 26840, 49780)",
    "1233911": "",
    "1233055": "",
    "1233001": ""
  }
}