{
  "id": 202171,
  "title": "Reshape + Transpose Trick by lafoss Explained",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/202171",
  "author_name": "",
  "post_date": "2020-12-08T18:00:54.169151800Z",
  "votes": 22,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all , <br>\nI would first like to thanks Kaggle and organizers for this lovely competition , although I was really shocked to see just 13 files amongst test and train when I just entered the competition . After two days of  Studying the competition , now its fairly clear of what is going on . </p>\n<p>Thanks to <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and his trick , I learned how to tile the big tiff files into something trainable .<br>\nIts a fairly simple and elegant trick but at first sight might look overwhelming and complicated , Since I know how difficult it can be deciphering it , I just thought to share the explanation for making it easy to others . I will walk you through the code one by one </p>\n<p>Step 1 : Reading the big tiff file</p>\n<blockquote>\n  <p>img = tiff.imread(os.path.join(DATA,index+'.tiff'))</p>\n</blockquote>\n<p>Step 2 : Checking if the Image is having three dimensions , if not squeeze the Image and rearrange axeses to get it to (H,W,C)</p>\n<blockquote>\n  <p>if len(img.shape) == 5:img = np.transpose(img.squeeze(), (1,2,0))<br>\n          mask = enc2mask(encs,(img.shape[1],img.shape[0]))</p>\n</blockquote>\n<p>Step 3 : Add padding on either sides to make the Image divisible into the given tiles properly . Its simple maths . If the height of Image is H (shape[0]) and we want the tile size (reduce<em>sz) to completely divide it , then what number should be substracted from H , its simple right (shape[0]%(reduce</em>sz)) and thus after subtracting this we can get pad value for Height . Similarly we can find pad value for width as well </p>\n<p><code>shape = img.shape\n  pad0 = (reduce*sz - shape[0]%(reduce*sz))%(reduce*sz)\n  pad1 = (reduce*sz - shape[1]%(reduce*sz))%(reduce*sz)</code></p>\n<p>After that we can use numpy to fill values of zero on either side of the Image</p>\n<p><code>img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],\n                    constant_values=0)\n        mask = np.pad(mask,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2]],\n                    constant_values=0)</code></p>\n<p>Step 4 : Resizing the Image to reduce it by 4(reduce) times</p>\n<p><code>img = cv2.resize(img,(img.shape[1]//reduce,img.shape[0]//reduce),\n                         interpolation = cv2.INTER_AREA)</code></p>\n<p>Step 5 : Now comes the tricky part which will need some explanation</p>\n<p>So the trick is based on how numpy arrays are stored in memory . Every numpy array be it any dimension for us is stored as a 1D array internally , meaning when we define a numpy array of shape let's say (512,512,3) , internally a contiguous memory location is alloted to this array and the are arranged as a sequence of numbers laid down as [ 1,2,3……512,1,2,3….512,1,2,3 ] and what we see on our end is just different views of the same 1D array . Now if we reshape the 3D array we just split the 1D array at a different place and get a different view . <br>\nThus if we reshape the original Image (H,W,C) into (number of tiles , H,W,C) our job is done right? That is the trick , the only thing we need to care about is we split at the right place .</p>\n<p><code>Here is how lafoss explains the trick :\nTo better understand what is going on, you should recall how the tensor is placed in the memory: everything is represented by a 1d array (without going into other complications), and the dimensionality is just an additional description for interpretation of this 1d array. So, for example let's consider a split of 1d array into tiles. The original array [0,1,2,3,4,5,6,7,8,9] after reshaping to (2,5) would look like [0,1,2,3,4|5,6,7,8,9]. Nothing happened with the way how the data is stored, you just put a separator showing that right now you interpret the data as a (2,5) shape array. Similar things are done for the images: u add separators to divide the image into tiles. The only problem right now is that the individual tiles are not contiguous in the memory. Therefore, you need to permute the dimensions, merge the dimensions for x and y indexes of tiles, and make everything contiguous. The last two steps are performed with reshape, and finally u have a tensor containing a list of tiles.</code></p>\n<h1>How is the splitting done ?</h1>\n<p>Now the 1D array might be arrange sequentially as H,W,C <br>\nWe break the height as H/tile_size(sz) and width as W/tile_size(sz) and thus the array is rehsaped to </p>\n<blockquote>\n  <p>(H/sz,sz,W/sz,sz,3)<br>\n  img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)</p>\n</blockquote>\n<p>Now we have created the tiles but they are contiguous in memory (that is not in right sequence) and thus we transpose the image to get </p>\n<blockquote>\n  <p>(H/sz,W/sz,sz,sz,3)<br>\n   img = img.transpose(0,2,1,3,4)</p>\n</blockquote>\n<p>Now all we need to do is to merge axis 0 and axis 1 to get to our tiles </p>\n<blockquote>\n  <p>(number of tiles , sz,sz,3)   <br>\n  Number of tiles = (H/sz) * (W/sz)<br>\n  img = img.reshape(-1,sz,sz,3)</p>\n</blockquote>\n<p>I hope this explanation helps people to understand it easily and allows the code to be used to any custom Image data they want in future</p>\n<p>Thanks for reading</p>",
  "messages": [
    {
      "id": "1106327",
      "postDate": "12/08/2020 18:00:54",
      "content": "<p>Hi all , <br>\nI would first like to thanks Kaggle and organizers for this lovely competition , although I was really shocked to see just 13 files amongst test and train when I just entered the competition . After two days of  Studying the competition , now its fairly clear of what is going on . </p>\n<p>Thanks to <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and his trick , I learned how to tile the big tiff files into something trainable .<br>\nIts a fairly simple and elegant trick but at first sight might look overwhelming and complicated , Since I know how difficult it can be deciphering it , I just thought to share the explanation for making it easy to others . I will walk you through the code one by one </p>\n<p>Step 1 : Reading the big tiff file</p>\n<blockquote>\n  <p>img = tiff.imread(os.path.join(DATA,index+'.tiff'))</p>\n</blockquote>\n<p>Step 2 : Checking if the Image is having three dimensions , if not squeeze the Image and rearrange axeses to get it to (H,W,C)</p>\n<blockquote>\n  <p>if len(img.shape) == 5:img = np.transpose(img.squeeze(), (1,2,0))<br>\n          mask = enc2mask(encs,(img.shape[1],img.shape[0]))</p>\n</blockquote>\n<p>Step 3 : Add padding on either sides to make the Image divisible into the given tiles properly . Its simple maths . If the height of Image is H (shape[0]) and we want the tile size (reduce<em>sz) to completely divide it , then what number should be substracted from H , its simple right (shape[0]%(reduce</em>sz)) and thus after subtracting this we can get pad value for Height . Similarly we can find pad value for width as well </p>\n<p><code>shape = img.shape\n  pad0 = (reduce*sz - shape[0]%(reduce*sz))%(reduce*sz)\n  pad1 = (reduce*sz - shape[1]%(reduce*sz))%(reduce*sz)</code></p>\n<p>After that we can use numpy to fill values of zero on either side of the Image</p>\n<p><code>img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],\n                    constant_values=0)\n        mask = np.pad(mask,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2]],\n                    constant_values=0)</code></p>\n<p>Step 4 : Resizing the Image to reduce it by 4(reduce) times</p>\n<p><code>img = cv2.resize(img,(img.shape[1]//reduce,img.shape[0]//reduce),\n                         interpolation = cv2.INTER_AREA)</code></p>\n<p>Step 5 : Now comes the tricky part which will need some explanation</p>\n<p>So the trick is based on how numpy arrays are stored in memory . Every numpy array be it any dimension for us is stored as a 1D array internally , meaning when we define a numpy array of shape let's say (512,512,3) , internally a contiguous memory location is alloted to this array and the are arranged as a sequence of numbers laid down as [ 1,2,3……512,1,2,3….512,1,2,3 ] and what we see on our end is just different views of the same 1D array . Now if we reshape the 3D array we just split the 1D array at a different place and get a different view . <br>\nThus if we reshape the original Image (H,W,C) into (number of tiles , H,W,C) our job is done right? That is the trick , the only thing we need to care about is we split at the right place .</p>\n<p><code>Here is how lafoss explains the trick :\nTo better understand what is going on, you should recall how the tensor is placed in the memory: everything is represented by a 1d array (without going into other complications), and the dimensionality is just an additional description for interpretation of this 1d array. So, for example let's consider a split of 1d array into tiles. The original array [0,1,2,3,4,5,6,7,8,9] after reshaping to (2,5) would look like [0,1,2,3,4|5,6,7,8,9]. Nothing happened with the way how the data is stored, you just put a separator showing that right now you interpret the data as a (2,5) shape array. Similar things are done for the images: u add separators to divide the image into tiles. The only problem right now is that the individual tiles are not contiguous in the memory. Therefore, you need to permute the dimensions, merge the dimensions for x and y indexes of tiles, and make everything contiguous. The last two steps are performed with reshape, and finally u have a tensor containing a list of tiles.</code></p>\n<h1>How is the splitting done ?</h1>\n<p>Now the 1D array might be arrange sequentially as H,W,C <br>\nWe break the height as H/tile_size(sz) and width as W/tile_size(sz) and thus the array is rehsaped to </p>\n<blockquote>\n  <p>(H/sz,sz,W/sz,sz,3)<br>\n  img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)</p>\n</blockquote>\n<p>Now we have created the tiles but they are contiguous in memory (that is not in right sequence) and thus we transpose the image to get </p>\n<blockquote>\n  <p>(H/sz,W/sz,sz,sz,3)<br>\n   img = img.transpose(0,2,1,3,4)</p>\n</blockquote>\n<p>Now all we need to do is to merge axis 0 and axis 1 to get to our tiles </p>\n<blockquote>\n  <p>(number of tiles , sz,sz,3)   <br>\n  Number of tiles = (H/sz) * (W/sz)<br>\n  img = img.reshape(-1,sz,sz,3)</p>\n</blockquote>\n<p>I hope this explanation helps people to understand it easily and allows the code to be used to any custom Image data they want in future</p>\n<p>Thanks for reading</p>",
      "rawMarkdown": "Hi all , \nI would first like to thanks Kaggle and organizers for this lovely competition , although I was really shocked to see just 13 files amongst test and train when I just entered the competition . After two days of  Studying the competition , now its fairly clear of what is going on . \n\nThanks to @iafoss and his trick , I learned how to tile the big tiff files into something trainable .\nIts a fairly simple and elegant trick but at first sight might look overwhelming and complicated , Since I know how difficult it can be deciphering it , I just thought to share the explanation for making it easy to others . I will walk you through the code one by one \n\nStep 1 : Reading the big tiff file\n\n> img = tiff.imread(os.path.join(DATA,index+'.tiff'))\n\nStep 2 : Checking if the Image is having three dimensions , if not squeeze the Image and rearrange axeses to get it to (H,W,C)\n\n>  if len(img.shape) == 5:img = np.transpose(img.squeeze(), (1,2,0))\n        mask = enc2mask(encs,(img.shape[1],img.shape[0]))\n\nStep 3 : Add padding on either sides to make the Image divisible into the given tiles properly . Its simple maths . If the height of Image is H (shape[0]) and we want the tile size (reduce*sz) to completely divide it , then what number should be substracted from H , its simple right (shape[0]%(reduce*sz)) and thus after subtracting this we can get pad value for Height . Similarly we can find pad value for width as well \n\n`shape = img.shape\n  pad0 = (reduce*sz - shape[0]%(reduce*sz))%(reduce*sz)\n  pad1 = (reduce*sz - shape[1]%(reduce*sz))%(reduce*sz)`\n\nAfter that we can use numpy to fill values of zero on either side of the Image\n\n`img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],\n                    constant_values=0)\n        mask = np.pad(mask,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2]],\n                    constant_values=0)`\n\nStep 4 : Resizing the Image to reduce it by 4(reduce) times\n\n`img = cv2.resize(img,(img.shape[1]//reduce,img.shape[0]//reduce),\n                         interpolation = cv2.INTER_AREA)`\n\nStep 5 : Now comes the tricky part which will need some explanation\n\nSo the trick is based on how numpy arrays are stored in memory . Every numpy array be it any dimension for us is stored as a 1D array internally , meaning when we define a numpy array of shape let's say (512,512,3) , internally a contiguous memory location is alloted to this array and the are arranged as a sequence of numbers laid down as [ 1,2,3......512,1,2,3....512,1,2,3 ] and what we see on our end is just different views of the same 1D array . Now if we reshape the 3D array we just split the 1D array at a different place and get a different view . \nThus if we reshape the original Image (H,W,C) into (number of tiles , H,W,C) our job is done right? That is the trick , the only thing we need to care about is we split at the right place .\n\n`Here is how lafoss explains the trick :\nTo better understand what is going on, you should recall how the tensor is placed in the memory: everything is represented by a 1d array (without going into other complications), and the dimensionality is just an additional description for interpretation of this 1d array. So, for example let's consider a split of 1d array into tiles. The original array [0,1,2,3,4,5,6,7,8,9] after reshaping to (2,5) would look like [0,1,2,3,4|5,6,7,8,9]. Nothing happened with the way how the data is stored, you just put a separator showing that right now you interpret the data as a (2,5) shape array. Similar things are done for the images: u add separators to divide the image into tiles. The only problem right now is that the individual tiles are not contiguous in the memory. Therefore, you need to permute the dimensions, merge the dimensions for x and y indexes of tiles, and make everything contiguous. The last two steps are performed with reshape, and finally u have a tensor containing a list of tiles.`\n\n# How is the splitting done ?\n\nNow the 1D array might be arrange sequentially as H,W,C \nWe break the height as H/tile_size(sz) and width as W/tile_size(sz) and thus the array is rehsaped to \n\n> (H/sz,sz,W/sz,sz,3)\nimg = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)\n\nNow we have created the tiles but they are contiguous in memory (that is not in right sequence) and thus we transpose the image to get \n\n> (H/sz,W/sz,sz,sz,3)\n img = img.transpose(0,2,1,3,4)\n\nNow all we need to do is to merge axis 0 and axis 1 to get to our tiles \n\n>  (number of tiles , sz,sz,3)   \nNumber of tiles = (H/sz) * (W/sz)\nimg = img.reshape(-1,sz,sz,3)\n\nI hope this explanation helps people to understand it easily and allows the code to be used to any custom Image data they want in future\n\nThanks for reading",
      "votes": null
    },
    {
      "id": "1106770",
      "postDate": "12/09/2020 05:37:13",
      "content": "<p>thanks for your explanation</p>",
      "rawMarkdown": "thanks for your explanation",
      "votes": null
    },
    {
      "id": "1109723",
      "postDate": "12/12/2020 02:10:32",
      "content": "<p>Very clear explanation 💯 </p>",
      "rawMarkdown": "Very clear explanation 💯",
      "votes": null
    },
    {
      "id": "1116596",
      "postDate": "12/17/2020 10:01:29",
      "content": "<p>Wow, so cool. So does the final <code>img = img.reshape(-1, sz, sz, 3)</code> bring us all tiles in the format of a huge \"batch\"? And if we want to store a local dataset, we can simply loop over batch axis and output each image? </p>\n<p>Thanks!</p>",
      "rawMarkdown": "Wow, so cool. So does the final `img = img.reshape(-1, sz, sz, 3)` bring us all tiles in the format of a huge \"batch\"? And if we want to store a local dataset, we can simply loop over batch axis and output each image? \n\nThanks!",
      "votes": null
    },
    {
      "id": "1228111",
      "postDate": "03/06/2021 06:08:42",
      "content": "<p>Cool! Thank you for your explanation!</p>",
      "rawMarkdown": "Cool! Thank you for your explanation!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1106770,
      "author_name": "tiandaye",
      "author_url": "",
      "post_date": "12/09/2020 05:37:13",
      "content": "<p>thanks for your explanation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1109723,
      "author_name": "rowhitswami",
      "author_url": "",
      "post_date": "12/12/2020 02:10:32",
      "content": "<p>Very clear explanation 💯 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1116596,
      "author_name": "fuckvenkatraman",
      "author_url": "",
      "post_date": "12/17/2020 10:01:29",
      "content": "<p>Wow, so cool. So does the final <code>img = img.reshape(-1, sz, sz, 3)</code> bring us all tiles in the format of a huge \"batch\"? And if we want to store a local dataset, we can simply loop over batch axis and output each image? </p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1228111,
      "author_name": "wangtyi",
      "author_url": "",
      "post_date": "03/06/2021 06:08:42",
      "content": "<p>Cool! Thank you for your explanation!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1106327": "Hi all , \nI would first like to thanks Kaggle and organizers for this lovely competition , although I was really shocked to see just 13 files amongst test and train when I just entered the competition . After two days of  Studying the competition , now its fairly clear of what is going on . \n\nThanks to @iafoss and his trick , I learned how to tile the big tiff files into something trainable .\nIts a fairly simple and elegant trick but at first sight might look overwhelming and complicated , Since I know how difficult it can be deciphering it , I just thought to share the explanation for making it easy to others . I will walk you through the code one by one \n\nStep 1 : Reading the big tiff file\n\n> img = tiff.imread(os.path.join(DATA,index+'.tiff'))\n\nStep 2 : Checking if the Image is having three dimensions , if not squeeze the Image and rearrange axeses to get it to (H,W,C)\n\n>  if len(img.shape) == 5:img = np.transpose(img.squeeze(), (1,2,0))\n        mask = enc2mask(encs,(img.shape[1],img.shape[0]))\n\nStep 3 : Add padding on either sides to make the Image divisible into the given tiles properly . Its simple maths . If the height of Image is H (shape[0]) and we want the tile size (reduce*sz) to completely divide it , then what number should be substracted from H , its simple right (shape[0]%(reduce*sz)) and thus after subtracting this we can get pad value for Height . Similarly we can find pad value for width as well \n\n`shape = img.shape\n  pad0 = (reduce*sz - shape[0]%(reduce*sz))%(reduce*sz)\n  pad1 = (reduce*sz - shape[1]%(reduce*sz))%(reduce*sz)`\n\nAfter that we can use numpy to fill values of zero on either side of the Image\n\n`img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],\n                    constant_values=0)\n        mask = np.pad(mask,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2]],\n                    constant_values=0)`\n\nStep 4 : Resizing the Image to reduce it by 4(reduce) times\n\n`img = cv2.resize(img,(img.shape[1]//reduce,img.shape[0]//reduce),\n                         interpolation = cv2.INTER_AREA)`\n\nStep 5 : Now comes the tricky part which will need some explanation\n\nSo the trick is based on how numpy arrays are stored in memory . Every numpy array be it any dimension for us is stored as a 1D array internally , meaning when we define a numpy array of shape let's say (512,512,3) , internally a contiguous memory location is alloted to this array and the are arranged as a sequence of numbers laid down as [ 1,2,3......512,1,2,3....512,1,2,3 ] and what we see on our end is just different views of the same 1D array . Now if we reshape the 3D array we just split the 1D array at a different place and get a different view . \nThus if we reshape the original Image (H,W,C) into (number of tiles , H,W,C) our job is done right? That is the trick , the only thing we need to care about is we split at the right place .\n\n`Here is how lafoss explains the trick :\nTo better understand what is going on, you should recall how the tensor is placed in the memory: everything is represented by a 1d array (without going into other complications), and the dimensionality is just an additional description for interpretation of this 1d array. So, for example let's consider a split of 1d array into tiles. The original array [0,1,2,3,4,5,6,7,8,9] after reshaping to (2,5) would look like [0,1,2,3,4|5,6,7,8,9]. Nothing happened with the way how the data is stored, you just put a separator showing that right now you interpret the data as a (2,5) shape array. Similar things are done for the images: u add separators to divide the image into tiles. The only problem right now is that the individual tiles are not contiguous in the memory. Therefore, you need to permute the dimensions, merge the dimensions for x and y indexes of tiles, and make everything contiguous. The last two steps are performed with reshape, and finally u have a tensor containing a list of tiles.`\n\n# How is the splitting done ?\n\nNow the 1D array might be arrange sequentially as H,W,C \nWe break the height as H/tile_size(sz) and width as W/tile_size(sz) and thus the array is rehsaped to \n\n> (H/sz,sz,W/sz,sz,3)\nimg = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)\n\nNow we have created the tiles but they are contiguous in memory (that is not in right sequence) and thus we transpose the image to get \n\n> (H/sz,W/sz,sz,sz,3)\n img = img.transpose(0,2,1,3,4)\n\nNow all we need to do is to merge axis 0 and axis 1 to get to our tiles \n\n>  (number of tiles , sz,sz,3)   \nNumber of tiles = (H/sz) * (W/sz)\nimg = img.reshape(-1,sz,sz,3)\n\nI hope this explanation helps people to understand it easily and allows the code to be used to any custom Image data they want in future\n\nThanks for reading",
    "1106770": "thanks for your explanation",
    "1109723": "Very clear explanation 💯",
    "1116596": "Wow, so cool. So does the final `img = img.reshape(-1, sz, sz, 3)` bring us all tiles in the format of a huge \"batch\"? And if we want to store a local dataset, we can simply loop over batch axis and output each image? \n\nThanks!",
    "1228111": "Cool! Thank you for your explanation!"
  },
  "source": "meta"
}