{
  "id": 248440,
  "title": "Batch Size Issue while using Spatial Images",
  "url": "/competitions/seti-breakthrough-listen/discussion/248440",
  "author_name": "Tarushi Pathak",
  "post_date": "2021-06-23T14:43:07.066000",
  "votes": 0,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello Everyone !</p>\n<p>I am running across an issue while training on Spatial Images. The model does start training but somewhere in the middle , fails to complete the process with the following error:</p>\n<pre><code>InvalidArgumentError: 2 root error(s) found.\n  (0) Invalid argument:  Incompatible shapes: [4,1] vs. [32,1]\n     [[node binary_crossentropy/logistic_loss/mul (defined at &lt;ipython-input-44-ae511a693fcc&gt;:11) ]]\n     [[assert_less_equal/All/_196]]\n  (1) Invalid argument:  Incompatible shapes: [4,1] vs. [32,1]\n     [[node binary_crossentropy/logistic_loss/mul (defined at &lt;ipython-input-44-ae511a693fcc&gt;:11) ]]\n0 successful operations.\n0 derived errors ignored. [Op:__inference_train_function_16980]\n\nFunction call stack:\ntrain_function -&gt; train_function\n</code></pre>\n<p>The below code helps in converting the numpy files to numpy matrices. The else condition in the get item function converts the images to their spatial form. </p>\n<pre><code>class CustomData(Sequence):\n    def __init__(self, x_set, y_set=None, batch_size=32 , channel = True):\n        self.x, self.y = x_set, y_set\n        self.batch_size = batch_size\n        self.is_train = False if y_set is None else True\n        self.channel = channel\n\n    def __len__(self):\n        return math.ceil(len(self.x) / self.batch_size)\n\n    def __getitem__(self, idx):\n        batch_ids = self.x[idx * self.batch_size: (idx + 1) * self.batch_size]\n        if self.y is not None:\n            batch_y = self.y[idx * self.batch_size: (idx + 1) * self.batch_size]\n\n        # taking channels 1, 3, and 5 only\n        if self.channel:\n            list_x = [np.load(path)[::2] for path in batch_ids]\n            batch_x = np.moveaxis(list_x,1,-1)\n            batch_x = batch_x.astype(np.float32) / 255\n        else:\n   #spatial images\n            list_x = [np.vstack(np.load(path)[::2].astype(np.float32)).transpose((1,0)) for path in batch_ids]\n            batch_x = np.resize(list_x,(self.batch_size,273,256,3)) / 255\n        if self.is_train:\n            return batch_x, batch_y\n        else:\n            return batch_x\n</code></pre>\n<p>TensorFlow Version = 2.4.1</p>",
  "messages": [
    {
      "id": 1365969,
      "postDate": "2021-06-26T10:13:05.747Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tarushi89\" target=\"_blank\">@tarushi89</a> ,<br>\nThe problem issue I see in this approach is that for the last batch it is not getting the same examples as all the other batches which is changing the dimension of the tensor your NN is expecting.<br>\nMy advise would be to either drop the last batch or select a batch size which perfectly divides your number of examples without leaving any remainder.</p>",
      "rawMarkdown": "Hi @tarushi89 ,\nThe problem issue I see in this approach is that for the last batch it is not getting the same examples as all the other batches which is changing the dimension of the tensor your NN is expecting.\nMy advise would be to either drop the last batch or select a batch size which perfectly divides your number of examples without leaving any remainder.",
      "votes": 1
    },
    {
      "id": 1364006,
      "postDate": "2021-06-24T14:31:04.087Z",
      "content": "<p>I am not using TF but my guess is you get a batch size of 4 in your last batch.  If your code has batch size hardcoded then you should skip the last batch if it is not full size.</p>",
      "rawMarkdown": "I am not using TF but my guess is you get a batch size of 4 in your last batch.  If your code has batch size hardcoded then you should skip the last batch if it is not full size.",
      "votes": 1,
      "replies": [
        {
          "id": 1371139,
          "postDate": "2021-06-30T17:39:51.203Z",
          "content": "<p>Hi ! Thanks , I hadn't thought of dropping the last batch. 😅 <br>\nBut anyway , I used Mithil's code and that helped me out.</p>",
          "rawMarkdown": "Hi ! Thanks , I hadn't thought of dropping the last batch. 😅 \nBut anyway , I used Mithil's code and that helped me out."
        }
      ]
    },
    {
      "id": 1363204,
      "postDate": "2021-06-24T03:02:47.470Z",
      "content": "<p>You can try this as a generator </p>\n<p>`def <strong>getitem</strong>(self, ids):</p>\n<pre><code>    batch_ids = self.idx[ids * self.batch_size:(ids + 1) * self.batch_size]\n    if self.y is not None:\n        batch_y = self.y[ids * self.batch_size: (ids + 1) * self.batch_size]`\n\n    list_x1 = np.array(\n        [np.load(id_to_path(x, self.is_train))[::2].reshape(3 * 273, 256) for x in batch_ids]).transpose(1, 2, 0)\n    list_x2 = np.array([np.zeros((3, 3 * 273, 256)) for x in batch_ids]).transpose(1, 2, 3, 0)\n    list_x2[0::] = list_x1\n    list_x2[1::] = list_x1\n    list_x2[2::] = list_x1\n    batch_x = np.transpose(list_x2, (3, 1, 2, 0))\n\n\n    if self.is_train:\n        return batch_x, batch_y\n    else:\n        return batch_x\n</code></pre>",
      "rawMarkdown": "You can try this as a generator \n\n\n`def __getitem__(self, ids):\n\n\t\tbatch_ids = self.idx[ids * self.batch_size:(ids + 1) * self.batch_size]\n\t\tif self.y is not None:\n\t\t\tbatch_y = self.y[ids * self.batch_size: (ids + 1) * self.batch_size]`\n\n\t\tlist_x1 = np.array(\n\t\t\t[np.load(id_to_path(x, self.is_train))[::2].reshape(3 * 273, 256) for x in batch_ids]).transpose(1, 2, 0)\n\t\tlist_x2 = np.array([np.zeros((3, 3 * 273, 256)) for x in batch_ids]).transpose(1, 2, 3, 0)\n\t\tlist_x2[0::] = list_x1\n\t\tlist_x2[1::] = list_x1\n\t\tlist_x2[2::] = list_x1\n\t\tbatch_x = np.transpose(list_x2, (3, 1, 2, 0))\n\t\t\n\n\t\tif self.is_train:\n\t\t\treturn batch_x, batch_y\n\t\telse:\n\t\t\treturn batch_x\n",
      "votes": 1,
      "replies": [
        {
          "id": 1363452,
          "postDate": "2021-06-24T06:40:16.853Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1363699,
          "postDate": "2021-06-24T09:28:02.210Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> ! Thanks , this worked ! However I had to reduce the batch size to 16 , as it was giving an OOM  error for batch size 32 . Is yours working with batch size 32? Another thing, are you using input shape as (819,256,3) ?</p>",
          "rawMarkdown": "Hi @mithilsalunkhe ! Thanks , this worked ! However I had to reduce the batch size to 16 , as it was giving an OOM  error for batch size 32 . Is yours working with batch size 32? Another thing, are you using input shape as (819,256,3) ?"
        },
        {
          "id": 1363778,
          "postDate": "2021-06-24T11:04:46.247Z",
          "content": "<p>Actually I resize my image to (512,1024,3) after this so I have to use batch size 12.<br>\nEdit - I become Discussion Expert With this comment </p>",
          "rawMarkdown": "Actually I resize my image to (512,1024,3) after this so I have to use batch size 12.\nEdit - I become Discussion Expert With this comment \n\n",
          "votes": 1
        },
        {
          "id": 1363827,
          "postDate": "2021-06-24T11:37:50.920Z",
          "content": "<p><a href=\"https://www.kaggle.com/tarushi89\" target=\"_blank\">@tarushi89</a> try tpu.<br>\nwith tpu, speed increase almost x20</p>",
          "rawMarkdown": "@tarushi89 try tpu.\nwith tpu, speed increase almost x20",
          "votes": 1
        },
        {
          "id": 1363871,
          "postDate": "2021-06-24T12:32:41.563Z",
          "content": "<p><a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a> thanks ! I'll do that. :)</p>",
          "rawMarkdown": "@assign thanks ! I'll do that. :)"
        },
        {
          "id": 1363880,
          "postDate": "2021-06-24T12:46:02.220Z",
          "content": "<p><a href=\"https://www.kaggle.com/as1125883\" target=\"_blank\">@as1125883</a> Generators don't work with TPU.</p>",
          "rawMarkdown": "@as1125883 Generators don't work with TPU."
        }
      ]
    },
    {
      "id": 1362626,
      "postDate": "2021-06-23T14:43:07.067Z",
      "content": "<p>Hello Everyone !</p>\n<p>I am running across an issue while training on Spatial Images. The model does start training but somewhere in the middle , fails to complete the process with the following error:</p>\n<pre><code>InvalidArgumentError: 2 root error(s) found.\n  (0) Invalid argument:  Incompatible shapes: [4,1] vs. [32,1]\n     [[node binary_crossentropy/logistic_loss/mul (defined at &lt;ipython-input-44-ae511a693fcc&gt;:11) ]]\n     [[assert_less_equal/All/_196]]\n  (1) Invalid argument:  Incompatible shapes: [4,1] vs. [32,1]\n     [[node binary_crossentropy/logistic_loss/mul (defined at &lt;ipython-input-44-ae511a693fcc&gt;:11) ]]\n0 successful operations.\n0 derived errors ignored. [Op:__inference_train_function_16980]\n\nFunction call stack:\ntrain_function -&gt; train_function\n</code></pre>\n<p>The below code helps in converting the numpy files to numpy matrices. The else condition in the get item function converts the images to their spatial form. </p>\n<pre><code>class CustomData(Sequence):\n    def __init__(self, x_set, y_set=None, batch_size=32 , channel = True):\n        self.x, self.y = x_set, y_set\n        self.batch_size = batch_size\n        self.is_train = False if y_set is None else True\n        self.channel = channel\n\n    def __len__(self):\n        return math.ceil(len(self.x) / self.batch_size)\n\n    def __getitem__(self, idx):\n        batch_ids = self.x[idx * self.batch_size: (idx + 1) * self.batch_size]\n        if self.y is not None:\n            batch_y = self.y[idx * self.batch_size: (idx + 1) * self.batch_size]\n\n        # taking channels 1, 3, and 5 only\n        if self.channel:\n            list_x = [np.load(path)[::2] for path in batch_ids]\n            batch_x = np.moveaxis(list_x,1,-1)\n            batch_x = batch_x.astype(np.float32) / 255\n        else:\n   #spatial images\n            list_x = [np.vstack(np.load(path)[::2].astype(np.float32)).transpose((1,0)) for path in batch_ids]\n            batch_x = np.resize(list_x,(self.batch_size,273,256,3)) / 255\n        if self.is_train:\n            return batch_x, batch_y\n        else:\n            return batch_x\n</code></pre>\n<p>TensorFlow Version = 2.4.1</p>",
      "rawMarkdown": "Hello Everyone !\n\nI am running across an issue while training on Spatial Images. The model does start training but somewhere in the middle , fails to complete the process with the following error:\n\n```\nInvalidArgumentError: 2 root error(s) found.\n  (0) Invalid argument:  Incompatible shapes: [4,1] vs. [32,1]\n\t [[node binary_crossentropy/logistic_loss/mul (defined at <ipython-input-44-ae511a693fcc>:11) ]]\n\t [[assert_less_equal/All/_196]]\n  (1) Invalid argument:  Incompatible shapes: [4,1] vs. [32,1]\n\t [[node binary_crossentropy/logistic_loss/mul (defined at <ipython-input-44-ae511a693fcc>:11) ]]\n0 successful operations.\n0 derived errors ignored. [Op:__inference_train_function_16980]\n\nFunction call stack:\ntrain_function -> train_function\n\n```\nThe below code helps in converting the numpy files to numpy matrices. The else condition in the get item function converts the images to their spatial form. \n\n```\nclass CustomData(Sequence):\n    def __init__(self, x_set, y_set=None, batch_size=32 , channel = True):\n        self.x, self.y = x_set, y_set\n        self.batch_size = batch_size\n        self.is_train = False if y_set is None else True\n        self.channel = channel\n    \n    def __len__(self):\n        return math.ceil(len(self.x) / self.batch_size)\n    \n    def __getitem__(self, idx):\n        batch_ids = self.x[idx * self.batch_size: (idx + 1) * self.batch_size]\n        if self.y is not None:\n            batch_y = self.y[idx * self.batch_size: (idx + 1) * self.batch_size]\n        \n        # taking channels 1, 3, and 5 only\n        if self.channel:\n            list_x = [np.load(path)[::2] for path in batch_ids]\n            batch_x = np.moveaxis(list_x,1,-1)\n            batch_x = batch_x.astype(np.float32) / 255\n        else:\n   #spatial images\n            list_x = [np.vstack(np.load(path)[::2].astype(np.float32)).transpose((1,0)) for path in batch_ids]\n            batch_x = np.resize(list_x,(self.batch_size,273,256,3)) / 255\n        if self.is_train:\n            return batch_x, batch_y\n        else:\n            return batch_x\n```\n\n\nTensorFlow Version = 2.4.1"
    }
  ],
  "comments": [
    {
      "id": 1365969,
      "author_name": "Manav",
      "author_url": "",
      "post_date": "2021-06-26T10:13:05.747000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tarushi89\" target=\"_blank\">@tarushi89</a> ,<br>\nThe problem issue I see in this approach is that for the last batch it is not getting the same examples as all the other batches which is changing the dimension of the tensor your NN is expecting.<br>\nMy advise would be to either drop the last batch or select a batch size which perfectly divides your number of examples without leaving any remainder.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1364006,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-06-24T14:31:04.087000",
      "content": "<p>I am not using TF but my guess is you get a batch size of 4 in your last batch.  If your code has batch size hardcoded then you should skip the last batch if it is not full size.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1371139,
          "author_name": "Tarushi Pathak",
          "author_url": "",
          "post_date": "2021-06-30T17:39:51.203000",
          "content": "<p>Hi ! Thanks , I hadn't thought of dropping the last batch. 😅 <br>\nBut anyway , I used Mithil's code and that helped me out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1363204,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-06-24T03:02:47.470000",
      "content": "<p>You can try this as a generator </p>\n<p>`def <strong>getitem</strong>(self, ids):</p>\n<pre><code>    batch_ids = self.idx[ids * self.batch_size:(ids + 1) * self.batch_size]\n    if self.y is not None:\n        batch_y = self.y[ids * self.batch_size: (ids + 1) * self.batch_size]`\n\n    list_x1 = np.array(\n        [np.load(id_to_path(x, self.is_train))[::2].reshape(3 * 273, 256) for x in batch_ids]).transpose(1, 2, 0)\n    list_x2 = np.array([np.zeros((3, 3 * 273, 256)) for x in batch_ids]).transpose(1, 2, 3, 0)\n    list_x2[0::] = list_x1\n    list_x2[1::] = list_x1\n    list_x2[2::] = list_x1\n    batch_x = np.transpose(list_x2, (3, 1, 2, 0))\n\n\n    if self.is_train:\n        return batch_x, batch_y\n    else:\n        return batch_x\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 1363452,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-06-24T06:40:16.853000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1363699,
          "author_name": "Tarushi Pathak",
          "author_url": "",
          "post_date": "2021-06-24T09:28:02.210000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> ! Thanks , this worked ! However I had to reduce the batch size to 16 , as it was giving an OOM  error for batch size 32 . Is yours working with batch size 32? Another thing, are you using input shape as (819,256,3) ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1363778,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-06-24T11:04:46.247000",
          "content": "<p>Actually I resize my image to (512,1024,3) after this so I have to use batch size 12.<br>\nEdit - I become Discussion Expert With this comment </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1363827,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-24T11:37:50.920000",
          "content": "<p><a href=\"https://www.kaggle.com/tarushi89\" target=\"_blank\">@tarushi89</a> try tpu.<br>\nwith tpu, speed increase almost x20</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1363871,
          "author_name": "Tarushi Pathak",
          "author_url": "",
          "post_date": "2021-06-24T12:32:41.563000",
          "content": "<p><a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a> thanks ! I'll do that. :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1363880,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-06-24T12:46:02.220000",
          "content": "<p><a href=\"https://www.kaggle.com/as1125883\" target=\"_blank\">@as1125883</a> Generators don't work with TPU.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1365969": "Hi @tarushi89 ,\nThe problem issue I see in this approach is that for the last batch it is not getting the same examples as all the other batches which is changing the dimension of the tensor your NN is expecting.\nMy advise would be to either drop the last batch or select a batch size which perfectly divides your number of examples without leaving any remainder.",
    "1364006": "I am not using TF but my guess is you get a batch size of 4 in your last batch.  If your code has batch size hardcoded then you should skip the last batch if it is not full size.",
    "1363204": "You can try this as a generator \n\n\n`def __getitem__(self, ids):\n\n\t\tbatch_ids = self.idx[ids * self.batch_size:(ids + 1) * self.batch_size]\n\t\tif self.y is not None:\n\t\t\tbatch_y = self.y[ids * self.batch_size: (ids + 1) * self.batch_size]`\n\n\t\tlist_x1 = np.array(\n\t\t\t[np.load(id_to_path(x, self.is_train))[::2].reshape(3 * 273, 256) for x in batch_ids]).transpose(1, 2, 0)\n\t\tlist_x2 = np.array([np.zeros((3, 3 * 273, 256)) for x in batch_ids]).transpose(1, 2, 3, 0)\n\t\tlist_x2[0::] = list_x1\n\t\tlist_x2[1::] = list_x1\n\t\tlist_x2[2::] = list_x1\n\t\tbatch_x = np.transpose(list_x2, (3, 1, 2, 0))\n\t\t\n\n\t\tif self.is_train:\n\t\t\treturn batch_x, batch_y\n\t\telse:\n\t\t\treturn batch_x\n",
    "1362626": "Hello Everyone !\n\nI am running across an issue while training on Spatial Images. The model does start training but somewhere in the middle , fails to complete the process with the following error:\n\n```\nInvalidArgumentError: 2 root error(s) found.\n  (0) Invalid argument:  Incompatible shapes: [4,1] vs. [32,1]\n\t [[node binary_crossentropy/logistic_loss/mul (defined at <ipython-input-44-ae511a693fcc>:11) ]]\n\t [[assert_less_equal/All/_196]]\n  (1) Invalid argument:  Incompatible shapes: [4,1] vs. [32,1]\n\t [[node binary_crossentropy/logistic_loss/mul (defined at <ipython-input-44-ae511a693fcc>:11) ]]\n0 successful operations.\n0 derived errors ignored. [Op:__inference_train_function_16980]\n\nFunction call stack:\ntrain_function -> train_function\n\n```\nThe below code helps in converting the numpy files to numpy matrices. The else condition in the get item function converts the images to their spatial form. \n\n```\nclass CustomData(Sequence):\n    def __init__(self, x_set, y_set=None, batch_size=32 , channel = True):\n        self.x, self.y = x_set, y_set\n        self.batch_size = batch_size\n        self.is_train = False if y_set is None else True\n        self.channel = channel\n    \n    def __len__(self):\n        return math.ceil(len(self.x) / self.batch_size)\n    \n    def __getitem__(self, idx):\n        batch_ids = self.x[idx * self.batch_size: (idx + 1) * self.batch_size]\n        if self.y is not None:\n            batch_y = self.y[idx * self.batch_size: (idx + 1) * self.batch_size]\n        \n        # taking channels 1, 3, and 5 only\n        if self.channel:\n            list_x = [np.load(path)[::2] for path in batch_ids]\n            batch_x = np.moveaxis(list_x,1,-1)\n            batch_x = batch_x.astype(np.float32) / 255\n        else:\n   #spatial images\n            list_x = [np.vstack(np.load(path)[::2].astype(np.float32)).transpose((1,0)) for path in batch_ids]\n            batch_x = np.resize(list_x,(self.batch_size,273,256,3)) / 255\n        if self.is_train:\n            return batch_x, batch_y\n        else:\n            return batch_x\n```\n\n\nTensorFlow Version = 2.4.1"
  }
}