{
  "id": 481133,
  "title": "need help with tensorflow 2.15.0 model debugging",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/481133",
  "author_name": "",
  "post_date": "2024-03-02T09:41:31.000183200Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am trying to use the latest environments but i am facing some weird issue, I have tried using a TPU which also had teh same error, the model I am experimenting with is the EfficientNet model from this <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">notebook</a>, the model has 6 output classes, but when predicting it generates wrong number of outputs</p>\n<p>for i, batch in enumerate(train_gen):<br>\n    print(f\"Batch {i+1}: Input shape: {batch[0].shape}, Target shape: {batch[1].shape}\")<br>\n    break</p>\n<p>with strategy.scope():<br>\n    model = build_model()<br>\n    model.fit( batch[0], batch[1],epochs = 1, verbose=1)<br>\n    oof = model.predict(batch[0], verbose=1)<br>\n    print(batch[0].shape,batch[1].shape)<br>\n    print(model.summary())</p>\n<p>(44, 128, 256, 8) (44, 6) are the respectives of batch[0] and batch[1], </p>\n<p>and the oof.shape is <br>\n(2,)<br>\nI nearly wasted 10 hours of GPU previous week trying to saving my notebook versions while experimenting, just to know it was all wasted, I encountered the same is occuring in TPU and when I am passing the datagenerator itself, it runs fine the entire training process and its throwing error in the last split</p>\n<p>I would appreciate your help in fixing this<br>\nThank you.</p>",
  "messages": [
    {
      "id": "2677566",
      "postDate": "03/02/2024 09:41:31",
      "content": "<p>I am trying to use the latest environments but i am facing some weird issue, I have tried using a TPU which also had teh same error, the model I am experimenting with is the EfficientNet model from this <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">notebook</a>, the model has 6 output classes, but when predicting it generates wrong number of outputs</p>\n<p>for i, batch in enumerate(train_gen):<br>\n    print(f\"Batch {i+1}: Input shape: {batch[0].shape}, Target shape: {batch[1].shape}\")<br>\n    break</p>\n<p>with strategy.scope():<br>\n    model = build_model()<br>\n    model.fit( batch[0], batch[1],epochs = 1, verbose=1)<br>\n    oof = model.predict(batch[0], verbose=1)<br>\n    print(batch[0].shape,batch[1].shape)<br>\n    print(model.summary())</p>\n<p>(44, 128, 256, 8) (44, 6) are the respectives of batch[0] and batch[1], </p>\n<p>and the oof.shape is <br>\n(2,)<br>\nI nearly wasted 10 hours of GPU previous week trying to saving my notebook versions while experimenting, just to know it was all wasted, I encountered the same is occuring in TPU and when I am passing the datagenerator itself, it runs fine the entire training process and its throwing error in the last split</p>\n<p>I would appreciate your help in fixing this<br>\nThank you.</p>",
      "rawMarkdown": "I am trying to use the latest environments but i am facing some weird issue, I have tried using a TPU which also had teh same error, the model I am experimenting with is the EfficientNet model from this [notebook](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43), the model has 6 output classes, but when predicting it generates wrong number of outputs\n\nfor i, batch in enumerate(train_gen):\n    print(f\"Batch {i+1}: Input shape: {batch[0].shape}, Target shape: {batch[1].shape}\")\n    break\n\nwith strategy.scope():\n    model = build_model()\n    model.fit( batch[0], batch[1],epochs = 1, verbose=1)\n    oof = model.predict(batch[0], verbose=1)\n    print(batch[0].shape,batch[1].shape)\n    print(model.summary())\n\n(44, 128, 256, 8) (44, 6) are the respectives of batch[0] and batch[1], \n\nand the oof.shape is \n(2,)\nI nearly wasted 10 hours of GPU previous week trying to saving my notebook versions while experimenting, just to know it was all wasted, I encountered the same is occuring in TPU and when I am passing the datagenerator itself, it runs fine the entire training process and its throwing error in the last split\n\nI would appreciate your help in fixing this\nThank you.",
      "votes": null
    },
    {
      "id": "2678585",
      "postDate": "03/02/2024 22:09:50",
      "content": "<p>I came across the same issue with TPU, I also tried to move to 2.15.0, so I could switch easily between Colab and Kaggle, but I got alot of headache.<br>\nI think there is a bug in 2.14 and 2.15 that doesn't register the GPUs properly, you see that error at the very beginning when you try to import tensorflow, that issue could be related, I am not sure.</p>",
      "rawMarkdown": "I came across the same issue with TPU, I also tried to move to 2.15.0, so I could switch easily between Colab and Kaggle, but I got alot of headache.\nI think there is a bug in 2.14 and 2.15 that doesn't register the GPUs properly, you see that error at the very beginning when you try to import tensorflow, that issue could be related, I am not sure.",
      "votes": null
    },
    {
      "id": "2679578",
      "postDate": "03/03/2024 15:52:34",
      "content": "<p>Thank you for the response, I also tried downgrading tf to 2.13.0 while using newer env, but all the dependency errors that I faced were too much to deal with, I moved to a pytorch based notebook for experimentation, which takes nearly twice as much as time for each epoch when compared with tf (just beginning of the week and already running of resources),</p>",
      "rawMarkdown": "Thank you for the response, I also tried downgrading tf to 2.13.0 while using newer env, but all the dependency errors that I faced were too much to deal with, I moved to a pytorch based notebook for experimentation, which takes nearly twice as much as time for each epoch when compared with tf (just beginning of the week and already running of resources),",
      "votes": null
    },
    {
      "id": "2679593",
      "postDate": "03/03/2024 16:09:31",
      "content": "<p>You could fork this <a href=\"https://www.kaggle.com/code/nartaa/features-head-starter\" target=\"_blank\">notebook</a>, it uses 2.13.0 by default, so no need to downgrade.</p>",
      "rawMarkdown": "You could fork this [notebook](https://www.kaggle.com/code/nartaa/features-head-starter), it uses 2.13.0 by default, so no need to downgrade.",
      "votes": null
    },
    {
      "id": "2679624",
      "postDate": "03/03/2024 16:36:35",
      "content": "<p>Hi, I already did, it was a really good notebook for you to be able to fit all 4 models in a single notebook, but it can be over whelming when trying to experiment with individual models, so I am focusing on simple notebooks for all the testing and experimentations :')</p>",
      "rawMarkdown": "Hi, I already did, it was a really good notebook for you to be able to fit all 4 models in a single notebook, but it can be over whelming when trying to experiment with individual models, so I am focusing on simple notebooks for all the testing and experimentations :')",
      "votes": null
    },
    {
      "id": "2679720",
      "postDate": "03/03/2024 17:34:52",
      "content": "<p>I hear ya 👍</p>",
      "rawMarkdown": "I hear ya 👍",
      "votes": null
    },
    {
      "id": "2680011",
      "postDate": "03/03/2024 22:37:25",
      "content": "<p>I encountered a similar problem and found a solution. Initially, my code looked like this:</p>\n<pre><code>with strategy.scope():\n    model = build_model()\nmodel.fit(train_gen, =1, =valid_gen)\noof = model.predict(valid_gen, =1)\n</code></pre>\n<p>However, the shape of <code>oof</code> turned out to be (1, 53), which likely corresponds to (1, number_of_batches).</p>\n<p>To address this, I modified my code as follows:</p>\n<pre><code>with strategy.scope():\n     = build_model()\n.fit(train_gen, verbose=, validation_data=valid_gen)\n\n = build_model()  # Rebuild the \n.load_weights(weight_path)\n\noof = .predict(valid_gen, verbose=)\n</code></pre>\n<p>This adjustment successfully changed the shape of <code>oof</code> to (3346, 6), which matches the expected (data_length, number_of_outputs). It appears that strategy.scope() can impact the prediction phase.</p>",
      "rawMarkdown": "I encountered a similar problem and found a solution. Initially, my code looked like this:\n```\nwith strategy.scope():\n    model = build_model()\nmodel.fit(train_gen, verbose=1, validation_data=valid_gen)\noof = model.predict(valid_gen, verbose=1)\n```\nHowever, the shape of `oof` turned out to be (1, 53), which likely corresponds to (1, number_of_batches).\n\nTo address this, I modified my code as follows:\n```\nwith strategy.scope():\n    model = build_model()\nmodel.fit(train_gen, verbose=1, validation_data=valid_gen)\n\nmodel = build_model()  # Rebuild the model\nmodel.load_weights(weight_path)\n\noof = model.predict(valid_gen, verbose=1)\n```\nThis adjustment successfully changed the shape of `oof` to (3346, 6), which matches the expected (data_length, number_of_outputs). It appears that strategy.scope() can impact the prediction phase.",
      "votes": null
    },
    {
      "id": "2680494",
      "postDate": "03/04/2024 08:33:32",
      "content": "<p>Thanks for helping out, it fixed the issue :')</p>",
      "rawMarkdown": "Thanks for helping out, it fixed the issue :')",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2678585,
      "author_name": "nartaa",
      "author_url": "",
      "post_date": "03/02/2024 22:09:50",
      "content": "<p>I came across the same issue with TPU, I also tried to move to 2.15.0, so I could switch easily between Colab and Kaggle, but I got alot of headache.<br>\nI think there is a bug in 2.14 and 2.15 that doesn't register the GPUs properly, you see that error at the very beginning when you try to import tensorflow, that issue could be related, I am not sure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2679578,
          "author_name": "arunsensei",
          "author_url": "",
          "post_date": "03/03/2024 15:52:34",
          "content": "<p>Thank you for the response, I also tried downgrading tf to 2.13.0 while using newer env, but all the dependency errors that I faced were too much to deal with, I moved to a pytorch based notebook for experimentation, which takes nearly twice as much as time for each epoch when compared with tf (just beginning of the week and already running of resources),</p>",
          "votes": null,
          "replies": [
            {
              "id": 2679593,
              "author_name": "nartaa",
              "author_url": "",
              "post_date": "03/03/2024 16:09:31",
              "content": "<p>You could fork this <a href=\"https://www.kaggle.com/code/nartaa/features-head-starter\" target=\"_blank\">notebook</a>, it uses 2.13.0 by default, so no need to downgrade.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2679624,
                  "author_name": "arunsensei",
                  "author_url": "",
                  "post_date": "03/03/2024 16:36:35",
                  "content": "<p>Hi, I already did, it was a really good notebook for you to be able to fit all 4 models in a single notebook, but it can be over whelming when trying to experiment with individual models, so I am focusing on simple notebooks for all the testing and experimentations :')</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2679720,
                      "author_name": "nartaa",
                      "author_url": "",
                      "post_date": "03/03/2024 17:34:52",
                      "content": "<p>I hear ya 👍</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2680011,
      "author_name": "spicatom0914",
      "author_url": "",
      "post_date": "03/03/2024 22:37:25",
      "content": "<p>I encountered a similar problem and found a solution. Initially, my code looked like this:</p>\n<pre><code>with strategy.scope():\n    model = build_model()\nmodel.fit(train_gen, =1, =valid_gen)\noof = model.predict(valid_gen, =1)\n</code></pre>\n<p>However, the shape of <code>oof</code> turned out to be (1, 53), which likely corresponds to (1, number_of_batches).</p>\n<p>To address this, I modified my code as follows:</p>\n<pre><code>with strategy.scope():\n     = build_model()\n.fit(train_gen, verbose=, validation_data=valid_gen)\n\n = build_model()  # Rebuild the \n.load_weights(weight_path)\n\noof = .predict(valid_gen, verbose=)\n</code></pre>\n<p>This adjustment successfully changed the shape of <code>oof</code> to (3346, 6), which matches the expected (data_length, number_of_outputs). It appears that strategy.scope() can impact the prediction phase.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2680494,
          "author_name": "arunsensei",
          "author_url": "",
          "post_date": "03/04/2024 08:33:32",
          "content": "<p>Thanks for helping out, it fixed the issue :')</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2677566": "I am trying to use the latest environments but i am facing some weird issue, I have tried using a TPU which also had teh same error, the model I am experimenting with is the EfficientNet model from this [notebook](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43), the model has 6 output classes, but when predicting it generates wrong number of outputs\n\nfor i, batch in enumerate(train_gen):\n    print(f\"Batch {i+1}: Input shape: {batch[0].shape}, Target shape: {batch[1].shape}\")\n    break\n\nwith strategy.scope():\n    model = build_model()\n    model.fit( batch[0], batch[1],epochs = 1, verbose=1)\n    oof = model.predict(batch[0], verbose=1)\n    print(batch[0].shape,batch[1].shape)\n    print(model.summary())\n\n(44, 128, 256, 8) (44, 6) are the respectives of batch[0] and batch[1], \n\nand the oof.shape is \n(2,)\nI nearly wasted 10 hours of GPU previous week trying to saving my notebook versions while experimenting, just to know it was all wasted, I encountered the same is occuring in TPU and when I am passing the datagenerator itself, it runs fine the entire training process and its throwing error in the last split\n\nI would appreciate your help in fixing this\nThank you.",
    "2678585": "I came across the same issue with TPU, I also tried to move to 2.15.0, so I could switch easily between Colab and Kaggle, but I got alot of headache.\nI think there is a bug in 2.14 and 2.15 that doesn't register the GPUs properly, you see that error at the very beginning when you try to import tensorflow, that issue could be related, I am not sure.",
    "2679578": "Thank you for the response, I also tried downgrading tf to 2.13.0 while using newer env, but all the dependency errors that I faced were too much to deal with, I moved to a pytorch based notebook for experimentation, which takes nearly twice as much as time for each epoch when compared with tf (just beginning of the week and already running of resources),",
    "2679593": "You could fork this [notebook](https://www.kaggle.com/code/nartaa/features-head-starter), it uses 2.13.0 by default, so no need to downgrade.",
    "2679624": "Hi, I already did, it was a really good notebook for you to be able to fit all 4 models in a single notebook, but it can be over whelming when trying to experiment with individual models, so I am focusing on simple notebooks for all the testing and experimentations :')",
    "2679720": "I hear ya 👍",
    "2680011": "I encountered a similar problem and found a solution. Initially, my code looked like this:\n```\nwith strategy.scope():\n    model = build_model()\nmodel.fit(train_gen, verbose=1, validation_data=valid_gen)\noof = model.predict(valid_gen, verbose=1)\n```\nHowever, the shape of `oof` turned out to be (1, 53), which likely corresponds to (1, number_of_batches).\n\nTo address this, I modified my code as follows:\n```\nwith strategy.scope():\n    model = build_model()\nmodel.fit(train_gen, verbose=1, validation_data=valid_gen)\n\nmodel = build_model()  # Rebuild the model\nmodel.load_weights(weight_path)\n\noof = model.predict(valid_gen, verbose=1)\n```\nThis adjustment successfully changed the shape of `oof` to (3346, 6), which matches the expected (data_length, number_of_outputs). It appears that strategy.scope() can impact the prediction phase.",
    "2680494": "Thanks for helping out, it fixed the issue :')"
  },
  "source": "meta"
}