{
  "id": 498693,
  "title": "[Bug] Submission time out - Fixed",
  "url": "/competitions/birdclef-2024/discussion/498693",
  "author_name": "",
  "post_date": "2024-04-29T09:02:58.822056700Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p><strong>UPDATE: FIXED (we had <code>cat submission.csv</code>) at the end.</strong></p>\n<p>We have been stuck on submission timeouts recently which we cannot explain.</p>\n<p>We have already optimized our inference speed quite a bit, and when we test in our kaggle notebook on the <code>unlabelled soundscapes</code>, we observe inference speeds of around <code>1.5s / 4 minute file</code>. This should take ~27.5 minutes in total for 1100 files. When we submit this, we obtain a notebook timeout error and we see that our notebook is still running ~2h+…</p>\n<p>We are using Torch data-loader <code>num_workers=8</code> with <code>prefetch_factor=2</code> and ONNX for inference speedup, with an efficientNet_b0 model as a baseline. We don't observe any memory leaks when testing with the unlabelled_soundscapes or other weird behaviour that could lead to our submission timeouts..</p>\n<p>What we have tried:</p>\n<ul>\n<li>Predicting randomly with torch random instead of giving it to the model. (This might mean something is going on with the data loading on the submission notebook, that is different from when we run the Kaggle notebook)</li>\n<li>Ensured Kaggle notebook is on CPU (Accelerator = None)</li>\n</ul>\n<p>What can cause this:</p>\n<ul>\n<li>Considerably slower hardware on submission runtime compared to notebook runtime.</li>\n<li>Different environment on submission notebook? Maybe different package versions / python version that can cause this?</li>\n</ul>\n<p>Any help is appreciated since we could not get any notebook to score yet!</p>",
  "messages": [
    {
      "id": "2782324",
      "postDate": "04/29/2024 09:02:58",
      "content": "<p><strong>UPDATE: FIXED (we had <code>cat submission.csv</code>) at the end.</strong></p>\n<p>We have been stuck on submission timeouts recently which we cannot explain.</p>\n<p>We have already optimized our inference speed quite a bit, and when we test in our kaggle notebook on the <code>unlabelled soundscapes</code>, we observe inference speeds of around <code>1.5s / 4 minute file</code>. This should take ~27.5 minutes in total for 1100 files. When we submit this, we obtain a notebook timeout error and we see that our notebook is still running ~2h+…</p>\n<p>We are using Torch data-loader <code>num_workers=8</code> with <code>prefetch_factor=2</code> and ONNX for inference speedup, with an efficientNet_b0 model as a baseline. We don't observe any memory leaks when testing with the unlabelled_soundscapes or other weird behaviour that could lead to our submission timeouts..</p>\n<p>What we have tried:</p>\n<ul>\n<li>Predicting randomly with torch random instead of giving it to the model. (This might mean something is going on with the data loading on the submission notebook, that is different from when we run the Kaggle notebook)</li>\n<li>Ensured Kaggle notebook is on CPU (Accelerator = None)</li>\n</ul>\n<p>What can cause this:</p>\n<ul>\n<li>Considerably slower hardware on submission runtime compared to notebook runtime.</li>\n<li>Different environment on submission notebook? Maybe different package versions / python version that can cause this?</li>\n</ul>\n<p>Any help is appreciated since we could not get any notebook to score yet!</p>",
      "rawMarkdown": "**UPDATE: FIXED (we had `cat submission.csv`) at the end.**\n\n\nWe have been stuck on submission timeouts recently which we cannot explain.\n\nWe have already optimized our inference speed quite a bit, and when we test in our kaggle notebook on the `unlabelled soundscapes`, we observe inference speeds of around `1.5s / 4 minute file`. This should take ~27.5 minutes in total for 1100 files. When we submit this, we obtain a notebook timeout error and we see that our notebook is still running ~2h+...\n\nWe are using Torch data-loader `num_workers=8` with `prefetch_factor=2` and ONNX for inference speedup, with an efficientNet_b0 model as a baseline. We don't observe any memory leaks when testing with the unlabelled_soundscapes or other weird behaviour that could lead to our submission timeouts..\n\nWhat we have tried:\n\n- Predicting randomly with torch random instead of giving it to the model. (This might mean something is going on with the data loading on the submission notebook, that is different from when we run the Kaggle notebook)\n- Ensured Kaggle notebook is on CPU (Accelerator = None)\n\nWhat can cause this:\n\n- Considerably slower hardware on submission runtime compared to notebook runtime.\n- Different environment on submission notebook? Maybe different package versions / python version that can cause this?\n\nAny help is appreciated since we could not get any notebook to score yet!",
      "votes": null
    },
    {
      "id": "2782332",
      "postDate": "04/29/2024 09:09:19",
      "content": "<p>Here you can see our Kaggle inference times on the <code>unlabelled_soundscapes</code>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8913321%2Fb528bee0b88989ea700181891b81ed23%2Finference.png?generation=1714381758161200&amp;alt=media\"></p>",
      "rawMarkdown": "Here you can see our Kaggle inference times on the `unlabelled_soundscapes`:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8913321%2Fb528bee0b88989ea700181891b81ed23%2Finference.png?generation=1714381758161200&alt=media)",
      "votes": null
    },
    {
      "id": "2782413",
      "postDate": "04/29/2024 09:48:10",
      "content": "<blockquote>\n  <p>test in our local kaggle notebook</p>\n</blockquote>\n<p>local or on kaggle?  You should test on Kaggle.</p>",
      "rawMarkdown": "> test in our local kaggle notebook\n\nlocal or on kaggle?  You should test on Kaggle.",
      "votes": null
    },
    {
      "id": "2782425",
      "postDate": "04/29/2024 09:52:54",
      "content": "<p>That was worded a bit confusingly indeed, I meant on Kaggle (removed the local keyword). Testing was done on Kaggle with a screenshot of our inference times above. Locally we even see faster inference times.</p>",
      "rawMarkdown": "That was worded a bit confusingly indeed, I meant on Kaggle (removed the local keyword). Testing was done on Kaggle with a screenshot of our inference times above. Locally we even see faster inference times.",
      "votes": null
    },
    {
      "id": "2782479",
      "postDate": "04/29/2024 10:24:23",
      "content": "<p>OK, you ran your code with 1100 unlabelled soundscapes on Kaggle, and it took 27 minutes to run?</p>",
      "rawMarkdown": "OK, you ran your code with 1100 unlabelled soundscapes on Kaggle, and it took 27 minutes to run?",
      "votes": null
    },
    {
      "id": "2782641",
      "postDate": "04/29/2024 12:00:53",
      "content": "<p>We figured it out. We had <code>!cat submission.csv</code> at the end of our notebook haha, which was fine for 10 samples but not for 1100. That caused the timeout for submission :)</p>",
      "rawMarkdown": "We figured it out. We had `!cat submission.csv` at the end of our notebook haha, which was fine for 10 samples but not for 1100. That caused the timeout for submission :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2782332,
      "author_name": "hugodeheer",
      "author_url": "",
      "post_date": "04/29/2024 09:09:19",
      "content": "<p>Here you can see our Kaggle inference times on the <code>unlabelled_soundscapes</code>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8913321%2Fb528bee0b88989ea700181891b81ed23%2Finference.png?generation=1714381758161200&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2782413,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/29/2024 09:48:10",
      "content": "<blockquote>\n  <p>test in our local kaggle notebook</p>\n</blockquote>\n<p>local or on kaggle?  You should test on Kaggle.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2782425,
          "author_name": "hugodeheer",
          "author_url": "",
          "post_date": "04/29/2024 09:52:54",
          "content": "<p>That was worded a bit confusingly indeed, I meant on Kaggle (removed the local keyword). Testing was done on Kaggle with a screenshot of our inference times above. Locally we even see faster inference times.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2782479,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "04/29/2024 10:24:23",
              "content": "<p>OK, you ran your code with 1100 unlabelled soundscapes on Kaggle, and it took 27 minutes to run?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2782641,
                  "author_name": "hugodeheer",
                  "author_url": "",
                  "post_date": "04/29/2024 12:00:53",
                  "content": "<p>We figured it out. We had <code>!cat submission.csv</code> at the end of our notebook haha, which was fine for 10 samples but not for 1100. That caused the timeout for submission :)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2782324": "**UPDATE: FIXED (we had `cat submission.csv`) at the end.**\n\n\nWe have been stuck on submission timeouts recently which we cannot explain.\n\nWe have already optimized our inference speed quite a bit, and when we test in our kaggle notebook on the `unlabelled soundscapes`, we observe inference speeds of around `1.5s / 4 minute file`. This should take ~27.5 minutes in total for 1100 files. When we submit this, we obtain a notebook timeout error and we see that our notebook is still running ~2h+...\n\nWe are using Torch data-loader `num_workers=8` with `prefetch_factor=2` and ONNX for inference speedup, with an efficientNet_b0 model as a baseline. We don't observe any memory leaks when testing with the unlabelled_soundscapes or other weird behaviour that could lead to our submission timeouts..\n\nWhat we have tried:\n\n- Predicting randomly with torch random instead of giving it to the model. (This might mean something is going on with the data loading on the submission notebook, that is different from when we run the Kaggle notebook)\n- Ensured Kaggle notebook is on CPU (Accelerator = None)\n\nWhat can cause this:\n\n- Considerably slower hardware on submission runtime compared to notebook runtime.\n- Different environment on submission notebook? Maybe different package versions / python version that can cause this?\n\nAny help is appreciated since we could not get any notebook to score yet!",
    "2782332": "Here you can see our Kaggle inference times on the `unlabelled_soundscapes`:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8913321%2Fb528bee0b88989ea700181891b81ed23%2Finference.png?generation=1714381758161200&alt=media)",
    "2782413": "> test in our local kaggle notebook\n\nlocal or on kaggle?  You should test on Kaggle.",
    "2782425": "That was worded a bit confusingly indeed, I meant on Kaggle (removed the local keyword). Testing was done on Kaggle with a screenshot of our inference times above. Locally we even see faster inference times.",
    "2782479": "OK, you ran your code with 1100 unlabelled soundscapes on Kaggle, and it took 27 minutes to run?",
    "2782641": "We figured it out. We had `!cat submission.csv` at the end of our notebook haha, which was fine for 10 samples but not for 1100. That caused the timeout for submission :)"
  },
  "source": "meta"
}