{
  "id": 396546,
  "title": "Notebook timeout error with pretraining",
  "url": "/competitions/birdclef-2023/discussion/396546",
  "author_name": "0x4RY4N",
  "post_date": "2023-03-22T03:24:28.680000",
  "votes": 15,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I tried submitting a notebook with an effnet b1 model pretrained on the 2021-22 dataset, but that notebook gave me a timeout error. In contrast, a notebook with no pretraining(just the imagenet weights) is getting scored fine. The difference is just in the training notebook; the inference notebook is essentially the same but with different weights.<br>\nBoth these notebooks had the same model arch, the same model size, and the same running time, so why am I getting this error?</p>",
  "messages": [
    {
      "id": 2191524,
      "postDate": "2023-03-22T03:24:28.680Z",
      "content": "<p>I tried submitting a notebook with an effnet b1 model pretrained on the 2021-22 dataset, but that notebook gave me a timeout error. In contrast, a notebook with no pretraining(just the imagenet weights) is getting scored fine. The difference is just in the training notebook; the inference notebook is essentially the same but with different weights.<br>\nBoth these notebooks had the same model arch, the same model size, and the same running time, so why am I getting this error?</p>",
      "rawMarkdown": "I tried submitting a notebook with an effnet b1 model pretrained on the 2021-22 dataset, but that notebook gave me a timeout error. In contrast, a notebook with no pretraining(just the imagenet weights) is getting scored fine. The difference is just in the training notebook; the inference notebook is essentially the same but with different weights.\nBoth these notebooks had the same model arch, the same model size, and the same running time, so why am I getting this error?",
      "votes": 15
    },
    {
      "id": 2196527,
      "postDate": "2023-03-25T13:22:30.927Z",
      "content": "<p>I debugged my inference notebook and found out that this is due to environment changes. Running the notebook pinned to the original environment, I get an inference time of ~20 seconds and an estimated submission time of 1 hour 5 minutes. The original pinned environment uses an AMD EPYC 7B12 CPU.</p>\n<p>When \"Always using the latest environment\",  I get an inference time of ~40 seconds and an estimated submission time of 2 hour 30 minutes. The latest environment uses an Intel(R) Xeon(R) CPU @ 2.20GHz which is slower than the AMD one.</p>\n<p>The unfortunate thing now is that the latest environment seems to always be used even when not selected.</p>\n<p><a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> can you please looked into this?</p>",
      "rawMarkdown": "I debugged my inference notebook and found out that this is due to environment changes. Running the notebook pinned to the original environment, I get an inference time of ~20 seconds and an estimated submission time of 1 hour 5 minutes. The original pinned environment uses an AMD EPYC 7B12 CPU.\n\nWhen \"Always using the latest environment\",  I get an inference time of ~40 seconds and an estimated submission time of 2 hour 30 minutes. The latest environment uses an Intel(R) Xeon(R) CPU @ 2.20GHz which is slower than the AMD one.\n\nThe unfortunate thing now is that the latest environment seems to always be used even when not selected.\n\n@maggiemd @stefankahl can you please looked into this?",
      "votes": 7,
      "replies": [
        {
          "id": 2196821,
          "postDate": "2023-03-25T17:24:05.737Z",
          "content": "<p>Thank you for finding this out! I have been debugging this week, with a lot of frustration and no luck. </p>\n<p>However, which CPU it uses was very random for me. Switching between \"Always using the latest environment\" and \"original pinned environment\" did not switch the CPU. I kept restarting the CPU kernel and when I got an AMD CPU (a lot of restarts) the difference between in interference time between the intel and AMD CPU was even larger for me. With the intel CPU it takes 2:27 min (8 hours for submission) and with the AMD CPU it took only 10 seconds! (33 min for submission). Which is a bit shocking.</p>\n<p>Are all the submissions run on the same hardware? Otherwise this could make it a bit unfair which CPU will be used for each submission.</p>",
          "rawMarkdown": "Thank you for finding this out! I have been debugging this week, with a lot of frustration and no luck. \n\nHowever, which CPU it uses was very random for me. Switching between \"Always using the latest environment\" and \"original pinned environment\" did not switch the CPU. I kept restarting the CPU kernel and when I got an AMD CPU (a lot of restarts) the difference between in interference time between the intel and AMD CPU was even larger for me. With the intel CPU it takes 2:27 min (8 hours for submission) and with the AMD CPU it took only 10 seconds! (33 min for submission). Which is a bit shocking.\n\nAre all the submissions run on the same hardware? Otherwise this could make it a bit unfair which CPU will be used for each submission.\n",
          "votes": 1
        },
        {
          "id": 2196831,
          "postDate": "2023-03-25T17:35:22.303Z",
          "content": "<p>Huge thanks for finding this; we'll take a look with the Kaggle folks.</p>",
          "rawMarkdown": "Huge thanks for finding this; we'll take a look with the Kaggle folks.",
          "votes": 4,
          "replies": [
            {
              "id": 2199593,
              "postDate": "2023-03-27T20:43:12.057Z",
              "content": "<p><a href=\"https://www.kaggle.com/menno1111\" target=\"_blank\">@menno1111</a> \"Latest environment\" refers to the \"Docker image\" that is used, not the hardware.<br>\nThe Docker image is basically just the collection of package versions installed and released together on Kaggle.</p>\n<p>The hardware is where you're seeing AMD vs. Intel CPU differences.<br>\nOur CPU machine hardware for regular sessions is variable to support the volume of sessions, however the hardware used for running your submissions (sync, bulk reruns etc.) is not variable, and is set to Intel Skylake.</p>",
              "rawMarkdown": "@menno1111 \"Latest environment\" refers to the \"Docker image\" that is used, not the hardware.\nThe Docker image is basically just the collection of package versions installed and released together on Kaggle.\n\nThe hardware is where you're seeing AMD vs. Intel CPU differences.\nOur CPU machine hardware for regular sessions is variable to support the volume of sessions, however the hardware used for running your submissions (sync, bulk reruns etc.) is not variable, and is set to Intel Skylake.",
              "votes": 8
            },
            {
              "id": 2200743,
              "postDate": "2023-03-28T19:01:49.940Z",
              "content": "<p>Thank you very much for your response! <br>\nIs the speed of the Intel Skylake CPU comparable to the Intel CPU we can use in our Kaggle environment? Or is this something that can not be disclosed.</p>",
              "rawMarkdown": "Thank you very much for your response! \nIs the speed of the Intel Skylake CPU comparable to the Intel CPU we can use in our Kaggle environment? Or is this something that can not be disclosed."
            },
            {
              "id": 2200759,
              "postDate": "2023-03-28T19:18:38.650Z",
              "content": "<p><a href=\"https://www.kaggle.com/menno1111\" target=\"_blank\">@menno1111</a> I believe the Intel CPU you'll get in CPU interactive sessions will usually be an Intel Broadwell, Skylake is a newer model so it should be at least as good on average.</p>",
              "rawMarkdown": "@menno1111 I believe the Intel CPU you'll get in CPU interactive sessions will usually be an Intel Broadwell, Skylake is a newer model so it should be at least as good on average.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2217414,
      "postDate": "2023-04-10T21:55:29.983Z",
      "content": "<p>I was having the same issue, but without any pre-training, so I think this is maybe not the root cause.  <br>\nMy inference time on the sample submission (ignoring the rest of the code execution) changed from 21 seconds (which submits fine) with one set of weights, to 80 seconds (which times out).   All of it done on the Kaggle platform, forked from <a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-training-w-cmap\" target=\"_blank\">Nischay's</a> Pytorch Lightning notebooks with the same environment and version of TIMM. </p>\n<p>My inference notebook is running OK for now.  I re-built the training data pipeline, removed some unnecessary data-type changes and redundant functions, and array transpositions, and made sure I had exactly the same process for training and inference.   Sorry can't be more specific, as I never actually managed to pin-point the root cause.</p>",
      "rawMarkdown": "I was having the same issue, but without any pre-training, so I think this is maybe not the root cause.  \nMy inference time on the sample submission (ignoring the rest of the code execution) changed from 21 seconds (which submits fine) with one set of weights, to 80 seconds (which times out).   All of it done on the Kaggle platform, forked from [Nischay's](https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-training-w-cmap) Pytorch Lightning notebooks with the same environment and version of TIMM. \n\nMy inference notebook is running OK for now.  I re-built the training data pipeline, removed some unnecessary data-type changes and redundant functions, and array transpositions, and made sure I had exactly the same process for training and inference.   Sorry can't be more specific, as I never actually managed to pin-point the root cause.",
      "votes": 1,
      "replies": [
        {
          "id": 2239998,
          "postDate": "2023-04-30T05:53:44.590Z",
          "content": "<p>after pretraining do you got any improvement as i can see no improvement in the score even my cv improves from 0.78 to 0.8</p>",
          "rawMarkdown": "after pretraining do you got any improvement as i can see no improvement in the score even my cv improves from 0.78 to 0.8"
        }
      ]
    },
    {
      "id": 2203735,
      "postDate": "2023-03-31T04:43:13.720Z",
      "content": "<p>Intel CPU [with pretraining]:  1:15min(2+ hours for submission ) Timed out<br>\nAMD [with pretraining]:  7sec (2+ hours of submission) Timed out<br>\nwith just imagenet weights and no pretraining on previous year data:  successful submission, no matter which cpu.<br>\nIs there any solution for it?</p>",
      "rawMarkdown": "Intel CPU [with pretraining]:  1:15min(2+ hours for submission ) Timed out\nAMD [with pretraining]:  7sec (2+ hours of submission) Timed out\nwith just imagenet weights and no pretraining on previous year data:  successful submission, no matter which cpu.\nIs there any solution for it?",
      "votes": 1,
      "replies": [
        {
          "id": 2204412,
          "postDate": "2023-03-31T14:41:23.053Z",
          "content": "<p>Exactly the same problem, but I have no idea how to solve it.</p>",
          "rawMarkdown": "Exactly the same problem, but I have no idea how to solve it."
        },
        {
          "id": 2205848,
          "postDate": "2023-04-02T02:10:19.467Z",
          "content": "<ul>\n<li><p>If possible, it might be helpful to run these models on a fixed machine that you have access to and thereby isolate whether the pretraining itself is causing models to run slower.</p></li>\n<li><p>For this kind of debugging often you can get away with training just one step, since you only care about model speed. This will let you iterate faster than waiting for a full pretraining run.</p></li>\n<li><p>Check whether there are benchmarking tools with per-op timings available in your library of choice. Benchmarking can help uncover what ops are costing the most time, and surface differences between models.</p></li>\n</ul>",
          "rawMarkdown": "* If possible, it might be helpful to run these models on a fixed machine that you have access to and thereby isolate whether the pretraining itself is causing models to run slower.\n\n* For this kind of debugging often you can get away with training just one step, since you only care about model speed. This will let you iterate faster than waiting for a full pretraining run.\n\n* Check whether there are benchmarking tools with per-op timings available in your library of choice. Benchmarking can help uncover what ops are costing the most time, and surface differences between models.",
          "replies": [
            {
              "id": 2205948,
              "postDate": "2023-04-02T05:54:22.500Z",
              "content": "<p>I have conducted some experiments recently, but still have not achieved any significant results. These experiments include: </p>\n<ul>\n<li>Directly using the pre-trained model commit without fine-tuning. </li>\n<li>Rewriting the inference code to exclude some potentially problematic packages. </li>\n<li>Running two models on local CPU and GPU respectively, but there was no significant difference in performance between pre-training and non-pre-training on local CPU, and the same was observed on GPU.<br>\nT<br>\nI will try to test the inference time of my model at each step on a notebook. Thank you for your reminder.</li>\n</ul>",
              "rawMarkdown": "I have conducted some experiments recently, but still have not achieved any significant results. These experiments include: \n- Directly using the pre-trained model commit without fine-tuning. \n- Rewriting the inference code to exclude some potentially problematic packages. \n- Running two models on local CPU and GPU respectively, but there was no significant difference in performance between pre-training and non-pre-training on local CPU, and the same was observed on GPU.\nT\nI will try to test the inference time of my model at each step on a notebook. Thank you for your reminder."
            },
            {
              "id": 2206288,
              "postDate": "2023-04-02T13:01:26.403Z",
              "content": "<p>i didn't understand if with AMD CPU it took 7-10 sec for an audio file(which is a test sample) then why it is taking more than 2 hrs for submission(which has approx 200 files) and that too only with pretrained models. I had also used around 200 train samples for prediction in interactive mode and it is working fine.</p>",
              "rawMarkdown": "i didn't understand if with AMD CPU it took 7-10 sec for an audio file(which is a test sample) then why it is taking more than 2 hrs for submission(which has approx 200 files) and that too only with pretrained models. I had also used around 200 train samples for prediction in interactive mode and it is working fine."
            },
            {
              "id": 2206307,
              "postDate": "2023-04-02T13:20:23.530Z",
              "content": "<p>Kaggle always uses Intel cpu to run your submission not AMD. While in interactive sessions，it depends on your choice.</p>",
              "rawMarkdown": "Kaggle always uses Intel cpu to run your submission not AMD. While in interactive sessions，it depends on your choice."
            }
          ]
        }
      ]
    },
    {
      "id": 2201435,
      "postDate": "2023-03-29T10:00:10.173Z",
      "content": "<p>Is there a solution to this problem?<br>\nI am having the same problem.</p>",
      "rawMarkdown": "Is there a solution to this problem?\nI am having the same problem.",
      "votes": 1
    },
    {
      "id": 2192545,
      "postDate": "2023-03-22T18:04:58.037Z",
      "content": "<p>Hi, are you using the library timm? if yes, have you checked the version that you trained the model and the one you are using to load the checkpoint? Once I had a problem like this and I have fixed it by changing the version of timm in the inference notebook</p>",
      "rawMarkdown": "Hi, are you using the library timm? if yes, have you checked the version that you trained the model and the one you are using to load the checkpoint? Once I had a problem like this and I have fixed it by changing the version of timm in the inference notebook",
      "votes": 1,
      "replies": [
        {
          "id": 2192625,
          "postDate": "2023-03-22T19:28:38.963Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2192627,
          "postDate": "2023-03-22T19:29:41.233Z",
          "content": "<p>Yes, both the notebooks have the same version of timm.</p>",
          "rawMarkdown": "Yes, both the notebooks have the same version of timm.",
          "votes": 1
        },
        {
          "id": 2205435,
          "postDate": "2023-04-01T14:48:08.090Z",
          "content": "<p>It's not a timm problem, i even tried torchvision.models instead of timm, and getting out of time. Whenever i use pretrained model on external data i got this problem</p>",
          "rawMarkdown": "It's not a timm problem, i even tried torchvision.models instead of timm, and getting out of time. Whenever i use pretrained model on external data i got this problem",
          "votes": 1
        }
      ]
    },
    {
      "id": 2234805,
      "postDate": "2023-04-25T13:58:10.983Z",
      "content": "<p>Let me add myself to the list of folks, whose submission is experiencing timeout. (3 tries last night, no luck :-) No reason for it as far as I can tell, even without GPU, it runs fairly quickly. Access to error logs will be super helpful.</p>",
      "rawMarkdown": "Let me add myself to the list of folks, whose submission is experiencing timeout. (3 tries last night, no luck :-) No reason for it as far as I can tell, even without GPU, it runs fairly quickly. Access to error logs will be super helpful."
    },
    {
      "id": 2212296,
      "postDate": "2023-04-06T17:13:47.540Z",
      "content": "<p>I am having the same problem, is there a solution to this problem?</p>",
      "rawMarkdown": "I am having the same problem, is there a solution to this problem?"
    },
    {
      "id": 2192158,
      "postDate": "2023-03-22T13:18:07.160Z",
      "content": "<p>I have the same problem as you! I am very confused why this is happening.</p>",
      "rawMarkdown": "I have the same problem as you! I am very confused why this is happening."
    },
    {
      "id": 2191773,
      "postDate": "2023-03-22T07:44:07.753Z",
      "content": "<p>Same problem here</p>",
      "rawMarkdown": "Same problem here",
      "replies": [
        {
          "id": 2205080,
          "postDate": "2023-04-01T08:22:35.803Z",
          "content": "<p>Your current ranking is quite high. Did you utilize a pretrained model? If yes, may I ask how you addressed this issue?</p>",
          "rawMarkdown": "Your current ranking is quite high. Did you utilize a pretrained model? If yes, may I ask how you addressed this issue?",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2196527,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2023-03-25T13:22:30.927000",
      "content": "<p>I debugged my inference notebook and found out that this is due to environment changes. Running the notebook pinned to the original environment, I get an inference time of ~20 seconds and an estimated submission time of 1 hour 5 minutes. The original pinned environment uses an AMD EPYC 7B12 CPU.</p>\n<p>When \"Always using the latest environment\",  I get an inference time of ~40 seconds and an estimated submission time of 2 hour 30 minutes. The latest environment uses an Intel(R) Xeon(R) CPU @ 2.20GHz which is slower than the AMD one.</p>\n<p>The unfortunate thing now is that the latest environment seems to always be used even when not selected.</p>\n<p><a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> can you please looked into this?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2196821,
          "author_name": "menno",
          "author_url": "",
          "post_date": "2023-03-25T17:24:05.737000",
          "content": "<p>Thank you for finding this out! I have been debugging this week, with a lot of frustration and no luck. </p>\n<p>However, which CPU it uses was very random for me. Switching between \"Always using the latest environment\" and \"original pinned environment\" did not switch the CPU. I kept restarting the CPU kernel and when I got an AMD CPU (a lot of restarts) the difference between in interference time between the intel and AMD CPU was even larger for me. With the intel CPU it takes 2:27 min (8 hours for submission) and with the AMD CPU it took only 10 seconds! (33 min for submission). Which is a bit shocking.</p>\n<p>Are all the submissions run on the same hardware? Otherwise this could make it a bit unfair which CPU will be used for each submission.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2196831,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2023-03-25T17:35:22.303000",
          "content": "<p>Huge thanks for finding this; we'll take a look with the Kaggle folks.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2199593,
              "author_name": "Dustin",
              "author_url": "",
              "post_date": "2023-03-27T20:43:12.057000",
              "content": "<p><a href=\"https://www.kaggle.com/menno1111\" target=\"_blank\">@menno1111</a> \"Latest environment\" refers to the \"Docker image\" that is used, not the hardware.<br>\nThe Docker image is basically just the collection of package versions installed and released together on Kaggle.</p>\n<p>The hardware is where you're seeing AMD vs. Intel CPU differences.<br>\nOur CPU machine hardware for regular sessions is variable to support the volume of sessions, however the hardware used for running your submissions (sync, bulk reruns etc.) is not variable, and is set to Intel Skylake.</p>",
              "votes": 8,
              "replies": []
            },
            {
              "id": 2200743,
              "author_name": "menno",
              "author_url": "",
              "post_date": "2023-03-28T19:01:49.940000",
              "content": "<p>Thank you very much for your response! <br>\nIs the speed of the Intel Skylake CPU comparable to the Intel CPU we can use in our Kaggle environment? Or is this something that can not be disclosed.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2200759,
              "author_name": "Dustin",
              "author_url": "",
              "post_date": "2023-03-28T19:18:38.650000",
              "content": "<p><a href=\"https://www.kaggle.com/menno1111\" target=\"_blank\">@menno1111</a> I believe the Intel CPU you'll get in CPU interactive sessions will usually be an Intel Broadwell, Skylake is a newer model so it should be at least as good on average.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2217414,
      "author_name": "Olly Powell",
      "author_url": "",
      "post_date": "2023-04-10T21:55:29.983000",
      "content": "<p>I was having the same issue, but without any pre-training, so I think this is maybe not the root cause.  <br>\nMy inference time on the sample submission (ignoring the rest of the code execution) changed from 21 seconds (which submits fine) with one set of weights, to 80 seconds (which times out).   All of it done on the Kaggle platform, forked from <a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-training-w-cmap\" target=\"_blank\">Nischay's</a> Pytorch Lightning notebooks with the same environment and version of TIMM. </p>\n<p>My inference notebook is running OK for now.  I re-built the training data pipeline, removed some unnecessary data-type changes and redundant functions, and array transpositions, and made sure I had exactly the same process for training and inference.   Sorry can't be more specific, as I never actually managed to pin-point the root cause.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2239998,
          "author_name": "Himanshu Nayal",
          "author_url": "",
          "post_date": "2023-04-30T05:53:44.590000",
          "content": "<p>after pretraining do you got any improvement as i can see no improvement in the score even my cv improves from 0.78 to 0.8</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2203735,
      "author_name": "Himanshu Nayal",
      "author_url": "",
      "post_date": "2023-03-31T04:43:13.720000",
      "content": "<p>Intel CPU [with pretraining]:  1:15min(2+ hours for submission ) Timed out<br>\nAMD [with pretraining]:  7sec (2+ hours of submission) Timed out<br>\nwith just imagenet weights and no pretraining on previous year data:  successful submission, no matter which cpu.<br>\nIs there any solution for it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2204412,
          "author_name": "Orion",
          "author_url": "",
          "post_date": "2023-03-31T14:41:23.053000",
          "content": "<p>Exactly the same problem, but I have no idea how to solve it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2205848,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2023-04-02T02:10:19.467000",
          "content": "<ul>\n<li><p>If possible, it might be helpful to run these models on a fixed machine that you have access to and thereby isolate whether the pretraining itself is causing models to run slower.</p></li>\n<li><p>For this kind of debugging often you can get away with training just one step, since you only care about model speed. This will let you iterate faster than waiting for a full pretraining run.</p></li>\n<li><p>Check whether there are benchmarking tools with per-op timings available in your library of choice. Benchmarking can help uncover what ops are costing the most time, and surface differences between models.</p></li>\n</ul>",
          "votes": 0,
          "replies": [
            {
              "id": 2205948,
              "author_name": "Orion",
              "author_url": "",
              "post_date": "2023-04-02T05:54:22.500000",
              "content": "<p>I have conducted some experiments recently, but still have not achieved any significant results. These experiments include: </p>\n<ul>\n<li>Directly using the pre-trained model commit without fine-tuning. </li>\n<li>Rewriting the inference code to exclude some potentially problematic packages. </li>\n<li>Running two models on local CPU and GPU respectively, but there was no significant difference in performance between pre-training and non-pre-training on local CPU, and the same was observed on GPU.<br>\nT<br>\nI will try to test the inference time of my model at each step on a notebook. Thank you for your reminder.</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2206288,
              "author_name": "Himanshu Nayal",
              "author_url": "",
              "post_date": "2023-04-02T13:01:26.403000",
              "content": "<p>i didn't understand if with AMD CPU it took 7-10 sec for an audio file(which is a test sample) then why it is taking more than 2 hrs for submission(which has approx 200 files) and that too only with pretrained models. I had also used around 200 train samples for prediction in interactive mode and it is working fine.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2206307,
              "author_name": "Orion",
              "author_url": "",
              "post_date": "2023-04-02T13:20:23.530000",
              "content": "<p>Kaggle always uses Intel cpu to run your submission not AMD. While in interactive sessions，it depends on your choice.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2201435,
      "author_name": "torahirod",
      "author_url": "",
      "post_date": "2023-03-29T10:00:10.173000",
      "content": "<p>Is there a solution to this problem?<br>\nI am having the same problem.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2192545,
      "author_name": "Paulo Junqueira",
      "author_url": "",
      "post_date": "2023-03-22T18:04:58.037000",
      "content": "<p>Hi, are you using the library timm? if yes, have you checked the version that you trained the model and the one you are using to load the checkpoint? Once I had a problem like this and I have fixed it by changing the version of timm in the inference notebook</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2192625,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-22T19:28:38.963000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2192627,
          "author_name": "0x4RY4N",
          "author_url": "",
          "post_date": "2023-03-22T19:29:41.233000",
          "content": "<p>Yes, both the notebooks have the same version of timm.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2205435,
          "author_name": "Himanshu Nayal",
          "author_url": "",
          "post_date": "2023-04-01T14:48:08.090000",
          "content": "<p>It's not a timm problem, i even tried torchvision.models instead of timm, and getting out of time. Whenever i use pretrained model on external data i got this problem</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2234805,
      "author_name": "Ashish Kumar",
      "author_url": "",
      "post_date": "2023-04-25T13:58:10.983000",
      "content": "<p>Let me add myself to the list of folks, whose submission is experiencing timeout. (3 tries last night, no luck :-) No reason for it as far as I can tell, even without GPU, it runs fairly quickly. Access to error logs will be super helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2212296,
      "author_name": "Sanjay Acharjee",
      "author_url": "",
      "post_date": "2023-04-06T17:13:47.540000",
      "content": "<p>I am having the same problem, is there a solution to this problem?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2192158,
      "author_name": "menno",
      "author_url": "",
      "post_date": "2023-03-22T13:18:07.160000",
      "content": "<p>I have the same problem as you! I am very confused why this is happening.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2191773,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2023-03-22T07:44:07.753000",
      "content": "<p>Same problem here</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2205080,
          "author_name": "Orion",
          "author_url": "",
          "post_date": "2023-04-01T08:22:35.803000",
          "content": "<p>Your current ranking is quite high. Did you utilize a pretrained model? If yes, may I ask how you addressed this issue?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2191524": "I tried submitting a notebook with an effnet b1 model pretrained on the 2021-22 dataset, but that notebook gave me a timeout error. In contrast, a notebook with no pretraining(just the imagenet weights) is getting scored fine. The difference is just in the training notebook; the inference notebook is essentially the same but with different weights.\nBoth these notebooks had the same model arch, the same model size, and the same running time, so why am I getting this error?",
    "2196527": "I debugged my inference notebook and found out that this is due to environment changes. Running the notebook pinned to the original environment, I get an inference time of ~20 seconds and an estimated submission time of 1 hour 5 minutes. The original pinned environment uses an AMD EPYC 7B12 CPU.\n\nWhen \"Always using the latest environment\",  I get an inference time of ~40 seconds and an estimated submission time of 2 hour 30 minutes. The latest environment uses an Intel(R) Xeon(R) CPU @ 2.20GHz which is slower than the AMD one.\n\nThe unfortunate thing now is that the latest environment seems to always be used even when not selected.\n\n@maggiemd @stefankahl can you please looked into this?",
    "2217414": "I was having the same issue, but without any pre-training, so I think this is maybe not the root cause.  \nMy inference time on the sample submission (ignoring the rest of the code execution) changed from 21 seconds (which submits fine) with one set of weights, to 80 seconds (which times out).   All of it done on the Kaggle platform, forked from [Nischay's](https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-training-w-cmap) Pytorch Lightning notebooks with the same environment and version of TIMM. \n\nMy inference notebook is running OK for now.  I re-built the training data pipeline, removed some unnecessary data-type changes and redundant functions, and array transpositions, and made sure I had exactly the same process for training and inference.   Sorry can't be more specific, as I never actually managed to pin-point the root cause.",
    "2203735": "Intel CPU [with pretraining]:  1:15min(2+ hours for submission ) Timed out\nAMD [with pretraining]:  7sec (2+ hours of submission) Timed out\nwith just imagenet weights and no pretraining on previous year data:  successful submission, no matter which cpu.\nIs there any solution for it?",
    "2201435": "Is there a solution to this problem?\nI am having the same problem.",
    "2192545": "Hi, are you using the library timm? if yes, have you checked the version that you trained the model and the one you are using to load the checkpoint? Once I had a problem like this and I have fixed it by changing the version of timm in the inference notebook",
    "2234805": "Let me add myself to the list of folks, whose submission is experiencing timeout. (3 tries last night, no luck :-) No reason for it as far as I can tell, even without GPU, it runs fairly quickly. Access to error logs will be super helpful.",
    "2212296": "I am having the same problem, is there a solution to this problem?",
    "2192158": "I have the same problem as you! I am very confused why this is happening.",
    "2191773": "Same problem here"
  }
}