{
  "id": 409331,
  "title": "Avoid Timeout by Rounding Your Model's Parameters",
  "url": "/competitions/birdclef-2023/discussion/409331",
  "author_name": "",
  "post_date": "2023-05-10T15:22:38.226950800Z",
  "votes": 16,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Today my submission suddenly timed out, even though it had been working well before.</p>\n<p>Reason: The parameters contain too many decimals or the numbers are just too small. This can cause problems for some CPUs, although it appears that AMD CPUs are less affected.</p>\n<p>Solution: Round your model's parameters, similar to <a href=\"https://github.com/Daniil-Osokin/lightweight-human-pose-estimation.pytorch/issues/32\" target=\"_blank\">this</a> issue. </p>\n<pre><code>params = list(model.parameters())\nprint('the length of parameters is', len(params))\nfor i in range(len(params)):\n    params[i].data = torch.round(params[i].data*10**4) / 10**4\n</code></pre>\n<p>In my case, before rounding the model's parameters, it took 60 seconds to process one audio clip, but after rounding, the processing time was reduced to approximately 4 seconds.</p>\n<p>Related discussions: <br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/396546\" target=\"_blank\">Notebook timeout error with pretraining</a> <a href=\"https://www.kaggle.com/aryankhatana\" target=\"_blank\">@aryankhatana</a><br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/407532\" target=\"_blank\">Is submission time estimation reliable?</a> <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a></p>\n<hr>\n<p>Update:<br>\nAnother related post: <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/401584\" target=\"_blank\">Fluctuating inference times, by factor of &gt; 20</a> <a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a></p>",
  "messages": [
    {
      "id": "2253995",
      "postDate": "05/10/2023 15:22:38",
      "content": "<p>Today my submission suddenly timed out, even though it had been working well before.</p>\n<p>Reason: The parameters contain too many decimals or the numbers are just too small. This can cause problems for some CPUs, although it appears that AMD CPUs are less affected.</p>\n<p>Solution: Round your model's parameters, similar to <a href=\"https://github.com/Daniil-Osokin/lightweight-human-pose-estimation.pytorch/issues/32\" target=\"_blank\">this</a> issue. </p>\n<pre><code>params = list(model.parameters())\nprint('the length of parameters is', len(params))\nfor i in range(len(params)):\n    params[i].data = torch.round(params[i].data*10**4) / 10**4\n</code></pre>\n<p>In my case, before rounding the model's parameters, it took 60 seconds to process one audio clip, but after rounding, the processing time was reduced to approximately 4 seconds.</p>\n<p>Related discussions: <br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/396546\" target=\"_blank\">Notebook timeout error with pretraining</a> <a href=\"https://www.kaggle.com/aryankhatana\" target=\"_blank\">@aryankhatana</a><br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/407532\" target=\"_blank\">Is submission time estimation reliable?</a> <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a></p>\n<hr>\n<p>Update:<br>\nAnother related post: <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/401584\" target=\"_blank\">Fluctuating inference times, by factor of &gt; 20</a> <a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a></p>",
      "rawMarkdown": "Today my submission suddenly timed out, even though it had been working well before.\n\nReason: The parameters contain too many decimals or the numbers are just too small. This can cause problems for some CPUs, although it appears that AMD CPUs are less affected.\n\nSolution: Round your model's parameters, similar to [this](https://github.com/Daniil-Osokin/lightweight-human-pose-estimation.pytorch/issues/32) issue. \n\n```\nparams = list(model.parameters())\nprint('the length of parameters is', len(params))\nfor i in range(len(params)):\n    params[i].data = torch.round(params[i].data*10**4) / 10**4\n```\nIn my case, before rounding the model's parameters, it took 60 seconds to process one audio clip, but after rounding, the processing time was reduced to approximately 4 seconds.\n\nRelated discussions: \n[Notebook timeout error with pretraining](https://www.kaggle.com/competitions/birdclef-2023/discussion/396546) @aryankhatana\n[Is submission time estimation reliable?](https://www.kaggle.com/competitions/birdclef-2023/discussion/407532) @leonshangguan\n\n---\nUpdate:\nAnother related post: [Fluctuating inference times, by factor of > 20](https://www.kaggle.com/competitions/birdclef-2023/discussion/401584) @ollypowell",
      "votes": null
    },
    {
      "id": "2254079",
      "postDate": "05/10/2023 16:37:06",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chrisqiu\" target=\"_blank\">@chrisqiu</a> ,<br>\nDoes this affect the model's performance in any way?<br>\nThanks.</p>",
      "rawMarkdown": "Hi @chrisqiu ,\nDoes this affect the model's performance in any way?\nThanks.",
      "votes": null
    },
    {
      "id": "2254421",
      "postDate": "05/11/2023 00:52:32",
      "content": "<p>Another related discussion is over here. <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/401584\" target=\"_blank\">Fluctuating inference times, by factor of &gt; 20</a><br>\nThe main issue here is that denormal values seem to have different behavior on different CPUs. Rounding parameters works but rounding to 1e-4 might have a big impact on the performance depending on the model.<br>\n<a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a> found out retraining with  <code>torch.set_flush_denormal(True)</code> works. Another fast and safe way I think should work is to discard denormals with <code>torch.set_flush_denormal(True)</code> combined with discarding extremely small parameters maybe &lt;1e-18 depending on your model. </p>",
      "rawMarkdown": "Another related discussion is over here. [Fluctuating inference times, by factor of > 20](https://www.kaggle.com/competitions/birdclef-2023/discussion/401584)\nThe main issue here is that denormal values seem to have different behavior on different CPUs. Rounding parameters works but rounding to 1e-4 might have a big impact on the performance depending on the model.\n@ollypowell found out retraining with  `torch.set_flush_denormal(True)` works. Another fast and safe way I think should work is to discard denormals with `torch.set_flush_denormal(True)` combined with discarding extremely small parameters maybe <1e-18 depending on your model.",
      "votes": null
    },
    {
      "id": "2254610",
      "postDate": "05/11/2023 05:37:43",
      "content": "<p>My LB score hasn't decreased, but it's difficult to determine the exact performance as the results are only given to two significant digits. <br>\nOne alternative method to test the method is by rounding the parameters during evaluation and observing the results of cross-validation. However, I have not yet attempted this approach.</p>",
      "rawMarkdown": "My LB score hasn't decreased, but it's difficult to determine the exact performance as the results are only given to two significant digits. \nOne alternative method to test the method is by rounding the parameters during evaluation and observing the results of cross-validation. However, I have not yet attempted this approach.",
      "votes": null
    },
    {
      "id": "2254614",
      "postDate": "05/11/2023 05:41:04",
      "content": "<p>Thank you for bringing this up! <br>\nI attempted to use <code>torch.set_flush_denormal(True)</code> during inference, but unfortunately, it did not work as expected. <br>\nHowever, I have not yet retrained the network with <code>torch.set_flush_denormal(True)</code> enabled, so that may be worth exploring in the future.</p>",
      "rawMarkdown": "Thank you for bringing this up! \nI attempted to use `torch.set_flush_denormal(True)` during inference, but unfortunately, it did not work as expected. \nHowever, I have not yet retrained the network with `torch.set_flush_denormal(True)` enabled, so that may be worth exploring in the future.",
      "votes": null
    },
    {
      "id": "2254658",
      "postDate": "05/11/2023 06:33:58",
      "content": "<p>Yeah <code>torch.set_flush_denormal(True)</code> doesn't really clean up all the denormal values like I mentioned in the discussion, it must be combined with manually flushing the small parameters that are almost denormal which is quite frustrating.</p>",
      "rawMarkdown": "Yeah `torch.set_flush_denormal(True)` doesn't really clean up all the denormal values like I mentioned in the discussion, it must be combined with manually flushing the small parameters that are almost denormal which is quite frustrating.",
      "votes": null
    },
    {
      "id": "2258687",
      "postDate": "05/14/2023 12:09:10",
      "content": "<p>Maybe you can check in your submissions and sort by score and see if the new submission is below previous submission or not.</p>",
      "rawMarkdown": "Maybe you can check in your submissions and sort by score and see if the new submission is below previous submission or not.",
      "votes": null
    },
    {
      "id": "2259055",
      "postDate": "05/14/2023 16:52:42",
      "content": "<p>This is a good idea 🤔 Although the total submissions per day are limited.</p>",
      "rawMarkdown": "This is a good idea 🤔 Although the total submissions per day are limited.",
      "votes": null
    },
    {
      "id": "2260771",
      "postDate": "05/15/2023 21:13:54",
      "content": "<p>Hey Everyone !<br>\nDid anyone face Notebook timeout error after 10 mins of run post submission ? It's written 120 mins in the documentation. I wonder, Why my submission timeouts after 10 mins ? I even incorporated the suggestions posted out here in my second submission but the result remains same. I wonder what's the catch ?</p>\n<p>Refer the attached image.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9201105%2Feed91ce81904b32a7c6f2b8a406c4a52%2FScreenshot%20from%202023-05-16%2002-40-54.png?generation=1684185216479448&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hey Everyone !\nDid anyone face Notebook timeout error after 10 mins of run post submission ? It's written 120 mins in the documentation. I wonder, Why my submission timeouts after 10 mins ? I even incorporated the suggestions posted out here in my second submission but the result remains same. I wonder what's the catch ?\n\nRefer the attached image.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9201105%2Feed91ce81904b32a7c6f2b8a406c4a52%2FScreenshot%20from%202023-05-16%2002-40-54.png?generation=1684185216479448&alt=media)",
      "votes": null
    },
    {
      "id": "2260837",
      "postDate": "05/15/2023 23:28:16",
      "content": "<p>60 secs to 4 secs per audio?? that's crazy! 💥💥</p>",
      "rawMarkdown": "60 secs to 4 secs per audio?? that's crazy! 💥💥",
      "votes": null
    },
    {
      "id": "2261244",
      "postDate": "05/16/2023 07:42:54",
      "content": "<p>Indeed! I think it is quite unusual for a single model to take 60 seconds to run on a CPU.</p>",
      "rawMarkdown": "Indeed! I think it is quite unusual for a single model to take 60 seconds to run on a CPU.",
      "votes": null
    },
    {
      "id": "2261247",
      "postDate": "05/16/2023 07:43:29",
      "content": "<p>This is strange 🤔 Maybe it is because of an error instead of timeout?</p>",
      "rawMarkdown": "This is strange 🤔 Maybe it is because of an error instead of timeout?",
      "votes": null
    },
    {
      "id": "2261294",
      "postDate": "05/16/2023 08:24:04",
      "content": "<p>I think it was a Kaggle bug</p>\n<p>I had strange Timeout at the same time you have (if you have made a post right after getting it)</p>",
      "rawMarkdown": "I think it was a Kaggle bug\n\nI had strange Timeout at the same time you have (if you have made a post right after getting it)",
      "votes": null
    },
    {
      "id": "2261716",
      "postDate": "05/16/2023 14:10:13",
      "content": "<p>Is it not necessary to load the params into the model? If necessary, how to load?</p>",
      "rawMarkdown": "Is it not necessary to load the params into the model? If necessary, how to load?",
      "votes": null
    },
    {
      "id": "2261751",
      "postDate": "05/16/2023 14:22:27",
      "content": "<p>Yeah <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a>, I posted rightaway after facing the bug ! It seems like t'was a kaggle bug ! It's atleast running for more than 10 mins, now ! Hope that the glitch's fixed now ! </p>",
      "rawMarkdown": "Yeah @vladimirsydor, I posted rightaway after facing the bug ! It seems like t'was a kaggle bug ! It's atleast running for more than 10 mins, now ! Hope that the glitch's fixed now !",
      "votes": null
    },
    {
      "id": "2261755",
      "postDate": "05/16/2023 14:23:14",
      "content": "<p>As mentioned, It seems like a kaggle bug that was persistent across submissions yesterday.</p>",
      "rawMarkdown": "As mentioned, It seems like a kaggle bug that was persistent across submissions yesterday.",
      "votes": null
    },
    {
      "id": "2269353",
      "postDate": "05/22/2023 12:12:48",
      "content": "<p>Yes, how do we get the list back into the state_dict format to reload the rounded parameters? Any help appreciated!</p>",
      "rawMarkdown": "Yes, how do we get the list back into the state_dict format to reload the rounded parameters? Any help appreciated!",
      "votes": null
    },
    {
      "id": "2269372",
      "postDate": "05/22/2023 12:27:24",
      "content": "<p>It is not necessary. Just mutate the object.</p>",
      "rawMarkdown": "It is not necessary. Just mutate the object.",
      "votes": null
    },
    {
      "id": "2269375",
      "postDate": "05/22/2023 12:41:14",
      "content": "<p>Hey Chris,<br>\nappreciate your fast answer! Nonetheless, I am still pretty new to all this - can you give me some additional information on what exactly you mean by mutating the object? <br>\nBest,<br>\nJan</p>",
      "rawMarkdown": "Hey Chris,\nappreciate your fast answer! Nonetheless, I am still pretty new to all this - can you give me some additional information on what exactly you mean by mutating the object? \nBest,\nJan",
      "votes": null
    },
    {
      "id": "2269380",
      "postDate": "05/22/2023 12:44:45",
      "content": "<p>Is this what you mean by mutation in this case? </p>\n<p><code>for param in model.parameters():\n     param.data = torch.round(param.data*10**10) / 10**10</code></p>",
      "rawMarkdown": "Is this what you mean by mutation in this case? \n\n`for param in model.parameters():\n     param.data = torch.round(param.data*10**10) / 10**10`",
      "votes": null
    },
    {
      "id": "2269394",
      "postDate": "05/22/2023 12:53:23",
      "content": "<p>Ok. <code>list(model.parameters())</code> returns a list of params. <br>\nEach param is a <code>torch.nn.parameter.Parameter</code> object.<br>\nYou can modify the object's <code>data</code> property directly.</p>",
      "rawMarkdown": "Ok. `list(model.parameters())` returns a list of params. \nEach param is a `torch.nn.parameter.Parameter` object.\nYou can modify the object's `data` property directly.",
      "votes": null
    },
    {
      "id": "2269452",
      "postDate": "05/22/2023 13:34:32",
      "content": "<p>Hey Chris, thank you!</p>",
      "rawMarkdown": "Hey Chris, thank you!",
      "votes": null
    },
    {
      "id": "2269712",
      "postDate": "05/22/2023 16:43:12",
      "content": "<p>I modify the parameters following this guide :<br>\n<a href=\"https://discuss.pytorch.org/t/overwrite-parameters-of-model-with-new-values/92952\" target=\"_blank\">https://discuss.pytorch.org/t/overwrite-parameters-of-model-with-new-values/92952</a><br>\nBut the rounding doesn't seem to affect my runtime <br>\nIt's becoming slower these past days </p>",
      "rawMarkdown": "I modify the parameters following this guide :\nhttps://discuss.pytorch.org/t/overwrite-parameters-of-model-with-new-values/92952\nBut the rounding doesn't seem to affect my runtime \nIt's becoming slower these past days",
      "votes": null
    },
    {
      "id": "2270084",
      "postDate": "05/23/2023 01:46:55",
      "content": "<p>I think rounding parameters might not necessarily speed up the inference. It can be used to solve unexpected time out. Like in my case, it is 'unusual' to take 60 seconds to inference one model. <br>\nBut depending on your situation, rounding parameters might not have significant impact.</p>",
      "rawMarkdown": "I think rounding parameters might not necessarily speed up the inference. It can be used to solve unexpected time out. Like in my case, it is 'unusual' to take 60 seconds to inference one model. \nBut depending on your situation, rounding parameters might not have significant impact.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2254079,
      "author_name": "shashwatraman",
      "author_url": "",
      "post_date": "05/10/2023 16:37:06",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chrisqiu\" target=\"_blank\">@chrisqiu</a> ,<br>\nDoes this affect the model's performance in any way?<br>\nThanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2254610,
          "author_name": "chrisqiu",
          "author_url": "",
          "post_date": "05/11/2023 05:37:43",
          "content": "<p>My LB score hasn't decreased, but it's difficult to determine the exact performance as the results are only given to two significant digits. <br>\nOne alternative method to test the method is by rounding the parameters during evaluation and observing the results of cross-validation. However, I have not yet attempted this approach.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2258687,
              "author_name": "salmanahmedtamu",
              "author_url": "",
              "post_date": "05/14/2023 12:09:10",
              "content": "<p>Maybe you can check in your submissions and sort by score and see if the new submission is below previous submission or not.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2259055,
                  "author_name": "chrisqiu",
                  "author_url": "",
                  "post_date": "05/14/2023 16:52:42",
                  "content": "<p>This is a good idea 🤔 Although the total submissions per day are limited.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2254421,
      "author_name": "lhanhsin",
      "author_url": "",
      "post_date": "05/11/2023 00:52:32",
      "content": "<p>Another related discussion is over here. <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/401584\" target=\"_blank\">Fluctuating inference times, by factor of &gt; 20</a><br>\nThe main issue here is that denormal values seem to have different behavior on different CPUs. Rounding parameters works but rounding to 1e-4 might have a big impact on the performance depending on the model.<br>\n<a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a> found out retraining with  <code>torch.set_flush_denormal(True)</code> works. Another fast and safe way I think should work is to discard denormals with <code>torch.set_flush_denormal(True)</code> combined with discarding extremely small parameters maybe &lt;1e-18 depending on your model. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2254614,
          "author_name": "chrisqiu",
          "author_url": "",
          "post_date": "05/11/2023 05:41:04",
          "content": "<p>Thank you for bringing this up! <br>\nI attempted to use <code>torch.set_flush_denormal(True)</code> during inference, but unfortunately, it did not work as expected. <br>\nHowever, I have not yet retrained the network with <code>torch.set_flush_denormal(True)</code> enabled, so that may be worth exploring in the future.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2254658,
              "author_name": "lhanhsin",
              "author_url": "",
              "post_date": "05/11/2023 06:33:58",
              "content": "<p>Yeah <code>torch.set_flush_denormal(True)</code> doesn't really clean up all the denormal values like I mentioned in the discussion, it must be combined with manually flushing the small parameters that are almost denormal which is quite frustrating.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2260771,
      "author_name": "suraj520",
      "author_url": "",
      "post_date": "05/15/2023 21:13:54",
      "content": "<p>Hey Everyone !<br>\nDid anyone face Notebook timeout error after 10 mins of run post submission ? It's written 120 mins in the documentation. I wonder, Why my submission timeouts after 10 mins ? I even incorporated the suggestions posted out here in my second submission but the result remains same. I wonder what's the catch ?</p>\n<p>Refer the attached image.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9201105%2Feed91ce81904b32a7c6f2b8a406c4a52%2FScreenshot%20from%202023-05-16%2002-40-54.png?generation=1684185216479448&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2261247,
          "author_name": "chrisqiu",
          "author_url": "",
          "post_date": "05/16/2023 07:43:29",
          "content": "<p>This is strange 🤔 Maybe it is because of an error instead of timeout?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2261755,
              "author_name": "suraj520",
              "author_url": "",
              "post_date": "05/16/2023 14:23:14",
              "content": "<p>As mentioned, It seems like a kaggle bug that was persistent across submissions yesterday.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2261294,
          "author_name": "vladimirsydor",
          "author_url": "",
          "post_date": "05/16/2023 08:24:04",
          "content": "<p>I think it was a Kaggle bug</p>\n<p>I had strange Timeout at the same time you have (if you have made a post right after getting it)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2261751,
              "author_name": "suraj520",
              "author_url": "",
              "post_date": "05/16/2023 14:22:27",
              "content": "<p>Yeah <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a>, I posted rightaway after facing the bug ! It seems like t'was a kaggle bug ! It's atleast running for more than 10 mins, now ! Hope that the glitch's fixed now ! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2260837,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "05/15/2023 23:28:16",
      "content": "<p>60 secs to 4 secs per audio?? that's crazy! 💥💥</p>",
      "votes": null,
      "replies": [
        {
          "id": 2261244,
          "author_name": "chrisqiu",
          "author_url": "",
          "post_date": "05/16/2023 07:42:54",
          "content": "<p>Indeed! I think it is quite unusual for a single model to take 60 seconds to run on a CPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2261716,
      "author_name": "shigemitsutomizawa",
      "author_url": "",
      "post_date": "05/16/2023 14:10:13",
      "content": "<p>Is it not necessary to load the params into the model? If necessary, how to load?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2269353,
          "author_name": "janbrederecke",
          "author_url": "",
          "post_date": "05/22/2023 12:12:48",
          "content": "<p>Yes, how do we get the list back into the state_dict format to reload the rounded parameters? Any help appreciated!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2269372,
              "author_name": "chrisqiu",
              "author_url": "",
              "post_date": "05/22/2023 12:27:24",
              "content": "<p>It is not necessary. Just mutate the object.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2269375,
                  "author_name": "janbrederecke",
                  "author_url": "",
                  "post_date": "05/22/2023 12:41:14",
                  "content": "<p>Hey Chris,<br>\nappreciate your fast answer! Nonetheless, I am still pretty new to all this - can you give me some additional information on what exactly you mean by mutating the object? <br>\nBest,<br>\nJan</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2269380,
                      "author_name": "janbrederecke",
                      "author_url": "",
                      "post_date": "05/22/2023 12:44:45",
                      "content": "<p>Is this what you mean by mutation in this case? </p>\n<p><code>for param in model.parameters():\n     param.data = torch.round(param.data*10**10) / 10**10</code></p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2269394,
                          "author_name": "chrisqiu",
                          "author_url": "",
                          "post_date": "05/22/2023 12:53:23",
                          "content": "<p>Ok. <code>list(model.parameters())</code> returns a list of params. <br>\nEach param is a <code>torch.nn.parameter.Parameter</code> object.<br>\nYou can modify the object's <code>data</code> property directly.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2269452,
                              "author_name": "janbrederecke",
                              "author_url": "",
                              "post_date": "05/22/2023 13:34:32",
                              "content": "<p>Hey Chris, thank you!</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2269712,
                                  "author_name": "nyleve",
                                  "author_url": "",
                                  "post_date": "05/22/2023 16:43:12",
                                  "content": "<p>I modify the parameters following this guide :<br>\n<a href=\"https://discuss.pytorch.org/t/overwrite-parameters-of-model-with-new-values/92952\" target=\"_blank\">https://discuss.pytorch.org/t/overwrite-parameters-of-model-with-new-values/92952</a><br>\nBut the rounding doesn't seem to affect my runtime <br>\nIt's becoming slower these past days </p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2270084,
                                      "author_name": "chrisqiu",
                                      "author_url": "",
                                      "post_date": "05/23/2023 01:46:55",
                                      "content": "<p>I think rounding parameters might not necessarily speed up the inference. It can be used to solve unexpected time out. Like in my case, it is 'unusual' to take 60 seconds to inference one model. <br>\nBut depending on your situation, rounding parameters might not have significant impact.</p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2253995": "Today my submission suddenly timed out, even though it had been working well before.\n\nReason: The parameters contain too many decimals or the numbers are just too small. This can cause problems for some CPUs, although it appears that AMD CPUs are less affected.\n\nSolution: Round your model's parameters, similar to [this](https://github.com/Daniil-Osokin/lightweight-human-pose-estimation.pytorch/issues/32) issue. \n\n```\nparams = list(model.parameters())\nprint('the length of parameters is', len(params))\nfor i in range(len(params)):\n    params[i].data = torch.round(params[i].data*10**4) / 10**4\n```\nIn my case, before rounding the model's parameters, it took 60 seconds to process one audio clip, but after rounding, the processing time was reduced to approximately 4 seconds.\n\nRelated discussions: \n[Notebook timeout error with pretraining](https://www.kaggle.com/competitions/birdclef-2023/discussion/396546) @aryankhatana\n[Is submission time estimation reliable?](https://www.kaggle.com/competitions/birdclef-2023/discussion/407532) @leonshangguan\n\n---\nUpdate:\nAnother related post: [Fluctuating inference times, by factor of > 20](https://www.kaggle.com/competitions/birdclef-2023/discussion/401584) @ollypowell",
    "2254079": "Hi @chrisqiu ,\nDoes this affect the model's performance in any way?\nThanks.",
    "2254421": "Another related discussion is over here. [Fluctuating inference times, by factor of > 20](https://www.kaggle.com/competitions/birdclef-2023/discussion/401584)\nThe main issue here is that denormal values seem to have different behavior on different CPUs. Rounding parameters works but rounding to 1e-4 might have a big impact on the performance depending on the model.\n@ollypowell found out retraining with  `torch.set_flush_denormal(True)` works. Another fast and safe way I think should work is to discard denormals with `torch.set_flush_denormal(True)` combined with discarding extremely small parameters maybe <1e-18 depending on your model.",
    "2254610": "My LB score hasn't decreased, but it's difficult to determine the exact performance as the results are only given to two significant digits. \nOne alternative method to test the method is by rounding the parameters during evaluation and observing the results of cross-validation. However, I have not yet attempted this approach.",
    "2254614": "Thank you for bringing this up! \nI attempted to use `torch.set_flush_denormal(True)` during inference, but unfortunately, it did not work as expected. \nHowever, I have not yet retrained the network with `torch.set_flush_denormal(True)` enabled, so that may be worth exploring in the future.",
    "2254658": "Yeah `torch.set_flush_denormal(True)` doesn't really clean up all the denormal values like I mentioned in the discussion, it must be combined with manually flushing the small parameters that are almost denormal which is quite frustrating.",
    "2258687": "Maybe you can check in your submissions and sort by score and see if the new submission is below previous submission or not.",
    "2259055": "This is a good idea 🤔 Although the total submissions per day are limited.",
    "2260771": "Hey Everyone !\nDid anyone face Notebook timeout error after 10 mins of run post submission ? It's written 120 mins in the documentation. I wonder, Why my submission timeouts after 10 mins ? I even incorporated the suggestions posted out here in my second submission but the result remains same. I wonder what's the catch ?\n\nRefer the attached image.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9201105%2Feed91ce81904b32a7c6f2b8a406c4a52%2FScreenshot%20from%202023-05-16%2002-40-54.png?generation=1684185216479448&alt=media)",
    "2260837": "60 secs to 4 secs per audio?? that's crazy! 💥💥",
    "2261244": "Indeed! I think it is quite unusual for a single model to take 60 seconds to run on a CPU.",
    "2261247": "This is strange 🤔 Maybe it is because of an error instead of timeout?",
    "2261294": "I think it was a Kaggle bug\n\nI had strange Timeout at the same time you have (if you have made a post right after getting it)",
    "2261716": "Is it not necessary to load the params into the model? If necessary, how to load?",
    "2261751": "Yeah @vladimirsydor, I posted rightaway after facing the bug ! It seems like t'was a kaggle bug ! It's atleast running for more than 10 mins, now ! Hope that the glitch's fixed now !",
    "2261755": "As mentioned, It seems like a kaggle bug that was persistent across submissions yesterday.",
    "2269353": "Yes, how do we get the list back into the state_dict format to reload the rounded parameters? Any help appreciated!",
    "2269372": "It is not necessary. Just mutate the object.",
    "2269375": "Hey Chris,\nappreciate your fast answer! Nonetheless, I am still pretty new to all this - can you give me some additional information on what exactly you mean by mutating the object? \nBest,\nJan",
    "2269380": "Is this what you mean by mutation in this case? \n\n`for param in model.parameters():\n     param.data = torch.round(param.data*10**10) / 10**10`",
    "2269394": "Ok. `list(model.parameters())` returns a list of params. \nEach param is a `torch.nn.parameter.Parameter` object.\nYou can modify the object's `data` property directly.",
    "2269452": "Hey Chris, thank you!",
    "2269712": "I modify the parameters following this guide :\nhttps://discuss.pytorch.org/t/overwrite-parameters-of-model-with-new-values/92952\nBut the rounding doesn't seem to affect my runtime \nIt's becoming slower these past days",
    "2270084": "I think rounding parameters might not necessarily speed up the inference. It can be used to solve unexpected time out. Like in my case, it is 'unusual' to take 60 seconds to inference one model. \nBut depending on your situation, rounding parameters might not have significant impact."
  },
  "source": "meta"
}