{
  "id": 401584,
  "title": "Fluctuating inference times, by factor of > 20",
  "url": "/competitions/birdclef-2023/discussion/401584",
  "author_name": "",
  "post_date": "2023-04-14T00:31:38.296169500Z",
  "votes": 2,
  "comment_count": 13,
  "views": 0,
  "content": "<p>There are a couple of other posts similar already, but the nature of this seems slightly different.  Or maybe we're all facing the same mysterious problem.</p>\n<p>I'm getting my submissions timing out because the evaluation time is varying enormously.  I'm timing the evaluation of the one 600 second clip in the test set, and not counting the rest of my code execution.  My observations so far are:</p>\n<ul>\n<li>I can now turn on and off the problem, simply by training for more or less epochs.  Not really a solution, but it does narrow down the problem a lot!</li>\n<li>Repeatable times with some training weights,  for example one set consistently takes 15 seconds with EfficientNetV2s.</li>\n<li>Inconsistent with other training weights, produced by the same notebook.  For example I've run a notebook with the inference step taking 9 seconds, but it failed submission.  I then forked it, made no changes at all, and now it takes 380 seconds every time.</li>\n<li>No changes, or small changes to the inference notebook in between problems. Consistent versions of Python environment, with both training and inference.  P100 GPU used for training, CPU only for inference of course.</li>\n<li>Both the training &amp; infer forked originally from <a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference\" target=\"_blank\">Nischay's</a> PyTorch Lightning starter, but I've made a lot of changes.</li>\n<li>It makes negligible difference if I run inference in smaller batches</li>\n<li>I tried different datatypes/normalisations but it made no difference, floats on [0,255] or [normalised on [-1,1].</li>\n<li>It takes only 3 seconds to create the images from the sound files, that's not the issue.</li>\n</ul>\n<p>It puzzles me that training weights have much to do with evaluation speed.</p>",
  "messages": [
    {
      "id": "2221066",
      "postDate": "04/14/2023 00:31:38",
      "content": "<p>There are a couple of other posts similar already, but the nature of this seems slightly different.  Or maybe we're all facing the same mysterious problem.</p>\n<p>I'm getting my submissions timing out because the evaluation time is varying enormously.  I'm timing the evaluation of the one 600 second clip in the test set, and not counting the rest of my code execution.  My observations so far are:</p>\n<ul>\n<li>I can now turn on and off the problem, simply by training for more or less epochs.  Not really a solution, but it does narrow down the problem a lot!</li>\n<li>Repeatable times with some training weights,  for example one set consistently takes 15 seconds with EfficientNetV2s.</li>\n<li>Inconsistent with other training weights, produced by the same notebook.  For example I've run a notebook with the inference step taking 9 seconds, but it failed submission.  I then forked it, made no changes at all, and now it takes 380 seconds every time.</li>\n<li>No changes, or small changes to the inference notebook in between problems. Consistent versions of Python environment, with both training and inference.  P100 GPU used for training, CPU only for inference of course.</li>\n<li>Both the training &amp; infer forked originally from <a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference\" target=\"_blank\">Nischay's</a> PyTorch Lightning starter, but I've made a lot of changes.</li>\n<li>It makes negligible difference if I run inference in smaller batches</li>\n<li>I tried different datatypes/normalisations but it made no difference, floats on [0,255] or [normalised on [-1,1].</li>\n<li>It takes only 3 seconds to create the images from the sound files, that's not the issue.</li>\n</ul>\n<p>It puzzles me that training weights have much to do with evaluation speed.</p>",
      "rawMarkdown": "There are a couple of other posts similar already, but the nature of this seems slightly different.  Or maybe we're all facing the same mysterious problem.\n\nI'm getting my submissions timing out because the evaluation time is varying enormously.  I'm timing the evaluation of the one 600 second clip in the test set, and not counting the rest of my code execution.  My observations so far are:\n\n- I can now turn on and off the problem, simply by training for more or less epochs.  Not really a solution, but it does narrow down the problem a lot!\n- Repeatable times with some training weights,  for example one set consistently takes 15 seconds with EfficientNetV2s.\n- Inconsistent with other training weights, produced by the same notebook.  For example I've run a notebook with the inference step taking 9 seconds, but it failed submission.  I then forked it, made no changes at all, and now it takes 380 seconds every time.\n- No changes, or small changes to the inference notebook in between problems. Consistent versions of Python environment, with both training and inference.  P100 GPU used for training, CPU only for inference of course.\n- Both the training & infer forked originally from [Nischay's](https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference) PyTorch Lightning starter, but I've made a lot of changes.\n- It makes negligible difference if I run inference in smaller batches\n- I tried different datatypes/normalisations but it made no difference, floats on [0,255] or [normalised on [-1,1].\n- It takes only 3 seconds to create the images from the sound files, that's not the issue.\n\nIt puzzles me that training weights have much to do with evaluation speed.",
      "votes": null
    },
    {
      "id": "2221134",
      "postDate": "04/14/2023 03:30:01",
      "content": "<p>For what it’s worth I’ve seen none of these issues. I used pytorch for everything and start with a pretrained tf_efficientnet_b0_ns from timm (similar in spirit but not forked from Nischay). All my submissions have taken about ~20 min, about the same time it takes to run over the test audio clip 200 times.</p>",
      "rawMarkdown": "For what it’s worth I’ve seen none of these issues. I used pytorch for everything and start with a pretrained tf_efficientnet_b0_ns from timm (similar in spirit but not forked from Nischay). All my submissions have taken about ~20 min, about the same time it takes to run over the test audio clip 200 times.",
      "votes": null
    },
    {
      "id": "2221175",
      "postDate": "04/14/2023 04:21:38",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/robbynevels\" target=\"_blank\">@robbynevels</a> .  If a bunch of people confirm the same thing I might make a fresh inference book with PyTorch and timm.  (I'm also using efficientnet from timm.   I can't see how it has much to do with the training notebook.  I can see a lot of folks putting a lot of time into troubleshooting this issue, but if it's changing randomly even with the same code, then it's not likely to be an easy thing to trace.</p>",
      "rawMarkdown": "Thanks @robbynevels .  If a bunch of people confirm the same thing I might make a fresh inference book with PyTorch and timm.  (I'm also using efficientnet from timm.   I can't see how it has much to do with the training notebook.  I can see a lot of folks putting a lot of time into troubleshooting this issue, but if it's changing randomly even with the same code, then it's not likely to be an easy thing to trace.",
      "votes": null
    },
    {
      "id": "2221222",
      "postDate": "04/14/2023 04:59:28",
      "content": "<p>I feel for you :/ I originally spent 3 days trying to figure out why my submission was scoring so low… only to realize I forgot to do \"model.eval()\" after loading the weights lol. Hopefully whatever issue y'all are having isn't that silly.</p>\n<p>If it helps here's a public copy of a previous submission I used (model is private). I just submitted and it took ~11 min: <a href=\"https://www.kaggle.com/code/robbynevels/example-submission-w-pytorch-effnet-joblib/notebook\" target=\"_blank\">https://www.kaggle.com/code/robbynevels/example-submission-w-pytorch-effnet-joblib/notebook</a></p>",
      "rawMarkdown": "I feel for you :/ I originally spent 3 days trying to figure out why my submission was scoring so low... only to realize I forgot to do \"model.eval()\" after loading the weights lol. Hopefully whatever issue y'all are having isn't that silly.\n\nIf it helps here's a public copy of a previous submission I used (model is private). I just submitted and it took ~11 min: https://www.kaggle.com/code/robbynevels/example-submission-w-pytorch-effnet-joblib/notebook",
      "votes": null
    },
    {
      "id": "2221363",
      "postDate": "04/14/2023 07:44:40",
      "content": "<p>Thanks for the help, and the link, it did help me to find a couple of mistakes in my code, but alas still not solved the main problem.   My submission would take about 21 hours at it's current speed!</p>",
      "rawMarkdown": "Thanks for the help, and the link, it did help me to find a couple of mistakes in my code, but alas still not solved the main problem.   My submission would take about 21 hours at it's current speed!",
      "votes": null
    },
    {
      "id": "2225263",
      "postDate": "04/18/2023 03:42:13",
      "content": "<p>I ended up forking your notebook, and running it on my model and model weights, and got an even worse result.  Like more than 20 minutes, for the one sound file!</p>\n<p>I suspect the heart of the issue is a datatype thing.  I notice you've left your melspecs as uint8.  Could you tell me did you also train that way?  My notebooks were throwing errors in both training and inference when I tried to use uint8.  My understanding is that EffNet expects floats from [0,255].</p>",
      "rawMarkdown": "I ended up forking your notebook, and running it on my model and model weights, and got an even worse result.  Like more than 20 minutes, for the one sound file!\n\nI suspect the heart of the issue is a datatype thing.  I notice you've left your melspecs as uint8.  Could you tell me did you also train that way?  My notebooks were throwing errors in both training and inference when I tried to use uint8.  My understanding is that EffNet expects floats from [0,255].",
      "votes": null
    },
    {
      "id": "2227371",
      "postDate": "04/19/2023 17:55:52",
      "content": "<p>That's bizarre! It's never taken anywhere close to that amount of time for me. Are you using another pretrained model from timms, or something else? I'd be extremely surprised if this had anything to do with datatypes. Could it be that you're accidentally running the model on the entire 10 min audio, instead of 5 sec crops? That could cause a lot of unnecessary processing and possibly disk paging.</p>\n<p>I precomputed the melspecs as uint8 because it compresses the dataset to about ~5 GB, making it possible to fit into a GPU during training, which really speeds stuff up (only about ~1 min per epoch). </p>\n<p>Before passing each melspec to the model, I convert it to floats by mapping the [0, 255] range to [-1, 1], since it's usually better to have input be centered around zero:</p>\n<pre><code>data = data / 127 - 1 \n# data.dtype is now float32 ranged [-1, 1]\n</code></pre>\n<p>However, I had no idea that effnet expects [0, 255]! I don't see that mentioned in the timms documentation and it's not obvious to me where that assumption is encoded in their implementation, but from googling it does indeed sound like that's how the original version from google brain worked. I'll experiment with keeping it at 0,255 instead during training and see if it converges faster. Thanks for letting me know!</p>",
      "rawMarkdown": "That's bizarre! It's never taken anywhere close to that amount of time for me. Are you using another pretrained model from timms, or something else? I'd be extremely surprised if this had anything to do with datatypes. Could it be that you're accidentally running the model on the entire 10 min audio, instead of 5 sec crops? That could cause a lot of unnecessary processing and possibly disk paging.\n\nI precomputed the melspecs as uint8 because it compresses the dataset to about ~5 GB, making it possible to fit into a GPU during training, which really speeds stuff up (only about ~1 min per epoch). \n\nBefore passing each melspec to the model, I convert it to floats by mapping the [0, 255] range to [-1, 1], since it's usually better to have input be centered around zero:\n\n```\ndata = data / 127 - 1 \n# data.dtype is now float32 ranged [-1, 1]\n```\n\nHowever, I had no idea that effnet expects [0, 255]! I don't see that mentioned in the timms documentation and it's not obvious to me where that assumption is encoded in their implementation, but from googling it does indeed sound like that's how the original version from google brain worked. I'll experiment with keeping it at 0,255 instead during training and see if it converges faster. Thanks for letting me know!",
      "votes": null
    },
    {
      "id": "2228923",
      "postDate": "04/20/2023 23:49:51",
      "content": "<p>Well I still think the root cause is something to do with the data pre-processing.  Thanks for the detailed response, there are some good hints in there.   I'm out of GPU allocation for the week, but will try again on Saturday.</p>\n<p>I may have been getting confused between standardizing for Ablumentations on [0,1] and Efficientnet on [-1,1].  And yes I'm using timm, but maybe that implementation Efficientnetcan't do [0,255].  It would be interesting to know if your model falls over if you tried that.</p>\n<p>It's still really surprising that any of this effects inference speed (as opposed to accuracy, training stability or convergence).  I wish I had an explanation for that.  It will be there under the hood, but I'd rather focus on other things right now!</p>",
      "rawMarkdown": "Well I still think the root cause is something to do with the data pre-processing.  Thanks for the detailed response, there are some good hints in there.   I'm out of GPU allocation for the week, but will try again on Saturday.\n\nI may have been getting confused between standardizing for Ablumentations on [0,1] and Efficientnet on [-1,1].  And yes I'm using timm, but maybe that implementation Efficientnetcan't do [0,255].  It would be interesting to know if your model falls over if you tried that.\n\nIt's still really surprising that any of this effects inference speed (as opposed to accuracy, training stability or convergence).  I wish I had an explanation for that.  It will be there under the hood, but I'd rather focus on other things right now!",
      "votes": null
    },
    {
      "id": "2228946",
      "postDate": "04/21/2023 00:45:41",
      "content": "<p>I retrained on data that was just [0, 255] floats and it didn't seem to have much affect on training or validation. I guess the model just learns to scale very quickly to the given range at the beginning of training.</p>\n<p>When you say you're out of GPU allocation -- that would only affect your training notebook, right? So you could still test/instrument the submission notebook. You might try profiling your code and seeing what part is taking the longest: <a href=\"https://pytorch.org/tutorials/recipes/recipes/profiler_recipe.html\" target=\"_blank\">https://pytorch.org/tutorials/recipes/recipes/profiler_recipe.html</a></p>",
      "rawMarkdown": "I retrained on data that was just [0, 255] floats and it didn't seem to have much affect on training or validation. I guess the model just learns to scale very quickly to the given range at the beginning of training.\n\nWhen you say you're out of GPU allocation -- that would only affect your training notebook, right? So you could still test/instrument the submission notebook. You might try profiling your code and seeing what part is taking the longest: https://pytorch.org/tutorials/recipes/recipes/profiler_recipe.html",
      "votes": null
    },
    {
      "id": "2231965",
      "postDate": "04/23/2023 21:40:12",
      "content": "<p>Thanks for trying on [255] that, It's great to have one more possibility eliminated!  Actually I don't think it is a datatype problem now, I'm back to training with images normalised on [-1,1] from now on.</p>\n<p>The problem is definitely with model evaluation.  The rest of the code execution is trivial.  But I have made one really useful observation: The problem only occurs if I train for too long.  I can get \"good\" weights, that evaluate in 15 seconds on the test sample, and submit OK, if I keep the number of epochs down.    So this points towards something going wrong with my loss function, or validation epoch end callback on training I think.   I'm using CrossEntropyLoss, but might go back to BCE loss.</p>",
      "rawMarkdown": "Thanks for trying on [255] that, It's great to have one more possibility eliminated!  Actually I don't think it is a datatype problem now, I'm back to training with images normalised on [-1,1] from now on.\n\nThe problem is definitely with model evaluation.  The rest of the code execution is trivial.  But I have made one really useful observation: The problem only occurs if I train for too long.  I can get \"good\" weights, that evaluate in 15 seconds on the test sample, and submit OK, if I keep the number of epochs down.    So this points towards something going wrong with my loss function, or validation epoch end callback on training I think.   I'm using CrossEntropyLoss, but might go back to BCE loss.",
      "votes": null
    },
    {
      "id": "2232483",
      "postDate": "04/24/2023 11:33:24",
      "content": "<p>If it's training weights that cause this issue, one thing that comes to mind is denormal values.</p>\n<p>You can check if there are denormal values with the answer in this issue.<br>\n<a href=\"https://github.com/microsoft/onnxruntime/issues/5472\" target=\"_blank\">https://github.com/microsoft/onnxruntime/issues/5472</a></p>",
      "rawMarkdown": "If it's training weights that cause this issue, one thing that comes to mind is denormal values.\n\nYou can check if there are denormal values with the answer in this issue.\nhttps://github.com/microsoft/onnxruntime/issues/5472",
      "votes": null
    },
    {
      "id": "2234091",
      "postDate": "04/24/2023 21:22:26",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/lhanhsin\" target=\"_blank\">@lhanhsin</a> ,   yeah that sounds like the same issue!  I think you're on the right track.   Actually ptrblck over on the Pytorch Discuss forum suggested this too, I set it on the inference notebook, it helped a little, (but only like a 20% speedup), but didn't occur to me to try it on the training one as well.  I'm re-running the training now to see if that helps.  Will know in a few hours!     Cheers.</p>",
      "rawMarkdown": "Thanks @lhanhsin ,   yeah that sounds like the same issue!  I think you're on the right track.   Actually ptrblck over on the Pytorch Discuss forum suggested this too, I set it on the inference notebook, it helped a little, (but only like a 20% speedup), but didn't occur to me to try it on the training one as well.  I'm re-running the training now to see if that helps.  Will know in a few hours!     Cheers.",
      "votes": null
    },
    {
      "id": "2234226",
      "postDate": "04/25/2023 03:13:08",
      "content": "<p>Yeah it might be worth trying, let me know if training with the flushing helps!</p>\n<p>By the way I ran a little experiment to see what <code>torch.set_flush_denormal(True)</code> actually does and the computing speed is still slow when the float is too small. I'm guessing it flushes denormal values when seen but doesn't prevent from computing such values? This kinda explains why there is partial improvement but still is slow.<br>\nA quick fix would be to manually flush out some parameters that you think are too small, perhaps &lt;1e-18.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F4427ac210dc4fdc7cfb83ccfd347473e%2F2023-04-25%20124151.jpg?generation=1682392336847987&amp;alt=media\" alt=\"torch.set_flush_denormal(True) behavior\"></p>",
      "rawMarkdown": "Yeah it might be worth trying, let me know if training with the flushing helps!\n\nBy the way I ran a little experiment to see what `torch.set_flush_denormal(True)` actually does and the computing speed is still slow when the float is too small. I'm guessing it flushes denormal values when seen but doesn't prevent from computing such values? This kinda explains why there is partial improvement but still is slow.\nA quick fix would be to manually flush out some parameters that you think are too small, perhaps <1e-18.\n![torch.set_flush_denormal(True) behavior](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F4427ac210dc4fdc7cfb83ccfd347473e%2F2023-04-25%20124151.jpg?generation=1682392336847987&alt=media)",
      "votes": null
    },
    {
      "id": "2235245",
      "postDate": "04/25/2023 21:20:52",
      "content": "<p>Well, I want to tentatively say that I think  <code>torch.set_flush_denormal(True)</code> in the training notebook has fixed the problem!   I've run a training notebook with 17 Epochs and the inference went fine.  (19 seconds for the 600second test clip, and it submitted OK).   </p>\n<p>I've had this problem come and go before, so not sure I'm ready to declare total victory 😂, but it's looking good so far anyway.  Thanks so much for the suggestion.</p>",
      "rawMarkdown": "Well, I want to tentatively say that I think  `torch.set_flush_denormal(True)` in the training notebook has fixed the problem!   I've run a training notebook with 17 Epochs and the inference went fine.  (19 seconds for the 600second test clip, and it submitted OK).   \n\nI've had this problem come and go before, so not sure I'm ready to declare total victory 😂, but it's looking good so far anyway.  Thanks so much for the suggestion.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2221134,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "04/14/2023 03:30:01",
      "content": "<p>For what it’s worth I’ve seen none of these issues. I used pytorch for everything and start with a pretrained tf_efficientnet_b0_ns from timm (similar in spirit but not forked from Nischay). All my submissions have taken about ~20 min, about the same time it takes to run over the test audio clip 200 times.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2221175,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "04/14/2023 04:21:38",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/robbynevels\" target=\"_blank\">@robbynevels</a> .  If a bunch of people confirm the same thing I might make a fresh inference book with PyTorch and timm.  (I'm also using efficientnet from timm.   I can't see how it has much to do with the training notebook.  I can see a lot of folks putting a lot of time into troubleshooting this issue, but if it's changing randomly even with the same code, then it's not likely to be an easy thing to trace.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2221222,
              "author_name": "robbynevels",
              "author_url": "",
              "post_date": "04/14/2023 04:59:28",
              "content": "<p>I feel for you :/ I originally spent 3 days trying to figure out why my submission was scoring so low… only to realize I forgot to do \"model.eval()\" after loading the weights lol. Hopefully whatever issue y'all are having isn't that silly.</p>\n<p>If it helps here's a public copy of a previous submission I used (model is private). I just submitted and it took ~11 min: <a href=\"https://www.kaggle.com/code/robbynevels/example-submission-w-pytorch-effnet-joblib/notebook\" target=\"_blank\">https://www.kaggle.com/code/robbynevels/example-submission-w-pytorch-effnet-joblib/notebook</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2221363,
                  "author_name": "ollypowell",
                  "author_url": "",
                  "post_date": "04/14/2023 07:44:40",
                  "content": "<p>Thanks for the help, and the link, it did help me to find a couple of mistakes in my code, but alas still not solved the main problem.   My submission would take about 21 hours at it's current speed!</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2225263,
                  "author_name": "ollypowell",
                  "author_url": "",
                  "post_date": "04/18/2023 03:42:13",
                  "content": "<p>I ended up forking your notebook, and running it on my model and model weights, and got an even worse result.  Like more than 20 minutes, for the one sound file!</p>\n<p>I suspect the heart of the issue is a datatype thing.  I notice you've left your melspecs as uint8.  Could you tell me did you also train that way?  My notebooks were throwing errors in both training and inference when I tried to use uint8.  My understanding is that EffNet expects floats from [0,255].</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2227371,
                      "author_name": "robbynevels",
                      "author_url": "",
                      "post_date": "04/19/2023 17:55:52",
                      "content": "<p>That's bizarre! It's never taken anywhere close to that amount of time for me. Are you using another pretrained model from timms, or something else? I'd be extremely surprised if this had anything to do with datatypes. Could it be that you're accidentally running the model on the entire 10 min audio, instead of 5 sec crops? That could cause a lot of unnecessary processing and possibly disk paging.</p>\n<p>I precomputed the melspecs as uint8 because it compresses the dataset to about ~5 GB, making it possible to fit into a GPU during training, which really speeds stuff up (only about ~1 min per epoch). </p>\n<p>Before passing each melspec to the model, I convert it to floats by mapping the [0, 255] range to [-1, 1], since it's usually better to have input be centered around zero:</p>\n<pre><code>data = data / 127 - 1 \n# data.dtype is now float32 ranged [-1, 1]\n</code></pre>\n<p>However, I had no idea that effnet expects [0, 255]! I don't see that mentioned in the timms documentation and it's not obvious to me where that assumption is encoded in their implementation, but from googling it does indeed sound like that's how the original version from google brain worked. I'll experiment with keeping it at 0,255 instead during training and see if it converges faster. Thanks for letting me know!</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2228923,
                          "author_name": "ollypowell",
                          "author_url": "",
                          "post_date": "04/20/2023 23:49:51",
                          "content": "<p>Well I still think the root cause is something to do with the data pre-processing.  Thanks for the detailed response, there are some good hints in there.   I'm out of GPU allocation for the week, but will try again on Saturday.</p>\n<p>I may have been getting confused between standardizing for Ablumentations on [0,1] and Efficientnet on [-1,1].  And yes I'm using timm, but maybe that implementation Efficientnetcan't do [0,255].  It would be interesting to know if your model falls over if you tried that.</p>\n<p>It's still really surprising that any of this effects inference speed (as opposed to accuracy, training stability or convergence).  I wish I had an explanation for that.  It will be there under the hood, but I'd rather focus on other things right now!</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2228946,
                              "author_name": "robbynevels",
                              "author_url": "",
                              "post_date": "04/21/2023 00:45:41",
                              "content": "<p>I retrained on data that was just [0, 255] floats and it didn't seem to have much affect on training or validation. I guess the model just learns to scale very quickly to the given range at the beginning of training.</p>\n<p>When you say you're out of GPU allocation -- that would only affect your training notebook, right? So you could still test/instrument the submission notebook. You might try profiling your code and seeing what part is taking the longest: <a href=\"https://pytorch.org/tutorials/recipes/recipes/profiler_recipe.html\" target=\"_blank\">https://pytorch.org/tutorials/recipes/recipes/profiler_recipe.html</a></p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2231965,
                                  "author_name": "ollypowell",
                                  "author_url": "",
                                  "post_date": "04/23/2023 21:40:12",
                                  "content": "<p>Thanks for trying on [255] that, It's great to have one more possibility eliminated!  Actually I don't think it is a datatype problem now, I'm back to training with images normalised on [-1,1] from now on.</p>\n<p>The problem is definitely with model evaluation.  The rest of the code execution is trivial.  But I have made one really useful observation: The problem only occurs if I train for too long.  I can get \"good\" weights, that evaluate in 15 seconds on the test sample, and submit OK, if I keep the number of epochs down.    So this points towards something going wrong with my loss function, or validation epoch end callback on training I think.   I'm using CrossEntropyLoss, but might go back to BCE loss.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2232483,
      "author_name": "lhanhsin",
      "author_url": "",
      "post_date": "04/24/2023 11:33:24",
      "content": "<p>If it's training weights that cause this issue, one thing that comes to mind is denormal values.</p>\n<p>You can check if there are denormal values with the answer in this issue.<br>\n<a href=\"https://github.com/microsoft/onnxruntime/issues/5472\" target=\"_blank\">https://github.com/microsoft/onnxruntime/issues/5472</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2234091,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "04/24/2023 21:22:26",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/lhanhsin\" target=\"_blank\">@lhanhsin</a> ,   yeah that sounds like the same issue!  I think you're on the right track.   Actually ptrblck over on the Pytorch Discuss forum suggested this too, I set it on the inference notebook, it helped a little, (but only like a 20% speedup), but didn't occur to me to try it on the training one as well.  I'm re-running the training now to see if that helps.  Will know in a few hours!     Cheers.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2234226,
              "author_name": "lhanhsin",
              "author_url": "",
              "post_date": "04/25/2023 03:13:08",
              "content": "<p>Yeah it might be worth trying, let me know if training with the flushing helps!</p>\n<p>By the way I ran a little experiment to see what <code>torch.set_flush_denormal(True)</code> actually does and the computing speed is still slow when the float is too small. I'm guessing it flushes denormal values when seen but doesn't prevent from computing such values? This kinda explains why there is partial improvement but still is slow.<br>\nA quick fix would be to manually flush out some parameters that you think are too small, perhaps &lt;1e-18.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F4427ac210dc4fdc7cfb83ccfd347473e%2F2023-04-25%20124151.jpg?generation=1682392336847987&amp;alt=media\" alt=\"torch.set_flush_denormal(True) behavior\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2235245,
                  "author_name": "ollypowell",
                  "author_url": "",
                  "post_date": "04/25/2023 21:20:52",
                  "content": "<p>Well, I want to tentatively say that I think  <code>torch.set_flush_denormal(True)</code> in the training notebook has fixed the problem!   I've run a training notebook with 17 Epochs and the inference went fine.  (19 seconds for the 600second test clip, and it submitted OK).   </p>\n<p>I've had this problem come and go before, so not sure I'm ready to declare total victory 😂, but it's looking good so far anyway.  Thanks so much for the suggestion.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2221066": "There are a couple of other posts similar already, but the nature of this seems slightly different.  Or maybe we're all facing the same mysterious problem.\n\nI'm getting my submissions timing out because the evaluation time is varying enormously.  I'm timing the evaluation of the one 600 second clip in the test set, and not counting the rest of my code execution.  My observations so far are:\n\n- I can now turn on and off the problem, simply by training for more or less epochs.  Not really a solution, but it does narrow down the problem a lot!\n- Repeatable times with some training weights,  for example one set consistently takes 15 seconds with EfficientNetV2s.\n- Inconsistent with other training weights, produced by the same notebook.  For example I've run a notebook with the inference step taking 9 seconds, but it failed submission.  I then forked it, made no changes at all, and now it takes 380 seconds every time.\n- No changes, or small changes to the inference notebook in between problems. Consistent versions of Python environment, with both training and inference.  P100 GPU used for training, CPU only for inference of course.\n- Both the training & infer forked originally from [Nischay's](https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference) PyTorch Lightning starter, but I've made a lot of changes.\n- It makes negligible difference if I run inference in smaller batches\n- I tried different datatypes/normalisations but it made no difference, floats on [0,255] or [normalised on [-1,1].\n- It takes only 3 seconds to create the images from the sound files, that's not the issue.\n\nIt puzzles me that training weights have much to do with evaluation speed.",
    "2221134": "For what it’s worth I’ve seen none of these issues. I used pytorch for everything and start with a pretrained tf_efficientnet_b0_ns from timm (similar in spirit but not forked from Nischay). All my submissions have taken about ~20 min, about the same time it takes to run over the test audio clip 200 times.",
    "2221175": "Thanks @robbynevels .  If a bunch of people confirm the same thing I might make a fresh inference book with PyTorch and timm.  (I'm also using efficientnet from timm.   I can't see how it has much to do with the training notebook.  I can see a lot of folks putting a lot of time into troubleshooting this issue, but if it's changing randomly even with the same code, then it's not likely to be an easy thing to trace.",
    "2221222": "I feel for you :/ I originally spent 3 days trying to figure out why my submission was scoring so low... only to realize I forgot to do \"model.eval()\" after loading the weights lol. Hopefully whatever issue y'all are having isn't that silly.\n\nIf it helps here's a public copy of a previous submission I used (model is private). I just submitted and it took ~11 min: https://www.kaggle.com/code/robbynevels/example-submission-w-pytorch-effnet-joblib/notebook",
    "2221363": "Thanks for the help, and the link, it did help me to find a couple of mistakes in my code, but alas still not solved the main problem.   My submission would take about 21 hours at it's current speed!",
    "2225263": "I ended up forking your notebook, and running it on my model and model weights, and got an even worse result.  Like more than 20 minutes, for the one sound file!\n\nI suspect the heart of the issue is a datatype thing.  I notice you've left your melspecs as uint8.  Could you tell me did you also train that way?  My notebooks were throwing errors in both training and inference when I tried to use uint8.  My understanding is that EffNet expects floats from [0,255].",
    "2227371": "That's bizarre! It's never taken anywhere close to that amount of time for me. Are you using another pretrained model from timms, or something else? I'd be extremely surprised if this had anything to do with datatypes. Could it be that you're accidentally running the model on the entire 10 min audio, instead of 5 sec crops? That could cause a lot of unnecessary processing and possibly disk paging.\n\nI precomputed the melspecs as uint8 because it compresses the dataset to about ~5 GB, making it possible to fit into a GPU during training, which really speeds stuff up (only about ~1 min per epoch). \n\nBefore passing each melspec to the model, I convert it to floats by mapping the [0, 255] range to [-1, 1], since it's usually better to have input be centered around zero:\n\n```\ndata = data / 127 - 1 \n# data.dtype is now float32 ranged [-1, 1]\n```\n\nHowever, I had no idea that effnet expects [0, 255]! I don't see that mentioned in the timms documentation and it's not obvious to me where that assumption is encoded in their implementation, but from googling it does indeed sound like that's how the original version from google brain worked. I'll experiment with keeping it at 0,255 instead during training and see if it converges faster. Thanks for letting me know!",
    "2228923": "Well I still think the root cause is something to do with the data pre-processing.  Thanks for the detailed response, there are some good hints in there.   I'm out of GPU allocation for the week, but will try again on Saturday.\n\nI may have been getting confused between standardizing for Ablumentations on [0,1] and Efficientnet on [-1,1].  And yes I'm using timm, but maybe that implementation Efficientnetcan't do [0,255].  It would be interesting to know if your model falls over if you tried that.\n\nIt's still really surprising that any of this effects inference speed (as opposed to accuracy, training stability or convergence).  I wish I had an explanation for that.  It will be there under the hood, but I'd rather focus on other things right now!",
    "2228946": "I retrained on data that was just [0, 255] floats and it didn't seem to have much affect on training or validation. I guess the model just learns to scale very quickly to the given range at the beginning of training.\n\nWhen you say you're out of GPU allocation -- that would only affect your training notebook, right? So you could still test/instrument the submission notebook. You might try profiling your code and seeing what part is taking the longest: https://pytorch.org/tutorials/recipes/recipes/profiler_recipe.html",
    "2231965": "Thanks for trying on [255] that, It's great to have one more possibility eliminated!  Actually I don't think it is a datatype problem now, I'm back to training with images normalised on [-1,1] from now on.\n\nThe problem is definitely with model evaluation.  The rest of the code execution is trivial.  But I have made one really useful observation: The problem only occurs if I train for too long.  I can get \"good\" weights, that evaluate in 15 seconds on the test sample, and submit OK, if I keep the number of epochs down.    So this points towards something going wrong with my loss function, or validation epoch end callback on training I think.   I'm using CrossEntropyLoss, but might go back to BCE loss.",
    "2232483": "If it's training weights that cause this issue, one thing that comes to mind is denormal values.\n\nYou can check if there are denormal values with the answer in this issue.\nhttps://github.com/microsoft/onnxruntime/issues/5472",
    "2234091": "Thanks @lhanhsin ,   yeah that sounds like the same issue!  I think you're on the right track.   Actually ptrblck over on the Pytorch Discuss forum suggested this too, I set it on the inference notebook, it helped a little, (but only like a 20% speedup), but didn't occur to me to try it on the training one as well.  I'm re-running the training now to see if that helps.  Will know in a few hours!     Cheers.",
    "2234226": "Yeah it might be worth trying, let me know if training with the flushing helps!\n\nBy the way I ran a little experiment to see what `torch.set_flush_denormal(True)` actually does and the computing speed is still slow when the float is too small. I'm guessing it flushes denormal values when seen but doesn't prevent from computing such values? This kinda explains why there is partial improvement but still is slow.\nA quick fix would be to manually flush out some parameters that you think are too small, perhaps <1e-18.\n![torch.set_flush_denormal(True) behavior](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4992459%2F4427ac210dc4fdc7cfb83ccfd347473e%2F2023-04-25%20124151.jpg?generation=1682392336847987&alt=media)",
    "2235245": "Well, I want to tentatively say that I think  `torch.set_flush_denormal(True)` in the training notebook has fixed the problem!   I've run a training notebook with 17 Epochs and the inference went fine.  (19 seconds for the 600second test clip, and it submitted OK).   \n\nI've had this problem come and go before, so not sure I'm ready to declare total victory 😂, but it's looking good so far anyway.  Thanks so much for the suggestion."
  },
  "source": "meta"
}