{
  "id": 499530,
  "title": "Question: Difference in prediction time for the same model",
  "url": "/competitions/birdclef-2024/discussion/499530",
  "author_name": "",
  "post_date": "2024-05-02T05:31:18.727445700Z",
  "votes": 3,
  "comment_count": 15,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16858852%2F2f388a81258434ff90e224610c5332bb%2Fcontents.png?generation=1714637198131654&amp;alt=media\">I am kaggle beginner. I need help.</p>\n<p>The prediction time differs greatly depending on the weights learned for the same model.</p>\n<p>In the CLEF2024 competition, when the model was submitted with the first learned weights, it was within 2 hours without any problem,<br>\nHowever, when the weights were improved using the optimizer's scheduler, the prediction time increased by a factor of 5.</p>\n<p>We investigated and found that CPU usage seemed to be increasing.<br>\nWhen we tried forecasting on the GPU, the difference in forecasting time between the previous and current learning weights disappeared.</p>\n<p>Since it is a save and load of state_dict(), nothing should have changed except the weights and bias values.<br>\nIs it possible for CPU usage to change in this situation?</p>",
  "messages": [
    {
      "id": "2788152",
      "postDate": "05/02/2024 05:31:18",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16858852%2F2f388a81258434ff90e224610c5332bb%2Fcontents.png?generation=1714637198131654&amp;alt=media\">I am kaggle beginner. I need help.</p>\n<p>The prediction time differs greatly depending on the weights learned for the same model.</p>\n<p>In the CLEF2024 competition, when the model was submitted with the first learned weights, it was within 2 hours without any problem,<br>\nHowever, when the weights were improved using the optimizer's scheduler, the prediction time increased by a factor of 5.</p>\n<p>We investigated and found that CPU usage seemed to be increasing.<br>\nWhen we tried forecasting on the GPU, the difference in forecasting time between the previous and current learning weights disappeared.</p>\n<p>Since it is a save and load of state_dict(), nothing should have changed except the weights and bias values.<br>\nIs it possible for CPU usage to change in this situation?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16858852%2F2f388a81258434ff90e224610c5332bb%2Fcontents.png?generation=1714637198131654&alt=media)I am kaggle beginner. I need help.\n\nThe prediction time differs greatly depending on the weights learned for the same model.\n\nIn the CLEF2024 competition, when the model was submitted with the first learned weights, it was within 2 hours without any problem,\nHowever, when the weights were improved using the optimizer's scheduler, the prediction time increased by a factor of 5.\n\nWe investigated and found that CPU usage seemed to be increasing.\nWhen we tried forecasting on the GPU, the difference in forecasting time between the previous and current learning weights disappeared.\n\nSince it is a save and load of state_dict(), nothing should have changed except the weights and bias values.\nIs it possible for CPU usage to change in this situation?",
      "votes": null
    },
    {
      "id": "2788746",
      "postDate": "05/02/2024 12:04:17",
      "content": "<p>It is normal for the prediction time to vary depending on the learned weights of the model. Optimization and weight updates can affect the model’s performance.</p>\n<p>When you see an increase in CPU usage, this situation can depend on several factors:</p>\n<p>Optimization Algorithms: Optimization algorithms can work at different speeds when updating weights. This can affect CPU usage.<br>\nData Size and Processing Load: The size of the training data and the processing load can affect CPU usage. Larger datasets or more complex processing loads can use the CPU more.<br>\nHyperparameter Settings: Hyperparameter settings (for example, learning rate, momentum, regularization term) can also affect CPU usage.<br>\nHardware Differences: Different hardware (CPU type, memory, disk speed, etc.) can trigger different CPU usage. When you see that GPU usage eliminates the difference in prediction time, it shows that the GPU benefits from parallel computing capabilities. The GPU can accelerate large matrix operations and help the model run faster.<br>\nSince the state_dict is saved and loaded, nothing should change except for the weights and bias values. However, this does not affect CPU usage. You should investigate other factors for changes in CPU usage.</p>\n<p>If you need more information about the change in CPU usage, it would be useful to check your hardware configuration and review the optimization steps. Also, continuing to run your model on the GPU can improve performance🥇</p>",
      "rawMarkdown": "It is normal for the prediction time to vary depending on the learned weights of the model. Optimization and weight updates can affect the model’s performance.\n\nWhen you see an increase in CPU usage, this situation can depend on several factors:\n\nOptimization Algorithms: Optimization algorithms can work at different speeds when updating weights. This can affect CPU usage.\nData Size and Processing Load: The size of the training data and the processing load can affect CPU usage. Larger datasets or more complex processing loads can use the CPU more.\nHyperparameter Settings: Hyperparameter settings (for example, learning rate, momentum, regularization term) can also affect CPU usage.\nHardware Differences: Different hardware (CPU type, memory, disk speed, etc.) can trigger different CPU usage. When you see that GPU usage eliminates the difference in prediction time, it shows that the GPU benefits from parallel computing capabilities. The GPU can accelerate large matrix operations and help the model run faster.\nSince the state_dict is saved and loaded, nothing should change except for the weights and bias values. However, this does not affect CPU usage. You should investigate other factors for changes in CPU usage.\n\nIf you need more information about the change in CPU usage, it would be useful to check your hardware configuration and review the optimization steps. Also, continuing to run your model on the GPU can improve performance🥇",
      "votes": null
    },
    {
      "id": "2788924",
      "postDate": "05/02/2024 13:32:20",
      "content": "<p>Thank you.<br>\nI understand that when learning, there can be a difference in learning time depending on the optimizer and other settings.</p>\n<p><strong>&gt;Since the state_dict is saved and loaded, nothing should change except for the weights and bias values.</strong><br>\nThe state_dict should only store the weights and bias information.<br>\nTherefore, I would expect the processing time to be the same for both learned parameters…</p>\n<p>By the way, I thought the first learned parameter had more zeros, so I investigated, but the second learned parameter had more zeros.</p>",
      "rawMarkdown": "Thank you.\nI understand that when learning, there can be a difference in learning time depending on the optimizer and other settings.\n\n**>Since the state_dict is saved and loaded, nothing should change except for the weights and bias values.**\nThe state_dict should only store the weights and bias information.\nTherefore, I would expect the processing time to be the same for both learned parameters...\n\nBy the way, I thought the first learned parameter had more zeros, so I investigated, but the second learned parameter had more zeros.",
      "votes": null
    },
    {
      "id": "2789239",
      "postDate": "05/02/2024 16:07:30",
      "content": "<p>You did more than retrainnig. You also added some metric computation with torch. Why do you assume that the slowdown is not due to the latter?</p>",
      "rawMarkdown": "You did more than retrainnig. You also added some metric computation with torch. Why do you assume that the slowdown is not due to the latter?",
      "votes": null
    },
    {
      "id": "2790614",
      "postDate": "05/03/2024 09:11:04",
      "content": "<p>Thank you.<br>\nSorry for the lack of explanation.</p>\n<p>I created separate notebooks for the study phase and the submission phase.<br>\nIn the learning phase, we added the evaluation metrics for the model, but in the evaluation phase in question, we did not change anything except the state_dict to be loaded.<br>\nWhile the model is the same and the size of parameters to be loaded is the same, there is a difference in processing time.</p>",
      "rawMarkdown": "Thank you.\nSorry for the lack of explanation.\n\nI created separate notebooks for the study phase and the submission phase.\nIn the learning phase, we added the evaluation metrics for the model, but in the evaluation phase in question, we did not change anything except the state_dict to be loaded.\nWhile the model is the same and the size of parameters to be loaded is the same, there is a difference in processing time.",
      "votes": null
    },
    {
      "id": "2790677",
      "postDate": "05/03/2024 09:45:28",
      "content": "<p>OK. This is the same code, the only change are the weights?</p>\n<p>Maybe you are close to RAM capacity or something, end depending on race conditions, you may get slowdown because you are using the compute resource at its limit?</p>",
      "rawMarkdown": "OK. This is the same code, the only change are the weights?\n\nMaybe you are close to RAM capacity or something, end depending on race conditions, you may get slowdown because you are using the compute resource at its limit?",
      "votes": null
    },
    {
      "id": "2791049",
      "postDate": "05/03/2024 13:00:16",
      "content": "<p>The only thing that has changed is the weight.</p>\n<p><strong>&gt;Maybe you are close to RAM capacity or something, end depending on race conditions, you may get slowdown because you are using the compute resource at its limit?</strong><br>\nThank you. I will look into the status of the various resources.</p>",
      "rawMarkdown": "The only thing that has changed is the weight.\n\n**>Maybe you are close to RAM capacity or something, end depending on race conditions, you may get slowdown because you are using the compute resource at its limit?**\nThank you. I will look into the status of the various resources.",
      "votes": null
    },
    {
      "id": "2792859",
      "postDate": "05/04/2024 12:40:14",
      "content": "<p>I had a similar situation. The inference time of a trained model was about 60 sec per 4-min record. Much too slow!</p>\n<p>I tried to figure out, why, and I ended up with a single PyTorch Conv2d layer. The inference time of this layer was about 3 sec, while the same layer initialized with random weights had inference time of 0.06 sec. No matter what I did, when the trained weights were used, the layer became 'slow'.</p>\n<p>The solution was to use <code>model.half().float()</code>, i.e. to round the weights of the model. Actually, I don't understand, why it works, because the type of weight tensor remains <code>float32</code>. Somehow, this significantly increased the speed of computations, while the results produced by the model are almost the same.</p>",
      "rawMarkdown": "I had a similar situation. The inference time of a trained model was about 60 sec per 4-min record. Much too slow!\n\nI tried to figure out, why, and I ended up with a single PyTorch Conv2d layer. The inference time of this layer was about 3 sec, while the same layer initialized with random weights had inference time of 0.06 sec. No matter what I did, when the trained weights were used, the layer became 'slow'.\n\nThe solution was to use `model.half().float()`, i.e. to round the weights of the model. Actually, I don't understand, why it works, because the type of weight tensor remains `float32`. Somehow, this significantly increased the speed of computations, while the results produced by the model are almost the same.",
      "votes": null
    },
    {
      "id": "2793082",
      "postDate": "05/04/2024 14:25:24",
      "content": "<p>Interesting. rounding your weights to fp16 precision helps. Are you using onnx or openvino? Maybe there is something with these that make them run FP16 when possible? I am speculating, I have absolutely no idea of why you speed your computation that way.</p>",
      "rawMarkdown": "Interesting. rounding your weights to fp16 precision helps. Are you using onnx or openvino? Maybe there is something with these that make them run FP16 when possible? I am speculating, I have absolutely no idea of why you speed your computation that way.",
      "votes": null
    },
    {
      "id": "2793162",
      "postDate": "05/04/2024 15:23:14",
      "content": "<p>If some of the model weights end up very small (e.g ~1e-30) during training, the inference time could become very large, it is because of the how CPU handles float computatitions in that range to maintain the required precision (that might be not true, but worth investigating, unfortunately I forgot the source of this information).</p>\n<p>Can you check the norm of the weights in the area around 0 before and after rounding?</p>\n<p>Most likely you'll be able to get a dependence in the chart \"lowest non-zero weight norm\"/inference_time</p>",
      "rawMarkdown": "If some of the model weights end up very small (e.g ~1e-30) during training, the inference time could become very large, it is because of the how CPU handles float computatitions in that range to maintain the required precision (that might be not true, but worth investigating, unfortunately I forgot the source of this information).\n\nCan you check the norm of the weights in the area around 0 before and after rounding?\n\nMost likely you'll be able to get a dependence in the chart \"lowest non-zero weight norm\"/inference_time",
      "votes": null
    },
    {
      "id": "2793192",
      "postDate": "05/04/2024 15:39:48",
      "content": "<p>What I'm reporting here is a pure PyTorch behaviour, without onnx/openvino.</p>\n<p>Just two layers: first initialized randomly, second - with <code>load_state_dict()</code>. I also tried to copy weight Parameter directly, or to convert it to numpy first, then back to tensor and Parameter. The result was the same: the weights trained on GPU made a layer extremely slow. I don't know exactly, but I suppose, there's some hidden data type conversion. However, I didn't find, where it happens.</p>",
      "rawMarkdown": "What I'm reporting here is a pure PyTorch behaviour, without onnx/openvino.\n\nJust two layers: first initialized randomly, second - with `load_state_dict()`. I also tried to copy weight Parameter directly, or to convert it to numpy first, then back to tensor and Parameter. The result was the same: the weights trained on GPU made a layer extremely slow. I don't know exactly, but I suppose, there's some hidden data type conversion. However, I didn't find, where it happens.",
      "votes": null
    },
    {
      "id": "2793257",
      "postDate": "05/04/2024 16:18:49",
      "content": "<p>Thank you, this may be the reason. I saw the weights like 1e-44, indeed, but didn't think about them in this way.</p>\n<p>While rounding the weights helps to solve the problem, it seems that adding L1 regularization should prevent it from even appearing.</p>",
      "rawMarkdown": "Thank you, this may be the reason. I saw the weights like 1e-44, indeed, but didn't think about them in this way.\n\nWhile rounding the weights helps to solve the problem, it seems that adding L1 regularization should prevent it from even appearing.",
      "votes": null
    },
    {
      "id": "2793781",
      "postDate": "05/04/2024 23:49:51",
      "content": "<p>Thank you!<br>\nI was surprised to find that the computation time can vary greatly depending on the weights for the same size.</p>\n<p>I will work on improving the model using the \"model.half().float()\" and other methods you mentioned.<br>\nFrom now on, I will work on learning the model while being aware of the value of the weights!</p>",
      "rawMarkdown": "Thank you!\nI was surprised to find that the computation time can vary greatly depending on the weights for the same size.\n\nI will work on improving the model using the \"model.half().float()\" and other methods you mentioned.\nFrom now on, I will work on learning the model while being aware of the value of the weights!",
      "votes": null
    },
    {
      "id": "2794415",
      "postDate": "05/05/2024 09:22:40",
      "content": "<p>Just let me know, if this actually helps, I'm curious about it.</p>",
      "rawMarkdown": "Just let me know, if this actually helps, I'm curious about it.",
      "votes": null
    },
    {
      "id": "2794634",
      "postDate": "05/05/2024 12:19:01",
      "content": "<p>The values of the front and rear weights were checked.<br>\nThe first learned parameter was 1e-8 with the smallest weight excluding 0, while the second learned parameter was e-45 with the smallest!</p>\n<p>Also, the second learned parameter was always less than 1e-40 in almost all layers.</p>\n<p>Using \"model.half().float()\", weights of the order of 1e-45 were reduced to the order of 1e-8, and the calculation was faster.<br>\nMin: 1.401298464324817e-45(Before using model.half().float())<br>\n↓↓<br>\nMin: 5.960464477539063e-08(After using model.half().float())</p>\n<p>Thank you all!<br>\nI am relieved that I found the cause!</p>",
      "rawMarkdown": "The values of the front and rear weights were checked.\nThe first learned parameter was 1e-8 with the smallest weight excluding 0, while the second learned parameter was e-45 with the smallest!\n\nAlso, the second learned parameter was always less than 1e-40 in almost all layers.\n\nUsing \"model.half().float()\", weights of the order of 1e-45 were reduced to the order of 1e-8, and the calculation was faster.\nMin: 1.401298464324817e-45(Before using model.half().float())\n↓↓\nMin: 5.960464477539063e-08(After using model.half().float())\n\n\nThank you all!\nI am relieved that I found the cause!",
      "votes": null
    },
    {
      "id": "2800067",
      "postDate": "05/08/2024 05:08:14",
      "content": "<p>Has anyone experimented training with 16bit precision?</p>",
      "rawMarkdown": "Has anyone experimented training with 16bit precision?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2788746,
      "author_name": "akifahiner",
      "author_url": "",
      "post_date": "05/02/2024 12:04:17",
      "content": "<p>It is normal for the prediction time to vary depending on the learned weights of the model. Optimization and weight updates can affect the model’s performance.</p>\n<p>When you see an increase in CPU usage, this situation can depend on several factors:</p>\n<p>Optimization Algorithms: Optimization algorithms can work at different speeds when updating weights. This can affect CPU usage.<br>\nData Size and Processing Load: The size of the training data and the processing load can affect CPU usage. Larger datasets or more complex processing loads can use the CPU more.<br>\nHyperparameter Settings: Hyperparameter settings (for example, learning rate, momentum, regularization term) can also affect CPU usage.<br>\nHardware Differences: Different hardware (CPU type, memory, disk speed, etc.) can trigger different CPU usage. When you see that GPU usage eliminates the difference in prediction time, it shows that the GPU benefits from parallel computing capabilities. The GPU can accelerate large matrix operations and help the model run faster.<br>\nSince the state_dict is saved and loaded, nothing should change except for the weights and bias values. However, this does not affect CPU usage. You should investigate other factors for changes in CPU usage.</p>\n<p>If you need more information about the change in CPU usage, it would be useful to check your hardware configuration and review the optimization steps. Also, continuing to run your model on the GPU can improve performance🥇</p>",
      "votes": null,
      "replies": [
        {
          "id": 2788924,
          "author_name": "madao13",
          "author_url": "",
          "post_date": "05/02/2024 13:32:20",
          "content": "<p>Thank you.<br>\nI understand that when learning, there can be a difference in learning time depending on the optimizer and other settings.</p>\n<p><strong>&gt;Since the state_dict is saved and loaded, nothing should change except for the weights and bias values.</strong><br>\nThe state_dict should only store the weights and bias information.<br>\nTherefore, I would expect the processing time to be the same for both learned parameters…</p>\n<p>By the way, I thought the first learned parameter had more zeros, so I investigated, but the second learned parameter had more zeros.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2789239,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/02/2024 16:07:30",
      "content": "<p>You did more than retrainnig. You also added some metric computation with torch. Why do you assume that the slowdown is not due to the latter?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2790614,
          "author_name": "madao13",
          "author_url": "",
          "post_date": "05/03/2024 09:11:04",
          "content": "<p>Thank you.<br>\nSorry for the lack of explanation.</p>\n<p>I created separate notebooks for the study phase and the submission phase.<br>\nIn the learning phase, we added the evaluation metrics for the model, but in the evaluation phase in question, we did not change anything except the state_dict to be loaded.<br>\nWhile the model is the same and the size of parameters to be loaded is the same, there is a difference in processing time.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2790677,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "05/03/2024 09:45:28",
              "content": "<p>OK. This is the same code, the only change are the weights?</p>\n<p>Maybe you are close to RAM capacity or something, end depending on race conditions, you may get slowdown because you are using the compute resource at its limit?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2791049,
                  "author_name": "madao13",
                  "author_url": "",
                  "post_date": "05/03/2024 13:00:16",
                  "content": "<p>The only thing that has changed is the weight.</p>\n<p><strong>&gt;Maybe you are close to RAM capacity or something, end depending on race conditions, you may get slowdown because you are using the compute resource at its limit?</strong><br>\nThank you. I will look into the status of the various resources.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2792859,
      "author_name": "kdmitrie",
      "author_url": "",
      "post_date": "05/04/2024 12:40:14",
      "content": "<p>I had a similar situation. The inference time of a trained model was about 60 sec per 4-min record. Much too slow!</p>\n<p>I tried to figure out, why, and I ended up with a single PyTorch Conv2d layer. The inference time of this layer was about 3 sec, while the same layer initialized with random weights had inference time of 0.06 sec. No matter what I did, when the trained weights were used, the layer became 'slow'.</p>\n<p>The solution was to use <code>model.half().float()</code>, i.e. to round the weights of the model. Actually, I don't understand, why it works, because the type of weight tensor remains <code>float32</code>. Somehow, this significantly increased the speed of computations, while the results produced by the model are almost the same.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2793082,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/04/2024 14:25:24",
          "content": "<p>Interesting. rounding your weights to fp16 precision helps. Are you using onnx or openvino? Maybe there is something with these that make them run FP16 when possible? I am speculating, I have absolutely no idea of why you speed your computation that way.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2793192,
              "author_name": "kdmitrie",
              "author_url": "",
              "post_date": "05/04/2024 15:39:48",
              "content": "<p>What I'm reporting here is a pure PyTorch behaviour, without onnx/openvino.</p>\n<p>Just two layers: first initialized randomly, second - with <code>load_state_dict()</code>. I also tried to copy weight Parameter directly, or to convert it to numpy first, then back to tensor and Parameter. The result was the same: the weights trained on GPU made a layer extremely slow. I don't know exactly, but I suppose, there's some hidden data type conversion. However, I didn't find, where it happens.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2793162,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "05/04/2024 15:23:14",
          "content": "<p>If some of the model weights end up very small (e.g ~1e-30) during training, the inference time could become very large, it is because of the how CPU handles float computatitions in that range to maintain the required precision (that might be not true, but worth investigating, unfortunately I forgot the source of this information).</p>\n<p>Can you check the norm of the weights in the area around 0 before and after rounding?</p>\n<p>Most likely you'll be able to get a dependence in the chart \"lowest non-zero weight norm\"/inference_time</p>",
          "votes": null,
          "replies": [
            {
              "id": 2793257,
              "author_name": "kdmitrie",
              "author_url": "",
              "post_date": "05/04/2024 16:18:49",
              "content": "<p>Thank you, this may be the reason. I saw the weights like 1e-44, indeed, but didn't think about them in this way.</p>\n<p>While rounding the weights helps to solve the problem, it seems that adding L1 regularization should prevent it from even appearing.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2793781,
                  "author_name": "madao13",
                  "author_url": "",
                  "post_date": "05/04/2024 23:49:51",
                  "content": "<p>Thank you!<br>\nI was surprised to find that the computation time can vary greatly depending on the weights for the same size.</p>\n<p>I will work on improving the model using the \"model.half().float()\" and other methods you mentioned.<br>\nFrom now on, I will work on learning the model while being aware of the value of the weights!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2794415,
                      "author_name": "kdmitrie",
                      "author_url": "",
                      "post_date": "05/05/2024 09:22:40",
                      "content": "<p>Just let me know, if this actually helps, I'm curious about it.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2794634,
                          "author_name": "madao13",
                          "author_url": "",
                          "post_date": "05/05/2024 12:19:01",
                          "content": "<p>The values of the front and rear weights were checked.<br>\nThe first learned parameter was 1e-8 with the smallest weight excluding 0, while the second learned parameter was e-45 with the smallest!</p>\n<p>Also, the second learned parameter was always less than 1e-40 in almost all layers.</p>\n<p>Using \"model.half().float()\", weights of the order of 1e-45 were reduced to the order of 1e-8, and the calculation was faster.<br>\nMin: 1.401298464324817e-45(Before using model.half().float())<br>\n↓↓<br>\nMin: 5.960464477539063e-08(After using model.half().float())</p>\n<p>Thank you all!<br>\nI am relieved that I found the cause!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2800067,
      "author_name": "srinath9s",
      "author_url": "",
      "post_date": "05/08/2024 05:08:14",
      "content": "<p>Has anyone experimented training with 16bit precision?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2788152": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16858852%2F2f388a81258434ff90e224610c5332bb%2Fcontents.png?generation=1714637198131654&alt=media)I am kaggle beginner. I need help.\n\nThe prediction time differs greatly depending on the weights learned for the same model.\n\nIn the CLEF2024 competition, when the model was submitted with the first learned weights, it was within 2 hours without any problem,\nHowever, when the weights were improved using the optimizer's scheduler, the prediction time increased by a factor of 5.\n\nWe investigated and found that CPU usage seemed to be increasing.\nWhen we tried forecasting on the GPU, the difference in forecasting time between the previous and current learning weights disappeared.\n\nSince it is a save and load of state_dict(), nothing should have changed except the weights and bias values.\nIs it possible for CPU usage to change in this situation?",
    "2788746": "It is normal for the prediction time to vary depending on the learned weights of the model. Optimization and weight updates can affect the model’s performance.\n\nWhen you see an increase in CPU usage, this situation can depend on several factors:\n\nOptimization Algorithms: Optimization algorithms can work at different speeds when updating weights. This can affect CPU usage.\nData Size and Processing Load: The size of the training data and the processing load can affect CPU usage. Larger datasets or more complex processing loads can use the CPU more.\nHyperparameter Settings: Hyperparameter settings (for example, learning rate, momentum, regularization term) can also affect CPU usage.\nHardware Differences: Different hardware (CPU type, memory, disk speed, etc.) can trigger different CPU usage. When you see that GPU usage eliminates the difference in prediction time, it shows that the GPU benefits from parallel computing capabilities. The GPU can accelerate large matrix operations and help the model run faster.\nSince the state_dict is saved and loaded, nothing should change except for the weights and bias values. However, this does not affect CPU usage. You should investigate other factors for changes in CPU usage.\n\nIf you need more information about the change in CPU usage, it would be useful to check your hardware configuration and review the optimization steps. Also, continuing to run your model on the GPU can improve performance🥇",
    "2788924": "Thank you.\nI understand that when learning, there can be a difference in learning time depending on the optimizer and other settings.\n\n**>Since the state_dict is saved and loaded, nothing should change except for the weights and bias values.**\nThe state_dict should only store the weights and bias information.\nTherefore, I would expect the processing time to be the same for both learned parameters...\n\nBy the way, I thought the first learned parameter had more zeros, so I investigated, but the second learned parameter had more zeros.",
    "2789239": "You did more than retrainnig. You also added some metric computation with torch. Why do you assume that the slowdown is not due to the latter?",
    "2790614": "Thank you.\nSorry for the lack of explanation.\n\nI created separate notebooks for the study phase and the submission phase.\nIn the learning phase, we added the evaluation metrics for the model, but in the evaluation phase in question, we did not change anything except the state_dict to be loaded.\nWhile the model is the same and the size of parameters to be loaded is the same, there is a difference in processing time.",
    "2790677": "OK. This is the same code, the only change are the weights?\n\nMaybe you are close to RAM capacity or something, end depending on race conditions, you may get slowdown because you are using the compute resource at its limit?",
    "2791049": "The only thing that has changed is the weight.\n\n**>Maybe you are close to RAM capacity or something, end depending on race conditions, you may get slowdown because you are using the compute resource at its limit?**\nThank you. I will look into the status of the various resources.",
    "2792859": "I had a similar situation. The inference time of a trained model was about 60 sec per 4-min record. Much too slow!\n\nI tried to figure out, why, and I ended up with a single PyTorch Conv2d layer. The inference time of this layer was about 3 sec, while the same layer initialized with random weights had inference time of 0.06 sec. No matter what I did, when the trained weights were used, the layer became 'slow'.\n\nThe solution was to use `model.half().float()`, i.e. to round the weights of the model. Actually, I don't understand, why it works, because the type of weight tensor remains `float32`. Somehow, this significantly increased the speed of computations, while the results produced by the model are almost the same.",
    "2793082": "Interesting. rounding your weights to fp16 precision helps. Are you using onnx or openvino? Maybe there is something with these that make them run FP16 when possible? I am speculating, I have absolutely no idea of why you speed your computation that way.",
    "2793162": "If some of the model weights end up very small (e.g ~1e-30) during training, the inference time could become very large, it is because of the how CPU handles float computatitions in that range to maintain the required precision (that might be not true, but worth investigating, unfortunately I forgot the source of this information).\n\nCan you check the norm of the weights in the area around 0 before and after rounding?\n\nMost likely you'll be able to get a dependence in the chart \"lowest non-zero weight norm\"/inference_time",
    "2793192": "What I'm reporting here is a pure PyTorch behaviour, without onnx/openvino.\n\nJust two layers: first initialized randomly, second - with `load_state_dict()`. I also tried to copy weight Parameter directly, or to convert it to numpy first, then back to tensor and Parameter. The result was the same: the weights trained on GPU made a layer extremely slow. I don't know exactly, but I suppose, there's some hidden data type conversion. However, I didn't find, where it happens.",
    "2793257": "Thank you, this may be the reason. I saw the weights like 1e-44, indeed, but didn't think about them in this way.\n\nWhile rounding the weights helps to solve the problem, it seems that adding L1 regularization should prevent it from even appearing.",
    "2793781": "Thank you!\nI was surprised to find that the computation time can vary greatly depending on the weights for the same size.\n\nI will work on improving the model using the \"model.half().float()\" and other methods you mentioned.\nFrom now on, I will work on learning the model while being aware of the value of the weights!",
    "2794415": "Just let me know, if this actually helps, I'm curious about it.",
    "2794634": "The values of the front and rear weights were checked.\nThe first learned parameter was 1e-8 with the smallest weight excluding 0, while the second learned parameter was e-45 with the smallest!\n\nAlso, the second learned parameter was always less than 1e-40 in almost all layers.\n\nUsing \"model.half().float()\", weights of the order of 1e-45 were reduced to the order of 1e-8, and the calculation was faster.\nMin: 1.401298464324817e-45(Before using model.half().float())\n↓↓\nMin: 5.960464477539063e-08(After using model.half().float())\n\n\nThank you all!\nI am relieved that I found the cause!",
    "2800067": "Has anyone experimented training with 16bit precision?"
  },
  "source": "meta"
}