{
  "id": 77931,
  "title": "Unreasonable memory usage ",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77931",
  "author_name": "",
  "post_date": "2019-01-17T19:01:47.707578600Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>has anyone accounted the same problem?\nI've a kernel that gives me the same results at each run (model made in PyTorch). \nBut in one rune it use 6 GB of ram memory during data transformation process, another time 11 GB. (I'm looking at bottom ribbon with cpu and ram information). I reset kernel before each run but memory consumption is completely unreasonable and random. </p>",
  "messages": [
    {
      "id": "457588",
      "postDate": "01/17/2019 19:01:47",
      "content": "<p>Hi,</p>\n\n<p>has anyone accounted the same problem?\nI've a kernel that gives me the same results at each run (model made in PyTorch). \nBut in one rune it use 6 GB of ram memory during data transformation process, another time 11 GB. (I'm looking at bottom ribbon with cpu and ram information). I reset kernel before each run but memory consumption is completely unreasonable and random. </p>",
      "rawMarkdown": "Hi,\n\nhas anyone accounted the same problem?\nI've a kernel that gives me the same results at each run (model made in PyTorch). \nBut in one rune it use 6 GB of ram memory during data transformation process, another time 11 GB. (I'm looking at bottom ribbon with cpu and ram information). I reset kernel before each run but memory consumption is completely unreasonable and random.",
      "votes": null
    },
    {
      "id": "457714",
      "postDate": "01/18/2019 01:13:42",
      "content": "<p>Well a lot of deep learning models are pretty memory dependent, but the extent to which this is apparent is based on your choice of hyperparameters. If you're working with full batch gradient descent on a huge network with highly dimensional data (as text-based data almost always is, unless you've invested a great deal of time into feature engineering) the memory usage will spike during training at certain points during each epoch. Think of it, not only do you have to have all of the data loaded, you also have to keep all engineered features loaded as well (which can be massive if you're talking about n-grams created without discrimination), the entire model itself will remain in memory as will any data needed during training (which fluctuates based on where it is in each epoch, which itself is random because batch creation is inherently random unless you specify otherwise). </p>\n\n<p>Without seeing your exact code or getting more details it would be hard to identify the problem to recommend a solution, but since you're still under the 16GB memory limit Kaggle sets, it shouldn't be a huge problem. That being said, I've noticed similar issues of weird random memory usage within kernels when there would be no logical reason for it, so I don't think that the data displayed by the kernels is very robust. </p>",
      "rawMarkdown": "Well a lot of deep learning models are pretty memory dependent, but the extent to which this is apparent is based on your choice of hyperparameters. If you're working with full batch gradient descent on a huge network with highly dimensional data (as text-based data almost always is, unless you've invested a great deal of time into feature engineering) the memory usage will spike during training at certain points during each epoch. Think of it, not only do you have to have all of the data loaded, you also have to keep all engineered features loaded as well (which can be massive if you're talking about n-grams created without discrimination), the entire model itself will remain in memory as will any data needed during training (which fluctuates based on where it is in each epoch, which itself is random because batch creation is inherently random unless you specify otherwise). \n\nWithout seeing your exact code or getting more details it would be hard to identify the problem to recommend a solution, but since you're still under the 16GB memory limit Kaggle sets, it shouldn't be a huge problem. That being said, I've noticed similar issues of weird random memory usage within kernels when there would be no logical reason for it, so I don't think that the data displayed by the kernels is very robust.",
      "votes": null
    },
    {
      "id": "457918",
      "postDate": "01/18/2019 10:49:39",
      "content": "<p>@DanielSzponar collect unused objects to free memory. Use memory efficient tricks such as using <code>numpy</code> rather than list. You may find this link helpful <a href=\"https://www.kaggle.com/product-feedback/72606#latest=429470\">https://www.kaggle.com/product-feedback/72606#latest=429470</a></p>",
      "rawMarkdown": "DanielSzponar collect unused objects to free memory. Use memory efficient tricks such as using `numpy` rather than list. You may find this link helpful https://www.kaggle.com/product-feedback/72606#latest=429470",
      "votes": null
    },
    {
      "id": "457924",
      "postDate": "01/18/2019 10:59:45",
      "content": "<p>Thanks a lot! Indeed i'm using list to apply functions over text. I'll try different implementation.</p>",
      "rawMarkdown": "Thanks a lot! Indeed i'm using list to apply functions over text. I'll try different implementation.",
      "votes": null
    },
    {
      "id": "458142",
      "postDate": "01/18/2019 22:07:28",
      "content": "<p>Maybe are you storing a tensor (with all its internal dependencies) instead of the value of a tensor?</p>\n\n<p>If, for example, you store iteratively the loss function output while it is still a tensor, your memory will explode, as it will contain all the backpropagation derivatives. To avoid that, when you store it, simply do something like <code>loss.data</code>, or even I think you can do <code>float(loss)</code> assuming the tensor <code>loss</code> has only one element.</p>\n\n<p>Hope this helps.</p>",
      "rawMarkdown": "Maybe are you storing a tensor (with all its internal dependencies) instead of the value of a tensor?\n\nIf, for example, you store iteratively the loss function output while it is still a tensor, your memory will explode, as it will contain all the backpropagation derivatives. To avoid that, when you store it, simply do something like `loss.data`, or even I think you can do `float(loss)` assuming the tensor `loss` has only one element.\n\nHope this helps.",
      "votes": null
    },
    {
      "id": "458350",
      "postDate": "01/19/2019 12:31:10",
      "content": "<p>Thanks. I'm using <code>loss.data</code> but It's really valuable explanations that I wasn't aware of.</p>",
      "rawMarkdown": "Thanks. I'm using ```loss.data``` but It's really valuable explanations that I wasn't aware of.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 457714,
      "author_name": "alecthekulak",
      "author_url": "",
      "post_date": "01/18/2019 01:13:42",
      "content": "<p>Well a lot of deep learning models are pretty memory dependent, but the extent to which this is apparent is based on your choice of hyperparameters. If you're working with full batch gradient descent on a huge network with highly dimensional data (as text-based data almost always is, unless you've invested a great deal of time into feature engineering) the memory usage will spike during training at certain points during each epoch. Think of it, not only do you have to have all of the data loaded, you also have to keep all engineered features loaded as well (which can be massive if you're talking about n-grams created without discrimination), the entire model itself will remain in memory as will any data needed during training (which fluctuates based on where it is in each epoch, which itself is random because batch creation is inherently random unless you specify otherwise). </p>\n\n<p>Without seeing your exact code or getting more details it would be hard to identify the problem to recommend a solution, but since you're still under the 16GB memory limit Kaggle sets, it shouldn't be a huge problem. That being said, I've noticed similar issues of weird random memory usage within kernels when there would be no logical reason for it, so I don't think that the data displayed by the kernels is very robust. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 457918,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "01/18/2019 10:49:39",
      "content": "<p>@DanielSzponar collect unused objects to free memory. Use memory efficient tricks such as using <code>numpy</code> rather than list. You may find this link helpful <a href=\"https://www.kaggle.com/product-feedback/72606#latest=429470\">https://www.kaggle.com/product-feedback/72606#latest=429470</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 457924,
          "author_name": "nicke1",
          "author_url": "",
          "post_date": "01/18/2019 10:59:45",
          "content": "<p>Thanks a lot! Indeed i'm using list to apply functions over text. I'll try different implementation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 458142,
      "author_name": "manuelsh",
      "author_url": "",
      "post_date": "01/18/2019 22:07:28",
      "content": "<p>Maybe are you storing a tensor (with all its internal dependencies) instead of the value of a tensor?</p>\n\n<p>If, for example, you store iteratively the loss function output while it is still a tensor, your memory will explode, as it will contain all the backpropagation derivatives. To avoid that, when you store it, simply do something like <code>loss.data</code>, or even I think you can do <code>float(loss)</code> assuming the tensor <code>loss</code> has only one element.</p>\n\n<p>Hope this helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 458350,
          "author_name": "nicke1",
          "author_url": "",
          "post_date": "01/19/2019 12:31:10",
          "content": "<p>Thanks. I'm using <code>loss.data</code> but It's really valuable explanations that I wasn't aware of.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "457588": "Hi,\n\nhas anyone accounted the same problem?\nI've a kernel that gives me the same results at each run (model made in PyTorch). \nBut in one rune it use 6 GB of ram memory during data transformation process, another time 11 GB. (I'm looking at bottom ribbon with cpu and ram information). I reset kernel before each run but memory consumption is completely unreasonable and random.",
    "457714": "Well a lot of deep learning models are pretty memory dependent, but the extent to which this is apparent is based on your choice of hyperparameters. If you're working with full batch gradient descent on a huge network with highly dimensional data (as text-based data almost always is, unless you've invested a great deal of time into feature engineering) the memory usage will spike during training at certain points during each epoch. Think of it, not only do you have to have all of the data loaded, you also have to keep all engineered features loaded as well (which can be massive if you're talking about n-grams created without discrimination), the entire model itself will remain in memory as will any data needed during training (which fluctuates based on where it is in each epoch, which itself is random because batch creation is inherently random unless you specify otherwise). \n\nWithout seeing your exact code or getting more details it would be hard to identify the problem to recommend a solution, but since you're still under the 16GB memory limit Kaggle sets, it shouldn't be a huge problem. That being said, I've noticed similar issues of weird random memory usage within kernels when there would be no logical reason for it, so I don't think that the data displayed by the kernels is very robust.",
    "457918": "DanielSzponar collect unused objects to free memory. Use memory efficient tricks such as using `numpy` rather than list. You may find this link helpful https://www.kaggle.com/product-feedback/72606#latest=429470",
    "457924": "Thanks a lot! Indeed i'm using list to apply functions over text. I'll try different implementation.",
    "458142": "Maybe are you storing a tensor (with all its internal dependencies) instead of the value of a tensor?\n\nIf, for example, you store iteratively the loss function output while it is still a tensor, your memory will explode, as it will contain all the backpropagation derivatives. To avoid that, when you store it, simply do something like `loss.data`, or even I think you can do `float(loss)` assuming the tensor `loss` has only one element.\n\nHope this helps.",
    "458350": "Thanks. I'm using ```loss.data``` but It's really valuable explanations that I wasn't aware of."
  },
  "source": "meta"
}