{
  "id": 111904,
  "title": "Some tricks to train faster (pytorch)",
  "url": "/competitions/understanding_cloud_organization/discussion/111904",
  "author_name": "Hanke Chen",
  "post_date": "2019-10-09T21:17:00.873000",
  "votes": 22,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>Some tricks to train faster (pytorch)</h1>\n\n<h2>Data Storage</h2>\n\n<p>If you have an SSD, then you can store your data on the SSD instead of HDD. Loading from <code>.npy</code> files can also be faster than from <code>.png</code> because it avoids CPU conversion. You can preprocess the dataset into some <code>.npy</code> matrix before training.</p>\n\n<h2>Code writing style</h2>\n\n<p>make sure you write code like this:\n<code>\ndef forward(x):\n  x = conv1(x)\n  x = conv2(x)\n  ...\n  return x\n</code>\nnot like this:\n<code>\ndef forward(x):\n  x1 = conv1(x)\n  x2 = conv2(x1)\n  ... return x34\n</code>\nif your variable do not need to be stored.\nAlso, using <code>nn.ReLU(inplace=True)</code> might reduce memory usage, but it might also introduce an extra step of gradient calculation.</p>\n\n<h2>Augmentation</h2>\n\n<p>Some augmentation library is faster than others. See: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1593699%2Faadd28a4c767ed1aaca953b5059b060a%2Ftmp.png?generation=1570654793513080&amp;alt=media\" alt=\"table from albunmentation\">\n(measured in img)\nChoose wisely.</p>\n\n<h2>Training Process</h2>\n\n<p>Check your code. Make sure you do not take data out of GPU to CPU if it is not necessary. Taking a mask out instead of your metrics can cost you a lot of time! (almost twice as slow)</p>\n\n<h2>Half Precision Training</h2>\n\n<p>Use APEX for FP16 training: <a href=\"https://github.com/NVIDIA/apex\">https://github.com/NVIDIA/apex</a>\nAPEX also supports multi-GPU parallel.\nIf you have trouble installing APEX, you can see my <a href=\"https://zhuanlan.zhihu.com/p/80386137\">知乎 if you read Chinese</a> or comment below.\n.\n.\n.\nHope it helps\nFeel free to correct me or give us some extra suggestions :)</p>",
  "messages": [
    {
      "id": 645149,
      "postDate": "2019-10-09T21:17:00.873Z",
      "content": "<h1>Some tricks to train faster (pytorch)</h1>\n\n<h2>Data Storage</h2>\n\n<p>If you have an SSD, then you can store your data on the SSD instead of HDD. Loading from <code>.npy</code> files can also be faster than from <code>.png</code> because it avoids CPU conversion. You can preprocess the dataset into some <code>.npy</code> matrix before training.</p>\n\n<h2>Code writing style</h2>\n\n<p>make sure you write code like this:\n<code>\ndef forward(x):\n  x = conv1(x)\n  x = conv2(x)\n  ...\n  return x\n</code>\nnot like this:\n<code>\ndef forward(x):\n  x1 = conv1(x)\n  x2 = conv2(x1)\n  ... return x34\n</code>\nif your variable do not need to be stored.\nAlso, using <code>nn.ReLU(inplace=True)</code> might reduce memory usage, but it might also introduce an extra step of gradient calculation.</p>\n\n<h2>Augmentation</h2>\n\n<p>Some augmentation library is faster than others. See: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1593699%2Faadd28a4c767ed1aaca953b5059b060a%2Ftmp.png?generation=1570654793513080&amp;alt=media\" alt=\"table from albunmentation\">\n(measured in img)\nChoose wisely.</p>\n\n<h2>Training Process</h2>\n\n<p>Check your code. Make sure you do not take data out of GPU to CPU if it is not necessary. Taking a mask out instead of your metrics can cost you a lot of time! (almost twice as slow)</p>\n\n<h2>Half Precision Training</h2>\n\n<p>Use APEX for FP16 training: <a href=\"https://github.com/NVIDIA/apex\">https://github.com/NVIDIA/apex</a>\nAPEX also supports multi-GPU parallel.\nIf you have trouble installing APEX, you can see my <a href=\"https://zhuanlan.zhihu.com/p/80386137\">知乎 if you read Chinese</a> or comment below.\n.\n.\n.\nHope it helps\nFeel free to correct me or give us some extra suggestions :)</p>",
      "rawMarkdown": "# Some tricks to train faster (pytorch)\n## Data Storage\nIf you have an SSD, then you can store your data on the SSD instead of HDD. Loading from `.npy` files can also be faster than from `.png` because it avoids CPU conversion. You can preprocess the dataset into some `.npy` matrix before training.\n\n## Code writing style\nmake sure you write code like this:\n```\ndef forward(x):\n  x = conv1(x)\n  x = conv2(x)\n  ...\n  return x\n```\nnot like this:\n```\ndef forward(x):\n  x1 = conv1(x)\n  x2 = conv2(x1)\n  ... return x34\n```\nif your variable do not need to be stored.\nAlso, using `nn.ReLU(inplace=True)` might reduce memory usage, but it might also introduce an extra step of gradient calculation.\n\n## Augmentation\nSome augmentation library is faster than others. See: \n![table from albunmentation](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1593699%2Faadd28a4c767ed1aaca953b5059b060a%2Ftmp.png?generation=1570654793513080&amp;alt=media)\n(measured in img)\nChoose wisely.\n\n## Training Process\nCheck your code. Make sure you do not take data out of GPU to CPU if it is not necessary. Taking a mask out instead of your metrics can cost you a lot of time! (almost twice as slow)\n\n## Half Precision Training\nUse APEX for FP16 training: https://github.com/NVIDIA/apex\nAPEX also supports multi-GPU parallel.\nIf you have trouble installing APEX, you can see my [知乎 if you read Chinese](https://zhuanlan.zhihu.com/p/80386137) or comment below.\n.\n.\n.\nHope it helps\nFeel free to correct me or give us some extra suggestions :)",
      "votes": 22
    },
    {
      "id": 913635,
      "postDate": "2020-07-03T10:23:35.223Z",
      "content": "<p>The data in the augmentation table means consuming time? So lesser is better? Thanks for clarification.</p>\n\n<p>I tried saving image in .npy file but I found the size is the largest compared to .png and .jpeg. The file size is ranked as follows.\n.jpeg &lt; .png &lt; .npy\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2663988%2F5e6072fe315a969b9762528f67823699%2FScreen%20Shot%202020-07-03%20at%208.25.27%20PM.png?generation=1593771967283912&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The data in the augmentation table means consuming time? So lesser is better? Thanks for clarification.\n\nI tried saving image in .npy file but I found the size is the largest compared to .png and .jpeg. The file size is ranked as follows.\n.jpeg &lt; .png &lt; .npy\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2663988%2F5e6072fe315a969b9762528f67823699%2FScreen%20Shot%202020-07-03%20at%208.25.27%20PM.png?generation=1593771967283912&amp;alt=media)\n",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 645649,
      "postDate": "2019-10-10T10:43:01.150Z",
      "content": "<p>Hi <a href=\"/kokecacao\">@kokecacao</a> , thanks for your sharing</p>",
      "rawMarkdown": "Hi @kokecacao , thanks for your sharing",
      "votes": 1
    },
    {
      "id": 645440,
      "postDate": "2019-10-10T06:05:36.370Z",
      "content": "<p>Very Helpful Tricks\nThanks <a href=\"/kokecacao\">@kokecacao</a> </p>",
      "rawMarkdown": "Very Helpful Tricks\nThanks @kokecacao ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 913635,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-03T10:23:35.223000",
      "content": "<p>The data in the augmentation table means consuming time? So lesser is better? Thanks for clarification.</p>\n\n<p>I tried saving image in .npy file but I found the size is the largest compared to .png and .jpeg. The file size is ranked as follows.\n.jpeg &lt; .png &lt; .npy\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2663988%2F5e6072fe315a969b9762528f67823699%2FScreen%20Shot%202020-07-03%20at%208.25.27%20PM.png?generation=1593771967283912&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 645649,
      "author_name": "cswwp",
      "author_url": "",
      "post_date": "2019-10-10T10:43:01.150000",
      "content": "<p>Hi <a href=\"/kokecacao\">@kokecacao</a> , thanks for your sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 645440,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-10-10T06:05:36.370000",
      "content": "<p>Very Helpful Tricks\nThanks <a href=\"/kokecacao\">@kokecacao</a> </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "645149": "# Some tricks to train faster (pytorch)\n## Data Storage\nIf you have an SSD, then you can store your data on the SSD instead of HDD. Loading from `.npy` files can also be faster than from `.png` because it avoids CPU conversion. You can preprocess the dataset into some `.npy` matrix before training.\n\n## Code writing style\nmake sure you write code like this:\n```\ndef forward(x):\n  x = conv1(x)\n  x = conv2(x)\n  ...\n  return x\n```\nnot like this:\n```\ndef forward(x):\n  x1 = conv1(x)\n  x2 = conv2(x1)\n  ... return x34\n```\nif your variable do not need to be stored.\nAlso, using `nn.ReLU(inplace=True)` might reduce memory usage, but it might also introduce an extra step of gradient calculation.\n\n## Augmentation\nSome augmentation library is faster than others. See: \n![table from albunmentation](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1593699%2Faadd28a4c767ed1aaca953b5059b060a%2Ftmp.png?generation=1570654793513080&amp;alt=media)\n(measured in img)\nChoose wisely.\n\n## Training Process\nCheck your code. Make sure you do not take data out of GPU to CPU if it is not necessary. Taking a mask out instead of your metrics can cost you a lot of time! (almost twice as slow)\n\n## Half Precision Training\nUse APEX for FP16 training: https://github.com/NVIDIA/apex\nAPEX also supports multi-GPU parallel.\nIf you have trouble installing APEX, you can see my [知乎 if you read Chinese](https://zhuanlan.zhihu.com/p/80386137) or comment below.\n.\n.\n.\nHope it helps\nFeel free to correct me or give us some extra suggestions :)",
    "913635": "The data in the augmentation table means consuming time? So lesser is better? Thanks for clarification.\n\nI tried saving image in .npy file but I found the size is the largest compared to .png and .jpeg. The file size is ranked as follows.\n.jpeg &lt; .png &lt; .npy\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2663988%2F5e6072fe315a969b9762528f67823699%2FScreen%20Shot%202020-07-03%20at%208.25.27%20PM.png?generation=1593771967283912&amp;alt=media)\n",
    "645649": "Hi @kokecacao , thanks for your sharing",
    "645440": "Very Helpful Tricks\nThanks @kokecacao "
  }
}