{
  "id": 152995,
  "title": "Solution for  ***Out-of-memory issues***",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/152995",
  "author_name": "",
  "post_date": "2020-05-22T15:54:20.674975400Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>We need to adjust below hyper parameter</p>\n\n<ul>\n<li><p>max_seq_length: Sequence lengths &lt;  512, Use a shorter max sequence length to save substantial memory. </p></li>\n<li><p>train_batch_size:  Memory usage directly proportional to batch size.</p></li>\n<li><p>Use BERT Base :  The BERT-Large model requires significantly more memory than BERT-Base.</p></li>\n<li><p>Optimizer: The default optimizer for BERT is Adam, which requires a lot of extra memory to store the m and v vectors. Try other optimizer can reduce memory usage,</p>\n\n<ul><li>RAM : TensorFlow GPU (minimum 12GB RAM) with TensorFlow 1.11.0:</li></ul></li>\n</ul>\n\n<p>System  BERT-Base\n Seq Length -  Max Batch Size\n     64         -        64\n    128        -         32\n   256         -         16\n   320        -          14\n   384         -         12\n   512        -           6\nBERT-Large <br>\n      Seq Length    -     Max Batch Size\n           64                 -           12\n          128        -            6\n          256        -             2\n          320        -            1\n          384        -            0\n          512        -            0</p>",
  "messages": [
    {
      "id": "857431",
      "postDate": "05/22/2020 15:54:20",
      "content": "<p>We need to adjust below hyper parameter</p>\n\n<ul>\n<li><p>max_seq_length: Sequence lengths &lt;  512, Use a shorter max sequence length to save substantial memory. </p></li>\n<li><p>train_batch_size:  Memory usage directly proportional to batch size.</p></li>\n<li><p>Use BERT Base :  The BERT-Large model requires significantly more memory than BERT-Base.</p></li>\n<li><p>Optimizer: The default optimizer for BERT is Adam, which requires a lot of extra memory to store the m and v vectors. Try other optimizer can reduce memory usage,</p>\n\n<ul><li>RAM : TensorFlow GPU (minimum 12GB RAM) with TensorFlow 1.11.0:</li></ul></li>\n</ul>\n\n<p>System  BERT-Base\n Seq Length -  Max Batch Size\n     64         -        64\n    128        -         32\n   256         -         16\n   320        -          14\n   384         -         12\n   512        -           6\nBERT-Large <br>\n      Seq Length    -     Max Batch Size\n           64                 -           12\n          128        -            6\n          256        -             2\n          320        -            1\n          384        -            0\n          512        -            0</p>",
      "rawMarkdown": "We need to adjust below hyper parameter\n\n- max_seq_length: Sequence lengths &lt;  512, Use a shorter max sequence length to save substantial memory. \n\n- train_batch_size:  Memory usage directly proportional to batch size.\n\n- Use BERT Base :  The BERT-Large model requires significantly more memory than BERT-Base.\n\n-  Optimizer: The default optimizer for BERT is Adam, which requires a lot of extra memory to store the m and v vectors. Try other optimizer can reduce memory usage,\n\n - RAM : TensorFlow GPU (minimum 12GB RAM) with TensorFlow 1.11.0:\n\nSystem  BERT-Base\n Seq Length\t-  Max Batch Size\n     64\t        -        64\n    128\t       -         32\n   256\t       -         16\n   320\t      -          14\n   384\t       -         12\n   512\t      -           6\nBERT-Large\t        \n      Seq Length\t-     Max Batch Size\n           64\t              -           12\n          128\t     -            6\n          256\t     -             2\n          320\t     -            1\n          384\t     -            0\n          512\t     -            0",
      "votes": null
    },
    {
      "id": "857723",
      "postDate": "05/22/2020 21:25:24",
      "content": "<p>great, thanks for sharing.</p>",
      "rawMarkdown": "great, thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 857723,
      "author_name": "rohitsingh9990",
      "author_url": "",
      "post_date": "05/22/2020 21:25:24",
      "content": "<p>great, thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "857431": "We need to adjust below hyper parameter\n\n- max_seq_length: Sequence lengths &lt;  512, Use a shorter max sequence length to save substantial memory. \n\n- train_batch_size:  Memory usage directly proportional to batch size.\n\n- Use BERT Base :  The BERT-Large model requires significantly more memory than BERT-Base.\n\n-  Optimizer: The default optimizer for BERT is Adam, which requires a lot of extra memory to store the m and v vectors. Try other optimizer can reduce memory usage,\n\n - RAM : TensorFlow GPU (minimum 12GB RAM) with TensorFlow 1.11.0:\n\nSystem  BERT-Base\n Seq Length\t-  Max Batch Size\n     64\t        -        64\n    128\t       -         32\n   256\t       -         16\n   320\t      -          14\n   384\t       -         12\n   512\t      -           6\nBERT-Large\t        \n      Seq Length\t-     Max Batch Size\n           64\t              -           12\n          128\t     -            6\n          256\t     -             2\n          320\t     -            1\n          384\t     -            0\n          512\t     -            0",
    "857723": "great, thanks for sharing."
  },
  "source": "meta"
}