{
  "id": 139128,
  "title": "Stater Pytorch TPU Google Colab",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/139128",
  "author_name": "Gopi Durgaprasad",
  "post_date": "2020-03-27T14:50:38.735000",
  "votes": 8,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi every one this kernal based on <a href=\"https://www.kaggle.com/abhishek\"></a><a href=\"/abhishek\">@abhishek</a>  keranl <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores-w-valid\">bert-multi-lingual-tpu-training-8-cores-w-valid</a>  .</p>\n\n<p>Same kernal setup in google colab free TPU </p>\n\n<p>Notebook : <a href=\"https://colab.research.google.com/drive/1CQMC6N6ZvbR0eB_hcEOutSlBKjNZjqOt\">https://colab.research.google.com/drive/1CQMC6N6ZvbR0eB_hcEOutSlBKjNZjqOt</a></p>\n\n<p>In this colab notebook setup three things :</p>\n\n<ol>\n<li>Download kaggle  dataset to google colab.</li>\n<li>Run the <a href=\"/abhishek\">@abhishek</a> sir model </li>\n<li>Upload trained weights to the kaggle dataset.</li>\n</ol>\n\n<h3>Inference</h3>\n\n<p><a href=\"https://www.kaggle.com/gopidurgaprasad/stater-pytorch-tpu-google-colab\">https://www.kaggle.com/gopidurgaprasad/stater-pytorch-tpu-google-colab</a></p>",
  "messages": [
    {
      "id": 788297,
      "postDate": "2020-03-27T14:50:38.737Z",
      "content": "<p>Hi every one this kernal based on <a href=\"https://www.kaggle.com/abhishek\"></a><a href=\"/abhishek\">@abhishek</a>  keranl <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores-w-valid\">bert-multi-lingual-tpu-training-8-cores-w-valid</a>  .</p>\n\n<p>Same kernal setup in google colab free TPU </p>\n\n<p>Notebook : <a href=\"https://colab.research.google.com/drive/1CQMC6N6ZvbR0eB_hcEOutSlBKjNZjqOt\">https://colab.research.google.com/drive/1CQMC6N6ZvbR0eB_hcEOutSlBKjNZjqOt</a></p>\n\n<p>In this colab notebook setup three things :</p>\n\n<ol>\n<li>Download kaggle  dataset to google colab.</li>\n<li>Run the <a href=\"/abhishek\">@abhishek</a> sir model </li>\n<li>Upload trained weights to the kaggle dataset.</li>\n</ol>\n\n<h3>Inference</h3>\n\n<p><a href=\"https://www.kaggle.com/gopidurgaprasad/stater-pytorch-tpu-google-colab\">https://www.kaggle.com/gopidurgaprasad/stater-pytorch-tpu-google-colab</a></p>",
      "rawMarkdown": "Hi every one this kernal based on [@abhishek](https://www.kaggle.com/abhishek)  keranl [bert-multi-lingual-tpu-training-8-cores-w-valid](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores-w-valid)  .\n\nSame kernal setup in google colab free TPU \n\nNotebook : https://colab.research.google.com/drive/1CQMC6N6ZvbR0eB_hcEOutSlBKjNZjqOt\n\nIn this colab notebook setup three things :\n\n1. Download kaggle  dataset to google colab.\n2. Run the @abhishek sir model \n3. Upload trained weights to the kaggle dataset.\n\n### Inference\n\nhttps://www.kaggle.com/gopidurgaprasad/stater-pytorch-tpu-google-colab",
      "votes": 8
    },
    {
      "id": 802511,
      "postDate": "2020-04-09T15:26:43.437Z",
      "content": "<p>Hi,\nThanks for this information.\nTPU is giving OOM error every time I run it in colab ?....any suggestions \nPS: I even tried BATCH_SIZE = 8..... but no luck!</p>",
      "rawMarkdown": "Hi,\nThanks for this information.\nTPU is giving OOM error every time I run it in colab ?....any suggestions \nPS: I even tried BATCH_SIZE = 8..... but no luck!",
      "replies": [
        {
          "id": 802515,
          "postDate": "2020-04-09T15:30:27.340Z",
          "content": "<p>Which batch size? Train or valid? \nTrain batch size is inside the code hard-coded, probably which is not changed.</p>",
          "rawMarkdown": "Which batch size? Train or valid? \nTrain batch size is inside the code hard-coded, probably which is not changed."
        },
        {
          "id": 803385,
          "postDate": "2020-04-10T13:06:06.067Z",
          "content": "<p><a href=\"/namanj27\">@namanj27</a> Hai, Check my colab notebook above... you need to get more ram. for that need to run the first line of code and wait for full the ram then colab ask you \"getting more ram\" in below fo the notebook\ncheck carefully.</p>",
          "rawMarkdown": "@namanj27 Hai, Check my colab notebook above... you need to get more ram. for that need to run the first line of code and wait for full the ram then colab ask you \"getting more ram\" in below fo the notebook\ncheck carefully."
        },
        {
          "id": 803530,
          "postDate": "2020-04-10T16:14:23.257Z",
          "content": "<p>Hi <a href=\"/gopidurgaprasad\">@gopidurgaprasad</a>, \nThe problem I m facing there is with TPU memory which is 8GBs, nothing to do with RAM ( Although I did increase RAM using while loop )\nIts g giving me this error: </p>\n\n<blockquote>\n  <p>ResourceExhaustedError: Ran out of memory in memory space hbm; used: YYY; limit: 7.48G.</p>\n</blockquote>",
          "rawMarkdown": "Hi @gopidurgaprasad, \nThe problem I m facing there is with TPU memory which is 8GBs, nothing to do with RAM ( Although I did increase RAM using while loop )\nIts g giving me this error: \n&gt; ResourceExhaustedError: Ran out of memory in memory space hbm; used: YYY; limit: 7.48G.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 789062,
      "postDate": "2020-03-28T10:32:29.600Z",
      "content": "<p>I was getting OOM error on colab. Did you change any parameter of abhishek's kernel?</p>",
      "rawMarkdown": "I was getting OOM error on colab. Did you change any parameter of abhishek's kernel?",
      "replies": [
        {
          "id": 789091,
          "postDate": "2020-03-28T11:29:48.267Z",
          "content": "<p><a href=\"/shayekh\">@shayekh</a> I am not change anything make sure increase the ram for that need to run the first block of code</p>",
          "rawMarkdown": "@shayekh I am not change anything make sure increase the ram for that need to run the first block of code"
        },
        {
          "id": 789098,
          "postDate": "2020-03-28T11:41:12.160Z",
          "content": "<p>\"make sure increase the ram\"\nso, you are using the paid version of colab.</p>",
          "rawMarkdown": "\"make sure increase the ram\"\nso, you are using the paid version of colab."
        },
        {
          "id": 789152,
          "postDate": "2020-03-28T12:57:15.130Z",
          "content": "<p>No ... just run first block of code then colab ask you want to high ram </p>",
          "rawMarkdown": "No ... just run first block of code then colab ask you want to high ram ",
          "votes": 1
        },
        {
          "id": 789231,
          "postDate": "2020-03-28T13:39:56.780Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 789304,
          "postDate": "2020-03-28T14:39:56.010Z",
          "content": "<p>Not working still getting error. Thanks. I decreased train and validation batch size to 16.</p>",
          "rawMarkdown": "Not working still getting error. Thanks. I decreased train and validation batch size to 16.",
          "votes": 1
        },
        {
          "id": 789455,
          "postDate": "2020-03-28T17:00:14.120Z",
          "content": "<p>run this code for getting more \"RAM\"\n<code>\na = []\nwhile(1):\n    a.append(\"1\")\n</code></p>",
          "rawMarkdown": "run this code for getting more \"RAM\"\n`\na = []\nwhile(1):\n    a.append(\"1\")\n`\n"
        },
        {
          "id": 789468,
          "postDate": "2020-03-28T17:08:10.860Z",
          "content": "<p>Ya man. Tried a lot. TPU memory is the issue.\nProbably bad luck now. Thanks.</p>",
          "rawMarkdown": "Ya man. Tried a lot. TPU memory is the issue.\nProbably bad luck now. Thanks.",
          "votes": 1
        }
      ]
    },
    {
      "id": 788312,
      "postDate": "2020-03-27T15:02:07.633Z",
      "content": "<p>Why did you upload trained weights to the kaggle dataset (third step)?</p>",
      "rawMarkdown": "Why did you upload trained weights to the kaggle dataset (third step)?",
      "replies": [
        {
          "id": 788339,
          "postDate": "2020-03-27T15:27:44.210Z",
          "content": "<p><a href=\"/hiramcho\">@hiramcho</a> Because we need setup inference/prediction for the test data using trained weights</p>",
          "rawMarkdown": "@hiramcho Because we need setup inference/prediction for the test data using trained weights",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 802511,
      "author_name": "Naman Jaswani",
      "author_url": "",
      "post_date": "2020-04-09T15:26:43.437000",
      "content": "<p>Hi,\nThanks for this information.\nTPU is giving OOM error every time I run it in colab ?....any suggestions \nPS: I even tried BATCH_SIZE = 8..... but no luck!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 802515,
          "author_name": "Shayekh Islam",
          "author_url": "",
          "post_date": "2020-04-09T15:30:27.340000",
          "content": "<p>Which batch size? Train or valid? \nTrain batch size is inside the code hard-coded, probably which is not changed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 803385,
          "author_name": "Gopi Durgaprasad",
          "author_url": "",
          "post_date": "2020-04-10T13:06:06.067000",
          "content": "<p><a href=\"/namanj27\">@namanj27</a> Hai, Check my colab notebook above... you need to get more ram. for that need to run the first line of code and wait for full the ram then colab ask you \"getting more ram\" in below fo the notebook\ncheck carefully.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 803530,
          "author_name": "Naman Jaswani",
          "author_url": "",
          "post_date": "2020-04-10T16:14:23.257000",
          "content": "<p>Hi <a href=\"/gopidurgaprasad\">@gopidurgaprasad</a>, \nThe problem I m facing there is with TPU memory which is 8GBs, nothing to do with RAM ( Although I did increase RAM using while loop )\nIts g giving me this error: </p>\n\n<blockquote>\n  <p>ResourceExhaustedError: Ran out of memory in memory space hbm; used: YYY; limit: 7.48G.</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 789062,
      "author_name": "Shayekh Islam",
      "author_url": "",
      "post_date": "2020-03-28T10:32:29.600000",
      "content": "<p>I was getting OOM error on colab. Did you change any parameter of abhishek's kernel?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 789091,
          "author_name": "Gopi Durgaprasad",
          "author_url": "",
          "post_date": "2020-03-28T11:29:48.267000",
          "content": "<p><a href=\"/shayekh\">@shayekh</a> I am not change anything make sure increase the ram for that need to run the first block of code</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 789098,
          "author_name": "Shayekh Islam",
          "author_url": "",
          "post_date": "2020-03-28T11:41:12.160000",
          "content": "<p>\"make sure increase the ram\"\nso, you are using the paid version of colab.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 789152,
          "author_name": "Gopi Durgaprasad",
          "author_url": "",
          "post_date": "2020-03-28T12:57:15.130000",
          "content": "<p>No ... just run first block of code then colab ask you want to high ram </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 789231,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-28T13:39:56.780000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 789304,
          "author_name": "Shayekh Islam",
          "author_url": "",
          "post_date": "2020-03-28T14:39:56.010000",
          "content": "<p>Not working still getting error. Thanks. I decreased train and validation batch size to 16.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 789455,
          "author_name": "Gopi Durgaprasad",
          "author_url": "",
          "post_date": "2020-03-28T17:00:14.120000",
          "content": "<p>run this code for getting more \"RAM\"\n<code>\na = []\nwhile(1):\n    a.append(\"1\")\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 789468,
          "author_name": "Shayekh Islam",
          "author_url": "",
          "post_date": "2020-03-28T17:08:10.860000",
          "content": "<p>Ya man. Tried a lot. TPU memory is the issue.\nProbably bad luck now. Thanks.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 788312,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-03-27T15:02:07.633000",
      "content": "<p>Why did you upload trained weights to the kaggle dataset (third step)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 788339,
          "author_name": "Gopi Durgaprasad",
          "author_url": "",
          "post_date": "2020-03-27T15:27:44.210000",
          "content": "<p><a href=\"/hiramcho\">@hiramcho</a> Because we need setup inference/prediction for the test data using trained weights</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "788297": "Hi every one this kernal based on [@abhishek](https://www.kaggle.com/abhishek)  keranl [bert-multi-lingual-tpu-training-8-cores-w-valid](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores-w-valid)  .\n\nSame kernal setup in google colab free TPU \n\nNotebook : https://colab.research.google.com/drive/1CQMC6N6ZvbR0eB_hcEOutSlBKjNZjqOt\n\nIn this colab notebook setup three things :\n\n1. Download kaggle  dataset to google colab.\n2. Run the @abhishek sir model \n3. Upload trained weights to the kaggle dataset.\n\n### Inference\n\nhttps://www.kaggle.com/gopidurgaprasad/stater-pytorch-tpu-google-colab",
    "802511": "Hi,\nThanks for this information.\nTPU is giving OOM error every time I run it in colab ?....any suggestions \nPS: I even tried BATCH_SIZE = 8..... but no luck!",
    "789062": "I was getting OOM error on colab. Did you change any parameter of abhishek's kernel?",
    "788312": "Why did you upload trained weights to the kaggle dataset (third step)?"
  }
}