{
  "id": 239750,
  "title": "Colab version TPU efficientNet v2 pipeline available",
  "url": "/competitions/bms-molecular-translation/discussion/239750",
  "author_name": "",
  "post_date": "2021-05-17T14:16:15.994635800Z",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Credit to Darien Schettler. The original notebook is<br>\n<a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs</a></p>\n<p>I just modified some dataset path so that It can be run on Colab.<br>\nIt seems much faster than efficientnet V1.</p>\n<p><a href=\"https://github.com/flydragon2018/bms_tpu_colab_demo/blob/main/_bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs_colab.ipynb\" target=\"_blank\">https://github.com/flydragon2018/bms_tpu_colab_demo/blob/main/_bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs_colab.ipynb</a></p>",
  "messages": [
    {
      "id": "1311630",
      "postDate": "05/17/2021 14:16:15",
      "content": "<p>Credit to Darien Schettler. The original notebook is<br>\n<a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs</a></p>\n<p>I just modified some dataset path so that It can be run on Colab.<br>\nIt seems much faster than efficientnet V1.</p>\n<p><a href=\"https://github.com/flydragon2018/bms_tpu_colab_demo/blob/main/_bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs_colab.ipynb\" target=\"_blank\">https://github.com/flydragon2018/bms_tpu_colab_demo/blob/main/_bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs_colab.ipynb</a></p>",
      "rawMarkdown": "Credit to Darien Schettler. The original notebook is\nhttps://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\n\nI just modified some dataset path so that It can be run on Colab.\nIt seems much faster than efficientnet V1.\n\nhttps://github.com/flydragon2018/bms_tpu_colab_demo/blob/main/_bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs_colab.ipynb",
      "votes": null
    },
    {
      "id": "1314252",
      "postDate": "05/19/2021 04:12:31",
      "content": "<p>I trained once, got LB 3.93, fine-tuned 2nd time,  NaN appears after 22 epochs.  Loaded previously good encoder/decoder, got LB 3.18.   However, I tried a few more times fine-tune training, NaN values appear soon.<br>\nI am not sure due to learning rate, or the mixed precision config. Any one would point out?  </p>",
      "rawMarkdown": "I trained once, got LB 3.93, fine-tuned 2nd time,  NaN appears after 22 epochs.  Loaded previously good encoder/decoder, got LB 3.18.   However, I tried a few more times fine-tune training, NaN values appear soon.\nI am not sure due to learning rate, or the mixed precision config. Any one would point out?",
      "votes": null
    },
    {
      "id": "1314742",
      "postDate": "05/19/2021 11:06:01",
      "content": "<p>efficientnet v2 B2, fine-tuned again, LB just improve to 3.05.  perhaps big models can do better?</p>",
      "rawMarkdown": "efficientnet v2 B2, fine-tuned again, LB just improve to 3.05.  perhaps big models can do better?",
      "votes": null
    },
    {
      "id": "1316196",
      "postDate": "05/20/2021 11:15:53",
      "content": "<p>Thanks for sharing the Colab version.<br>\nBy the way, it seems that fine-tuning the model again will improve the result.<br>\nDo you use the same dataset to fine-tune?</p>",
      "rawMarkdown": "Thanks for sharing the Colab version.\nBy the way, it seems that fine-tuning the model again will improve the result.\nDo you use the same dataset to fine-tune?",
      "votes": null
    },
    {
      "id": "1316448",
      "postDate": "05/20/2021 15:03:38",
      "content": "<p>same dataset.</p>\n<p>in case gcs_path does not work,  try to get the new one from kaggle and replace old. <br>\nIt happened to me.</p>",
      "rawMarkdown": "same dataset.\n\nin case gcs_path does not work,  try to get the new one from kaggle and replace old. \nIt happened to me.",
      "votes": null
    },
    {
      "id": "1316658",
      "postDate": "05/20/2021 18:52:00",
      "content": "<p>Did you try using LAMB with large batch size? I think that can help in training a higher number of epochs.<br>\nYou can refer this discussion:<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/240316\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/240316</a></p>",
      "rawMarkdown": "Did you try using LAMB with large batch size? I think that can help in training a higher number of epochs.\nYou can refer this discussion:\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/240316",
      "votes": null
    },
    {
      "id": "1316799",
      "postDate": "05/20/2021 22:20:49",
      "content": "<p>Thanks for sharing. Maybe try it later.</p>",
      "rawMarkdown": "Thanks for sharing. Maybe try it later.",
      "votes": null
    },
    {
      "id": "1316848",
      "postDate": "05/21/2021 00:05:17",
      "content": "<p>Got it, thanks.</p>",
      "rawMarkdown": "Got it, thanks.",
      "votes": null
    },
    {
      "id": "1317324",
      "postDate": "05/21/2021 10:21:21",
      "content": "<p>I tried to train from beginning using LAMB. however, NaN appears after acc reach 98.x%.  Temporarily I would like to push  Adam further. Perhaps the mixed-precision problem</p>",
      "rawMarkdown": "I tried to train from beginning using LAMB. however, NaN appears after acc reach 98.x%.  Temporarily I would like to push  Adam further. Perhaps the mixed-precision problem",
      "votes": null
    },
    {
      "id": "1322870",
      "postDate": "05/25/2021 18:47:40",
      "content": "<p>What's the batch size you used with LAMB ?</p>",
      "rawMarkdown": "What's the batch size you used with LAMB ?",
      "votes": null
    },
    {
      "id": "1322945",
      "postDate": "05/25/2021 20:11:17",
      "content": "<ol>\n<li>same as Adam.  perhaps model is too big. Now I switch from efficientnetv2-L to M.  However Colab is slower than Kaggle, also Colab is interactive, so It is time consuming.   </li>\n</ol>",
      "rawMarkdown": "64. same as Adam.  perhaps model is too big. Now I switch from efficientnetv2-L to M.  However Colab is slower than Kaggle, also Colab is interactive, so It is time consuming.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1314252,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "05/19/2021 04:12:31",
      "content": "<p>I trained once, got LB 3.93, fine-tuned 2nd time,  NaN appears after 22 epochs.  Loaded previously good encoder/decoder, got LB 3.18.   However, I tried a few more times fine-tune training, NaN values appear soon.<br>\nI am not sure due to learning rate, or the mixed precision config. Any one would point out?  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1314742,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "05/19/2021 11:06:01",
      "content": "<p>efficientnet v2 B2, fine-tuned again, LB just improve to 3.05.  perhaps big models can do better?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1316196,
          "author_name": "kuohsintu",
          "author_url": "",
          "post_date": "05/20/2021 11:15:53",
          "content": "<p>Thanks for sharing the Colab version.<br>\nBy the way, it seems that fine-tuning the model again will improve the result.<br>\nDo you use the same dataset to fine-tune?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316448,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "05/20/2021 15:03:38",
          "content": "<p>same dataset.</p>\n<p>in case gcs_path does not work,  try to get the new one from kaggle and replace old. <br>\nIt happened to me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316848,
          "author_name": "kuohsintu",
          "author_url": "",
          "post_date": "05/21/2021 00:05:17",
          "content": "<p>Got it, thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1316658,
      "author_name": "vikrant06",
      "author_url": "",
      "post_date": "05/20/2021 18:52:00",
      "content": "<p>Did you try using LAMB with large batch size? I think that can help in training a higher number of epochs.<br>\nYou can refer this discussion:<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/240316\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/240316</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1316799,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "05/20/2021 22:20:49",
          "content": "<p>Thanks for sharing. Maybe try it later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1317324,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "05/21/2021 10:21:21",
          "content": "<p>I tried to train from beginning using LAMB. however, NaN appears after acc reach 98.x%.  Temporarily I would like to push  Adam further. Perhaps the mixed-precision problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1322870,
          "author_name": "vikrant06",
          "author_url": "",
          "post_date": "05/25/2021 18:47:40",
          "content": "<p>What's the batch size you used with LAMB ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1322945,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "05/25/2021 20:11:17",
          "content": "<ol>\n<li>same as Adam.  perhaps model is too big. Now I switch from efficientnetv2-L to M.  However Colab is slower than Kaggle, also Colab is interactive, so It is time consuming.   </li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1311630": "Credit to Darien Schettler. The original notebook is\nhttps://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\n\nI just modified some dataset path so that It can be run on Colab.\nIt seems much faster than efficientnet V1.\n\nhttps://github.com/flydragon2018/bms_tpu_colab_demo/blob/main/_bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs_colab.ipynb",
    "1314252": "I trained once, got LB 3.93, fine-tuned 2nd time,  NaN appears after 22 epochs.  Loaded previously good encoder/decoder, got LB 3.18.   However, I tried a few more times fine-tune training, NaN values appear soon.\nI am not sure due to learning rate, or the mixed precision config. Any one would point out?",
    "1314742": "efficientnet v2 B2, fine-tuned again, LB just improve to 3.05.  perhaps big models can do better?",
    "1316196": "Thanks for sharing the Colab version.\nBy the way, it seems that fine-tuning the model again will improve the result.\nDo you use the same dataset to fine-tune?",
    "1316448": "same dataset.\n\nin case gcs_path does not work,  try to get the new one from kaggle and replace old. \nIt happened to me.",
    "1316658": "Did you try using LAMB with large batch size? I think that can help in training a higher number of epochs.\nYou can refer this discussion:\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/240316",
    "1316799": "Thanks for sharing. Maybe try it later.",
    "1316848": "Got it, thanks.",
    "1317324": "I tried to train from beginning using LAMB. however, NaN appears after acc reach 98.x%.  Temporarily I would like to push  Adam further. Perhaps the mixed-precision problem",
    "1322870": "What's the batch size you used with LAMB ?",
    "1322945": "64. same as Adam.  perhaps model is too big. Now I switch from efficientnetv2-L to M.  However Colab is slower than Kaggle, also Colab is interactive, so It is time consuming."
  },
  "source": "meta"
}