{
  "id": 238989,
  "title": "what is the difference between Kaggle TPU VM and Colab TPU VM",
  "url": "/competitions/bms-molecular-translation/discussion/238989",
  "author_name": "",
  "post_date": "2021-05-14T08:15:23.477594200Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I use the pretrained encoder/decoder for fine tune, it works on Kaggle, not far away with pretrained.<br>\nhowever, when I do same thing on Colab TPU,  it load the weights correctly, however the corresponding  val_loss and val_lsd are much bigger?</p>",
  "messages": [
    {
      "id": "1307053",
      "postDate": "05/14/2021 08:15:23",
      "content": "<p>I use the pretrained encoder/decoder for fine tune, it works on Kaggle, not far away with pretrained.<br>\nhowever, when I do same thing on Colab TPU,  it load the weights correctly, however the corresponding  val_loss and val_lsd are much bigger?</p>",
      "rawMarkdown": "I use the pretrained encoder/decoder for fine tune, it works on Kaggle, not far away with pretrained.\nhowever, when I do same thing on Colab TPU,  it load the weights correctly, however the corresponding  val_loss and val_lsd are much bigger?",
      "votes": null
    },
    {
      "id": "1307781",
      "postDate": "05/14/2021 16:49:28",
      "content": "<p>Yes there is an explanation. First of all metrics usually depends on batch size. Not much but it happens. <br>\nSecond: usually batch size is defined using number of cores i.e.:<br>\nBATCH_SIZE = 8 * strategy.num_replicas_in_sync<br>\nNow kaggle tpu has 8 cores while colab tpu only one. Therefore if you use same batch size - val loss should be approximately the same.</p>",
      "rawMarkdown": "Yes there is an explanation. First of all metrics usually depends on batch size. Not much but it happens. \nSecond: usually batch size is defined using number of cores i.e.:\nBATCH_SIZE = 8 * strategy.num_replicas_in_sync\nNow kaggle tpu has 8 cores while colab tpu only one. Therefore if you use same batch size - val loss should be approximately the same.",
      "votes": null
    },
    {
      "id": "1307893",
      "postDate": "05/14/2021 18:26:59",
      "content": "<p>I guess colab also uses 8 replicas.</p>",
      "rawMarkdown": "I guess colab also uses 8 replicas.",
      "votes": null
    },
    {
      "id": "1307918",
      "postDate": "05/14/2021 19:09:11",
      "content": "<p>They both have 8 cores.  But Kaggle has TPU v3 (16 Go/core) while Colab has TPU v2 (8 Go/core)</p>",
      "rawMarkdown": "They both have 8 cores.  But Kaggle has TPU v3 (16 Go/core) while Colab has TPU v2 (8 Go/core)",
      "votes": null
    },
    {
      "id": "1309112",
      "postDate": "05/15/2021 17:11:05",
      "content": "<p>First Thanks for all reply.    Colab TPU converges slowly,  and  NaN  appears frequently and thus training breaks.</p>",
      "rawMarkdown": "First Thanks for all reply.    Colab TPU converges slowly,  and  NaN  appears frequently and thus training breaks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1307781,
      "author_name": "wrrosa",
      "author_url": "",
      "post_date": "05/14/2021 16:49:28",
      "content": "<p>Yes there is an explanation. First of all metrics usually depends on batch size. Not much but it happens. <br>\nSecond: usually batch size is defined using number of cores i.e.:<br>\nBATCH_SIZE = 8 * strategy.num_replicas_in_sync<br>\nNow kaggle tpu has 8 cores while colab tpu only one. Therefore if you use same batch size - val loss should be approximately the same.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1307893,
          "author_name": "sorkun",
          "author_url": "",
          "post_date": "05/14/2021 18:26:59",
          "content": "<p>I guess colab also uses 8 replicas.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1307918,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/14/2021 19:09:11",
          "content": "<p>They both have 8 cores.  But Kaggle has TPU v3 (16 Go/core) while Colab has TPU v2 (8 Go/core)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1309112,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "05/15/2021 17:11:05",
      "content": "<p>First Thanks for all reply.    Colab TPU converges slowly,  and  NaN  appears frequently and thus training breaks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1307053": "I use the pretrained encoder/decoder for fine tune, it works on Kaggle, not far away with pretrained.\nhowever, when I do same thing on Colab TPU,  it load the weights correctly, however the corresponding  val_loss and val_lsd are much bigger?",
    "1307781": "Yes there is an explanation. First of all metrics usually depends on batch size. Not much but it happens. \nSecond: usually batch size is defined using number of cores i.e.:\nBATCH_SIZE = 8 * strategy.num_replicas_in_sync\nNow kaggle tpu has 8 cores while colab tpu only one. Therefore if you use same batch size - val loss should be approximately the same.",
    "1307893": "I guess colab also uses 8 replicas.",
    "1307918": "They both have 8 cores.  But Kaggle has TPU v3 (16 Go/core) while Colab has TPU v2 (8 Go/core)",
    "1309112": "First Thanks for all reply.    Colab TPU converges slowly,  and  NaN  appears frequently and thus training breaks."
  },
  "source": "meta"
}