{
  "id": 220724,
  "title": "A request to kaggle team!!!",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/220724",
  "author_name": "Raghawendra Singh",
  "post_date": "2021-02-19T09:56:09.876000",
  "votes": 10,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Kaggle provides 3 type of notebook:</p>\n<ol>\n<li>CPU notebook with 18GB of RAM.</li>\n<li>GPU notebook with 16 GB of RAM.</li>\n<li>TPU notebook with 18GB of RAM.</li>\n</ol>\n<p>Configuration of first 2 notebooks is very nice, but 3rd notebook lacks sufficient RAM to fully utilize the capacity of all 8 core of TPU. Google colab pro gives 36GB of RAM for its TPU notebook that I find better then Kaggle notebook. Competitions like this one requires preprocessing and augmentation of Large images that requires memory proportional to batch size. Because of this constraint we use either low batch size or single TPU core for training. Low batch size is not good for accuracy so we mostly use single core of TPU for training and remaining 7 cores remains unutilized. My request to you is to give more RAM to TPU notebook so that we can fully utilize it.</p>",
  "messages": [
    {
      "id": 1210264,
      "postDate": "2021-02-19T09:56:09.877Z",
      "content": "<p>Kaggle provides 3 type of notebook:</p>\n<ol>\n<li>CPU notebook with 18GB of RAM.</li>\n<li>GPU notebook with 16 GB of RAM.</li>\n<li>TPU notebook with 18GB of RAM.</li>\n</ol>\n<p>Configuration of first 2 notebooks is very nice, but 3rd notebook lacks sufficient RAM to fully utilize the capacity of all 8 core of TPU. Google colab pro gives 36GB of RAM for its TPU notebook that I find better then Kaggle notebook. Competitions like this one requires preprocessing and augmentation of Large images that requires memory proportional to batch size. Because of this constraint we use either low batch size or single TPU core for training. Low batch size is not good for accuracy so we mostly use single core of TPU for training and remaining 7 cores remains unutilized. My request to you is to give more RAM to TPU notebook so that we can fully utilize it.</p>",
      "rawMarkdown": "Kaggle provides 3 type of notebook:\n1. CPU notebook with 18GB of RAM.\n2. GPU notebook with 16 GB of RAM.\n3. TPU notebook with 18GB of RAM.\n\nConfiguration of first 2 notebooks is very nice, but 3rd notebook lacks sufficient RAM to fully utilize the capacity of all 8 core of TPU. Google colab pro gives 36GB of RAM for its TPU notebook that I find better then Kaggle notebook. Competitions like this one requires preprocessing and augmentation of Large images that requires memory proportional to batch size. Because of this constraint we use either low batch size or single TPU core for training. Low batch size is not good for accuracy so we mostly use single core of TPU for training and remaining 7 cores remains unutilized. My request to you is to give more RAM to TPU notebook so that we can fully utilize it.",
      "votes": 9
    },
    {
      "id": 1212152,
      "postDate": "2021-02-20T23:38:29.893Z",
      "content": "<p>When you enable TPU, your code does not use the local CPU RAM. Instead it uses a special CPU RAM connected with the TPU. My guess it that you get 256GB+ of RAM with TPU.</p>",
      "rawMarkdown": "When you enable TPU, your code does not use the local CPU RAM. Instead it uses a special CPU RAM connected with the TPU. My guess it that you get 256GB+ of RAM with TPU.",
      "votes": 4,
      "replies": [
        {
          "id": 1212291,
          "postDate": "2021-02-21T04:37:21.723Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. In pytorch, Model training happens on TPU RAM, but its dataloader,with augmentations code written in numpy and opencv, uses local CPU RAM. For a decent image size about 640X640 with a batch size of 16, a single instance of dataloader takes about 2.5-3.0 GB of CPU RAM. So training simultaneously on 8 core seems impossible with 8 instances of distributed dataloader with limited Local CPU RAM in kaggle notebook. I am able to use all 8 cores on Google colab pro TPU notebook since it gives 36 GB of CPU RAM. I think all the people upvoted this topic are facing similar issue. If you know its solution please let us know.</p>",
          "rawMarkdown": "Thanks @cdeotte. In pytorch, Model training happens on TPU RAM, but its dataloader,with augmentations code written in numpy and opencv, uses local CPU RAM. For a decent image size about 640X640 with a batch size of 16, a single instance of dataloader takes about 2.5-3.0 GB of CPU RAM. So training simultaneously on 8 core seems impossible with 8 instances of distributed dataloader with limited Local CPU RAM in kaggle notebook. I am able to use all 8 cores on Google colab pro TPU notebook since it gives 36 GB of CPU RAM. I think all the people upvoted this topic are facing similar issue. If you know its solution please let us know."
        },
        {
          "id": 1212331,
          "postDate": "2021-02-21T05:58:54.297Z",
          "content": "<p>Kaggle's Martin Gorner <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191#759409\" target=\"_blank\">here</a> says:</p>\n<blockquote>\n  <p>A small implementation detail is that even though tf.data.Datset runs on the \"CPU\", it is not the CPU of your Kaggle VM. There is another VM with cycles to spare which has the TPU attached to it through a PCI link. If you use tf.data.Dataset.prefetch(AUTO), you can get a lot of data transformations for \"free\" while the TPU is doing its forward and backward pass.</p>\n</blockquote>\n<p>When asked \"Do you know how large is this additional CPU?\", he responded:</p>\n<blockquote>\n  <p>Not precisely but I know it's a lot. This has been dimensioned to make sure the data-hungry TPUv3 can always be fed with enough data. </p>\n</blockquote>\n<p>So, we know that the TPUv3-8 has 128GB \"VRAM\" (similar to the VRAM using 4xV100-32GB GPUs). But we don't know how much CPU RAM the TPU has. Based on performance, I would guess that the TPU has a dedicated CPU with 256GB+ RAM and 20+ cores.</p>\n<p>If you use <strong>TensorFlow</strong>'s dataloader <code>tf.data.Dataset</code>, then your augmentations take place on the TPU CPU 256GB+ and the training takes place on the TPU 128GB \"VRAM\". If you use <strong>PyTorch</strong>, then you may use Kaggle's limited CPU that only has 16GB RAM and 2 cores.</p>\n<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> Can you shed more light on the situation?</p>",
          "rawMarkdown": "Kaggle's Martin Gorner [here][1] says:\n>A small implementation detail is that even though tf.data.Datset runs on the \"CPU\", it is not the CPU of your Kaggle VM. There is another VM with cycles to spare which has the TPU attached to it through a PCI link. If you use tf.data.Dataset.prefetch(AUTO), you can get a lot of data transformations for \"free\" while the TPU is doing its forward and backward pass.\n\nWhen asked \"Do you know how large is this additional CPU?\", he responded:\n>Not precisely but I know it's a lot. This has been dimensioned to make sure the data-hungry TPUv3 can always be fed with enough data. \n\nSo, we know that the TPUv3-8 has 128GB \"VRAM\" (similar to the VRAM using 4xV100-32GB GPUs). But we don't know how much CPU RAM the TPU has. Based on performance, I would guess that the TPU has a dedicated CPU with 256GB+ RAM and 20+ cores.\n\nIf you use **TensorFlow**'s dataloader `tf.data.Dataset`, then your augmentations take place on the TPU CPU 256GB+ and the training takes place on the TPU 128GB \"VRAM\". If you use **PyTorch**, then you may use Kaggle's limited CPU that only has 16GB RAM and 2 cores.\n\n@mgornergoogle Can you shed more light on the situation?\n\n[1]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191#759409",
          "votes": 6
        },
        {
          "id": 1212360,
          "postDate": "2021-02-21T06:44:29.850Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, It seems that pytorch is not mature enough to fully utilize TPU's.</p>",
          "rawMarkdown": "Thanks @cdeotte, It seems that pytorch is not mature enough to fully utilize TPU's."
        },
        {
          "id": 1212697,
          "postDate": "2021-02-21T13:46:07.467Z",
          "content": "<p>have you ever sucessfully could run in 8TPU cores in pytorch</p>",
          "rawMarkdown": "have you ever sucessfully could run in 8TPU cores in pytorch"
        },
        {
          "id": 1212719,
          "postDate": "2021-02-21T14:05:43.723Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> Yes, but with very small batch size per core on Kaggle notebook for Image competitions. Just set a small batch size and nprocs=8 in CFG class of following notebook, it will work but training will be very noisy.<br>\n<a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1</a><br>\nIn text processing dataloaders does not requirs lots of RAM, so in text competition:<br>\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification\" target=\"_blank\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification</a><br>\nyou can find many notebook that uses all 8 core successfully.<br>\nIn google colab pro, with 36 GB of CPU RAM, I managed to use decent batch size per core, but it provides TPU V2 having 8 GB VRAM per core So it also not good for training large models due to small batch size per core.<br>\nAnother solution I tried is to replace BatchNorm with SyncBatchNorm, but it seem that pytorch XLA does not support it. </p>",
          "rawMarkdown": "@morizin Yes, but with very small batch size per core on Kaggle notebook for Image competitions. Just set a small batch size and nprocs=8 in CFG class of following notebook, it will work but training will be very noisy.\nhttps://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1\nIn text processing dataloaders does not requirs lots of RAM, so in text competition:\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification\nyou can find many notebook that uses all 8 core successfully.\nIn google colab pro, with 36 GB of CPU RAM, I managed to use decent batch size per core, but it provides TPU V2 having 8 GB VRAM per core So it also not good for training large models due to small batch size per core.\nAnother solution I tried is to replace BatchNorm with SyncBatchNorm, but it seem that pytorch XLA does not support it. ",
          "votes": 1
        },
        {
          "id": 1213144,
          "postDate": "2021-02-21T21:43:35.233Z",
          "content": "<p>Have you tried to reduce augmentations?</p>",
          "rawMarkdown": "Have you tried to reduce augmentations?"
        },
        {
          "id": 1213637,
          "postDate": "2021-02-22T08:40:11.340Z",
          "content": "<p>It's worth bearing in mind, PyTorch on TPUs work very well in NLP competitions since there is little/no transformation of the data needed by the CPU.</p>\n<p>Image datasets require more work by the CPU in PyTorch, so this is where it may be advantageous to hop to TF if you want to use TPUs. I think maybe for the next image comp I'm going to solely try TF and see how I get on… 😅</p>",
          "rawMarkdown": "It's worth bearing in mind, PyTorch on TPUs work very well in NLP competitions since there is little/no transformation of the data needed by the CPU.\n\nImage datasets require more work by the CPU in PyTorch, so this is where it may be advantageous to hop to TF if you want to use TPUs. I think maybe for the next image comp I'm going to solely try TF and see how I get on... 😅",
          "votes": 1
        },
        {
          "id": 1213744,
          "postDate": "2021-02-22T10:04:51.940Z",
          "content": "<p>You're right. Unfortunately there's less pretrained models available for TF. For example, there's no ResNet200D implementation on TF (with pretrained weights).</p>",
          "rawMarkdown": "You're right. Unfortunately there's less pretrained models available for TF. For example, there's no ResNet200D implementation on TF (with pretrained weights)."
        },
        {
          "id": 1217125,
          "postDate": "2021-02-24T19:16:45.597Z",
          "content": "<p>I confirm <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s statement is accurate:</p>\n<blockquote>\n  <p>If you use TensorFlow's dataloader tf.data.Dataset, then your augmentations take place on the TPU CPU 256GB+ and the training takes place on the TPU 128GB \"VRAM\". If you use PyTorch, then you may use Kaggle's limited CPU that only has 16GB RAM and 2 cores.</p>\n</blockquote>",
          "rawMarkdown": "I confirm @cdeotte's statement is accurate:\n> If you use TensorFlow's dataloader tf.data.Dataset, then your augmentations take place on the TPU CPU 256GB+ and the training takes place on the TPU 128GB \"VRAM\". If you use PyTorch, then you may use Kaggle's limited CPU that only has 16GB RAM and 2 cores.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1225065,
      "postDate": "2021-03-03T09:25:56.460Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1212119,
      "postDate": "2021-02-20T22:26:03.467Z",
      "content": "<p>Thanks for this informations </p>",
      "rawMarkdown": "Thanks for this informations "
    }
  ],
  "comments": [
    {
      "id": 1212152,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-02-20T23:38:29.893000",
      "content": "<p>When you enable TPU, your code does not use the local CPU RAM. Instead it uses a special CPU RAM connected with the TPU. My guess it that you get 256GB+ of RAM with TPU.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1212291,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2021-02-21T04:37:21.723000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. In pytorch, Model training happens on TPU RAM, but its dataloader,with augmentations code written in numpy and opencv, uses local CPU RAM. For a decent image size about 640X640 with a batch size of 16, a single instance of dataloader takes about 2.5-3.0 GB of CPU RAM. So training simultaneously on 8 core seems impossible with 8 instances of distributed dataloader with limited Local CPU RAM in kaggle notebook. I am able to use all 8 cores on Google colab pro TPU notebook since it gives 36 GB of CPU RAM. I think all the people upvoted this topic are facing similar issue. If you know its solution please let us know.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1212331,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-02-21T05:58:54.297000",
          "content": "<p>Kaggle's Martin Gorner <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191#759409\" target=\"_blank\">here</a> says:</p>\n<blockquote>\n  <p>A small implementation detail is that even though tf.data.Datset runs on the \"CPU\", it is not the CPU of your Kaggle VM. There is another VM with cycles to spare which has the TPU attached to it through a PCI link. If you use tf.data.Dataset.prefetch(AUTO), you can get a lot of data transformations for \"free\" while the TPU is doing its forward and backward pass.</p>\n</blockquote>\n<p>When asked \"Do you know how large is this additional CPU?\", he responded:</p>\n<blockquote>\n  <p>Not precisely but I know it's a lot. This has been dimensioned to make sure the data-hungry TPUv3 can always be fed with enough data. </p>\n</blockquote>\n<p>So, we know that the TPUv3-8 has 128GB \"VRAM\" (similar to the VRAM using 4xV100-32GB GPUs). But we don't know how much CPU RAM the TPU has. Based on performance, I would guess that the TPU has a dedicated CPU with 256GB+ RAM and 20+ cores.</p>\n<p>If you use <strong>TensorFlow</strong>'s dataloader <code>tf.data.Dataset</code>, then your augmentations take place on the TPU CPU 256GB+ and the training takes place on the TPU 128GB \"VRAM\". If you use <strong>PyTorch</strong>, then you may use Kaggle's limited CPU that only has 16GB RAM and 2 cores.</p>\n<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> Can you shed more light on the situation?</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1212360,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2021-02-21T06:44:29.850000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, It seems that pytorch is not mature enough to fully utilize TPU's.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1212697,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-02-21T13:46:07.467000",
          "content": "<p>have you ever sucessfully could run in 8TPU cores in pytorch</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1212719,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2021-02-21T14:05:43.723000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> Yes, but with very small batch size per core on Kaggle notebook for Image competitions. Just set a small batch size and nprocs=8 in CFG class of following notebook, it will work but training will be very noisy.<br>\n<a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1</a><br>\nIn text processing dataloaders does not requirs lots of RAM, so in text competition:<br>\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification\" target=\"_blank\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification</a><br>\nyou can find many notebook that uses all 8 core successfully.<br>\nIn google colab pro, with 36 GB of CPU RAM, I managed to use decent batch size per core, but it provides TPU V2 having 8 GB VRAM per core So it also not good for training large models due to small batch size per core.<br>\nAnother solution I tried is to replace BatchNorm with SyncBatchNorm, but it seem that pytorch XLA does not support it. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1213144,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2021-02-21T21:43:35.233000",
          "content": "<p>Have you tried to reduce augmentations?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1213637,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2021-02-22T08:40:11.340000",
          "content": "<p>It's worth bearing in mind, PyTorch on TPUs work very well in NLP competitions since there is little/no transformation of the data needed by the CPU.</p>\n<p>Image datasets require more work by the CPU in PyTorch, so this is where it may be advantageous to hop to TF if you want to use TPUs. I think maybe for the next image comp I'm going to solely try TF and see how I get on… 😅</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1213744,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2021-02-22T10:04:51.940000",
          "content": "<p>You're right. Unfortunately there's less pretrained models available for TF. For example, there's no ResNet200D implementation on TF (with pretrained weights).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1217125,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2021-02-24T19:16:45.597000",
          "content": "<p>I confirm <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s statement is accurate:</p>\n<blockquote>\n  <p>If you use TensorFlow's dataloader tf.data.Dataset, then your augmentations take place on the TPU CPU 256GB+ and the training takes place on the TPU 128GB \"VRAM\". If you use PyTorch, then you may use Kaggle's limited CPU that only has 16GB RAM and 2 cores.</p>\n</blockquote>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1225065,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-03T09:25:56.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1212119,
      "author_name": "Miloud Belarebia",
      "author_url": "",
      "post_date": "2021-02-20T22:26:03.467000",
      "content": "<p>Thanks for this informations </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1210264": "Kaggle provides 3 type of notebook:\n1. CPU notebook with 18GB of RAM.\n2. GPU notebook with 16 GB of RAM.\n3. TPU notebook with 18GB of RAM.\n\nConfiguration of first 2 notebooks is very nice, but 3rd notebook lacks sufficient RAM to fully utilize the capacity of all 8 core of TPU. Google colab pro gives 36GB of RAM for its TPU notebook that I find better then Kaggle notebook. Competitions like this one requires preprocessing and augmentation of Large images that requires memory proportional to batch size. Because of this constraint we use either low batch size or single TPU core for training. Low batch size is not good for accuracy so we mostly use single core of TPU for training and remaining 7 cores remains unutilized. My request to you is to give more RAM to TPU notebook so that we can fully utilize it.",
    "1212152": "When you enable TPU, your code does not use the local CPU RAM. Instead it uses a special CPU RAM connected with the TPU. My guess it that you get 256GB+ of RAM with TPU.",
    "1225065": "",
    "1212119": "Thanks for this informations "
  }
}