{
  "id": 166353,
  "title": "How are people without multiple GPUs supposed to participate?",
  "url": "/competitions/landmark-retrieval-2020/discussion/166353",
  "author_name": "",
  "post_date": "2020-07-12T17:00:48.908194200Z",
  "votes": 9,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I have a 8GB 2070 GPU. Current estimate training time for 1.5M images and batch size of 4 (has to be low because of memory) is 19 hours per epoch. This is using ResNet50 backbone with a single FC hidden layer and output classes of 81k (yes triplet loss would be faster due to less parameters but ArcFace loss seems like the best approach). That means to train a model for 10 epoch, would need 8 days... . We only have 30 days. Maybe I am new to kaggle challanges, but surely there wouldn't be enough time to iterate a good model to beat the baseline, unless someone has 4 or 8 gpus to do distributed training. Seems like it is unfairly favour people with access to multiple GPUs.  Is anyone else facing the same issue? What hardware you guys have? Are other challenges also so hardware heavy?</p>",
  "messages": [
    {
      "id": "926363",
      "postDate": "07/12/2020 17:00:48",
      "content": "<p>I have a 8GB 2070 GPU. Current estimate training time for 1.5M images and batch size of 4 (has to be low because of memory) is 19 hours per epoch. This is using ResNet50 backbone with a single FC hidden layer and output classes of 81k (yes triplet loss would be faster due to less parameters but ArcFace loss seems like the best approach). That means to train a model for 10 epoch, would need 8 days... . We only have 30 days. Maybe I am new to kaggle challanges, but surely there wouldn't be enough time to iterate a good model to beat the baseline, unless someone has 4 or 8 gpus to do distributed training. Seems like it is unfairly favour people with access to multiple GPUs.  Is anyone else facing the same issue? What hardware you guys have? Are other challenges also so hardware heavy?</p>",
      "rawMarkdown": "I have a 8GB 2070 GPU. Current estimate training time for 1.5M images and batch size of 4 (has to be low because of memory) is 19 hours per epoch. This is using ResNet50 backbone with a single FC hidden layer and output classes of 81k (yes triplet loss would be faster due to less parameters but ArcFace loss seems like the best approach). That means to train a model for 10 epoch, would need 8 days... . We only have 30 days. Maybe I am new to kaggle challanges, but surely there wouldn't be enough time to iterate a good model to beat the baseline, unless someone has 4 or 8 gpus to do distributed training. Seems like it is unfairly favour people with access to multiple GPUs.  Is anyone else facing the same issue? What hardware you guys have? Are other challenges also so hardware heavy?",
      "votes": null
    },
    {
      "id": "926676",
      "postDate": "07/12/2020 21:36:23",
      "content": "<p>Have you tried using the free TPUs? The ones on Kaggle have 128GB of RAM and have the computation performance equivalent to 4xV100 (32GB) or 8xP100 (16GB). Since they are valued at 8$/h, with 30h free usage per week Kaggle is essentially give away 240$/week of free compute credits.</p>",
      "rawMarkdown": "Have you tried using the free TPUs? The ones on Kaggle have 128GB of RAM and have the computation performance equivalent to 4xV100 (32GB) or 8xP100 (16GB). Since they are valued at 8$/h, with 30h free usage per week Kaggle is essentially give away 240$/week of free compute credits.",
      "votes": null
    },
    {
      "id": "927107",
      "postDate": "07/13/2020 07:53:49",
      "content": "<p>yes its really sick, im trying to donwload images and using colab and gcloud</p>",
      "rawMarkdown": "yes its really sick, im trying to donwload images and using colab and gcloud",
      "votes": null
    },
    {
      "id": "927743",
      "postDate": "07/13/2020 15:10:11",
      "content": "<p>If you have a US credit card, you can consider colab pro with Google drive upgrade. This way, for around $10+5 USD a month you have much more cloud storage (which can be directly accessed from colab), 32gb Ram, and 24h sessions, and much better GPUs (e.g. v100)</p>",
      "rawMarkdown": "If you have a US credit card, you can consider colab pro with Google drive upgrade. This way, for around $10+5 USD a month you have much more cloud storage (which can be directly accessed from colab), 32gb Ram, and 24h sessions, and much better GPUs (e.g. v100)",
      "votes": null
    },
    {
      "id": "930573",
      "postDate": "07/15/2020 14:53:48",
      "content": "<p>No, Kaggle takes a fee from the clients. You pay for it with labor toward someone else's research project, unless you never plan to compete.</p>",
      "rawMarkdown": "No, Kaggle takes a fee from the clients. You pay for it with labor toward someone else's research project, unless you never plan to compete.",
      "votes": null
    },
    {
      "id": "930615",
      "postDate": "07/15/2020 15:43:02",
      "content": "<p>Great suggestion thanks. I finally got it to work on the TPU. Turns out serialization to shards gives an increase of 4x training speed, and TPU gives me another 4x increase, which is good but not great. I can't seem to get MXU high. Mine is around 2-3% and idle time is about 60-80%. I tried optimized it following this link: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443</a> by doing more step per TPU call. However although the idle time is less, I don't see an increase in training speed. Do you have any experience with optimizing TPU training? Any tips?</p>",
      "rawMarkdown": "Great suggestion thanks. I finally got it to work on the TPU. Turns out serialization to shards gives an increase of 4x training speed, and TPU gives me another 4x increase, which is good but not great. I can't seem to get MXU high. Mine is around 2-3% and idle time is about 60-80%. I tried optimized it following this link: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443 by doing more step per TPU call. However although the idle time is less, I don't see an increase in training speed. Do you have any experience with optimizing TPU training? Any tips?",
      "votes": null
    },
    {
      "id": "931222",
      "postDate": "07/16/2020 04:24:32",
      "content": "<p>I use Colab Pro and get P100. Do you regularly get v100? </p>",
      "rawMarkdown": "I use Colab Pro and get P100. Do you regularly get v100?",
      "votes": null
    },
    {
      "id": "931230",
      "postDate": "07/16/2020 04:34:26",
      "content": "<p>I guess it really depends. I've had various GPUs so I don't remember exactly.</p>",
      "rawMarkdown": "I guess it really depends. I've had various GPUs so I don't remember exactly.",
      "votes": null
    },
    {
      "id": "931630",
      "postDate": "07/16/2020 10:31:33",
      "content": "<p>It still won't be enough. DELG was trained with 30 P100 GPU days. </p>",
      "rawMarkdown": "It still won't be enough. DELG was trained with 30 P100 GPU days.",
      "votes": null
    },
    {
      "id": "931811",
      "postDate": "07/16/2020 13:20:00",
      "content": "<p>Although it's impossible to get a 100% mxu, 20% is definitely achievable from personal experience. Have you tried increasing the batch size?</p>",
      "rawMarkdown": "Although it's impossible to get a 100% mxu, 20% is definitely achievable from personal experience. Have you tried increasing the batch size?",
      "votes": null
    },
    {
      "id": "931817",
      "postDate": "07/16/2020 13:21:40",
      "content": "<p>Also I'm not sure if the idle time only pertains to the actual tpu usage, or also the pod cpu. If you are doing heavy cpu based computation before every iteration that might actually decrease your actual tpu usage efficiency.</p>",
      "rawMarkdown": "Also I'm not sure if the idle time only pertains to the actual tpu usage, or also the pod cpu. If you are doing heavy cpu based computation before every iteration that might actually decrease your actual tpu usage efficiency.",
      "votes": null
    },
    {
      "id": "932028",
      "postDate": "07/16/2020 16:27:30",
      "content": "<p>Batch size can't get much larger, since each core only has 16G memory. I could try with smaller model and see if the mxu improves</p>",
      "rawMarkdown": "Batch size can't get much larger, since each core only has 16G memory. I could try with smaller model and see if the mxu improves",
      "votes": null
    },
    {
      "id": "933740",
      "postDate": "07/18/2020 01:30:20",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> Do you know how to restore checkpoint in TPU? TPU does not have local system, and loading weights from model under tpu_strategy directly from google storage seems to cause the code to freeze.</p>",
      "rawMarkdown": "xhlulu Do you know how to restore checkpoint in TPU? TPU does not have local system, and loading weights from model under tpu_strategy directly from google storage seems to cause the code to freeze.",
      "votes": null
    },
    {
      "id": "956719",
      "postDate": "08/03/2020 18:06:29",
      "content": "<p><a href=\"/suruili\">@suruili</a> How did you get the model to work with a TPU?  I have been trying for a while with no luck.  A notebook would be greatly appreciated.</p>",
      "rawMarkdown": "suruili How did you get the model to work with a TPU?  I have been trying for a while with no luck.  A notebook would be greatly appreciated.",
      "votes": null
    },
    {
      "id": "956875",
      "postDate": "08/03/2020 21:24:57",
      "content": "<p>I published a notebook here: <a href=\"https://www.kaggle.com/suruili/arcface-gem-train-on-tpu\">https://www.kaggle.com/suruili/arcface-gem-train-on-tpu</a></p>",
      "rawMarkdown": "I published a notebook here: https://www.kaggle.com/suruili/arcface-gem-train-on-tpu",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 926676,
      "author_name": "xhlulu",
      "author_url": "",
      "post_date": "07/12/2020 21:36:23",
      "content": "<p>Have you tried using the free TPUs? The ones on Kaggle have 128GB of RAM and have the computation performance equivalent to 4xV100 (32GB) or 8xP100 (16GB). Since they are valued at 8$/h, with 30h free usage per week Kaggle is essentially give away 240$/week of free compute credits.</p>",
      "votes": null,
      "replies": [
        {
          "id": 930573,
          "author_name": "realabyszero",
          "author_url": "",
          "post_date": "07/15/2020 14:53:48",
          "content": "<p>No, Kaggle takes a fee from the clients. You pay for it with labor toward someone else's research project, unless you never plan to compete.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930615,
          "author_name": "suruili",
          "author_url": "",
          "post_date": "07/15/2020 15:43:02",
          "content": "<p>Great suggestion thanks. I finally got it to work on the TPU. Turns out serialization to shards gives an increase of 4x training speed, and TPU gives me another 4x increase, which is good but not great. I can't seem to get MXU high. Mine is around 2-3% and idle time is about 60-80%. I tried optimized it following this link: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443</a> by doing more step per TPU call. However although the idle time is less, I don't see an increase in training speed. Do you have any experience with optimizing TPU training? Any tips?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931811,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "07/16/2020 13:20:00",
          "content": "<p>Although it's impossible to get a 100% mxu, 20% is definitely achievable from personal experience. Have you tried increasing the batch size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931817,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "07/16/2020 13:21:40",
          "content": "<p>Also I'm not sure if the idle time only pertains to the actual tpu usage, or also the pod cpu. If you are doing heavy cpu based computation before every iteration that might actually decrease your actual tpu usage efficiency.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932028,
          "author_name": "suruili",
          "author_url": "",
          "post_date": "07/16/2020 16:27:30",
          "content": "<p>Batch size can't get much larger, since each core only has 16G memory. I could try with smaller model and see if the mxu improves</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933740,
          "author_name": "suruili",
          "author_url": "",
          "post_date": "07/18/2020 01:30:20",
          "content": "<p><a href=\"/xhlulu\">@xhlulu</a> Do you know how to restore checkpoint in TPU? TPU does not have local system, and loading weights from model under tpu_strategy directly from google storage seems to cause the code to freeze.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956719,
          "author_name": "fatehaliyev",
          "author_url": "",
          "post_date": "08/03/2020 18:06:29",
          "content": "<p><a href=\"/suruili\">@suruili</a> How did you get the model to work with a TPU?  I have been trying for a while with no luck.  A notebook would be greatly appreciated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956875,
          "author_name": "suruili",
          "author_url": "",
          "post_date": "08/03/2020 21:24:57",
          "content": "<p>I published a notebook here: <a href=\"https://www.kaggle.com/suruili/arcface-gem-train-on-tpu\">https://www.kaggle.com/suruili/arcface-gem-train-on-tpu</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 927107,
      "author_name": "enric1296",
      "author_url": "",
      "post_date": "07/13/2020 07:53:49",
      "content": "<p>yes its really sick, im trying to donwload images and using colab and gcloud</p>",
      "votes": null,
      "replies": [
        {
          "id": 927743,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "07/13/2020 15:10:11",
          "content": "<p>If you have a US credit card, you can consider colab pro with Google drive upgrade. This way, for around $10+5 USD a month you have much more cloud storage (which can be directly accessed from colab), 32gb Ram, and 24h sessions, and much better GPUs (e.g. v100)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931222,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "07/16/2020 04:24:32",
          "content": "<p>I use Colab Pro and get P100. Do you regularly get v100? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931230,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "07/16/2020 04:34:26",
          "content": "<p>I guess it really depends. I've had various GPUs so I don't remember exactly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931630,
          "author_name": "suruili",
          "author_url": "",
          "post_date": "07/16/2020 10:31:33",
          "content": "<p>It still won't be enough. DELG was trained with 30 P100 GPU days. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "926363": "I have a 8GB 2070 GPU. Current estimate training time for 1.5M images and batch size of 4 (has to be low because of memory) is 19 hours per epoch. This is using ResNet50 backbone with a single FC hidden layer and output classes of 81k (yes triplet loss would be faster due to less parameters but ArcFace loss seems like the best approach). That means to train a model for 10 epoch, would need 8 days... . We only have 30 days. Maybe I am new to kaggle challanges, but surely there wouldn't be enough time to iterate a good model to beat the baseline, unless someone has 4 or 8 gpus to do distributed training. Seems like it is unfairly favour people with access to multiple GPUs.  Is anyone else facing the same issue? What hardware you guys have? Are other challenges also so hardware heavy?",
    "926676": "Have you tried using the free TPUs? The ones on Kaggle have 128GB of RAM and have the computation performance equivalent to 4xV100 (32GB) or 8xP100 (16GB). Since they are valued at 8$/h, with 30h free usage per week Kaggle is essentially give away 240$/week of free compute credits.",
    "927107": "yes its really sick, im trying to donwload images and using colab and gcloud",
    "927743": "If you have a US credit card, you can consider colab pro with Google drive upgrade. This way, for around $10+5 USD a month you have much more cloud storage (which can be directly accessed from colab), 32gb Ram, and 24h sessions, and much better GPUs (e.g. v100)",
    "930573": "No, Kaggle takes a fee from the clients. You pay for it with labor toward someone else's research project, unless you never plan to compete.",
    "930615": "Great suggestion thanks. I finally got it to work on the TPU. Turns out serialization to shards gives an increase of 4x training speed, and TPU gives me another 4x increase, which is good but not great. I can't seem to get MXU high. Mine is around 2-3% and idle time is about 60-80%. I tried optimized it following this link: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443 by doing more step per TPU call. However although the idle time is less, I don't see an increase in training speed. Do you have any experience with optimizing TPU training? Any tips?",
    "931222": "I use Colab Pro and get P100. Do you regularly get v100?",
    "931230": "I guess it really depends. I've had various GPUs so I don't remember exactly.",
    "931630": "It still won't be enough. DELG was trained with 30 P100 GPU days.",
    "931811": "Although it's impossible to get a 100% mxu, 20% is definitely achievable from personal experience. Have you tried increasing the batch size?",
    "931817": "Also I'm not sure if the idle time only pertains to the actual tpu usage, or also the pod cpu. If you are doing heavy cpu based computation before every iteration that might actually decrease your actual tpu usage efficiency.",
    "932028": "Batch size can't get much larger, since each core only has 16G memory. I could try with smaller model and see if the mxu improves",
    "933740": "xhlulu Do you know how to restore checkpoint in TPU? TPU does not have local system, and loading weights from model under tpu_strategy directly from google storage seems to cause the code to freeze.",
    "956719": "suruili How did you get the model to work with a TPU?  I have been trying for a while with no luck.  A notebook would be greatly appreciated.",
    "956875": "I published a notebook here: https://www.kaggle.com/suruili/arcface-gem-train-on-tpu"
  },
  "source": "meta"
}