{
  "id": 267035,
  "title": "Tips for training your own model on Kaggle Kernel ",
  "url": "/competitions/landmark-recognition-2021/discussion/267035",
  "author_name": "",
  "post_date": "2021-08-21T12:25:42.758752800Z",
  "votes": 24,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Given the size of the competition dataset, training your own model on Kaggle Kernels may pose a challenge. Here are a few tips for starters that allow training a model using Kaggle Kernels:</p>\n<ol>\n<li>Image size - use smaller resolution, like 256</li>\n<li>Model size - use a smaller version of the model, e.g. EfficientNet B0</li>\n<li>Sample data - you don't need to use all of the data in one epoch. You can always sample data and  change see</li>\n<li>Beware of the <a href=\"https://github.com/pytorch/pytorch/issues/13246\" target=\"_blank\">memory leak</a> of pytorch dataloaders. I am still working on this - in anyone has an idea that is confirmed and works please let me know</li>\n</ol>\n<p>With the above settings (image size 256, Effnet B0, 0.75 fraction data used), I managed to train one epoch in around 4 hours using Kaggle kernel. I still have not fixed memory leak issue, and my kernel died during second epoch because of that, but this permitting you should be able to squeeze 2 epochs within 9 hours limit. </p>\n<p>The model trained for 3 epoch with above settings reached a 0.101 on LB without any post-processing (only global features).</p>\n<p>Best solutions from last year suggest around 10 epochs of training of the entire dataset, so with 2 weeks quota on GPU, you might be able to reach competitive levels.</p>",
  "messages": [
    {
      "id": "1484559",
      "postDate": "08/21/2021 12:25:42",
      "content": "<p>Given the size of the competition dataset, training your own model on Kaggle Kernels may pose a challenge. Here are a few tips for starters that allow training a model using Kaggle Kernels:</p>\n<ol>\n<li>Image size - use smaller resolution, like 256</li>\n<li>Model size - use a smaller version of the model, e.g. EfficientNet B0</li>\n<li>Sample data - you don't need to use all of the data in one epoch. You can always sample data and  change see</li>\n<li>Beware of the <a href=\"https://github.com/pytorch/pytorch/issues/13246\" target=\"_blank\">memory leak</a> of pytorch dataloaders. I am still working on this - in anyone has an idea that is confirmed and works please let me know</li>\n</ol>\n<p>With the above settings (image size 256, Effnet B0, 0.75 fraction data used), I managed to train one epoch in around 4 hours using Kaggle kernel. I still have not fixed memory leak issue, and my kernel died during second epoch because of that, but this permitting you should be able to squeeze 2 epochs within 9 hours limit. </p>\n<p>The model trained for 3 epoch with above settings reached a 0.101 on LB without any post-processing (only global features).</p>\n<p>Best solutions from last year suggest around 10 epochs of training of the entire dataset, so with 2 weeks quota on GPU, you might be able to reach competitive levels.</p>",
      "rawMarkdown": "Given the size of the competition dataset, training your own model on Kaggle Kernels may pose a challenge. Here are a few tips for starters that allow training a model using Kaggle Kernels:\n1. Image size - use smaller resolution, like 256\n2. Model size - use a smaller version of the model, e.g. EfficientNet B0\n3. Sample data - you don't need to use all of the data in one epoch. You can always sample data and  change see\n4. Beware of the [memory leak](https://github.com/pytorch/pytorch/issues/13246) of pytorch dataloaders. I am still working on this - in anyone has an idea that is confirmed and works please let me know\n\nWith the above settings (image size 256, Effnet B0, 0.75 fraction data used), I managed to train one epoch in around 4 hours using Kaggle kernel. I still have not fixed memory leak issue, and my kernel died during second epoch because of that, but this permitting you should be able to squeeze 2 epochs within 9 hours limit. \n\nThe model trained for 3 epoch with above settings reached a 0.101 on LB without any post-processing (only global features).\n\nBest solutions from last year suggest around 10 epochs of training of the entire dataset, so with 2 weeks quota on GPU, you might be able to reach competitive levels.",
      "votes": null
    },
    {
      "id": "1484564",
      "postDate": "08/21/2021 12:32:39",
      "content": "<p>If I'm not mistaken there is a memory leak in torch dataloader if you set <code>pin_memory=True</code></p>",
      "rawMarkdown": "If I'm not mistaken there is a memory leak in torch dataloader if you set `pin_memory=True`",
      "votes": null
    },
    {
      "id": "1484579",
      "postDate": "08/21/2021 12:45:14",
      "content": "<p>Thanks for the reply. I am actually seeing the memory leak even though I use <code>pin_memory=False</code></p>",
      "rawMarkdown": "Thanks for the reply. I am actually seeing the memory leak even though I use `pin_memory=False`",
      "votes": null
    },
    {
      "id": "1484584",
      "postDate": "08/21/2021 12:49:21",
      "content": "<p>Ah, that's ah shame :/<br>\nFor me setting <code>pin_memory=False</code> solved the issue</p>",
      "rawMarkdown": "Ah, that's ah shame :/\nFor me setting `pin_memory=False` solved the issue",
      "votes": null
    },
    {
      "id": "1484618",
      "postDate": "08/21/2021 13:09:39",
      "content": "<p>No worries, thanks for trying to help</p>",
      "rawMarkdown": "No worries, thanks for trying to help",
      "votes": null
    },
    {
      "id": "1484688",
      "postDate": "08/21/2021 14:01:21",
      "content": "<p>I would suggest training on a TPU. Using EfficientNetV2-M model (21M params) and 384x348 images an epoch takes just 35 minutes while using the whole train dataset, allowing you to train for 10 epochs in one run!</p>",
      "rawMarkdown": "I would suggest training on a TPU. Using EfficientNetV2-M model (21M params) and 384x348 images an epoch takes just 35 minutes while using the whole train dataset, allowing you to train for 10 epochs in one run!",
      "votes": null
    },
    {
      "id": "1484707",
      "postDate": "08/21/2021 14:14:32",
      "content": "<p>Are you using torch or tensorflow?</p>",
      "rawMarkdown": "Are you using torch or tensorflow?",
      "votes": null
    },
    {
      "id": "1484955",
      "postDate": "08/21/2021 17:23:46",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> and <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> for sharing this information!</p>",
      "rawMarkdown": "Thanks @markwijkhuizen and @narsil for sharing this information!",
      "votes": null
    },
    {
      "id": "1485079",
      "postDate": "08/21/2021 18:55:26",
      "content": "<p>are you using cloud storage bucket ? </p>",
      "rawMarkdown": "are you using cloud storage bucket ?",
      "votes": null
    },
    {
      "id": "1485104",
      "postDate": "08/21/2021 19:31:56",
      "content": "<p>I am using Tensorflow and the datasets are in TFRecord format. As the dataset is huge the dataset needed to be split into 3 parts.</p>",
      "rawMarkdown": "I am using Tensorflow and the datasets are in TFRecord format. As the dataset is huge the dataset needed to be split into 3 parts.",
      "votes": null
    },
    {
      "id": "1485162",
      "postDate": "08/21/2021 21:04:24",
      "content": "<p>Thanks for the tips <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>! At the end of the post, you suggest to start training a network, save the weights and resume training in another notebook, right?</p>",
      "rawMarkdown": "Thanks for the tips @narsil! At the end of the post, you suggest to start training a network, save the weights and resume training in another notebook, right?",
      "votes": null
    },
    {
      "id": "1485283",
      "postDate": "08/22/2021 00:16:17",
      "content": "<p>is TFRecords with local working ? or using GCS ? </p>",
      "rawMarkdown": "is TFRecords with local working ? or using GCS ?",
      "votes": null
    },
    {
      "id": "1485490",
      "postDate": "08/22/2021 06:35:13",
      "content": "<p>That is correct</p>",
      "rawMarkdown": "That is correct",
      "votes": null
    },
    {
      "id": "1485954",
      "postDate": "08/22/2021 14:29:04",
      "content": "<p>The TFRecords are stored on Google Cloud, training is done using TPU's provided by Kaggle. I will publish them somewhere next week including the notebook to create them.</p>",
      "rawMarkdown": "The TFRecords are stored on Google Cloud, training is done using TPU's provided by Kaggle. I will publish them somewhere next week including the notebook to create them.",
      "votes": null
    },
    {
      "id": "1487488",
      "postDate": "08/23/2021 16:49:05",
      "content": "<p>I prefer pytorch to tensorflow. Tried to make TPU work with pytorch but got bunch of exceptions / errors. I did some research  - my experience was not unique. Right now I feel there is some unnecessary tradeoff between using pytorch and TPU -&gt; seems there is space for a very successful library , if anyone is up for the work :) </p>",
      "rawMarkdown": "I prefer pytorch to tensorflow. Tried to make TPU work with pytorch but got bunch of exceptions / errors. I did some research  - my experience was not unique. Right now I feel there is some unnecessary tradeoff between using pytorch and TPU -> seems there is space for a very successful library , if anyone is up for the work :)",
      "votes": null
    },
    {
      "id": "1488060",
      "postDate": "08/24/2021 04:06:09",
      "content": "<p>Same here. Have you tried pytorch lightning? I haven't used it yet, but I've seen a lot of people use it for multi-GPU and TPU training.</p>",
      "rawMarkdown": "Same here. Have you tried pytorch lightning? I haven't used it yet, but I've seen a lot of people use it for multi-GPU and TPU training.",
      "votes": null
    },
    {
      "id": "1488157",
      "postDate": "08/24/2021 06:10:33",
      "content": "<p>The errors I got were exactly on Pytorch Lightning</p>",
      "rawMarkdown": "The errors I got were exactly on Pytorch Lightning",
      "votes": null
    },
    {
      "id": "1490428",
      "postDate": "08/25/2021 16:14:01",
      "content": "<p>I see. Did you get the score you mentioned above (0.101) by following a retrieval approach or simply applying softmax to the last layer of B0? I trained B0 with similar hyperparameters as you did, but for more epochs. When I tried a pure classification approach, it got a score of 0. </p>",
      "rawMarkdown": "I see. Did you get the score you mentioned above (0.101) by following a retrieval approach or simply applying softmax to the last layer of B0? I trained B0 with similar hyperparameters as you did, but for more epochs. When I tried a pure classification approach, it got a score of 0.",
      "votes": null
    },
    {
      "id": "1491073",
      "postDate": "08/26/2021 06:06:58",
      "content": "<p>Please have a look at the example notebook I prepared here <a href=\"https://www.kaggle.com/narsil/host-baseline-2020\" target=\"_blank\">https://www.kaggle.com/narsil/host-baseline-2020</a></p>\n<p>To get 0.101 I followed the same logic, just with different models and no post-processing (DELF)</p>",
      "rawMarkdown": "Please have a look at the example notebook I prepared here https://www.kaggle.com/narsil/host-baseline-2020\n\nTo get 0.101 I followed the same logic, just with different models and no post-processing (DELF)",
      "votes": null
    },
    {
      "id": "1494699",
      "postDate": "08/29/2021 01:49:00",
      "content": "<p>Referring to memory leak problem: there is a quite great summary of the methods to handle this issue <a href=\"https://github.com/pytorch/pytorch/issues/13246#issuecomment-905703662\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/13246#issuecomment-905703662</a><br>\n(use <code>np.array</code> instead of built-in python objects in your <code>Dataset</code> class and save strings with <code>np.string_</code> type - I think, these are the most common cases)<br>\nPlus, I would recommend to use latest torch version: I don't face this issue once I switched to the <code>1.9</code></p>",
      "rawMarkdown": "Referring to memory leak problem: there is a quite great summary of the methods to handle this issue https://github.com/pytorch/pytorch/issues/13246#issuecomment-905703662\n(use `np.array` instead of built-in python objects in your `Dataset` class and save strings with `np.string_` type - I think, these are the most common cases)\nPlus, I would recommend to use latest torch version: I don't face this issue once I switched to the `1.9`",
      "votes": null
    },
    {
      "id": "1496128",
      "postDate": "08/30/2021 06:15:19",
      "content": "<p>This is a great summary <a href=\"https://www.kaggle.com/podidiving\" target=\"_blank\">@podidiving</a> , thanks so much</p>",
      "rawMarkdown": "This is a great summary @podidiving , thanks so much",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1484564,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "08/21/2021 12:32:39",
      "content": "<p>If I'm not mistaken there is a memory leak in torch dataloader if you set <code>pin_memory=True</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1484579,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/21/2021 12:45:14",
          "content": "<p>Thanks for the reply. I am actually seeing the memory leak even though I use <code>pin_memory=False</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1484584,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "08/21/2021 12:49:21",
          "content": "<p>Ah, that's ah shame :/<br>\nFor me setting <code>pin_memory=False</code> solved the issue</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1484618,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/21/2021 13:09:39",
          "content": "<p>No worries, thanks for trying to help</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1484688,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "08/21/2021 14:01:21",
      "content": "<p>I would suggest training on a TPU. Using EfficientNetV2-M model (21M params) and 384x348 images an epoch takes just 35 minutes while using the whole train dataset, allowing you to train for 10 epochs in one run!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1484707,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/21/2021 14:14:32",
          "content": "<p>Are you using torch or tensorflow?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1484955,
          "author_name": "saurabhbagchi",
          "author_url": "",
          "post_date": "08/21/2021 17:23:46",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> and <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> for sharing this information!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1485079,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "08/21/2021 18:55:26",
          "content": "<p>are you using cloud storage bucket ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1485104,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "08/21/2021 19:31:56",
          "content": "<p>I am using Tensorflow and the datasets are in TFRecord format. As the dataset is huge the dataset needed to be split into 3 parts.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1485283,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "08/22/2021 00:16:17",
          "content": "<p>is TFRecords with local working ? or using GCS ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1485954,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "08/22/2021 14:29:04",
          "content": "<p>The TFRecords are stored on Google Cloud, training is done using TPU's provided by Kaggle. I will publish them somewhere next week including the notebook to create them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1487488,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/23/2021 16:49:05",
          "content": "<p>I prefer pytorch to tensorflow. Tried to make TPU work with pytorch but got bunch of exceptions / errors. I did some research  - my experience was not unique. Right now I feel there is some unnecessary tradeoff between using pytorch and TPU -&gt; seems there is space for a very successful library , if anyone is up for the work :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1488060,
          "author_name": "novice03",
          "author_url": "",
          "post_date": "08/24/2021 04:06:09",
          "content": "<p>Same here. Have you tried pytorch lightning? I haven't used it yet, but I've seen a lot of people use it for multi-GPU and TPU training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1488157,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/24/2021 06:10:33",
          "content": "<p>The errors I got were exactly on Pytorch Lightning</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1490428,
          "author_name": "novice03",
          "author_url": "",
          "post_date": "08/25/2021 16:14:01",
          "content": "<p>I see. Did you get the score you mentioned above (0.101) by following a retrieval approach or simply applying softmax to the last layer of B0? I trained B0 with similar hyperparameters as you did, but for more epochs. When I tried a pure classification approach, it got a score of 0. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1491073,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/26/2021 06:06:58",
          "content": "<p>Please have a look at the example notebook I prepared here <a href=\"https://www.kaggle.com/narsil/host-baseline-2020\" target=\"_blank\">https://www.kaggle.com/narsil/host-baseline-2020</a></p>\n<p>To get 0.101 I followed the same logic, just with different models and no post-processing (DELF)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1485162,
      "author_name": "leventelippenszky",
      "author_url": "",
      "post_date": "08/21/2021 21:04:24",
      "content": "<p>Thanks for the tips <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>! At the end of the post, you suggest to start training a network, save the weights and resume training in another notebook, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1485490,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/22/2021 06:35:13",
          "content": "<p>That is correct</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1494699,
      "author_name": "podidiving",
      "author_url": "",
      "post_date": "08/29/2021 01:49:00",
      "content": "<p>Referring to memory leak problem: there is a quite great summary of the methods to handle this issue <a href=\"https://github.com/pytorch/pytorch/issues/13246#issuecomment-905703662\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/13246#issuecomment-905703662</a><br>\n(use <code>np.array</code> instead of built-in python objects in your <code>Dataset</code> class and save strings with <code>np.string_</code> type - I think, these are the most common cases)<br>\nPlus, I would recommend to use latest torch version: I don't face this issue once I switched to the <code>1.9</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1496128,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "08/30/2021 06:15:19",
          "content": "<p>This is a great summary <a href=\"https://www.kaggle.com/podidiving\" target=\"_blank\">@podidiving</a> , thanks so much</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1484559": "Given the size of the competition dataset, training your own model on Kaggle Kernels may pose a challenge. Here are a few tips for starters that allow training a model using Kaggle Kernels:\n1. Image size - use smaller resolution, like 256\n2. Model size - use a smaller version of the model, e.g. EfficientNet B0\n3. Sample data - you don't need to use all of the data in one epoch. You can always sample data and  change see\n4. Beware of the [memory leak](https://github.com/pytorch/pytorch/issues/13246) of pytorch dataloaders. I am still working on this - in anyone has an idea that is confirmed and works please let me know\n\nWith the above settings (image size 256, Effnet B0, 0.75 fraction data used), I managed to train one epoch in around 4 hours using Kaggle kernel. I still have not fixed memory leak issue, and my kernel died during second epoch because of that, but this permitting you should be able to squeeze 2 epochs within 9 hours limit. \n\nThe model trained for 3 epoch with above settings reached a 0.101 on LB without any post-processing (only global features).\n\nBest solutions from last year suggest around 10 epochs of training of the entire dataset, so with 2 weeks quota on GPU, you might be able to reach competitive levels.",
    "1484564": "If I'm not mistaken there is a memory leak in torch dataloader if you set `pin_memory=True`",
    "1484579": "Thanks for the reply. I am actually seeing the memory leak even though I use `pin_memory=False`",
    "1484584": "Ah, that's ah shame :/\nFor me setting `pin_memory=False` solved the issue",
    "1484618": "No worries, thanks for trying to help",
    "1484688": "I would suggest training on a TPU. Using EfficientNetV2-M model (21M params) and 384x348 images an epoch takes just 35 minutes while using the whole train dataset, allowing you to train for 10 epochs in one run!",
    "1484707": "Are you using torch or tensorflow?",
    "1484955": "Thanks @markwijkhuizen and @narsil for sharing this information!",
    "1485079": "are you using cloud storage bucket ?",
    "1485104": "I am using Tensorflow and the datasets are in TFRecord format. As the dataset is huge the dataset needed to be split into 3 parts.",
    "1485162": "Thanks for the tips @narsil! At the end of the post, you suggest to start training a network, save the weights and resume training in another notebook, right?",
    "1485283": "is TFRecords with local working ? or using GCS ?",
    "1485490": "That is correct",
    "1485954": "The TFRecords are stored on Google Cloud, training is done using TPU's provided by Kaggle. I will publish them somewhere next week including the notebook to create them.",
    "1487488": "I prefer pytorch to tensorflow. Tried to make TPU work with pytorch but got bunch of exceptions / errors. I did some research  - my experience was not unique. Right now I feel there is some unnecessary tradeoff between using pytorch and TPU -> seems there is space for a very successful library , if anyone is up for the work :)",
    "1488060": "Same here. Have you tried pytorch lightning? I haven't used it yet, but I've seen a lot of people use it for multi-GPU and TPU training.",
    "1488157": "The errors I got were exactly on Pytorch Lightning",
    "1490428": "I see. Did you get the score you mentioned above (0.101) by following a retrieval approach or simply applying softmax to the last layer of B0? I trained B0 with similar hyperparameters as you did, but for more epochs. When I tried a pure classification approach, it got a score of 0.",
    "1491073": "Please have a look at the example notebook I prepared here https://www.kaggle.com/narsil/host-baseline-2020\n\nTo get 0.101 I followed the same logic, just with different models and no post-processing (DELF)",
    "1494699": "Referring to memory leak problem: there is a quite great summary of the methods to handle this issue https://github.com/pytorch/pytorch/issues/13246#issuecomment-905703662\n(use `np.array` instead of built-in python objects in your `Dataset` class and save strings with `np.string_` type - I think, these are the most common cases)\nPlus, I would recommend to use latest torch version: I don't face this issue once I switched to the `1.9`",
    "1496128": "This is a great summary @podidiving , thanks so much"
  },
  "source": "meta"
}