{
  "id": 388202,
  "title": "Tips for speeding up training",
  "url": "/competitions/nfl-player-contact-detection/discussion/388202",
  "author_name": "",
  "post_date": "2023-02-16T12:52:58.446748600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm experimenting with training a model based on the following notebook: <a href=\"https://www.kaggle.com/code/royalacecat/training-nfl-2-5d-cnn-lb-0-671-with-tta/notebook\" target=\"_blank\">notebook</a> (using Efficientnet B3 as the underlying model). </p>\n<p>Now, I don't own any hardware capable of doing this myself, so I have played around with various VMs on Google Cloud. With my current setup of an NVIDIA A100 GPU (40GB) with 12 CPUs (the max allowed for this GPU), I'm getting a training time per epoch of about 5 hours. Is this basically how long models take for this rather compute-intensive competition, or is there something I still can do to significantly speed up training?</p>",
  "messages": [
    {
      "id": "2147162",
      "postDate": "02/16/2023 12:52:58",
      "content": "<p>I'm experimenting with training a model based on the following notebook: <a href=\"https://www.kaggle.com/code/royalacecat/training-nfl-2-5d-cnn-lb-0-671-with-tta/notebook\" target=\"_blank\">notebook</a> (using Efficientnet B3 as the underlying model). </p>\n<p>Now, I don't own any hardware capable of doing this myself, so I have played around with various VMs on Google Cloud. With my current setup of an NVIDIA A100 GPU (40GB) with 12 CPUs (the max allowed for this GPU), I'm getting a training time per epoch of about 5 hours. Is this basically how long models take for this rather compute-intensive competition, or is there something I still can do to significantly speed up training?</p>",
      "rawMarkdown": "I'm experimenting with training a model based on the following notebook: [notebook](https://www.kaggle.com/code/royalacecat/training-nfl-2-5d-cnn-lb-0-671-with-tta/notebook) (using Efficientnet B3 as the underlying model). \n\nNow, I don't own any hardware capable of doing this myself, so I have played around with various VMs on Google Cloud. With my current setup of an NVIDIA A100 GPU (40GB) with 12 CPUs (the max allowed for this GPU), I'm getting a training time per epoch of about 5 hours. Is this basically how long models take for this rather compute-intensive competition, or is there something I still can do to significantly speed up training?",
      "votes": null
    },
    {
      "id": "2149318",
      "postDate": "02/18/2023 06:56:27",
      "content": "<p>I'm using A4000 GPU with 8 CPUs, and it's around 17 minutes per epoch. Try a simpler approach, and sample training data to decrease training time</p>",
      "rawMarkdown": "I'm using A4000 GPU with 8 CPUs, and it's around 17 minutes per epoch. Try a simpler approach, and sample training data to decrease training time",
      "votes": null
    },
    {
      "id": "2149537",
      "postDate": "02/18/2023 12:18:27",
      "content": "<p>Aha! 17 minutes is definitely quite a difference indeed! 👍(and seeing as you are apparently 33rd in the competition at the moment of writing, it seems to produce very good results as well!)</p>",
      "rawMarkdown": "Aha! 17 minutes is definitely quite a difference indeed! 👍(and seeing as you are apparently 33rd in the competition at the moment of writing, it seems to produce very good results as well!)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2149318,
      "author_name": "samir95",
      "author_url": "",
      "post_date": "02/18/2023 06:56:27",
      "content": "<p>I'm using A4000 GPU with 8 CPUs, and it's around 17 minutes per epoch. Try a simpler approach, and sample training data to decrease training time</p>",
      "votes": null,
      "replies": [
        {
          "id": 2149537,
          "author_name": "thomasbrekkunnvik",
          "author_url": "",
          "post_date": "02/18/2023 12:18:27",
          "content": "<p>Aha! 17 minutes is definitely quite a difference indeed! 👍(and seeing as you are apparently 33rd in the competition at the moment of writing, it seems to produce very good results as well!)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2147162": "I'm experimenting with training a model based on the following notebook: [notebook](https://www.kaggle.com/code/royalacecat/training-nfl-2-5d-cnn-lb-0-671-with-tta/notebook) (using Efficientnet B3 as the underlying model). \n\nNow, I don't own any hardware capable of doing this myself, so I have played around with various VMs on Google Cloud. With my current setup of an NVIDIA A100 GPU (40GB) with 12 CPUs (the max allowed for this GPU), I'm getting a training time per epoch of about 5 hours. Is this basically how long models take for this rather compute-intensive competition, or is there something I still can do to significantly speed up training?",
    "2149318": "I'm using A4000 GPU with 8 CPUs, and it's around 17 minutes per epoch. Try a simpler approach, and sample training data to decrease training time",
    "2149537": "Aha! 17 minutes is definitely quite a difference indeed! 👍(and seeing as you are apparently 33rd in the competition at the moment of writing, it seems to produce very good results as well!)"
  },
  "source": "meta"
}