{
  "id": 267430,
  "title": "How many epochs?",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/267430",
  "author_name": "",
  "post_date": "2021-08-23T07:46:58.831706900Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Due to the large amount of data, a question arises. How many epochs have you been training for? and how long does it take you?</p>",
  "messages": [
    {
      "id": "1486754",
      "postDate": "08/23/2021 07:46:58",
      "content": "<p>Due to the large amount of data, a question arises. How many epochs have you been training for? and how long does it take you?</p>",
      "rawMarkdown": "Due to the large amount of data, a question arises. How many epochs have you been training for? and how long does it take you?",
      "votes": null
    },
    {
      "id": "1486831",
      "postDate": "08/23/2021 08:32:58",
      "content": "<p>15-20 epochs in my case</p>",
      "rawMarkdown": "15-20 epochs in my case",
      "votes": null
    },
    {
      "id": "1487341",
      "postDate": "08/23/2021 15:04:06",
      "content": "<p>I'm stopping at 10 at the moment, but best score/loss is usually 9th or 10th epoch, which leads me to believe I can train longer. While metrics move favorably, they also move very asymptotically, which is time consuming. By converting to float16 and storing the entire dataset in ram, I'm able to do an epoch in a under three min on a single 3090. My hope was to run two experiments simultaneously, but unfortunately around third or forth fold, some weird thing happens which causes either cuda to freakout or pytorch's dataloader to segfault, so I've been running experiments serially for now.</p>",
      "rawMarkdown": "I'm stopping at 10 at the moment, but best score/loss is usually 9th or 10th epoch, which leads me to believe I can train longer. While metrics move favorably, they also move very asymptotically, which is time consuming. By converting to float16 and storing the entire dataset in ram, I'm able to do an epoch in a under three min on a single 3090. My hope was to run two experiments simultaneously, but unfortunately around third or forth fold, some weird thing happens which causes either cuda to freakout or pytorch's dataloader to segfault, so I've been running experiments serially for now.",
      "votes": null
    },
    {
      "id": "1487412",
      "postDate": "08/23/2021 16:05:49",
      "content": "<p>docker could help to run experiments simultaneously</p>",
      "rawMarkdown": "docker could help to run experiments simultaneously",
      "votes": null
    },
    {
      "id": "1487525",
      "postDate": "08/23/2021 17:12:00",
      "content": "<p>its usually depends on how big your dataset is the most formal epochs is 50 <br>\nbut jut in case you have smaller data set it can go to 100-500<br>\nbut for bigger dataset 20-25 is fine for me </p>",
      "rawMarkdown": "its usually depends on how big your dataset is the most formal epochs is 50 \nbut jut in case you have smaller data set it can go to 100-500\nbut for bigger dataset 20-25 is fine for me",
      "votes": null
    },
    {
      "id": "1489433",
      "postDate": "08/25/2021 02:09:10",
      "content": "<blockquote>\n  <p>By converting to float16 and storing the entire dataset in ram, I'm able to do an epoch in a under three min on a single 3090.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> That's impressive!! Could you elaborate more on this?</p>",
      "rawMarkdown": "> By converting to float16 and storing the entire dataset in ram, I'm able to do an epoch in a under three min on a single 3090.\n\n@authman That's impressive!! Could you elaborate more on this?",
      "votes": null
    },
    {
      "id": "1561191",
      "postDate": "10/27/2021 12:15:48",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1486831,
      "author_name": "zarif98sjs",
      "author_url": "",
      "post_date": "08/23/2021 08:32:58",
      "content": "<p>15-20 epochs in my case</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1487341,
      "author_name": "authman",
      "author_url": "",
      "post_date": "08/23/2021 15:04:06",
      "content": "<p>I'm stopping at 10 at the moment, but best score/loss is usually 9th or 10th epoch, which leads me to believe I can train longer. While metrics move favorably, they also move very asymptotically, which is time consuming. By converting to float16 and storing the entire dataset in ram, I'm able to do an epoch in a under three min on a single 3090. My hope was to run two experiments simultaneously, but unfortunately around third or forth fold, some weird thing happens which causes either cuda to freakout or pytorch's dataloader to segfault, so I've been running experiments serially for now.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1487412,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "08/23/2021 16:05:49",
          "content": "<p>docker could help to run experiments simultaneously</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1489433,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "08/25/2021 02:09:10",
          "content": "<blockquote>\n  <p>By converting to float16 and storing the entire dataset in ram, I'm able to do an epoch in a under three min on a single 3090.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> That's impressive!! Could you elaborate more on this?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1487525,
      "author_name": "shardulmehetar",
      "author_url": "",
      "post_date": "08/23/2021 17:12:00",
      "content": "<p>its usually depends on how big your dataset is the most formal epochs is 50 <br>\nbut jut in case you have smaller data set it can go to 100-500<br>\nbut for bigger dataset 20-25 is fine for me </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1561191,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 12:15:48",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1486754": "Due to the large amount of data, a question arises. How many epochs have you been training for? and how long does it take you?",
    "1486831": "15-20 epochs in my case",
    "1487341": "I'm stopping at 10 at the moment, but best score/loss is usually 9th or 10th epoch, which leads me to believe I can train longer. While metrics move favorably, they also move very asymptotically, which is time consuming. By converting to float16 and storing the entire dataset in ram, I'm able to do an epoch in a under three min on a single 3090. My hope was to run two experiments simultaneously, but unfortunately around third or forth fold, some weird thing happens which causes either cuda to freakout or pytorch's dataloader to segfault, so I've been running experiments serially for now.",
    "1487412": "docker could help to run experiments simultaneously",
    "1487525": "its usually depends on how big your dataset is the most formal epochs is 50 \nbut jut in case you have smaller data set it can go to 100-500\nbut for bigger dataset 20-25 is fine for me",
    "1489433": "> By converting to float16 and storing the entire dataset in ram, I'm able to do an epoch in a under three min on a single 3090.\n\n@authman That's impressive!! Could you elaborate more on this?",
    "1561191": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}