{
  "id": 261414,
  "title": "Big data, How do Kagglers handle it?",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/261414",
  "author_name": "Harshit Sati",
  "post_date": "2021-08-04T12:55:14.041000",
  "votes": 9,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Hey I noticed after enrolling that I have to deal with a 72GB data, how do kagglers deal with such a  big dataset?</p>",
  "messages": [
    {
      "id": 1450957,
      "postDate": "2021-08-05T08:17:29.857Z",
      "content": "<p>Hiiiiiiiiiiii :)</p>\n<p>Step 1: Start by building a small network and use a very small sample of data (I used 60,000 out of 560,000 and it did very good as a baseline in the LB).</p>\n<p>Unfortunately, you can't use more than ~60,000 within the Kaggle environment. You need to be weary of 2 things:</p>\n<ol>\n<li>How long does it take to commit: Kaggle shuts down any notebook after 9 hrs of running</li>\n<li>How much GPU you have left in the quota: this is a big one too. Don't forget to shut down your GPU after you commit a notebook, or else you'll consume double.</li>\n</ol>\n<p>Step 2: Improve on the \"small sample\"</p>\n<p>Step 3: Try Google Colab or a local device using GPU. I have an NVIDIA Quadro RTX 5000 on my workstation and it took ~15 hrs to train on the entire dataset. Remember that it takes ~ 2-3 hrs (on GPU) for the submission to be created as well.</p>\n<p>Good luck! 😁</p>",
      "rawMarkdown": "Hiiiiiiiiiiii :)\n\nStep 1: Start by building a small network and use a very small sample of data (I used 60,000 out of 560,000 and it did very good as a baseline in the LB).\n\nUnfortunately, you can't use more than ~60,000 within the Kaggle environment. You need to be weary of 2 things:\n1. How long does it take to commit: Kaggle shuts down any notebook after 9 hrs of running\n2. How much GPU you have left in the quota: this is a big one too. Don't forget to shut down your GPU after you commit a notebook, or else you'll consume double.\n\nStep 2: Improve on the \"small sample\"\n\nStep 3: Try Google Colab or a local device using GPU. I have an NVIDIA Quadro RTX 5000 on my workstation and it took ~15 hrs to train on the entire dataset. Remember that it takes ~ 2-3 hrs (on GPU) for the submission to be created as well.\n\nGood luck! 😁",
      "votes": 9,
      "replies": [
        {
          "id": 1456716,
          "postDate": "2021-08-07T04:17:34.133Z",
          "content": "<p>Whoa NVIDIA Quadro RTX 5000 🙌, thank you for the tips, I'll start with a short dataset right away :D</p>",
          "rawMarkdown": "Whoa NVIDIA Quadro RTX 5000 🙌, thank you for the tips, I'll start with a short dataset right away :D",
          "votes": 1
        }
      ]
    },
    {
      "id": 1446700,
      "postDate": "2021-08-04T12:55:14.040Z",
      "content": "<p>Hey I noticed after enrolling that I have to deal with a 72GB data, how do kagglers deal with such a  big dataset?</p>",
      "rawMarkdown": "Hey I noticed after enrolling that I have to deal with a 72GB data, how do kagglers deal with such a  big dataset?",
      "votes": 9
    },
    {
      "id": 1460516,
      "postDate": "2021-08-08T23:34:17.160Z",
      "content": "<p>Using TPUs and TFRecords I can train a single epoch in about 4m. See the discussion <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/261721\" target=\"_blank\">Things I've tried and how they worked out</a> for more details. </p>",
      "rawMarkdown": "Using TPUs and TFRecords I can train a single epoch in about 4m. See the discussion [Things I've tried and how they worked out](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/261721) for more details. ",
      "votes": 3,
      "replies": [
        {
          "id": 1463519,
          "postDate": "2021-08-10T08:13:55.273Z",
          "content": "<p>thank you I'll check it out for sure!</p>",
          "rawMarkdown": "thank you I'll check it out for sure!",
          "votes": 1
        },
        {
          "id": 1465333,
          "postDate": "2021-08-11T02:32:07.997Z",
          "content": "<p>BTW, I recently share some of the code and made the TRrecords public so you can more easily try this out. </p>",
          "rawMarkdown": "BTW, I recently share some of the code and made the TRrecords public so you can more easily try this out. ",
          "votes": 2
        },
        {
          "id": 1467159,
          "postDate": "2021-08-11T20:11:18.860Z",
          "content": "<p>Thak you for sharing</p>",
          "rawMarkdown": "Thak you for sharing"
        }
      ]
    },
    {
      "id": 1464182,
      "postDate": "2021-08-10T13:25:32.057Z",
      "content": "<p>Another thing I would point out -- don't feel obligated to throw models like EfficientnetB7 at the problem just yet. Those take way too long to train and will hinder your ability to explore new ways of handling the data. My current LB position is purely based on EfficientnetB0.</p>",
      "rawMarkdown": "Another thing I would point out -- don't feel obligated to throw models like EfficientnetB7 at the problem just yet. Those take way too long to train and will hinder your ability to explore new ways of handling the data. My current LB position is purely based on EfficientnetB0.",
      "votes": 4,
      "replies": [
        {
          "id": 1465742,
          "postDate": "2021-08-11T06:55:54.147Z",
          "content": "<p>whoa thank you for the heads up! </p>",
          "rawMarkdown": "whoa thank you for the heads up! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1447386,
      "postDate": "2021-08-04T14:47:37.520Z",
      "content": "<p>As a beginner I have been reading some of the codes on the \"Code\" tab above. Some of the notebooks contain data loading processes in which you might get a hang of it. This code does contain \"data processing\" section:<br>\n<a href=\"https://www.kaggle.com/xuzongniubi/g2net-efficientnet-b7-baseline-training\" target=\"_blank\">https://www.kaggle.com/xuzongniubi/g2net-efficientnet-b7-baseline-training</a></p>\n<p>Hope this helps:)</p>",
      "rawMarkdown": "As a beginner I have been reading some of the codes on the \"Code\" tab above. Some of the notebooks contain data loading processes in which you might get a hang of it. This code does contain \"data processing\" section:\nhttps://www.kaggle.com/xuzongniubi/g2net-efficientnet-b7-baseline-training\n\nHope this helps:)\n",
      "votes": 1,
      "replies": [
        {
          "id": 1456721,
          "postDate": "2021-08-07T04:20:13.153Z",
          "content": "<p>Thank you, sure I'll check it out!</p>",
          "rawMarkdown": "Thank you, sure I'll check it out!"
        }
      ]
    },
    {
      "id": 1451067,
      "postDate": "2021-08-05T08:52:30.120Z",
      "content": "<p>I think this is different for everyone. For me it's generally about experimenting quickly. That means I:</p>\n<ol>\n<li>try to optimise my code as much as possible. There's a related post with some nice tips and links over here <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/254775\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/254775</a></li>\n<li>prioritise my training around looking for a solution which starts off well and keeps getting better (as opposed to a solution that starts slowly/does badly at first but is great in the long run). I do this by plotting a lot of stuff inside of an epoch and killing it early if it's not doing well. Sure I'm limiting the space of solutions but this lets me try out a lot of models quickly so hopefully that makes up for it</li>\n</ol>\n<p>edit</p>\n<ol>\n<li>oh another one (one I'm terrible at), don't chase the leaderboard, try and figure out a strategy which will work over the course of a competition. For example I'm only submitting a single fold right now and have only done that to make sure it works as I expect. That way I'm not wasting time on something which isn't going to be a viable model</li>\n</ol>",
      "rawMarkdown": "I think this is different for everyone. For me it's generally about experimenting quickly. That means I:\n1. try to optimise my code as much as possible. There's a related post with some nice tips and links over here https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/254775\n2. prioritise my training around looking for a solution which starts off well and keeps getting better (as opposed to a solution that starts slowly/does badly at first but is great in the long run). I do this by plotting a lot of stuff inside of an epoch and killing it early if it's not doing well. Sure I'm limiting the space of solutions but this lets me try out a lot of models quickly so hopefully that makes up for it\n\nedit\n3. oh another one (one I'm terrible at), don't chase the leaderboard, try and figure out a strategy which will work over the course of a competition. For example I'm only submitting a single fold right now and have only done that to make sure it works as I expect. That way I'm not wasting time on something which isn't going to be a viable model",
      "votes": 2,
      "replies": [
        {
          "id": 1456712,
          "postDate": "2021-08-07T04:14:43.757Z",
          "content": "<p>Great tip, noted, thank you so much for answering!! :)</p>",
          "rawMarkdown": "Great tip, noted, thank you so much for answering!! :)"
        }
      ]
    },
    {
      "id": 1447807,
      "postDate": "2021-08-04T15:58:28.377Z",
      "content": "<p>Start with a tiny network and create a strong baseline. Look how far you can push your CV/LB before trying out something big.</p>",
      "rawMarkdown": "Start with a tiny network and create a strong baseline. Look how far you can push your CV/LB before trying out something big.",
      "votes": 2,
      "replies": [
        {
          "id": 1456719,
          "postDate": "2021-08-07T04:19:11.280Z",
          "content": "<p>do you mind telling this noob what CV/LB is? 😅 </p>",
          "rawMarkdown": "do you mind telling this noob what CV/LB is? 😅 "
        },
        {
          "id": 1456810,
          "postDate": "2021-08-07T05:16:26.700Z",
          "content": "<p>CV - Cross Validation. How well does your model perform on unseen data from training set.<br>\nLB - Leaderboard.</p>",
          "rawMarkdown": "CV - Cross Validation. How well does your model perform on unseen data from training set.\nLB - Leaderboard."
        },
        {
          "id": 1456835,
          "postDate": "2021-08-07T05:35:11.843Z",
          "content": "<p>Ah yes thank you</p>",
          "rawMarkdown": "Ah yes thank you"
        }
      ]
    },
    {
      "id": 1561206,
      "postDate": "2021-10-27T12:19:39.483Z",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
    },
    {
      "id": 1451838,
      "postDate": "2021-08-05T12:54:56.097Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 1456709,
          "postDate": "2021-08-07T04:12:35.890Z",
          "content": "<p>Noted, thank you so much!</p>",
          "rawMarkdown": "Noted, thank you so much!"
        }
      ]
    },
    {
      "id": 1447380,
      "postDate": "2021-08-04T14:46:29.080Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1450957,
      "author_name": "Andrada",
      "author_url": "",
      "post_date": "2021-08-05T08:17:29.857000",
      "content": "<p>Hiiiiiiiiiiii :)</p>\n<p>Step 1: Start by building a small network and use a very small sample of data (I used 60,000 out of 560,000 and it did very good as a baseline in the LB).</p>\n<p>Unfortunately, you can't use more than ~60,000 within the Kaggle environment. You need to be weary of 2 things:</p>\n<ol>\n<li>How long does it take to commit: Kaggle shuts down any notebook after 9 hrs of running</li>\n<li>How much GPU you have left in the quota: this is a big one too. Don't forget to shut down your GPU after you commit a notebook, or else you'll consume double.</li>\n</ol>\n<p>Step 2: Improve on the \"small sample\"</p>\n<p>Step 3: Try Google Colab or a local device using GPU. I have an NVIDIA Quadro RTX 5000 on my workstation and it took ~15 hrs to train on the entire dataset. Remember that it takes ~ 2-3 hrs (on GPU) for the submission to be created as well.</p>\n<p>Good luck! 😁</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1456716,
          "author_name": "Harshit Sati",
          "author_url": "",
          "post_date": "2021-08-07T04:17:34.133000",
          "content": "<p>Whoa NVIDIA Quadro RTX 5000 🙌, thank you for the tips, I'll start with a short dataset right away :D</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1460516,
      "author_name": "Fractal Feelings",
      "author_url": "",
      "post_date": "2021-08-08T23:34:17.160000",
      "content": "<p>Using TPUs and TFRecords I can train a single epoch in about 4m. See the discussion <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/261721\" target=\"_blank\">Things I've tried and how they worked out</a> for more details. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1463519,
          "author_name": "Harshit Sati",
          "author_url": "",
          "post_date": "2021-08-10T08:13:55.273000",
          "content": "<p>thank you I'll check it out for sure!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1465333,
          "author_name": "Fractal Feelings",
          "author_url": "",
          "post_date": "2021-08-11T02:32:07.997000",
          "content": "<p>BTW, I recently share some of the code and made the TRrecords public so you can more easily try this out. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1467159,
          "author_name": "Jurajicek",
          "author_url": "",
          "post_date": "2021-08-11T20:11:18.860000",
          "content": "<p>Thak you for sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1464182,
      "author_name": "Mark Tenenholtz",
      "author_url": "",
      "post_date": "2021-08-10T13:25:32.057000",
      "content": "<p>Another thing I would point out -- don't feel obligated to throw models like EfficientnetB7 at the problem just yet. Those take way too long to train and will hinder your ability to explore new ways of handling the data. My current LB position is purely based on EfficientnetB0.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1465742,
          "author_name": "Harshit Sati",
          "author_url": "",
          "post_date": "2021-08-11T06:55:54.147000",
          "content": "<p>whoa thank you for the heads up! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1447386,
      "author_name": "T Doug K",
      "author_url": "",
      "post_date": "2021-08-04T14:47:37.520000",
      "content": "<p>As a beginner I have been reading some of the codes on the \"Code\" tab above. Some of the notebooks contain data loading processes in which you might get a hang of it. This code does contain \"data processing\" section:<br>\n<a href=\"https://www.kaggle.com/xuzongniubi/g2net-efficientnet-b7-baseline-training\" target=\"_blank\">https://www.kaggle.com/xuzongniubi/g2net-efficientnet-b7-baseline-training</a></p>\n<p>Hope this helps:)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1456721,
          "author_name": "Harshit Sati",
          "author_url": "",
          "post_date": "2021-08-07T04:20:13.153000",
          "content": "<p>Thank you, sure I'll check it out!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1451067,
      "author_name": "Hamish",
      "author_url": "",
      "post_date": "2021-08-05T08:52:30.120000",
      "content": "<p>I think this is different for everyone. For me it's generally about experimenting quickly. That means I:</p>\n<ol>\n<li>try to optimise my code as much as possible. There's a related post with some nice tips and links over here <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/254775\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/254775</a></li>\n<li>prioritise my training around looking for a solution which starts off well and keeps getting better (as opposed to a solution that starts slowly/does badly at first but is great in the long run). I do this by plotting a lot of stuff inside of an epoch and killing it early if it's not doing well. Sure I'm limiting the space of solutions but this lets me try out a lot of models quickly so hopefully that makes up for it</li>\n</ol>\n<p>edit</p>\n<ol>\n<li>oh another one (one I'm terrible at), don't chase the leaderboard, try and figure out a strategy which will work over the course of a competition. For example I'm only submitting a single fold right now and have only done that to make sure it works as I expect. That way I'm not wasting time on something which isn't going to be a viable model</li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 1456712,
          "author_name": "Harshit Sati",
          "author_url": "",
          "post_date": "2021-08-07T04:14:43.757000",
          "content": "<p>Great tip, noted, thank you so much for answering!! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1447807,
      "author_name": "Harsh Patel",
      "author_url": "",
      "post_date": "2021-08-04T15:58:28.377000",
      "content": "<p>Start with a tiny network and create a strong baseline. Look how far you can push your CV/LB before trying out something big.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1456719,
          "author_name": "Harshit Sati",
          "author_url": "",
          "post_date": "2021-08-07T04:19:11.280000",
          "content": "<p>do you mind telling this noob what CV/LB is? 😅 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1456810,
          "author_name": "Harsh Patel",
          "author_url": "",
          "post_date": "2021-08-07T05:16:26.700000",
          "content": "<p>CV - Cross Validation. How well does your model perform on unseen data from training set.<br>\nLB - Leaderboard.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1456835,
          "author_name": "Harshit Sati",
          "author_url": "",
          "post_date": "2021-08-07T05:35:11.843000",
          "content": "<p>Ah yes thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1561206,
      "author_name": "ChristopherZerafa",
      "author_url": "",
      "post_date": "2021-10-27T12:19:39.483000",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1451838,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-05T12:54:56.097000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 1456709,
          "author_name": "Harshit Sati",
          "author_url": "",
          "post_date": "2021-08-07T04:12:35.890000",
          "content": "<p>Noted, thank you so much!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1447380,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-04T14:46:29.080000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1450957": "Hiiiiiiiiiiii :)\n\nStep 1: Start by building a small network and use a very small sample of data (I used 60,000 out of 560,000 and it did very good as a baseline in the LB).\n\nUnfortunately, you can't use more than ~60,000 within the Kaggle environment. You need to be weary of 2 things:\n1. How long does it take to commit: Kaggle shuts down any notebook after 9 hrs of running\n2. How much GPU you have left in the quota: this is a big one too. Don't forget to shut down your GPU after you commit a notebook, or else you'll consume double.\n\nStep 2: Improve on the \"small sample\"\n\nStep 3: Try Google Colab or a local device using GPU. I have an NVIDIA Quadro RTX 5000 on my workstation and it took ~15 hrs to train on the entire dataset. Remember that it takes ~ 2-3 hrs (on GPU) for the submission to be created as well.\n\nGood luck! 😁",
    "1446700": "Hey I noticed after enrolling that I have to deal with a 72GB data, how do kagglers deal with such a  big dataset?",
    "1460516": "Using TPUs and TFRecords I can train a single epoch in about 4m. See the discussion [Things I've tried and how they worked out](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/261721) for more details. ",
    "1464182": "Another thing I would point out -- don't feel obligated to throw models like EfficientnetB7 at the problem just yet. Those take way too long to train and will hinder your ability to explore new ways of handling the data. My current LB position is purely based on EfficientnetB0.",
    "1447386": "As a beginner I have been reading some of the codes on the \"Code\" tab above. Some of the notebooks contain data loading processes in which you might get a hang of it. This code does contain \"data processing\" section:\nhttps://www.kaggle.com/xuzongniubi/g2net-efficientnet-b7-baseline-training\n\nHope this helps:)\n",
    "1451067": "I think this is different for everyone. For me it's generally about experimenting quickly. That means I:\n1. try to optimise my code as much as possible. There's a related post with some nice tips and links over here https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/254775\n2. prioritise my training around looking for a solution which starts off well and keeps getting better (as opposed to a solution that starts slowly/does badly at first but is great in the long run). I do this by plotting a lot of stuff inside of an epoch and killing it early if it's not doing well. Sure I'm limiting the space of solutions but this lets me try out a lot of models quickly so hopefully that makes up for it\n\nedit\n3. oh another one (one I'm terrible at), don't chase the leaderboard, try and figure out a strategy which will work over the course of a competition. For example I'm only submitting a single fold right now and have only done that to make sure it works as I expect. That way I'm not wasting time on something which isn't going to be a viable model",
    "1447807": "Start with a tiny network and create a strong baseline. Look how far you can push your CV/LB before trying out something big.",
    "1561206": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1451838": "",
    "1447380": ""
  }
}