{
  "id": 221421,
  "title": "How to handle Large datasets",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/221421",
  "author_name": "Ikram Ul Haq",
  "post_date": "2021-02-22T16:44:33.900000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi,<br>\nHope you all are doing well. I was wondering if you people can advise based on your experience that how can one deal with these large datasets?<br>\nThis running competition contains 158 GBs of data. handling that without GPU is not valuable and Kaggle currently providing me 49 hours of GPU for a month. Is there anything I can do about it?<br>\nThank you.</p>",
  "messages": [
    {
      "id": 1214197,
      "postDate": "2021-02-22T16:44:33.900Z",
      "content": "<p>Hi,<br>\nHope you all are doing well. I was wondering if you people can advise based on your experience that how can one deal with these large datasets?<br>\nThis running competition contains 158 GBs of data. handling that without GPU is not valuable and Kaggle currently providing me 49 hours of GPU for a month. Is there anything I can do about it?<br>\nThank you.</p>",
      "rawMarkdown": "Hi,\nHope you all are doing well. I was wondering if you people can advise based on your experience that how can one deal with these large datasets?\nThis running competition contains 158 GBs of data. handling that without GPU is not valuable and Kaggle currently providing me 49 hours of GPU for a month. Is there anything I can do about it?\nThank you.",
      "votes": 1
    },
    {
      "id": 1217003,
      "postDate": "2021-02-24T17:36:39.097Z",
      "content": "<p>Maybe don't train on the full resolution images. I think if you try to inference on the full res images later, you might run into problems on Kaggle?</p>",
      "rawMarkdown": "Maybe don't train on the full resolution images. I think if you try to inference on the full res images later, you might run into problems on Kaggle?\n"
    },
    {
      "id": 1214767,
      "postDate": "2021-02-23T05:42:52.083Z",
      "content": "<p>I've been only working on this competition via Kaggle kernels so far, my approach is to create a smaller sample dataset (10GB) and use it for experimentation. I'm planning to release it as soon as I get a newsworthy score :) </p>",
      "rawMarkdown": "I've been only working on this competition via Kaggle kernels so far, my approach is to create a smaller sample dataset (10GB) and use it for experimentation. I'm planning to release it as soon as I get a newsworthy score :) "
    },
    {
      "id": 1214347,
      "postDate": "2021-02-22T19:15:45.613Z",
      "content": "<p>Some kind souls have shared data sets containing segments and segmented images.  If Kaggle is your only compute than you will want to use those shared data sets.  I am building my own set of single cell images and only about 1/2 thru the 21K of training images after around 30 hours on my dual GPU Ubuntu machine.  It would be a real pain to build that set of images in 9 hour increments on a kaggle GPU session.</p>\n<p>Once you have single cell images and the segments most other stuff to be done can be run in 9 hour sessions on GPU.  The competition has a long time yet to run - with good planning 49 hours per week is capable of getting decent result.  In TensorFlow it's very easy to convert a 40 hour training run into smaller pieces that fit the 9 hour limit.</p>\n<p>Your skill set may complement someone who has lots of compute but less skills - so create a good ad for your self in the \"wanting team\" discussion post.   You may find folks needing your skills.  Your profile suggests you would be a great partner for someone like myself - I have 4 dual GPU Ubuntu machines and a 10TB server but my Python skill set is still elementary school level :)  </p>\n<p>While a team only gets 5 submissions per day, a team of two has double the compute hours to play with.    Getting good team mates is however a bit of a problem.  Lots promise to be busy on the competition but end up not putting in the work to make a team worth the effort.  </p>",
      "rawMarkdown": "Some kind souls have shared data sets containing segments and segmented images.  If Kaggle is your only compute than you will want to use those shared data sets.  I am building my own set of single cell images and only about 1/2 thru the 21K of training images after around 30 hours on my dual GPU Ubuntu machine.  It would be a real pain to build that set of images in 9 hour increments on a kaggle GPU session.\n\nOnce you have single cell images and the segments most other stuff to be done can be run in 9 hour sessions on GPU.  The competition has a long time yet to run - with good planning 49 hours per week is capable of getting decent result.  In TensorFlow it's very easy to convert a 40 hour training run into smaller pieces that fit the 9 hour limit.\n\nYour skill set may complement someone who has lots of compute but less skills - so create a good ad for your self in the \"wanting team\" discussion post.   You may find folks needing your skills.  Your profile suggests you would be a great partner for someone like myself - I have 4 dual GPU Ubuntu machines and a 10TB server but my Python skill set is still elementary school level :)  \n\nWhile a team only gets 5 submissions per day, a team of two has double the compute hours to play with.    Getting good team mates is however a bit of a problem.  Lots promise to be busy on the competition but end up not putting in the work to make a team worth the effort.  ",
      "replies": [
        {
          "id": 1216842,
          "postDate": "2021-02-24T14:24:52.590Z",
          "content": "<p>You had me at 4 dual gpu machines , i have one single 3 gpu machine. Are you looking to team up in the near future?</p>",
          "rawMarkdown": "You had me at 4 dual gpu machines , i have one single 3 gpu machine. Are you looking to team up in the near future?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1217003,
      "author_name": "Alexander Riedel",
      "author_url": "",
      "post_date": "2021-02-24T17:36:39.097000",
      "content": "<p>Maybe don't train on the full resolution images. I think if you try to inference on the full res images later, you might run into problems on Kaggle?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1214767,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-02-23T05:42:52.083000",
      "content": "<p>I've been only working on this competition via Kaggle kernels so far, my approach is to create a smaller sample dataset (10GB) and use it for experimentation. I'm planning to release it as soon as I get a newsworthy score :) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1214347,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-02-22T19:15:45.613000",
      "content": "<p>Some kind souls have shared data sets containing segments and segmented images.  If Kaggle is your only compute than you will want to use those shared data sets.  I am building my own set of single cell images and only about 1/2 thru the 21K of training images after around 30 hours on my dual GPU Ubuntu machine.  It would be a real pain to build that set of images in 9 hour increments on a kaggle GPU session.</p>\n<p>Once you have single cell images and the segments most other stuff to be done can be run in 9 hour sessions on GPU.  The competition has a long time yet to run - with good planning 49 hours per week is capable of getting decent result.  In TensorFlow it's very easy to convert a 40 hour training run into smaller pieces that fit the 9 hour limit.</p>\n<p>Your skill set may complement someone who has lots of compute but less skills - so create a good ad for your self in the \"wanting team\" discussion post.   You may find folks needing your skills.  Your profile suggests you would be a great partner for someone like myself - I have 4 dual GPU Ubuntu machines and a 10TB server but my Python skill set is still elementary school level :)  </p>\n<p>While a team only gets 5 submissions per day, a team of two has double the compute hours to play with.    Getting good team mates is however a bit of a problem.  Lots promise to be busy on the competition but end up not putting in the work to make a team worth the effort.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1216842,
          "author_name": "Felipe Bivort Haiek",
          "author_url": "",
          "post_date": "2021-02-24T14:24:52.590000",
          "content": "<p>You had me at 4 dual gpu machines , i have one single 3 gpu machine. Are you looking to team up in the near future?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1214197": "Hi,\nHope you all are doing well. I was wondering if you people can advise based on your experience that how can one deal with these large datasets?\nThis running competition contains 158 GBs of data. handling that without GPU is not valuable and Kaggle currently providing me 49 hours of GPU for a month. Is there anything I can do about it?\nThank you.",
    "1217003": "Maybe don't train on the full resolution images. I think if you try to inference on the full res images later, you might run into problems on Kaggle?\n",
    "1214767": "I've been only working on this competition via Kaggle kernels so far, my approach is to create a smaller sample dataset (10GB) and use it for experimentation. I'm planning to release it as soon as I get a newsworthy score :) ",
    "1214347": "Some kind souls have shared data sets containing segments and segmented images.  If Kaggle is your only compute than you will want to use those shared data sets.  I am building my own set of single cell images and only about 1/2 thru the 21K of training images after around 30 hours on my dual GPU Ubuntu machine.  It would be a real pain to build that set of images in 9 hour increments on a kaggle GPU session.\n\nOnce you have single cell images and the segments most other stuff to be done can be run in 9 hour sessions on GPU.  The competition has a long time yet to run - with good planning 49 hours per week is capable of getting decent result.  In TensorFlow it's very easy to convert a 40 hour training run into smaller pieces that fit the 9 hour limit.\n\nYour skill set may complement someone who has lots of compute but less skills - so create a good ad for your self in the \"wanting team\" discussion post.   You may find folks needing your skills.  Your profile suggests you would be a great partner for someone like myself - I have 4 dual GPU Ubuntu machines and a 10TB server but my Python skill set is still elementary school level :)  \n\nWhile a team only gets 5 submissions per day, a team of two has double the compute hours to play with.    Getting good team mates is however a bit of a problem.  Lots promise to be busy on the competition but end up not putting in the work to make a team worth the effort.  "
  }
}