{
  "id": 55754,
  "title": "How to handle the amount of data ?",
  "url": "/competitions/trackml-particle-identification/discussion/55754",
  "author_name": "",
  "post_date": "2018-05-01T10:59:16.978406100Z",
  "votes": 6,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I am extremely interessted by this competition, but the amount of data is just so huge..\nI wanted to know if you know ways to be able to process all that without needing a crazy computer. Moreover will this competition propose some account and/or discount  for cloud computing ?</p>",
  "messages": [
    {
      "id": "321462",
      "postDate": "05/01/2018 10:59:16",
      "content": "<p>I am extremely interessted by this competition, but the amount of data is just so huge..\nI wanted to know if you know ways to be able to process all that without needing a crazy computer. Moreover will this competition propose some account and/or discount  for cloud computing ?</p>",
      "rawMarkdown": "I am extremely interessted by this competition, but the amount of data is just so huge..\nI wanted to know if you know ways to be able to process all that without needing a crazy computer. Moreover will this competition propose some account and/or discount  for cloud computing ?",
      "votes": null
    },
    {
      "id": "321551",
      "postDate": "05/01/2018 14:33:28",
      "content": "<p>Unfortunately, we are not able to provide GCP credits for this competition. </p>",
      "rawMarkdown": "Unfortunately, we are not able to provide GCP credits for this competition.",
      "votes": null
    },
    {
      "id": "321554",
      "postDate": "05/01/2018 14:46:20",
      "content": "<p>Floydhub gives you trail credit upon signing up. </p>",
      "rawMarkdown": "Floydhub gives you trail credit upon signing up.",
      "votes": null
    },
    {
      "id": "321774",
      "postDate": "05/01/2018 21:30:59",
      "content": "<p>We would advise to start with train_sample with 100 events to get a first feeling. We provided so much data because we could; it does not mean that only algorithm/code able to swallow all this data will have a good score.</p>",
      "rawMarkdown": "We would advise to start with train_sample with 100 events to get a first feeling. We provided so much data because we could; it does not mean that only algorithm/code able to swallow all this data will have a good score.",
      "votes": null
    },
    {
      "id": "321775",
      "postDate": "05/01/2018 21:39:39",
      "content": "<p>This is a really great point. Maybe worth creating a new forum topic advising people that they can / should start with a smaller amount of data.</p>",
      "rawMarkdown": "This is a really great point. Maybe worth creating a new forum topic advising people that they can / should start with a smaller amount of data.",
      "votes": null
    },
    {
      "id": "321776",
      "postDate": "05/01/2018 21:40:08",
      "content": "<p>I bet people will be less intimidated with getting started.</p>",
      "rawMarkdown": "I bet people will be less intimidated with getting started.",
      "votes": null
    },
    {
      "id": "321805",
      "postDate": "05/01/2018 23:45:13",
      "content": "<p>A small sample size would be good to build your model as well.</p>",
      "rawMarkdown": "A small sample size would be good to build your model as well.",
      "votes": null
    },
    {
      "id": "321810",
      "postDate": "05/02/2018 00:09:30",
      "content": "<p>Even the training set completely eats up the 17.2 GB if it's loaded into an online session. Not that it'd be a good idea to, just playing around and it happened to me. The trackml-library tools are powerful!<img src=\"https://image.ibb.co/fsckTn/Screen_Shot_2018_05_01_at_4_35_35_PM.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Even the training set completely eats up the 17.2 GB if it's loaded into an online session. Not that it'd be a good idea to, just playing around and it happened to me. The trackml-library tools are powerful!![enter image description here][1]\n\n\n  [1]: https://image.ibb.co/fsckTn/Screen_Shot_2018_05_01_at_4_35_35_PM.png",
      "votes": null
    },
    {
      "id": "322082",
      "postDate": "05/02/2018 11:56:18",
      "content": "<p>I kinda thought that a person who didn't use all the data didn't really have any chance to have a ~really good score</p>",
      "rawMarkdown": "I kinda thought that a person who didn't use all the data didn't really have any chance to have a ~really good score",
      "votes": null
    },
    {
      "id": "325564",
      "postDate": "05/08/2018 14:50:14",
      "content": "<p>Indeed. I was quite intimidated in the beginning. </p>",
      "rawMarkdown": "Indeed. I was quite intimidated in the beginning.",
      "votes": null
    },
    {
      "id": "326295",
      "postDate": "05/09/2018 14:21:59",
      "content": "<p>Look up learning curves for models, there will be a point were the amount of data used to train your model starts to give very little improvement</p>",
      "rawMarkdown": "Look up learning curves for models, there will be a point were the amount of data used to train your model starts to give very little improvement",
      "votes": null
    },
    {
      "id": "326334",
      "postDate": "05/09/2018 15:23:01",
      "content": "<p>It would be very helpful if the data was on a Google Cloud bucket, S3 bucket or something similar that we could use to copy the to our own cloud storage. Something similar was done for the TSA competition.</p>",
      "rawMarkdown": "It would be very helpful if the data was on a Google Cloud bucket, S3 bucket or something similar that we could use to copy the to our own cloud storage. Something similar was done for the TSA competition.",
      "votes": null
    },
    {
      "id": "326338",
      "postDate": "05/09/2018 15:27:35",
      "content": "<p>I have been using FloydHub for competitions that involved smaller datasets but their current data import capability makes it very difficult to work with large datasets, especially if you have a slow internet connection. I am going to try spin up an ec2 instance later, download the dataset to it and try upload to FloydHub from there.</p>",
      "rawMarkdown": "I have been using FloydHub for competitions that involved smaller datasets but their current data import capability makes it very difficult to work with large datasets, especially if you have a slow internet connection. I am going to try spin up an ec2 instance later, download the dataset to it and try upload to FloydHub from there.",
      "votes": null
    },
    {
      "id": "328575",
      "postDate": "05/14/2018 16:28:24",
      "content": "<p>@Jack  Can you let me know how the data upload to Floydhub works for you?</p>",
      "rawMarkdown": "Jack  Can you let me know how the data upload to Floydhub works for you?",
      "votes": null
    },
    {
      "id": "328749",
      "postDate": "05/15/2018 02:08:18",
      "content": "<p>I ended up not going with Floyd Hub since it required paying a extra $15 per month for storage and have instead decided to take the leap and build my own machine for DL.</p>",
      "rawMarkdown": "I ended up not going with Floyd Hub since it required paying a extra $15 per month for storage and have instead decided to take the leap and build my own machine for DL.",
      "votes": null
    },
    {
      "id": "337974",
      "postDate": "06/04/2018 06:42:16",
      "content": "<p>I loaded the CSV data into a Google Cloud bucket.  I then put the data into a Big Query dataset.  Does anyone want information on how to do this?\n<img src=\"http://8strong.com/trackml/dataset.png\" alt=\"dataset in GCP Big Query\"></p>",
      "rawMarkdown": "I loaded the CSV data into a Google Cloud bucket.  I then put the data into a Big Query dataset.  Does anyone want information on how to do this?\n![dataset in GCP Big Query][1]\n  [1]: http://8strong.com/trackml/dataset.png",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 321551,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/01/2018 14:33:28",
      "content": "<p>Unfortunately, we are not able to provide GCP credits for this competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 321554,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "05/01/2018 14:46:20",
      "content": "<p>Floydhub gives you trail credit upon signing up. </p>",
      "votes": null,
      "replies": [
        {
          "id": 326338,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "05/09/2018 15:27:35",
          "content": "<p>I have been using FloydHub for competitions that involved smaller datasets but their current data import capability makes it very difficult to work with large datasets, especially if you have a slow internet connection. I am going to try spin up an ec2 instance later, download the dataset to it and try upload to FloydHub from there.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328575,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "05/14/2018 16:28:24",
          "content": "<p>@Jack  Can you let me know how the data upload to Floydhub works for you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328749,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "05/15/2018 02:08:18",
          "content": "<p>I ended up not going with Floyd Hub since it required paying a extra $15 per month for storage and have instead decided to take the leap and build my own machine for DL.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321774,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "05/01/2018 21:30:59",
      "content": "<p>We would advise to start with train_sample with 100 events to get a first feeling. We provided so much data because we could; it does not mean that only algorithm/code able to swallow all this data will have a good score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 321775,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "05/01/2018 21:39:39",
          "content": "<p>This is a really great point. Maybe worth creating a new forum topic advising people that they can / should start with a smaller amount of data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321776,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "05/01/2018 21:40:08",
          "content": "<p>I bet people will be less intimidated with getting started.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321810,
          "author_name": "alleyvolta",
          "author_url": "",
          "post_date": "05/02/2018 00:09:30",
          "content": "<p>Even the training set completely eats up the 17.2 GB if it's loaded into an online session. Not that it'd be a good idea to, just playing around and it happened to me. The trackml-library tools are powerful!<img src=\"https://image.ibb.co/fsckTn/Screen_Shot_2018_05_01_at_4_35_35_PM.png\" alt=\"enter image description here\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325564,
          "author_name": "sarnavim",
          "author_url": "",
          "post_date": "05/08/2018 14:50:14",
          "content": "<p>Indeed. I was quite intimidated in the beginning. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321805,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "05/01/2018 23:45:13",
      "content": "<p>A small sample size would be good to build your model as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 322082,
          "author_name": "thomashirtz",
          "author_url": "",
          "post_date": "05/02/2018 11:56:18",
          "content": "<p>I kinda thought that a person who didn't use all the data didn't really have any chance to have a ~really good score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326295,
          "author_name": "eroticzombie",
          "author_url": "",
          "post_date": "05/09/2018 14:21:59",
          "content": "<p>Look up learning curves for models, there will be a point were the amount of data used to train your model starts to give very little improvement</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326334,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "05/09/2018 15:23:01",
      "content": "<p>It would be very helpful if the data was on a Google Cloud bucket, S3 bucket or something similar that we could use to copy the to our own cloud storage. Something similar was done for the TSA competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 337974,
      "author_name": "chceight",
      "author_url": "",
      "post_date": "06/04/2018 06:42:16",
      "content": "<p>I loaded the CSV data into a Google Cloud bucket.  I then put the data into a Big Query dataset.  Does anyone want information on how to do this?\n<img src=\"http://8strong.com/trackml/dataset.png\" alt=\"dataset in GCP Big Query\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "321462": "I am extremely interessted by this competition, but the amount of data is just so huge..\nI wanted to know if you know ways to be able to process all that without needing a crazy computer. Moreover will this competition propose some account and/or discount  for cloud computing ?",
    "321551": "Unfortunately, we are not able to provide GCP credits for this competition.",
    "321554": "Floydhub gives you trail credit upon signing up.",
    "321774": "We would advise to start with train_sample with 100 events to get a first feeling. We provided so much data because we could; it does not mean that only algorithm/code able to swallow all this data will have a good score.",
    "321775": "This is a really great point. Maybe worth creating a new forum topic advising people that they can / should start with a smaller amount of data.",
    "321776": "I bet people will be less intimidated with getting started.",
    "321805": "A small sample size would be good to build your model as well.",
    "321810": "Even the training set completely eats up the 17.2 GB if it's loaded into an online session. Not that it'd be a good idea to, just playing around and it happened to me. The trackml-library tools are powerful!![enter image description here][1]\n\n\n  [1]: https://image.ibb.co/fsckTn/Screen_Shot_2018_05_01_at_4_35_35_PM.png",
    "322082": "I kinda thought that a person who didn't use all the data didn't really have any chance to have a ~really good score",
    "325564": "Indeed. I was quite intimidated in the beginning.",
    "326295": "Look up learning curves for models, there will be a point were the amount of data used to train your model starts to give very little improvement",
    "326334": "It would be very helpful if the data was on a Google Cloud bucket, S3 bucket or something similar that we could use to copy the to our own cloud storage. Something similar was done for the TSA competition.",
    "326338": "I have been using FloydHub for competitions that involved smaller datasets but their current data import capability makes it very difficult to work with large datasets, especially if you have a slow internet connection. I am going to try spin up an ec2 instance later, download the dataset to it and try upload to FloydHub from there.",
    "328575": "Jack  Can you let me know how the data upload to Floydhub works for you?",
    "328749": "I ended up not going with Floyd Hub since it required paying a extra $15 per month for storage and have instead decided to take the leap and build my own machine for DL.",
    "337974": "I loaded the CSV data into a Google Cloud bucket.  I then put the data into a Big Query dataset.  Does anyone want information on how to do this?\n![dataset in GCP Big Query][1]\n  [1]: http://8strong.com/trackml/dataset.png"
  },
  "source": "meta"
}