{
  "id": 24128,
  "title": "Huge data size",
  "url": "/competitions/outbrain-click-prediction/discussion/24128",
  "author_name": "",
  "post_date": "2016-10-06T07:18:47.747Z",
  "votes": 3,
  "comment_count": 8,
  "views": 2086,
  "content": "<p>Hello,\nThe data size seems to be really to big to fit in one's laptop, are you all working on clusters? (AWS ....)\nThanks </p>",
  "messages": [
    {
      "id": "137944",
      "postDate": "10/06/2016 07:18:47",
      "content": "<p>Hello,\nThe data size seems to be really to big to fit in one's laptop, are you all working on clusters? (AWS ....)\nThanks </p>",
      "rawMarkdown": "Hello,\r\nThe data size seems to be really to big to fit in one's laptop, are you all working on clusters? (AWS ....)\r\nThanks",
      "votes": null
    },
    {
      "id": "137962",
      "postDate": "10/06/2016 09:15:26",
      "content": "<p>Hi Ferodia. It's certainly possible to do some restricted analysis on a laptop. You can train models on small subsets of the training rows or restrict yourself to a small number of features. Training many of these and ensembling them is something that I've found to work well in previous competitions with large data sets.</p>\n\n<p>Better hardware will let you include more data and/or train models quicker. It's obviously advantageous, but not usually essential to post a decent score.</p>",
      "rawMarkdown": "Hi Ferodia. It's certainly possible to do some restricted analysis on a laptop. You can train models on small subsets of the training rows or restrict yourself to a small number of features. Training many of these and ensembling them is something that I've found to work well in previous competitions with large data sets.\r\n\r\nBetter hardware will let you include more data and/or train models quicker. It's obviously advantageous, but not usually essential to post a decent score.",
      "votes": null
    },
    {
      "id": "138695",
      "postDate": "10/10/2016 12:46:27",
      "content": "<p>I wonder what will be the best approach to downscale the dataset on my private laptop?</p>",
      "rawMarkdown": "I wonder what will be the best approach to downscale the dataset on my private laptop?",
      "votes": null
    },
    {
      "id": "139219",
      "postDate": "10/12/2016 23:10:53",
      "content": "<p>I'm using an AWS spot instance with 4 cores and 30GB memory. At 4c per hour its less than my coffee bill for the time spent on the competition. </p>",
      "rawMarkdown": "I'm using an AWS spot instance with 4 cores and 30GB memory. At 4c per hour its less than my coffee bill for the time spent on the competition.",
      "votes": null
    },
    {
      "id": "139236",
      "postDate": "10/13/2016 01:36:30",
      "content": "<p>I am able to do some data preprocessing on my laptop with 16G ram and 128G SSD. It is also possible to train model with all data (not including page_views.csv) with some trick such as mini-batch sgd or hash trick. The training speed largely depends on your CPU and I/O.</p>",
      "rawMarkdown": "I am able to do some data preprocessing on my laptop with 16G ram and 128G SSD. It is also possible to train model with all data (not including page_views.csv) with some trick such as mini-batch sgd or hash trick. The training speed largely depends on your CPU and I/O.",
      "votes": null
    },
    {
      "id": "139249",
      "postDate": "10/13/2016 03:28:41",
      "content": "<p>When your data is too huge, you can do your analysis on a sample of the data (the sample should represent your whole universe) using techniques like stratified sampling.</p>",
      "rawMarkdown": "When your data is too huge, you can do your analysis on a sample of the data (the sample should represent your whole universe) using techniques like stratified sampling.",
      "votes": null
    },
    {
      "id": "139365",
      "postDate": "10/13/2016 18:03:22",
      "content": "<p>Hi everyone, thank you all for your replies, I think I will go for a sample to do data analysis locally, then use AWS spot instance for training the model on all the data :) </p>",
      "rawMarkdown": "Hi everyone, thank you all for your replies, I think I will go for a sample to do data analysis locally, then use AWS spot instance for training the model on all the data :)",
      "votes": null
    },
    {
      "id": "139484",
      "postDate": "10/14/2016 12:59:34",
      "content": "<p>these are some interesting tricks</p>",
      "rawMarkdown": "these are some interesting tricks",
      "votes": null
    },
    {
      "id": "142306",
      "postDate": "11/01/2016 12:38:07",
      "content": "<p>There is also possibility to use signing bonuses</p>\n\n<ul>\n<li><p>$300 from Google Cloud if you haven't used it previously </p></li>\n<li><p>$200 from Microsoft Azure if you haven't used it previously</p></li>\n</ul>",
      "rawMarkdown": "There is also possibility to use signing bonuses\r\n\r\n - $300 from Google Cloud if you haven't used it previously \r\n\r\n - $200 from Microsoft Azure if you haven't used it previously",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 137962,
      "author_name": "joconnor",
      "author_url": "",
      "post_date": "10/06/2016 09:15:26",
      "content": "<p>Hi Ferodia. It's certainly possible to do some restricted analysis on a laptop. You can train models on small subsets of the training rows or restrict yourself to a small number of features. Training many of these and ensembling them is something that I've found to work well in previous competitions with large data sets.</p>\n\n<p>Better hardware will let you include more data and/or train models quicker. It's obviously advantageous, but not usually essential to post a decent score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138695,
      "author_name": "some01",
      "author_url": "",
      "post_date": "10/10/2016 12:46:27",
      "content": "<p>I wonder what will be the best approach to downscale the dataset on my private laptop?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139219,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "10/12/2016 23:10:53",
      "content": "<p>I'm using an AWS spot instance with 4 cores and 30GB memory. At 4c per hour its less than my coffee bill for the time spent on the competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139236,
      "author_name": "qqgeogor",
      "author_url": "",
      "post_date": "10/13/2016 01:36:30",
      "content": "<p>I am able to do some data preprocessing on my laptop with 16G ram and 128G SSD. It is also possible to train model with all data (not including page_views.csv) with some trick such as mini-batch sgd or hash trick. The training speed largely depends on your CPU and I/O.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139249,
      "author_name": "",
      "author_url": "",
      "post_date": "10/13/2016 03:28:41",
      "content": "<p>When your data is too huge, you can do your analysis on a sample of the data (the sample should represent your whole universe) using techniques like stratified sampling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139365,
      "author_name": "ferodia",
      "author_url": "",
      "post_date": "10/13/2016 18:03:22",
      "content": "<p>Hi everyone, thank you all for your replies, I think I will go for a sample to do data analysis locally, then use AWS spot instance for training the model on all the data :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 139484,
      "author_name": "bennie",
      "author_url": "",
      "post_date": "10/14/2016 12:59:34",
      "content": "<p>these are some interesting tricks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 142306,
      "author_name": "adamszalucha",
      "author_url": "",
      "post_date": "11/01/2016 12:38:07",
      "content": "<p>There is also possibility to use signing bonuses</p>\n\n<ul>\n<li><p>$300 from Google Cloud if you haven't used it previously </p></li>\n<li><p>$200 from Microsoft Azure if you haven't used it previously</p></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "137944": "Hello,\r\nThe data size seems to be really to big to fit in one's laptop, are you all working on clusters? (AWS ....)\r\nThanks",
    "137962": "Hi Ferodia. It's certainly possible to do some restricted analysis on a laptop. You can train models on small subsets of the training rows or restrict yourself to a small number of features. Training many of these and ensembling them is something that I've found to work well in previous competitions with large data sets.\r\n\r\nBetter hardware will let you include more data and/or train models quicker. It's obviously advantageous, but not usually essential to post a decent score.",
    "138695": "I wonder what will be the best approach to downscale the dataset on my private laptop?",
    "139219": "I'm using an AWS spot instance with 4 cores and 30GB memory. At 4c per hour its less than my coffee bill for the time spent on the competition.",
    "139236": "I am able to do some data preprocessing on my laptop with 16G ram and 128G SSD. It is also possible to train model with all data (not including page_views.csv) with some trick such as mini-batch sgd or hash trick. The training speed largely depends on your CPU and I/O.",
    "139249": "When your data is too huge, you can do your analysis on a sample of the data (the sample should represent your whole universe) using techniques like stratified sampling.",
    "139365": "Hi everyone, thank you all for your replies, I think I will go for a sample to do data analysis locally, then use AWS spot instance for training the model on all the data :)",
    "139484": "these are some interesting tricks",
    "142306": "There is also possibility to use signing bonuses\r\n\r\n - $300 from Google Cloud if you haven't used it previously \r\n\r\n - $200 from Microsoft Azure if you haven't used it previously"
  },
  "source": "meta"
}