{
  "id": 14972,
  "title": "How are you handling the massive amounts of data?",
  "url": "/competitions/avito-context-ad-clicks/discussion/14972",
  "author_name": "",
  "post_date": "2015-07-01T08:58:28.353000",
  "votes": 1,
  "comment_count": 8,
  "views": 2238,
  "content": "<p>I have used the PostGreSQL import script that was mentioned in another thread, but I find that running queries is&nbsp;painstakingly slow. How are you guys processing the massive amounts of data?</p>",
  "messages": [
    {
      "id": 83066,
      "postDate": "2015-07-01T09:35:48.803Z",
      "content": "<p>I've only just started but I've been using Perl</p>",
      "votes": 1
    },
    {
      "id": 83063,
      "postDate": "2015-07-01T08:58:28.353Z",
      "content": "<p>I have used the PostGreSQL import script that was mentioned in another thread, but I find that running queries is&nbsp;painstakingly slow. How are you guys processing the massive amounts of data?</p>",
      "votes": 1
    },
    {
      "id": 83162,
      "postDate": "2015-07-02T02:34:15.397Z",
      "content": "<p>I R one could use package ff or ffbase to do this stuff. Then proccess all using chunks. It works beautifully.</p>",
      "votes": 2
    },
    {
      "id": 86200,
      "postDate": "2015-07-20T21:20:16.620Z",
      "content": "<p>I've been using hadoop to merge and transform. Check out AWS spot. It's cheap alternative. </p>",
      "rawMarkdown": "I've been using hadoop to merge and transform. Check out AWS spot. It's cheap alternative. "
    },
    {
      "id": 83120,
      "postDate": "2015-07-01T18:34:00.067Z",
      "content": "<p>[quote=MAS;83112]</p>\n<p>When you just randomly sample your dataset, you might end up with a subset that is mostly, if not purely, &nbsp;of class zero.&nbsp;</p>\n<p>[/quote]</p>\n\n<p>If I made the sub-sampled dataset very small then it could have significantly different statistical values than the entire dataset.&nbsp; I'm still using millions of rows though, so I believe the class label sparsity won't be an issue (since there are still 10's of thousands of clicks in it) if I don't have too many features.</p>"
    },
    {
      "id": 83112,
      "postDate": "2015-07-01T17:53:04.760Z",
      "content": "<p>When you just randomly sample your dataset, you might end up with a subset that is mostly, if not purely, &nbsp;of class zero.&nbsp;</p>"
    },
    {
      "id": 83108,
      "postDate": "2015-07-01T16:52:56.147Z",
      "content": "<p>I'm not using all of it.&nbsp; I'm testing things out on a smaller random subsample of the training set I made.&nbsp; The cross-validation vs. training set size levels off quite a bit for what I'm currently doing.&nbsp; In the end I'll take what I'm doing and retrain on the full dataset before predicting the test set.&nbsp; I guess this is technically overfitting since I'm not cross validating my feature selection process (I'm selecting features on the subset I'm analyzing), but I think (hope) that this will not have a large effect.&nbsp; What do you guys think?</p>"
    },
    {
      "id": 83274,
      "postDate": "2015-07-02T20:41:38.857Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 83157,
      "postDate": "2015-07-02T00:22:50.310Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 83066,
      "author_name": "Bluefool",
      "author_url": "",
      "post_date": "2015-07-01T09:35:48.803000",
      "content": "<p>I've only just started but I've been using Perl</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 83162,
      "author_name": "Leustagos",
      "author_url": "",
      "post_date": "2015-07-02T02:34:15.397000",
      "content": "<p>I R one could use package ff or ffbase to do this stuff. Then proccess all using chunks. It works beautifully.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 86200,
      "author_name": "RN",
      "author_url": "",
      "post_date": "2015-07-20T21:20:16.620000",
      "content": "<p>I've been using hadoop to merge and transform. Check out AWS spot. It's cheap alternative. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 83120,
      "author_name": "J Kolb",
      "author_url": "",
      "post_date": "2015-07-01T18:34:00.067000",
      "content": "<p>[quote=MAS;83112]</p>\n<p>When you just randomly sample your dataset, you might end up with a subset that is mostly, if not purely, &nbsp;of class zero.&nbsp;</p>\n<p>[/quote]</p>\n\n<p>If I made the sub-sampled dataset very small then it could have significantly different statistical values than the entire dataset.&nbsp; I'm still using millions of rows though, so I believe the class label sparsity won't be an issue (since there are still 10's of thousands of clicks in it) if I don't have too many features.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 83112,
      "author_name": "MAS",
      "author_url": "",
      "post_date": "2015-07-01T17:53:04.760000",
      "content": "<p>When you just randomly sample your dataset, you might end up with a subset that is mostly, if not purely, &nbsp;of class zero.&nbsp;</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 83108,
      "author_name": "J Kolb",
      "author_url": "",
      "post_date": "2015-07-01T16:52:56.147000",
      "content": "<p>I'm not using all of it.&nbsp; I'm testing things out on a smaller random subsample of the training set I made.&nbsp; The cross-validation vs. training set size levels off quite a bit for what I'm currently doing.&nbsp; In the end I'll take what I'm doing and retrain on the full dataset before predicting the test set.&nbsp; I guess this is technically overfitting since I'm not cross validating my feature selection process (I'm selecting features on the subset I'm analyzing), but I think (hope) that this will not have a large effect.&nbsp; What do you guys think?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 83274,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-07-02T20:41:38.857000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 83157,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-07-02T00:22:50.310000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "83066": "",
    "83063": "",
    "83162": "",
    "86200": "I've been using hadoop to merge and transform. Check out AWS spot. It's cheap alternative. ",
    "83120": "",
    "83112": "",
    "83108": "",
    "83274": "",
    "83157": ""
  }
}