{
  "id": 55107,
  "title": "Dataset size, memory limitations and building a balanced training set",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55107",
  "author_name": "",
  "post_date": "2018-04-22T06:34:23.029798100Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The main reason I spend some time with these competitions and using some of the available datasets for predictive model building, is that, as I don't do a huge amount of model building at work, and what I do always tends to be on the same type of data, it's an opportunity to learn some new things that might help me out later.</p>\n\n<p>With this 7 GB dataset, loading the whole lot into R on my 16 GB early 2011 MBP was causing some problems. I had been working with cut-down versions of the training set using the <code>nrows = x</code> argument and that was working fine.</p>\n\n<p>However, I wanted to build a balanced dataset, which, given the rarity of the <code>is_attributed == 1</code>, meant I really wanted to import every instance from the whole training set, and use that as to build my reduced, balanced dataset.</p>\n\n<p>A search for how to do that <a href=\"https://www.rdocumentation.org/packages/sqldf/versions/0.4-11/topics/read.csv.sql\">led me to the <code>sqldf</code> package</a>. What a useful find. Now, I can just import the wanted rows with a simple sql query. Only the results of the query are then processed by R, not the whole dataset, solving that first big memory use problem.</p>\n\n<p><code>allAttributed &lt;- read.csv.sql(\"train.csv\", \n                            sql = \"SELECT * FROM file WHERE `is_attributed` = 1\")</code></p>\n\n<p>Another dataset, another useful R lesson learned and filed away for future use!</p>",
  "messages": [
    {
      "id": "317653",
      "postDate": "04/22/2018 06:34:23",
      "content": "<p>The main reason I spend some time with these competitions and using some of the available datasets for predictive model building, is that, as I don't do a huge amount of model building at work, and what I do always tends to be on the same type of data, it's an opportunity to learn some new things that might help me out later.</p>\n\n<p>With this 7 GB dataset, loading the whole lot into R on my 16 GB early 2011 MBP was causing some problems. I had been working with cut-down versions of the training set using the <code>nrows = x</code> argument and that was working fine.</p>\n\n<p>However, I wanted to build a balanced dataset, which, given the rarity of the <code>is_attributed == 1</code>, meant I really wanted to import every instance from the whole training set, and use that as to build my reduced, balanced dataset.</p>\n\n<p>A search for how to do that <a href=\"https://www.rdocumentation.org/packages/sqldf/versions/0.4-11/topics/read.csv.sql\">led me to the <code>sqldf</code> package</a>. What a useful find. Now, I can just import the wanted rows with a simple sql query. Only the results of the query are then processed by R, not the whole dataset, solving that first big memory use problem.</p>\n\n<p><code>allAttributed &lt;- read.csv.sql(\"train.csv\", \n                            sql = \"SELECT * FROM file WHERE `is_attributed` = 1\")</code></p>\n\n<p>Another dataset, another useful R lesson learned and filed away for future use!</p>",
      "rawMarkdown": "The main reason I spend some time with these competitions and using some of the available datasets for predictive model building, is that, as I don't do a huge amount of model building at work, and what I do always tends to be on the same type of data, it's an opportunity to learn some new things that might help me out later.\n\nWith this 7 GB dataset, loading the whole lot into R on my 16 GB early 2011 MBP was causing some problems. I had been working with cut-down versions of the training set using the `nrows = x` argument and that was working fine.\n\nHowever, I wanted to build a balanced dataset, which, given the rarity of the `is_attributed == 1`, meant I really wanted to import every instance from the whole training set, and use that as to build my reduced, balanced dataset.\n\nA search for how to do that [led me to the `sqldf` package][1]. What a useful find. Now, I can just import the wanted rows with a simple sql query. Only the results of the query are then processed by R, not the whole dataset, solving that first big memory use problem.\n\n``allAttributed &lt;- read.csv.sql(\"train.csv\", \n                            sql = \"SELECT * FROM file WHERE `is_attributed` = 1\")``\n\nAnother dataset, another useful R lesson learned and filed away for future use!\n\n[1]: https://www.rdocumentation.org/packages/sqldf/versions/0.4-11/topics/read.csv.sql",
      "votes": null
    },
    {
      "id": "728570",
      "postDate": "01/24/2020 22:56:53",
      "content": "<p>Thanks for sharing it! </p>",
      "rawMarkdown": "Thanks for sharing it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 728570,
      "author_name": "carloscosta88",
      "author_url": "",
      "post_date": "01/24/2020 22:56:53",
      "content": "<p>Thanks for sharing it! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "317653": "The main reason I spend some time with these competitions and using some of the available datasets for predictive model building, is that, as I don't do a huge amount of model building at work, and what I do always tends to be on the same type of data, it's an opportunity to learn some new things that might help me out later.\n\nWith this 7 GB dataset, loading the whole lot into R on my 16 GB early 2011 MBP was causing some problems. I had been working with cut-down versions of the training set using the `nrows = x` argument and that was working fine.\n\nHowever, I wanted to build a balanced dataset, which, given the rarity of the `is_attributed == 1`, meant I really wanted to import every instance from the whole training set, and use that as to build my reduced, balanced dataset.\n\nA search for how to do that [led me to the `sqldf` package][1]. What a useful find. Now, I can just import the wanted rows with a simple sql query. Only the results of the query are then processed by R, not the whole dataset, solving that first big memory use problem.\n\n``allAttributed &lt;- read.csv.sql(\"train.csv\", \n                            sql = \"SELECT * FROM file WHERE `is_attributed` = 1\")``\n\nAnother dataset, another useful R lesson learned and filed away for future use!\n\n[1]: https://www.rdocumentation.org/packages/sqldf/versions/0.4-11/topics/read.csv.sql",
    "728570": "Thanks for sharing it!"
  },
  "source": "meta"
}