{
  "id": 51455,
  "title": "Need Advice on handling large volume of data",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51455",
  "author_name": "",
  "post_date": "2018-03-09T03:27:31.197402900Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi Everyone ,\n   I need some advice on handling the large data volume.\n   What I a planing to do is use 2 parameters in read_csv.\n   skiprows &amp; n_rows.</p>\n\n<p>For skiprow i will set a counter and loop over the data set and n_rows i will read the required number of rows. \n  I then intend to do whatever processing is required and feed  to the model.\n Then increment skip rows and read the next set of rows. </p>\n\n<p>I would like some advise if this approach is feasible.\n Also if it is there any to get the number of records in the sv file</p>",
  "messages": [
    {
      "id": "293024",
      "postDate": "03/09/2018 03:27:31",
      "content": "<p>Hi Everyone ,\n   I need some advice on handling the large data volume.\n   What I a planing to do is use 2 parameters in read_csv.\n   skiprows &amp; n_rows.</p>\n\n<p>For skiprow i will set a counter and loop over the data set and n_rows i will read the required number of rows. \n  I then intend to do whatever processing is required and feed  to the model.\n Then increment skip rows and read the next set of rows. </p>\n\n<p>I would like some advise if this approach is feasible.\n Also if it is there any to get the number of records in the sv file</p>",
      "rawMarkdown": "Hi Everyone ,\n   I need some advice on handling the large data volume.\n   What I a planing to do is use 2 parameters in read_csv.\n   skiprows &amp; n_rows.\n   \n   For skiprow i will set a counter and loop over the data set and n_rows i will read the required number of rows. \n  I then intend to do whatever processing is required and feed  to the model.\n Then increment skip rows and read the next set of rows. \n\n I would like some advise if this approach is feasible.\n Also if it is there any to get the number of records in the sv file",
      "votes": null
    },
    {
      "id": "293050",
      "postDate": "03/09/2018 04:40:51",
      "content": "<p>You may easily achieve this by processing the CSV file in chunks. Panda's <code>read_csv()</code> function has a <em>chunksize</em> parameter for doing so which generates an iterator.</p>\n\n<pre><code>chunkSize = n_rows\nfor chunk in pd.read_csv(filename, chunksize=chunkSize):\n    process(chunk)\n</code></pre>",
      "rawMarkdown": "You may easily achieve this by processing the CSV file in chunks. Panda's `read_csv()` function has a *chunksize* parameter for doing so which generates an iterator.\n\n    chunkSize = n_rows\n    for chunk in pd.read_csv(filename, chunksize=chunkSize):\n        process(chunk)",
      "votes": null
    },
    {
      "id": "293591",
      "postDate": "03/10/2018 07:31:35",
      "content": "<p>0-9308567=11.6,\n9308568-68941877=11.7,\n68941878-131886952=11.8,\n131886953-184903889=11.9,\nif you would like to use skiprows &amp; nrows.</p>",
      "rawMarkdown": "0-9308567=11.6,\n9308568-68941877=11.7,\n68941878-131886952=11.8,\n131886953-184903889=11.9,\nif you would like to use skiprows &amp; nrows.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 293050,
      "author_name": "jeabat",
      "author_url": "",
      "post_date": "03/09/2018 04:40:51",
      "content": "<p>You may easily achieve this by processing the CSV file in chunks. Panda's <code>read_csv()</code> function has a <em>chunksize</em> parameter for doing so which generates an iterator.</p>\n\n<pre><code>chunkSize = n_rows\nfor chunk in pd.read_csv(filename, chunksize=chunkSize):\n    process(chunk)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 293591,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "03/10/2018 07:31:35",
      "content": "<p>0-9308567=11.6,\n9308568-68941877=11.7,\n68941878-131886952=11.8,\n131886953-184903889=11.9,\nif you would like to use skiprows &amp; nrows.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "293024": "Hi Everyone ,\n   I need some advice on handling the large data volume.\n   What I a planing to do is use 2 parameters in read_csv.\n   skiprows &amp; n_rows.\n   \n   For skiprow i will set a counter and loop over the data set and n_rows i will read the required number of rows. \n  I then intend to do whatever processing is required and feed  to the model.\n Then increment skip rows and read the next set of rows. \n\n I would like some advise if this approach is feasible.\n Also if it is there any to get the number of records in the sv file",
    "293050": "You may easily achieve this by processing the CSV file in chunks. Panda's `read_csv()` function has a *chunksize* parameter for doing so which generates an iterator.\n\n    chunkSize = n_rows\n    for chunk in pd.read_csv(filename, chunksize=chunkSize):\n        process(chunk)",
    "293591": "0-9308567=11.6,\n9308568-68941877=11.7,\n68941878-131886952=11.8,\n131886953-184903889=11.9,\nif you would like to use skiprows &amp; nrows."
  },
  "source": "meta"
}