{
  "id": 52167,
  "title": "Quick tip for subsampling across time when reading the data",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52167",
  "author_name": "",
  "post_date": "2018-03-16T19:24:45.464344200Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>It's useful to be able to randomly sample lines from the data at read time -- you don't have to load it all into memory and can also capture the full time range instead of taking rows in a consecutive time range (the data is ordered by time). With python/pandas if you attempt this by creating a random sample of skiprows from a range, generating the random sample can be pretty slow.</p>\n\n<p>Instead, you can exploit the time ordering of the data by reading lines with a fixed step size based on the sample size you want. This way you can quickly draw a random sample that still mirrors the  time distribution of the full data set. </p>\n\n<p>Here's code for that:      </p>\n\n<pre><code>def get_skiprows(total_rows, sample_size):\n    inc = total_rows // sample_size\n    return [row for row in range(1, total_rows) if row % inc != 0]\n\ndf = pd.read_csv('train.csv',\n                 skiprows=get_skiprows(total_rows,sample_size), \n                 parse_dates=['click_time','attributed_time'])\n</code></pre>",
  "messages": [
    {
      "id": "297325",
      "postDate": "03/16/2018 19:24:45",
      "content": "<p>It's useful to be able to randomly sample lines from the data at read time -- you don't have to load it all into memory and can also capture the full time range instead of taking rows in a consecutive time range (the data is ordered by time). With python/pandas if you attempt this by creating a random sample of skiprows from a range, generating the random sample can be pretty slow.</p>\n\n<p>Instead, you can exploit the time ordering of the data by reading lines with a fixed step size based on the sample size you want. This way you can quickly draw a random sample that still mirrors the  time distribution of the full data set. </p>\n\n<p>Here's code for that:      </p>\n\n<pre><code>def get_skiprows(total_rows, sample_size):\n    inc = total_rows // sample_size\n    return [row for row in range(1, total_rows) if row % inc != 0]\n\ndf = pd.read_csv('train.csv',\n                 skiprows=get_skiprows(total_rows,sample_size), \n                 parse_dates=['click_time','attributed_time'])\n</code></pre>",
      "rawMarkdown": "It's useful to be able to randomly sample lines from the data at read time -- you don't have to load it all into memory and can also capture the full time range instead of taking rows in a consecutive time range (the data is ordered by time). With python/pandas if you attempt this by creating a random sample of skiprows from a range, generating the random sample can be pretty slow.\n\nInstead, you can exploit the time ordering of the data by reading lines with a fixed step size based on the sample size you want. This way you can quickly draw a random sample that still mirrors the  time distribution of the full data set. \n\nHere's code for that:      \n\n    def get_skiprows(total_rows, sample_size):\n        inc = total_rows // sample_size\n        return [row for row in range(1, total_rows) if row % inc != 0]\n     \n    df = pd.read_csv('train.csv',\n                     skiprows=get_skiprows(total_rows,sample_size), \n                     parse_dates=['click_time','attributed_time'])",
      "votes": null
    },
    {
      "id": "298230",
      "postDate": "03/19/2018 04:53:36",
      "content": "<p>Didn't know that we can use a function in skiprows. \nSo, to select only odd rows, we can write get_skiprows like:- \ndef get_skiprows(total_rows):\n       return [rows for row in range(1, total_rows,2)] </p>",
      "rawMarkdown": "Didn't know that we can use a function in skiprows. \nSo, to select only odd rows, we can write get_skiprows like:- \ndef get_skiprows(total_rows):\n       return [rows for row in range(1, total_rows,2)]",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 298230,
      "author_name": "princeatul",
      "author_url": "",
      "post_date": "03/19/2018 04:53:36",
      "content": "<p>Didn't know that we can use a function in skiprows. \nSo, to select only odd rows, we can write get_skiprows like:- \ndef get_skiprows(total_rows):\n       return [rows for row in range(1, total_rows,2)] </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "297325": "It's useful to be able to randomly sample lines from the data at read time -- you don't have to load it all into memory and can also capture the full time range instead of taking rows in a consecutive time range (the data is ordered by time). With python/pandas if you attempt this by creating a random sample of skiprows from a range, generating the random sample can be pretty slow.\n\nInstead, you can exploit the time ordering of the data by reading lines with a fixed step size based on the sample size you want. This way you can quickly draw a random sample that still mirrors the  time distribution of the full data set. \n\nHere's code for that:      \n\n    def get_skiprows(total_rows, sample_size):\n        inc = total_rows // sample_size\n        return [row for row in range(1, total_rows) if row % inc != 0]\n     \n    df = pd.read_csv('train.csv',\n                     skiprows=get_skiprows(total_rows,sample_size), \n                     parse_dates=['click_time','attributed_time'])",
    "298230": "Didn't know that we can use a function in skiprows. \nSo, to select only odd rows, we can write get_skiprows like:- \ndef get_skiprows(total_rows):\n       return [rows for row in range(1, total_rows,2)]"
  },
  "source": "meta"
}