{
  "id": 51809,
  "title": "Random selection of all the train set 🎲",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51809",
  "author_name": "",
  "post_date": "2018-03-13T10:12:50.172118700Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone ✋,</p>\n\n<p>I actually work on feature engineering to improve my model but when I tried to work with click_time column I found out that when I load the train set it only takes a short period (not even more than an hour). So <strong>I developp some code which select random lines</strong> from de train set (it's a little longer : for me <strong>80 seconds for 1% of the train dataset</strong>) </p>\n\n<p>I also find that there are approximately <strong>184 000 000 lines</strong> in the train dataset.</p>\n\n<p>For the code you need <strong>pandas</strong>, <strong>dask</strong> (that load realy quickly the data) and <strong>time</strong> if you want to see how many time the code last :</p>\n\n<pre><code>import pandas as pd\nimport dask.dataframe as dd\nimport time\n\nstart_time = time.time()\nprint('[{}] Start to load data'.format(time.time() - start_time))\npath = 'input/'\ntrain = dd.read_csv(path + \"train.csv\")\nfreq = 0.01 # Frequency of train dataset you want\ntrain = train.random_split([freq , 1-freq])[0]\ntrain = train.compute()\nprint('[{}] Finished to load data'.format(time.time() - start_time))\n</code></pre>\n\n<p>I hope it will help you for the competition ! 👍</p>",
  "messages": [
    {
      "id": "295230",
      "postDate": "03/13/2018 10:12:50",
      "content": "<p>Hi everyone ✋,</p>\n\n<p>I actually work on feature engineering to improve my model but when I tried to work with click_time column I found out that when I load the train set it only takes a short period (not even more than an hour). So <strong>I developp some code which select random lines</strong> from de train set (it's a little longer : for me <strong>80 seconds for 1% of the train dataset</strong>) </p>\n\n<p>I also find that there are approximately <strong>184 000 000 lines</strong> in the train dataset.</p>\n\n<p>For the code you need <strong>pandas</strong>, <strong>dask</strong> (that load realy quickly the data) and <strong>time</strong> if you want to see how many time the code last :</p>\n\n<pre><code>import pandas as pd\nimport dask.dataframe as dd\nimport time\n\nstart_time = time.time()\nprint('[{}] Start to load data'.format(time.time() - start_time))\npath = 'input/'\ntrain = dd.read_csv(path + \"train.csv\")\nfreq = 0.01 # Frequency of train dataset you want\ntrain = train.random_split([freq , 1-freq])[0]\ntrain = train.compute()\nprint('[{}] Finished to load data'.format(time.time() - start_time))\n</code></pre>\n\n<p>I hope it will help you for the competition ! 👍</p>",
      "rawMarkdown": "Hi everyone ✋,\n\nI actually work on feature engineering to improve my model but when I tried to work with click_time column I found out that when I load the train set it only takes a short period (not even more than an hour). So **I developp some code which select random lines** from de train set (it's a little longer : for me **80 seconds for 1% of the train dataset**) \n\nI also find that there are approximately **184 000 000 lines** in the train dataset.\n\nFor the code you need **pandas**, **dask** (that load realy quickly the data) and **time** if you want to see how many time the code last :\n\n    import pandas as pd\n    import dask.dataframe as dd\n    import time\n    \n    start_time = time.time()\n    print('[{}] Start to load data'.format(time.time() - start_time))\n    path = 'input/'\n    train = dd.read_csv(path + \"train.csv\")\n    freq = 0.01 # Frequency of train dataset you want\n    train = train.random_split([freq , 1-freq])[0]\n    train = train.compute()\n    print('[{}] Finished to load data'.format(time.time() - start_time))\n\n\nI hope it will help you for the competition ! 👍",
      "votes": null
    },
    {
      "id": "295236",
      "postDate": "03/13/2018 10:22:00",
      "content": "<p>For everyone using Dask: \nDo you find it really speeds things up when working on a single, moderate machine? i.e a Macbook with 16GM ram and the like? </p>",
      "rawMarkdown": "For everyone using Dask: \nDo you find it really speeds things up when working on a single, moderate machine? i.e a Macbook with 16GM ram and the like?",
      "votes": null
    },
    {
      "id": "295508",
      "postDate": "03/13/2018 18:54:11",
      "content": "<p>Great post!</p>",
      "rawMarkdown": "Great post!",
      "votes": null
    },
    {
      "id": "297888",
      "postDate": "03/18/2018 08:22:47",
      "content": "<p>Great, the method helped alot</p>",
      "rawMarkdown": "Great, the method helped alot",
      "votes": null
    },
    {
      "id": "298079",
      "postDate": "03/18/2018 18:21:26",
      "content": "<p>I'm glad it helped you ! </p>",
      "rawMarkdown": "I'm glad it helped you !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 295236,
      "author_name": "danofer",
      "author_url": "",
      "post_date": "03/13/2018 10:22:00",
      "content": "<p>For everyone using Dask: \nDo you find it really speeds things up when working on a single, moderate machine? i.e a Macbook with 16GM ram and the like? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 295508,
      "author_name": "mah0128",
      "author_url": "",
      "post_date": "03/13/2018 18:54:11",
      "content": "<p>Great post!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 297888,
      "author_name": "starcrab",
      "author_url": "",
      "post_date": "03/18/2018 08:22:47",
      "content": "<p>Great, the method helped alot</p>",
      "votes": null,
      "replies": [
        {
          "id": 298079,
          "author_name": "nathanlauga",
          "author_url": "",
          "post_date": "03/18/2018 18:21:26",
          "content": "<p>I'm glad it helped you ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "295230": "Hi everyone ✋,\n\nI actually work on feature engineering to improve my model but when I tried to work with click_time column I found out that when I load the train set it only takes a short period (not even more than an hour). So **I developp some code which select random lines** from de train set (it's a little longer : for me **80 seconds for 1% of the train dataset**) \n\nI also find that there are approximately **184 000 000 lines** in the train dataset.\n\nFor the code you need **pandas**, **dask** (that load realy quickly the data) and **time** if you want to see how many time the code last :\n\n    import pandas as pd\n    import dask.dataframe as dd\n    import time\n    \n    start_time = time.time()\n    print('[{}] Start to load data'.format(time.time() - start_time))\n    path = 'input/'\n    train = dd.read_csv(path + \"train.csv\")\n    freq = 0.01 # Frequency of train dataset you want\n    train = train.random_split([freq , 1-freq])[0]\n    train = train.compute()\n    print('[{}] Finished to load data'.format(time.time() - start_time))\n\n\nI hope it will help you for the competition ! 👍",
    "295236": "For everyone using Dask: \nDo you find it really speeds things up when working on a single, moderate machine? i.e a Macbook with 16GM ram and the like?",
    "295508": "Great post!",
    "297888": "Great, the method helped alot",
    "298079": "I'm glad it helped you !"
  },
  "source": "meta"
}