{
  "id": 51823,
  "title": "Shrinking the data set using Python and SQLite",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51823",
  "author_name": "",
  "post_date": "2018-03-13T13:11:22.739742400Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As already suggested by Konrad Banachewicz, Asparuh Hristov and others, we can shrink the train full data set by removing the attributed_time column and reducing click_time to hour and days value only.</p>\n\n<p>Konrad Banachewicz has already posted a very fast way to change click_time column by removing year and month from each row. He has used command line for this. Link :- <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347</a></p>\n\n<p>I wanted to achieve the above-mentioned shrinkage by using python and also store it so that I can do my explorations on it. I have used python and SQLite to do this. </p>\n\n<h2>Code for the same is :-</h2>\n\n<pre><code>import sqlite3\nimport re\nimport time\nimport os\nos.chdir(\"your directory path\") your directory path  \n# defining pattern to identify time and date from click_time\npattern = r'\\s*\\d+-\\d+-(\\d+)\\s+(\\d+):\\d+:\\d+'\npattern_regex = re.compile(pattern)\n\n# connecting and creating the SQLite database\nconn = sqlite3.connect('talkingdata.db')\nc = conn.cursor()\n# creating table and defining column. I have divided click_time into click_time_hr and            click_time_day\nc.execute('''CREATE TABLE IF NOT EXISTS train_full (\n            ip INTEGER,\n            app INTEGER,\n            device INTEGER,\n            os INTEGER,\n            channel INTEGER,\n            click_time_day INTEGER,\n            click_time_hr INTEGER,\n            is_attributed INTEGER)''')\n\n# reading the file line by line\ni = 0\nstart = time.time()\nfor line in open(\"train.csv\"):\n    if i == 0:\n        i += 1\n        continue\n    csv_row = line.split(',')\n    #finding date and hr from click_time\n    result = pattern_regex.findall(csv_row[5])\n    i += 1\n    print(\"Inserting line no :\" + str(i))\n    c.execute('''INSERT INTO train_full (ip, app, device, os, channel, click_time_day, click_time_hr, is_attributed) VALUES(?,?,?,?,?,?,?,?)''',(csv_row[0], csv_row[1], csv_row[2], csv_row[3], csv_row[4],result[0][0],result[0][1],csv_row[7]))\n\nconn.commit()\nc.close()\nconn.close()\n\nend = time.time()\nprint(\"total time taken \")\nprint(end - start)\n\"2553.701733112335\"\n</code></pre>\n\n<p>This took around 42 minutes and now the train data set is stored in sqlite database. I can do my exploration there. By removing print and \"i\" counter from the loop we can save some time, but i like a progress indication.</p>\n\n<p>Let me know your views on this and what all can be done here to make it faster. </p>",
  "messages": [
    {
      "id": "295304",
      "postDate": "03/13/2018 13:11:22",
      "content": "<p>As already suggested by Konrad Banachewicz, Asparuh Hristov and others, we can shrink the train full data set by removing the attributed_time column and reducing click_time to hour and days value only.</p>\n\n<p>Konrad Banachewicz has already posted a very fast way to change click_time column by removing year and month from each row. He has used command line for this. Link :- <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347</a></p>\n\n<p>I wanted to achieve the above-mentioned shrinkage by using python and also store it so that I can do my explorations on it. I have used python and SQLite to do this. </p>\n\n<h2>Code for the same is :-</h2>\n\n<pre><code>import sqlite3\nimport re\nimport time\nimport os\nos.chdir(\"your directory path\") your directory path  \n# defining pattern to identify time and date from click_time\npattern = r'\\s*\\d+-\\d+-(\\d+)\\s+(\\d+):\\d+:\\d+'\npattern_regex = re.compile(pattern)\n\n# connecting and creating the SQLite database\nconn = sqlite3.connect('talkingdata.db')\nc = conn.cursor()\n# creating table and defining column. I have divided click_time into click_time_hr and            click_time_day\nc.execute('''CREATE TABLE IF NOT EXISTS train_full (\n            ip INTEGER,\n            app INTEGER,\n            device INTEGER,\n            os INTEGER,\n            channel INTEGER,\n            click_time_day INTEGER,\n            click_time_hr INTEGER,\n            is_attributed INTEGER)''')\n\n# reading the file line by line\ni = 0\nstart = time.time()\nfor line in open(\"train.csv\"):\n    if i == 0:\n        i += 1\n        continue\n    csv_row = line.split(',')\n    #finding date and hr from click_time\n    result = pattern_regex.findall(csv_row[5])\n    i += 1\n    print(\"Inserting line no :\" + str(i))\n    c.execute('''INSERT INTO train_full (ip, app, device, os, channel, click_time_day, click_time_hr, is_attributed) VALUES(?,?,?,?,?,?,?,?)''',(csv_row[0], csv_row[1], csv_row[2], csv_row[3], csv_row[4],result[0][0],result[0][1],csv_row[7]))\n\nconn.commit()\nc.close()\nconn.close()\n\nend = time.time()\nprint(\"total time taken \")\nprint(end - start)\n\"2553.701733112335\"\n</code></pre>\n\n<p>This took around 42 minutes and now the train data set is stored in sqlite database. I can do my exploration there. By removing print and \"i\" counter from the loop we can save some time, but i like a progress indication.</p>\n\n<p>Let me know your views on this and what all can be done here to make it faster. </p>",
      "rawMarkdown": "As already suggested by Konrad Banachewicz, Asparuh Hristov and others, we can shrink the train full data set by removing the attributed_time column and reducing click_time to hour and days value only.\n\nKonrad Banachewicz has already posted a very fast way to change click_time column by removing year and month from each row. He has used command line for this. Link :- https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347\n\nI wanted to achieve the above-mentioned shrinkage by using python and also store it so that I can do my explorations on it. I have used python and SQLite to do this. \n\n## Code for the same is :- \n\n    import sqlite3\n    import re\n    import time\n    import os\n    os.chdir(\"your directory path\") your directory path  \n    # defining pattern to identify time and date from click_time\n    pattern = r'\\s*\\d+-\\d+-(\\d+)\\s+(\\d+):\\d+:\\d+'\n    pattern_regex = re.compile(pattern)\n\n    # connecting and creating the SQLite database\n    conn = sqlite3.connect('talkingdata.db')\n    c = conn.cursor()\n    # creating table and defining column. I have divided click_time into click_time_hr and            click_time_day\n    c.execute('''CREATE TABLE IF NOT EXISTS train_full (\n                ip INTEGER,\n                app INTEGER,\n                device INTEGER,\n                os INTEGER,\n                channel INTEGER,\n                click_time_day INTEGER,\n                click_time_hr INTEGER,\n                is_attributed INTEGER)''')\n\n    # reading the file line by line\n    i = 0\n    start = time.time()\n    for line in open(\"train.csv\"):\n        if i == 0:\n            i += 1\n            continue\n        csv_row = line.split(',')\n        #finding date and hr from click_time\n        result = pattern_regex.findall(csv_row[5])\n        i += 1\n        print(\"Inserting line no :\" + str(i))\n        c.execute('''INSERT INTO train_full (ip, app, device, os, channel, click_time_day, click_time_hr, is_attributed) VALUES(?,?,?,?,?,?,?,?)''',(csv_row[0], csv_row[1], csv_row[2], csv_row[3], csv_row[4],result[0][0],result[0][1],csv_row[7]))\n\n    conn.commit()\n    c.close()\n    conn.close()\n\n    end = time.time()\n    print(\"total time taken \")\n    print(end - start)\n    \"2553.701733112335\"\n\nThis took around 42 minutes and now the train data set is stored in sqlite database. I can do my exploration there. By removing print and \"i\" counter from the loop we can save some time, but i like a progress indication.\n\nLet me know your views on this and what all can be done here to make it faster.",
      "votes": null
    },
    {
      "id": "295698",
      "postDate": "03/14/2018 02:58:26",
      "content": "<p>Can you put this on public kernel and produce a lite version of the dataset? That will be a good showcase of the concept.</p>",
      "rawMarkdown": "Can you put this on public kernel and produce a lite version of the dataset? That will be a good showcase of the concept.",
      "votes": null
    },
    {
      "id": "295707",
      "postDate": "03/14/2018 03:17:15",
      "content": "<p>Please check this kernel :- \n<a href=\"https://www.kaggle.com/princeatul/shrinking-the-data-using-py3-on-4-or-8-gb-laptop\">https://www.kaggle.com/princeatul/shrinking-the-data-using-py3-on-4-or-8-gb-laptop</a></p>\n\n<p>I have used train_sample.csv here because of kernel constraint. It can be changed to train.csv while working on the local system.</p>",
      "rawMarkdown": "Please check this kernel :- \nhttps://www.kaggle.com/princeatul/shrinking-the-data-using-py3-on-4-or-8-gb-laptop\n\nI have used train_sample.csv here because of kernel constraint. It can be changed to train.csv while working on the local system.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 295698,
      "author_name": "muhammadalfiansyah",
      "author_url": "",
      "post_date": "03/14/2018 02:58:26",
      "content": "<p>Can you put this on public kernel and produce a lite version of the dataset? That will be a good showcase of the concept.</p>",
      "votes": null,
      "replies": [
        {
          "id": 295707,
          "author_name": "princeatul",
          "author_url": "",
          "post_date": "03/14/2018 03:17:15",
          "content": "<p>Please check this kernel :- \n<a href=\"https://www.kaggle.com/princeatul/shrinking-the-data-using-py3-on-4-or-8-gb-laptop\">https://www.kaggle.com/princeatul/shrinking-the-data-using-py3-on-4-or-8-gb-laptop</a></p>\n\n<p>I have used train_sample.csv here because of kernel constraint. It can be changed to train.csv while working on the local system.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "295304": "As already suggested by Konrad Banachewicz, Asparuh Hristov and others, we can shrink the train full data set by removing the attributed_time column and reducing click_time to hour and days value only.\n\nKonrad Banachewicz has already posted a very fast way to change click_time column by removing year and month from each row. He has used command line for this. Link :- https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347\n\nI wanted to achieve the above-mentioned shrinkage by using python and also store it so that I can do my explorations on it. I have used python and SQLite to do this. \n\n## Code for the same is :- \n\n    import sqlite3\n    import re\n    import time\n    import os\n    os.chdir(\"your directory path\") your directory path  \n    # defining pattern to identify time and date from click_time\n    pattern = r'\\s*\\d+-\\d+-(\\d+)\\s+(\\d+):\\d+:\\d+'\n    pattern_regex = re.compile(pattern)\n\n    # connecting and creating the SQLite database\n    conn = sqlite3.connect('talkingdata.db')\n    c = conn.cursor()\n    # creating table and defining column. I have divided click_time into click_time_hr and            click_time_day\n    c.execute('''CREATE TABLE IF NOT EXISTS train_full (\n                ip INTEGER,\n                app INTEGER,\n                device INTEGER,\n                os INTEGER,\n                channel INTEGER,\n                click_time_day INTEGER,\n                click_time_hr INTEGER,\n                is_attributed INTEGER)''')\n\n    # reading the file line by line\n    i = 0\n    start = time.time()\n    for line in open(\"train.csv\"):\n        if i == 0:\n            i += 1\n            continue\n        csv_row = line.split(',')\n        #finding date and hr from click_time\n        result = pattern_regex.findall(csv_row[5])\n        i += 1\n        print(\"Inserting line no :\" + str(i))\n        c.execute('''INSERT INTO train_full (ip, app, device, os, channel, click_time_day, click_time_hr, is_attributed) VALUES(?,?,?,?,?,?,?,?)''',(csv_row[0], csv_row[1], csv_row[2], csv_row[3], csv_row[4],result[0][0],result[0][1],csv_row[7]))\n\n    conn.commit()\n    c.close()\n    conn.close()\n\n    end = time.time()\n    print(\"total time taken \")\n    print(end - start)\n    \"2553.701733112335\"\n\nThis took around 42 minutes and now the train data set is stored in sqlite database. I can do my exploration there. By removing print and \"i\" counter from the loop we can save some time, but i like a progress indication.\n\nLet me know your views on this and what all can be done here to make it faster.",
    "295698": "Can you put this on public kernel and produce a lite version of the dataset? That will be a good showcase of the concept.",
    "295707": "Please check this kernel :- \nhttps://www.kaggle.com/princeatul/shrinking-the-data-using-py3-on-4-or-8-gb-laptop\n\nI have used train_sample.csv here because of kernel constraint. It can be changed to train.csv while working on the local system."
  },
  "source": "meta"
}