{
  "id": 55666,
  "title": "how to creat histogram feature?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55666",
  "author_name": "",
  "post_date": "2018-04-30T13:02:49.017211100Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>i am not so good in using pandas dataframe object. How I can create histogram feature like the one shown in the diagram below? thanks!</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321042/9330/df_histogram.png\" alt=\"enter image description here\"></p>",
  "messages": [
    {
      "id": "321042",
      "postDate": "04/30/2018 13:02:49",
      "content": "<p>i am not so good in using pandas dataframe object. How I can create histogram feature like the one shown in the diagram below? thanks!</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321042/9330/df_histogram.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "i am not so good in using pandas dataframe object. How I can create histogram feature like the one shown in the diagram below? thanks!\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/321042/9330/df_histogram.png",
      "votes": null
    },
    {
      "id": "321047",
      "postDate": "04/30/2018 13:20:49",
      "content": "<p>just add the time_index  col to df (something like df['click_time'].dt.hour) and join</p>",
      "rawMarkdown": "just add the time_index  col to df (something like df['click_time'].dt.hour) and join",
      "votes": null
    },
    {
      "id": "321055",
      "postDate": "04/30/2018 13:34:51",
      "content": "<p>Got it! Thanks</p>\n\n<p>create an index key \"df['t] = df (something like df['click_time'].dt.hour) \" in both dataframe and then  df = df.merge(g, on=['ip', 'app', 'device', 'os', 'channel','t'], how='left')</p>",
      "rawMarkdown": "Got it! Thanks\n\ncreate an index key \"df['t] = df (something like df['click_time'].dt.hour) \" in both dataframe and then  df = df.merge(g, on=['ip', 'app', 'device', 'os', 'channel','t'], how='left')",
      "votes": null
    },
    {
      "id": "321056",
      "postDate": "04/30/2018 13:39:32",
      "content": "<p>df['click_time'] = pd.to_datetime(df['click_time'])</p>\n\n<p>df.index = df['click_time']</p>\n\n<p>g = df.groupby(by=['ip','app', 'device','os','channel'])['is_attributed'].resample('1H').count().reset_index().rename(index=str, columns={'is_attributed': 'ip_app_device_os_channel%s_count' % '1H'})</p>\n\n<p>df = df.merge(g, on=['ip','app', 'device','os','channel'], how='left')</p>\n\n<p>However to run this on the whole dataset takes a very long time and takes up a lot of RAM</p>",
      "rawMarkdown": "df['click_time'] = pd.to_datetime(df['click_time'])\n\ndf.index = df['click_time']\n\ng = df.groupby(by=['ip','app', 'device','os','channel'])['is_attributed'].resample('1H').count().reset_index().rename(index=str, columns={'is_attributed': 'ip_app_device_os_channel%s_count' % '1H'})\n\ndf = df.merge(g, on=['ip','app', 'device','os','channel'], how='left')\n\nHowever to run this on the whole dataset takes a very long time and takes up a lot of RAM",
      "votes": null
    },
    {
      "id": "321066",
      "postDate": "04/30/2018 14:18:03",
      "content": "<p>Good to see you here!  </p>\n\n<p><a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">This notebook</a> contains various examples of features like the one you want to implement.  More important, it abstracts a bit how to build them.</p>",
      "rawMarkdown": "Good to see you here!  \n\n[This notebook][1] contains various examples of features like the one you want to implement.  More important, it abstracts a bit how to build them.\n\n\n  [1]: https://www.kaggle.com/nanomathias/feature-engineering-importance-testing",
      "votes": null
    },
    {
      "id": "321941",
      "postDate": "05/02/2018 07:21:40",
      "content": "<p>df.groupby(by=['ip','app', 'device','os','channel'])['is_attributed'].rolling('1H').count().reset_index(['ip','app', 'device','os','channel']).rename(index=str, columns={'is_attributed': 'ip_app_device_os_channel%s_count' % '1H'}</p>\n\n<p>i think use rolling may be a better way.</p>",
      "rawMarkdown": "df.groupby(by=['ip','app', 'device','os','channel'])['is_attributed'].rolling('1H').count().reset_index(['ip','app', 'device','os','channel']).rename(index=str, columns={'is_attributed': 'ip_app_device_os_channel%s_count' % '1H'}\n\ni think use rolling may be a better way.",
      "votes": null
    },
    {
      "id": "322238",
      "postDate": "05/02/2018 15:46:41",
      "content": "<p>merging step is exploding, and the process is getting killed. Any workarounds ?</p>",
      "rawMarkdown": "merging step is exploding, and the process is getting killed. Any workarounds ?",
      "votes": null
    },
    {
      "id": "322246",
      "postDate": "05/02/2018 15:59:26",
      "content": "<p>this wasn't exploding when merging dfs with normal indices. Is it cz of the datetime index?</p>",
      "rawMarkdown": "this wasn't exploding when merging dfs with normal indices. Is it cz of the datetime index?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 321047,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "04/30/2018 13:20:49",
      "content": "<p>just add the time_index  col to df (something like df['click_time'].dt.hour) and join</p>",
      "votes": null,
      "replies": [
        {
          "id": 321055,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/30/2018 13:34:51",
          "content": "<p>Got it! Thanks</p>\n\n<p>create an index key \"df['t] = df (something like df['click_time'].dt.hour) \" in both dataframe and then  df = df.merge(g, on=['ip', 'app', 'device', 'os', 'channel','t'], how='left')</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321056,
      "author_name": "dicksonchin93",
      "author_url": "",
      "post_date": "04/30/2018 13:39:32",
      "content": "<p>df['click_time'] = pd.to_datetime(df['click_time'])</p>\n\n<p>df.index = df['click_time']</p>\n\n<p>g = df.groupby(by=['ip','app', 'device','os','channel'])['is_attributed'].resample('1H').count().reset_index().rename(index=str, columns={'is_attributed': 'ip_app_device_os_channel%s_count' % '1H'})</p>\n\n<p>df = df.merge(g, on=['ip','app', 'device','os','channel'], how='left')</p>\n\n<p>However to run this on the whole dataset takes a very long time and takes up a lot of RAM</p>",
      "votes": null,
      "replies": [
        {
          "id": 322238,
          "author_name": "rrqqmm",
          "author_url": "",
          "post_date": "05/02/2018 15:46:41",
          "content": "<p>merging step is exploding, and the process is getting killed. Any workarounds ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322246,
          "author_name": "rrqqmm",
          "author_url": "",
          "post_date": "05/02/2018 15:59:26",
          "content": "<p>this wasn't exploding when merging dfs with normal indices. Is it cz of the datetime index?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321066,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/30/2018 14:18:03",
      "content": "<p>Good to see you here!  </p>\n\n<p><a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">This notebook</a> contains various examples of features like the one you want to implement.  More important, it abstracts a bit how to build them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 321941,
      "author_name": "marvinxu",
      "author_url": "",
      "post_date": "05/02/2018 07:21:40",
      "content": "<p>df.groupby(by=['ip','app', 'device','os','channel'])['is_attributed'].rolling('1H').count().reset_index(['ip','app', 'device','os','channel']).rename(index=str, columns={'is_attributed': 'ip_app_device_os_channel%s_count' % '1H'}</p>\n\n<p>i think use rolling may be a better way.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "321042": "i am not so good in using pandas dataframe object. How I can create histogram feature like the one shown in the diagram below? thanks!\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/321042/9330/df_histogram.png",
    "321047": "just add the time_index  col to df (something like df['click_time'].dt.hour) and join",
    "321055": "Got it! Thanks\n\ncreate an index key \"df['t] = df (something like df['click_time'].dt.hour) \" in both dataframe and then  df = df.merge(g, on=['ip', 'app', 'device', 'os', 'channel','t'], how='left')",
    "321056": "df['click_time'] = pd.to_datetime(df['click_time'])\n\ndf.index = df['click_time']\n\ng = df.groupby(by=['ip','app', 'device','os','channel'])['is_attributed'].resample('1H').count().reset_index().rename(index=str, columns={'is_attributed': 'ip_app_device_os_channel%s_count' % '1H'})\n\ndf = df.merge(g, on=['ip','app', 'device','os','channel'], how='left')\n\nHowever to run this on the whole dataset takes a very long time and takes up a lot of RAM",
    "321066": "Good to see you here!  \n\n[This notebook][1] contains various examples of features like the one you want to implement.  More important, it abstracts a bit how to build them.\n\n\n  [1]: https://www.kaggle.com/nanomathias/feature-engineering-importance-testing",
    "321941": "df.groupby(by=['ip','app', 'device','os','channel'])['is_attributed'].rolling('1H').count().reset_index(['ip','app', 'device','os','channel']).rename(index=str, columns={'is_attributed': 'ip_app_device_os_channel%s_count' % '1H'}\n\ni think use rolling may be a better way.",
    "322238": "merging step is exploding, and the process is getting killed. Any workarounds ?",
    "322246": "this wasn't exploding when merging dfs with normal indices. Is it cz of the datetime index?"
  },
  "source": "meta"
}