{
  "id": 53188,
  "title": "Should we exhaust all the possibilities of the combination for features ?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53188",
  "author_name": "",
  "post_date": "2018-03-28T03:40:22.327024900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<pre><code>dict_cnt = {'ip':              'channel',\\\n            'ip_day':          'channel',\\\n            'ip_hour':         'channel',\\\n            'ip_wday':         'channel',\\\n            'ip_day_hour':     'channel',\\\n            'ip_wday_hour':    'channel',\\\n            'ip_app_os':       'channel',\\\n            'ip_hour_channel': 'os',\\\n            'ip_hour_os':      'channel',\\\n            'ip_hour_app':     'channel'\\\n              }\n\nfor key, value in dict_cnt.iteritems():\n    print 'key:', key, '; ', 'value:', value \n    new_col_name = key.split('_') + [value]\n    gp = df[key.split('_') + [value]].groupby(by=key.split('_'))[[value]].count().reset_index().rename(index=str, columns={value: new_col_name})\n    df = df.merge(gp, on=key.split('_'), how='left')\n    df[value] = df[value].astype('uint16')\n    del gp; gc.collect()\n</code></pre>\n\n<p>I find that many kernals take the groupby feature as main features. Is there any potential or necessities to  enumerate all the possibilities of the combination ? It seems that it will cost too much computation resource. Any good ideas to get the cross stats features ?</p>",
  "messages": [
    {
      "id": "304815",
      "postDate": "03/28/2018 03:40:22",
      "content": "<pre><code>dict_cnt = {'ip':              'channel',\\\n            'ip_day':          'channel',\\\n            'ip_hour':         'channel',\\\n            'ip_wday':         'channel',\\\n            'ip_day_hour':     'channel',\\\n            'ip_wday_hour':    'channel',\\\n            'ip_app_os':       'channel',\\\n            'ip_hour_channel': 'os',\\\n            'ip_hour_os':      'channel',\\\n            'ip_hour_app':     'channel'\\\n              }\n\nfor key, value in dict_cnt.iteritems():\n    print 'key:', key, '; ', 'value:', value \n    new_col_name = key.split('_') + [value]\n    gp = df[key.split('_') + [value]].groupby(by=key.split('_'))[[value]].count().reset_index().rename(index=str, columns={value: new_col_name})\n    df = df.merge(gp, on=key.split('_'), how='left')\n    df[value] = df[value].astype('uint16')\n    del gp; gc.collect()\n</code></pre>\n\n<p>I find that many kernals take the groupby feature as main features. Is there any potential or necessities to  enumerate all the possibilities of the combination ? It seems that it will cost too much computation resource. Any good ideas to get the cross stats features ?</p>",
      "rawMarkdown": "dict_cnt = {'ip':              'channel',\\\n                'ip_day':          'channel',\\\n                'ip_hour':         'channel',\\\n                'ip_wday':         'channel',\\\n                'ip_day_hour':     'channel',\\\n                'ip_wday_hour':    'channel',\\\n                'ip_app_os':       'channel',\\\n                'ip_hour_channel': 'os',\\\n                'ip_hour_os':      'channel',\\\n                'ip_hour_app':     'channel'\\\n                  }\n    \n    for key, value in dict_cnt.iteritems():\n        print 'key:', key, '; ', 'value:', value \n        new_col_name = key.split('_') + [value]\n        gp = df[key.split('_') + [value]].groupby(by=key.split('_'))[[value]].count().reset_index().rename(index=str, columns={value: new_col_name})\n        df = df.merge(gp, on=key.split('_'), how='left')\n        df[value] = df[value].astype('uint16')\n        del gp; gc.collect()\n\nI find that many kernals take the groupby feature as main features. Is there any potential or necessities to  enumerate all the possibilities of the combination ? It seems that it will cost too much computation resource. Any good ideas to get the cross stats features ?",
      "votes": null
    },
    {
      "id": "307254",
      "postDate": "04/01/2018 07:26:11",
      "content": "<p>I will just copy paste my comment from another topic which seems relevant here as well:</p>\n\n<p>\"Surely, just experimenting will result in finding some good features and you will practice coding. Just trial and error. Try this combination, see the result, try that check if the LB score increases. Var on ip+os+attribute_time per hour? Day of the week as a feature where there are only 4 days in the whole data? Sure, just throw everything...</p>\n\n<p>I will go against all those things. Again - you will become better coder, but I dont think you will become better and feature engineering. My advice is - try to think what makes sense in real world ! What are the existing features, what is the problem, what do you want to predict? How would you do it if you had 1 shot? First try those things, and then of course a bit of experimenatition wouldn't hurt, but it should be supplementary and not instead of logical thinking and understanding the real business problem.\"</p>",
      "rawMarkdown": "I will just copy paste my comment from another topic which seems relevant here as well:\n\n\"Surely, just experimenting will result in finding some good features and you will practice coding. Just trial and error. Try this combination, see the result, try that check if the LB score increases. Var on ip+os+attribute_time per hour? Day of the week as a feature where there are only 4 days in the whole data? Sure, just throw everything...\n\nI will go against all those things. Again - you will become better coder, but I dont think you will become better and feature engineering. My advice is - try to think what makes sense in real world ! What are the existing features, what is the problem, what do you want to predict? How would you do it if you had 1 shot? First try those things, and then of course a bit of experimenatition wouldn't hurt, but it should be supplementary and not instead of logical thinking and understanding the real business problem.\"",
      "votes": null
    },
    {
      "id": "307357",
      "postDate": "04/01/2018 12:55:41",
      "content": "<p>well said!</p>",
      "rawMarkdown": "well said!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 307254,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "04/01/2018 07:26:11",
      "content": "<p>I will just copy paste my comment from another topic which seems relevant here as well:</p>\n\n<p>\"Surely, just experimenting will result in finding some good features and you will practice coding. Just trial and error. Try this combination, see the result, try that check if the LB score increases. Var on ip+os+attribute_time per hour? Day of the week as a feature where there are only 4 days in the whole data? Sure, just throw everything...</p>\n\n<p>I will go against all those things. Again - you will become better coder, but I dont think you will become better and feature engineering. My advice is - try to think what makes sense in real world ! What are the existing features, what is the problem, what do you want to predict? How would you do it if you had 1 shot? First try those things, and then of course a bit of experimenatition wouldn't hurt, but it should be supplementary and not instead of logical thinking and understanding the real business problem.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 307357,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "04/01/2018 12:55:41",
          "content": "<p>well said!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "304815": "dict_cnt = {'ip':              'channel',\\\n                'ip_day':          'channel',\\\n                'ip_hour':         'channel',\\\n                'ip_wday':         'channel',\\\n                'ip_day_hour':     'channel',\\\n                'ip_wday_hour':    'channel',\\\n                'ip_app_os':       'channel',\\\n                'ip_hour_channel': 'os',\\\n                'ip_hour_os':      'channel',\\\n                'ip_hour_app':     'channel'\\\n                  }\n    \n    for key, value in dict_cnt.iteritems():\n        print 'key:', key, '; ', 'value:', value \n        new_col_name = key.split('_') + [value]\n        gp = df[key.split('_') + [value]].groupby(by=key.split('_'))[[value]].count().reset_index().rename(index=str, columns={value: new_col_name})\n        df = df.merge(gp, on=key.split('_'), how='left')\n        df[value] = df[value].astype('uint16')\n        del gp; gc.collect()\n\nI find that many kernals take the groupby feature as main features. Is there any potential or necessities to  enumerate all the possibilities of the combination ? It seems that it will cost too much computation resource. Any good ideas to get the cross stats features ?",
    "307254": "I will just copy paste my comment from another topic which seems relevant here as well:\n\n\"Surely, just experimenting will result in finding some good features and you will practice coding. Just trial and error. Try this combination, see the result, try that check if the LB score increases. Var on ip+os+attribute_time per hour? Day of the week as a feature where there are only 4 days in the whole data? Sure, just throw everything...\n\nI will go against all those things. Again - you will become better coder, but I dont think you will become better and feature engineering. My advice is - try to think what makes sense in real world ! What are the existing features, what is the problem, what do you want to predict? How would you do it if you had 1 shot? First try those things, and then of course a bit of experimenatition wouldn't hurt, but it should be supplementary and not instead of logical thinking and understanding the real business problem.\"",
    "307357": "well said!"
  },
  "source": "meta"
}