{"cells":[{"metadata":{"_uuid":"649a6d7dc84f3b47e58762dc8739529b20b458fd"},"cell_type":"markdown","source":"# Automated feature engineering using Featuretools\n\n## What's Featuretools:\n\n* a python library to perform automated feature engineering.\n* based on \"Deep Feature Synthesis\" paper/ research\n* Documentation: https://docs.featuretools.com/\n* Source code: https://github.com/Featuretools/featuretools\n* Other examples: https://www.featuretools.com/demos\n\n## Deep Feature Synthesis\n* Paper: http://www.jmaxkanter.com/static/papers/DSAA_DSM_2015.pdf\n* Article: https://www.featurelabs.com/blog/deep-feature-synthesis/\n* DFS works with the structured transactional and relational datasets\n* Across datasets features are derived by using primitive mathematical operations\n* New features are composed from using derived features (hence \"Deep\")\n"},{"metadata":{"_cell_guid":"95dffc98-9214-46d1-a646-7c7319e7c9d0","_uuid":"0acae3074fbc18e36ea405f626b744f23784a7ec","collapsed":true,"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport featuretools as ft\n\nfrom datetime import datetime\n\nfrom featuretools.primitives import *","execution_count":4,"outputs":[]},{"metadata":{"_uuid":"46a6e96cb043881dd2dc7fb34b1f594a881c56cc"},"cell_type":"markdown","source":"# 1. Load data"},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","collapsed":true,"trusted":true},"cell_type":"code","source":"dtypes = {\n    'ip': 'uint32',\n    'app': 'uint16',\n    'device': 'uint16',\n    'os': 'uint16',\n    'channel': 'uint16',\n    'is_attributed': 'uint8'\n}\nto_read = ['ip', 'app', 'device', 'os', 'channel', 'click_time', 'is_attributed']\nto_parse = ['click_time']","execution_count":5,"outputs":[]},{"metadata":{"_cell_guid":"12076134-0487-4c47-8648-1c15be8c0ce0","_uuid":"5c5bfc7d0b463699375b90ec831216bccaa2833f","collapsed":true,"trusted":true},"cell_type":"code","source":"df = pd.read_csv('../input/train_sample.csv', usecols=to_read, dtype=dtypes, parse_dates=to_parse)\ndf['id'] = df.index","execution_count":6,"outputs":[]},{"metadata":{"_uuid":"112a0f7fc59f0e14091ac1ac37a2791e30a17867"},"cell_type":"markdown","source":"# 2. Prepare data"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# Create an entity set, a collection of entities (tables) and their relationships\nes = ft.EntitySet(id='clicks')\n\n# Create an entity \"clicks\" based on pandas dataframe and add it to the entity set\nes = es.entity_from_dataframe(\n    entity_id='clicks',\n    dataframe=df,\n    index='id',\n    time_index='click_time',\n    variable_types={\n        # We need to set proper types so that Featuretools won't treat them as numericals\n        'ip': ft.variable_types.Categorical,\n        'app': ft.variable_types.Categorical,\n        'device': ft.variable_types.Categorical,\n        'os': ft.variable_types.Categorical,\n        'channel': ft.variable_types.Categorical,\n        'is_attributed': ft.variable_types.Boolean,\n    }\n)\n\n# We can create new enities based on information we have, e.g. for ips or apps. We “normalize” the entity and extract a new one, this automatically adds a relationship between them\nes = es.normalize_entity(base_entity_id='clicks', new_entity_id='ip', index='ip')\nes = es.normalize_entity(base_entity_id='clicks', new_entity_id='app', index='app')\nes = es.normalize_entity(base_entity_id='clicks', new_entity_id='device', index='device')\nes = es.normalize_entity(base_entity_id='clicks', new_entity_id='channel', index='channel')\nes = es.normalize_entity(base_entity_id='clicks', new_entity_id='os', index='os')\n\n# How our entityset looks like:\nes","execution_count":7,"outputs":[]},{"metadata":{"_uuid":"94f8af0dbffbd6f744dea42f6c76be99dd7269de"},"cell_type":"markdown","source":"# 3. Create features"},{"metadata":{"_cell_guid":"0e606d66-c55c-4d58-9089-694f690f69c7","_uuid":"c6e0fff5980f68dd640fb003b83461a3b7a5a620","trusted":true},"cell_type":"code","source":"# Run Deep Feature Synthesis for app as a target entity (features will be create for each app)\nfeature_matrix, feature_defs = ft.dfs(\n    entityset=es,\n    target_entity='app'\n)\n\n# List of created features:\nfeature_defs","execution_count":9,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"864e84b5bd9423c07d95ca87903e532dfef88a2a"},"cell_type":"code","source":"# The features values\nfeature_matrix.head()","execution_count":11,"outputs":[]},{"metadata":{"_uuid":"8ab7db9c649e68ce90fbfb799daad1239ca62b0b"},"cell_type":"markdown","source":"# 4. Feature primitives\n\n* The units/ building blocks of Featuretools\n* Computations applied to raw datasets to create new features\n* Constrains the input and output data types\n* Two types of primitives: aggregation and transform\n\n### Aggregation vs Transform Primitive:\n\n1. Aggregation primitives: These primitives take related instances as an input and output a single value. They are applied across a parent-child relationship in an entity set. E.g: Count, Sum, AvgTimeBetween.\n\n2. Transform primitives: These primitives take one or more variables from an entity as an input and output a new variable for that entity. They are applied to a single entity. E.g: Hour, TimeSincePrevious, Absolute.\n\n3. Custom primitives: You can define your own aggregation and transform primitives"},{"metadata":{"trusted":true,"_uuid":"bba42cbdc26828aa6386d297199dfb64e4abd11a"},"cell_type":"code","source":"# Create feature with your own primitives\nfeature_matrix, feature_defs = ft.dfs(\n    entityset=es,\n    target_entity='app',\n    trans_primitives=[Hour],\n    agg_primitives=[PercentTrue, Mode]\n)\n\n# List of created features:\nfeature_defs","execution_count":18,"outputs":[]},{"metadata":{"_uuid":"81addeb9f67d353413a4e9623db5c6bba911ef09"},"cell_type":"markdown","source":"# 5. Handling time\n\n* Featuretools designed to take time into consideration\n* Entities have a column (time index) that specifies the point in time when data in that row became available\n* Cutoff Time specifies the time to calculate features. Only data prior to this time will be used.\n* Training window specifies the time to calculate features. Only data after this time will be used."},{"metadata":{"trusted":true,"_uuid":"4533a66895496bc898784f9db5dfe03c5996e18d"},"cell_type":"code","source":"# Tell Featuretools to add time when entity was last seen \nes.add_last_time_indexes()\n    \ntrain_cutoff_time = datetime.datetime(2017, 11, 8, 17, 0)\ntrain_training_window = ft.Timedelta(\"1 day\")\n\nfeature_matrix, feature_defs = ft.dfs(\n    entityset=es,\n    target_entity='app',\n    cutoff_time=train_cutoff_time,\n    training_window=train_training_window,\n)\n\nfeature_matrix.head()","execution_count":14,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}