{"nbformat": 4, "nbformat_minor": 1, "cells": [{"cell_type": "markdown", "metadata": {"_cell_guid": "04add1a5-d78d-4f07-8b7e-521dc29ca352", "_uuid": "57eb8343717eeae2b2a3c2f300b8a2acfacbbcb6"}, "source": ["\"***Feature Engineering is an art***\". Someone once said to spend more time on deriving more and **meaningful** features. Why did I bold meaningful? Because infomation fed to the model must make sense in whatever form it is given. Creating loads of features having no sense is of no use at all. To add more features, an added advantage would be understanding the real world situation at hand or holding sme (subject ,atter experience).\n", "\n", "On this second kernel of mine let's see how we can garner more features from the existing ones. This kernel did not run properly upon execution. So I have decided to incorporate memory reduction techniques adopted from my previous kernel mentioned below.\n", "\n", "To see my first kernel on how to reduce memory effectively [SEE THIS KERNEL](https://www.kaggle.com/jeru666/memory-reduction-and-data-insights/notebook)\n", "\n", "## Loading libraries and data"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "1f145073-4e4a-4403-9955-8d9e7c6594ee", "_uuid": "1ffcb3f110e5b228060feab970e2d82f13f12cbf"}, "source": ["import numpy as np # linear algebra\n", "import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n", "import seaborn as sns\n", "from matplotlib import pyplot\n", "\n", "# Input data files are available in the \"../input/\" directory.\n", "# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n", "\n", "from subprocess import check_output\n", "\n", "#df_members = pd.read_csv('../input/members_v3.csv')\n", "df_transactions = pd.read_csv('../input/transactions.csv')"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "78884905-e1f9-437b-b652-5443be7242a6", "_uuid": "b5bf5c597f15baa468ac5095807dbd0b82f5d169"}, "source": ["Have a quick look at the head to see what features we can create!"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "1781cd2c-a4b3-4814-923b-f32bb626e274", "_uuid": "d5e1b90fa0bf2ed89cc26f6651d5aeede8a28d3b"}, "source": ["print(df_transactions.shape)\n", "df_transactions.head()"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "a6d82aac-41c9-491d-bdbe-fa5acc3dbded", "_uuid": "4e26632e76bbd6b4978bcd96cd3954aee763bb2a"}, "source": ["## Memory Reduction\n", "As stated earlier, let us first reduce memory wherever  possible. But first how much memory does df_train dataframe consume?"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "5d659637-5071-4679-95fb-f7a8b6d0279f", "_uuid": "7ec9273634b04ec622c0efdd13e1339fffcf9de8"}, "source": ["mem = df_transactions.memory_usage(index=True).sum()\n", "print(mem/ 1024**2,\" MB\")"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "d6517fcb-84b0-4d37-81a5-1176db47a0b7", "_uuid": "11d3f4382126075516f5363b63fecc14dfc048f4"}, "source": ["The following functions check whether a column's datatype can be reduced based on the maximum and minimum value present in that column"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "08d9241d-31fe-4a63-96b7-45de09d34b7f", "_uuid": "43c0295ac2953a79fdeeedac48841b2865ada02f"}, "source": ["def change_datatype(df):\n", "    int_cols = list(df.select_dtypes(include=['int']).columns)\n", "    for col in int_cols:\n", "        if ((np.max(df[col]) <= 127) and(np.min(df[col] >= -128))):\n", "            df[col] = df[col].astype(np.int8)\n", "        elif ((np.max(df[col]) <= 32767) and(np.min(df[col] >= -32768))):\n", "            df[col] = df[col].astype(np.int16)\n", "        elif ((np.max(df[col]) <= 2147483647) and(np.min(df[col] >= -2147483648))):\n", "            df[col] = df[col].astype(np.int32)\n", "        else:\n", "            df[col] = df[col].astype(np.int64)\n", "\n", "change_datatype(df_transactions)\n", "\n", "def change_datatype_float(df):\n", "    float_cols = list(df.select_dtypes(include=['float']).columns)\n", "    for col in float_cols:\n", "        df[col] = df[col].astype(np.float32)\n", "        \n", "change_datatype_float(df_transactions)\n", "\n", "mem = df_transactions.memory_usage(index=True).sum()\n", "print(mem/ 1024**2,\" MB\")"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "eeda3d9f-cfda-4845-bcca-274c5a4fd2e1", "_uuid": "64900519a42d790c0904f22dc46edf8936641fe8"}, "source": ["We have reduced memory of **transactions** dataframe from 1.4 GB to ~500 MB. Now performing the same for **members** dataframe."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "48b24f50-7720-4cd0-ae3f-dc15a50e5b39", "_uuid": "52f5217c9b872f86a69ab47238f900d2d3e7a568"}, "source": ["#--- Members dataframe\n", "mem = df_members.memory_usage(index=True).sum()\n", "print(mem/ 1024**2,\" MB\")\n", "\n", "change_datatype(df_members)\n", "change_datatype_float(df_members)\n", "\n", "#--- Recheck memory of Members dataframe\n", "mem = df_members.memory_usage(index=True).sum()\n", "print(mem/ 1024**2,\" MB\")"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "551cb855-f9a4-4332-a9c2-b4c5b8a2b4a0", "_uuid": "fcd205ec1ce23e56915d5600500fc73eb03992a0"}, "source": ["We have reduced memory almost by 50%. Let us see the column types to notice the changes:"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "294840f9-b36b-430e-b933-1c3b8cd95bea", "_uuid": "143e31746ae5c0079b865b3637826e4f61ea66dc"}, "source": ["print(df_transactions.dtypes, '\\n')\n", "print(df_members.dtypes)"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "c6b80232-903e-41fa-905f-c28c83f01553", "_uuid": "1fa1c49e8e74e61ce0622985f225e8754cc8e50b"}, "source": ["## Transactions dataframe\n", "\n", "Now let us create new features!!\n", "\n", "Before creating features let us keep a count of the number of columns we have at the moment:"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "adaee137-11bc-4c2a-be20-ed28482c8ec6", "_uuid": "9897342b38645400a7f3d80f17c51ebbdf601832"}, "source": ["len(df_transactions.columns)"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "554600c2-7a48-4aba-b763-f9a75cec2263", "_uuid": "30e0eb26e250bbd631e8ee4c8cbdc856720a6f8d"}, "source": ["## Feature 1 : ***discount***\n", "We can create a **discount** column to see how much discount was offered to the customer."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "cef17194-5d65-42d7-bad7-50a29953bd35", "_uuid": "23c6ac3310392eeb284cc6c4ff6e4d1d6055e961"}, "source": ["df_transactions['discount'] = df_transactions['plan_list_price'] - df_transactions['actual_amount_paid']\n", "\n", "df_transactions['discount'].unique()"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "fea81b41-a0b3-4355-b480-f66a61ffe8af", "_uuid": "616e524787cf52314eea6b9f6973d7907c688cd3"}, "source": ["## Feature 2 : ***is_discount***\n", "Let us create another column **is_column** to check whether the customer has availed any discount or not. \n", "\n", "Why this feature? Oh come on you now why!"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "abf434de-8c1a-4944-9d89-02cd157fe9d7", "_uuid": "c0402a972d434a0737c08cff8e18002031769752"}, "source": ["df_transactions['is_discount'] = df_transactions.discount.apply(lambda x: 1 if x > 0 else 0)\n", "print(df_transactions['is_discount'].head())\n", "print(df_transactions['is_discount'].unique())"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "41a2831a-5724-4b74-9abb-c3b7e9c5412b", "_uuid": "4dd613f50607d61a3a820e975d4fc4910157a8b8"}, "source": ["## Feature 3 : ***amount_per_day***\n", "A new column featuring amount per-day can be added."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "cac6d21b-03ba-4d90-b9dd-50e0441e1f3f", "_uuid": "26b150e27aeeceb5c04f673acee2b6f8ea960a2e"}, "source": ["df_transactions['amt_per_day'] = df_transactions['actual_amount_paid'] / df_transactions['payment_plan_days']\n", "df_transactions['amt_per_day'].head()"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "7be993d0-9cb4-42c3-b9c7-94e2c4e34a9a", "_uuid": "edff38d7bd9fcbe4690130273fffbb147f8bdeb1"}, "source": ["Now we have two date columns :\n", "* transaction_date\t\n", "* membership_expire_date\n", "\n", "Let us see if we can extract some features from them!"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "d2ecbb05-3d4b-498d-a65b-e23eb9aef539", "_uuid": "f22eb1e1f13563375e75a21401b5dd6f1a0f2b6c"}, "source": ["date_cols = ['transaction_date', 'membership_expire_date']\n", "print(df_transactions[date_cols].dtypes)\n"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "1dda5069-0d9a-4b21-9b74-a8613766946c", "_uuid": "3bc83a28fec3b4e07d7dd24dc48d6005aab554a5"}, "source": ["Both the date columns are of **integer** type. We have to convert them to type **datetime**.\n"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "4a54abf5-a220-4865-84d5-8df2f6256e0e", "_uuid": "8bc4060e28e76c4cd3745f54f7945aa48e725f27"}, "source": ["Converting date columns from integer to datetime:"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "c43884ae-b919-41b6-b736-ed2c4548ba39", "_uuid": "c1148b39cc0058cc19068995df0ebe5dd6688424"}, "source": ["for col in date_cols:\n", "    df_transactions[col] = pd.to_datetime(df_transactions[col], format='%Y%m%d')\n", "    \n", "df_transactions.head()"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "53340a12-3716-4167-8a41-e9cfa6e9c3fb", "_uuid": "713adeaa092e2c67896f015d2d7115f65aac7b31"}, "source": ["## Feature 4 : ***membership_duration***\n", "\n", "The difference between **transaction_date** and\t**membership_expire_date** would give us membership duration.\n", "\n", "Here we  find the differnce between these two columns in terms of days and months and later preserve the result as type integer."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "cd6bdb34-2cd1-45c5-96c1-ff6a929affe5", "_uuid": "409501abae9bda483feacb3f63c4f060c14fdafb"}, "source": ["#--- difference in days ---\n", "df_transactions['membership_duration'] = df_transactions.membership_expire_date - df_transactions.transaction_date\n", "df_transactions['membership_duration'] = df_transactions['membership_duration'] / np.timedelta64(1, 'D')\n", "df_transactions['membership_duration'] = df_transactions['membership_duration'].astype(int)\n", "\n", " \n", "#---difference in months ---\n", "#df_transactions['membership_duration_M'] = (df_transactions.membership_expire_date - df_transactions.transaction_date)/ np.timedelta64(1, 'M')\n", "#df_transactions['membership_duration_M'] = round(df_transactions['membership_duration_M']).astype(int)\n", "#df_transactions['membership_duration_M'].head()\n"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "16910f7f-c7af-458d-96d4-49cf47caca70", "_uuid": "6564c6f568cf4bc81e06be5221c6a9dbc15a2482"}, "source": ["Let us check the number of columns now:"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "b2399ebc-c480-4abf-ac9d-7d5d54d63176", "_uuid": "ab299f81d66233662435d13be2b37fc713f203d3"}, "source": ["len(df_transactions.columns)"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "60a2bd6b-3f13-4a6c-9db1-2abb18f165b3", "_uuid": "347884a42795373ac154d26b6d83df8cb72c524f"}, "source": ["Now that we have created 5 more columns, we have increased the memory consumption as well. So we will run the previous functions again to keep the memory in check."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "f96696c9-5aed-4d07-9768-55d4eb250330", "_uuid": "e513d243b2c2b21388c4f4685c4cde821cf66a20"}, "source": ["change_datatype(df_transactions)\n", "change_datatype_float(df_transactions)"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "4d3a4c66-42c1-4881-a1fd-953407163d9b", "_uuid": "0b58ea4075d5429b27fc679b45a2b6b25b3f1374"}, "source": ["## Members dataframe\n", "Now let us see the members.csv file"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "0cc34307-dc73-4756-86db-9d2612785c9e", "_uuid": "05868ff4df5bdb962134dc2837748091c9e4a182"}, "source": ["df_members.head()"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "ecaaedc0-bc47-4cb5-9b8a-a2433adc7692", "_uuid": "1b1d2d404e32d52f22712ab6e04327363f939fed"}, "source": ["#--- Number of columns \n", "len(df_members.columns)"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "9bbc9103-f6d0-4278-8098-5106e29e9125", "_uuid": "35ab19921336ae10bd7b0f4efdf9ad371024754e"}, "source": ["We will have to convert the date columns as before:"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "440dad4d-9cf3-4387-9e71-bf47ed953ea8", "_uuid": "2a2fffe903d1e43b37440e972e0745bf14e2ceb4"}, "source": ["date_cols = ['registration_init_time', 'expiration_date']\n", "\n", "for col in date_cols:\n", "    df_members[col] = pd.to_datetime(df_members[col], format='%Y%m%d')"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "10c8e6f1-17fe-4f03-bf10-14bc0ed49294", "_uuid": "23e846a27e40cf26e90d5d8161bb95a0c752d440"}, "source": ["## Feature 5 : ***registration_duration***"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "0dd428f7-784a-473d-8df9-6943d23c5a34", "_uuid": "5b7ad23fad7431e78d72dbee15a7703713ec5a88"}, "source": ["#--- difference in days ---\n", "df_members['registration_duration'] = df_members.expiration_date - df_members.registration_init_time\n", "df_members['registration_duration'] = df_members['registration_duration'] / np.timedelta64(1, 'D')\n", "df_members['registration_duration'] = df_members['registration_duration'].astype(int)\n", "\n", "#---difference in months ---\n", "#df_members['registration_duration_M'] = (df_members.expiration_date - df_members.registration_init_time)/ np.timedelta64(1, 'M')\n", "#df_members['registration_duration_M'] = round(df_members['registration_duration_M']).astype(int)\n"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "5ce13735-91a6-4a91-a079-0b1a83f0523e", "_uuid": "be51bdde7976432bd806167e4f9811a927aa471e"}, "source": ["#--- Reducing and checking memory again ---\n", "change_datatype(df_members)\n", "change_datatype_float(df_members)\n", "\n", "#--- Recheck memory of Members dataframe\n", "mem = df_members.memory_usage(index=True).sum()\n", "print(mem/ 1024**2,\" MB\")"]}, {"cell_type": "markdown", "metadata": {"collapsed": true, "_cell_guid": "c49286a9-f002-46c2-887c-916bbb207231", "_uuid": "27b1201066b0bbbff2f70afea8da4caa88b2fca1"}, "source": ["## Merging dataframes (*transactions* and *members*)\n", "Now let us combine both both the dataframes(**transactions** and **members**) based on **msno** and see if we can create anything more."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "317c336c-74c4-497e-bc9b-e9d1ceae80da", "_uuid": "22161502fa0a50c0d44961287b46c1cf68f35b78"}, "source": ["#-- merging the two dataframes---\n", "df_comb = pd.merge(df_transactions, df_members, on='msno', how='inner')\n", "\n", "#--- deleting the dataframes to save memory\n", "del df_transactions\n", "del df_members\n", "\n", "df_comb.head()"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "80d46061-9277-4828-b6c6-ab424836d43c", "_uuid": "83faf8f1b5620f3950b1a26ff574679bf88f01f8"}, "source": ["#df_comb = df_comb.drop('msno', 1)\n", "mem = df_comb.memory_usage(index=True).sum()\n", "print(\"Memory consumed by training set  :   {} MB\" .format(mem/ 1024**2))"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "5eec481f-17fa-4463-932e-e7283716dab1", "_uuid": "870922d955d85370cded888835ea962d8736f45f"}, "source": ["## Feature 6 : ***reg_mem_duration***\n", "\n", "After merging both the dataframes, we observe that the  average **registration_duration** is much longer than that of **membership_duration**.\n", "\n", "So it is highly likely that customers renew their membership on a monthly or quarterly basis.\n", "\n", "We can create another column stating the difference between the registration and membership duration. I don't know if it makes sense, but let's create it."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "45910c68-3f23-4d02-9f96-59a2962890c2", "_uuid": "604f239e1ef057ae2b917fe79bc3a5958d2f3c09"}, "source": ["df_comb['reg_mem_duration'] = df_comb['registration_duration'] - df_comb['membership_duration']\n", "#df_comb['reg_mem_duration_M'] = df_comb['registration_duration_M'] - df_comb['membership_duration_M']\n", "df_comb.head()"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "9ebb0480-2221-492d-b935-e744e972d4da", "_uuid": "7af9956e747ea41f7ea89acbe6577171f6754890"}, "source": ["## Feature 7 : ***autorenew_&_not_cancel***\n", "\n", "A binary feature to see whether mebers have** auto renewed** and **not cancelled** at the same time:\n", "* **auto_renew** = 1 and\n", "* **is_cancel** = 0"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "40a890b4-875b-4f4c-841c-79a4346038bf", "_uuid": "37358c1ec55c4e649880e52846bf5b64cd13df57"}, "source": ["df_comb['autorenew_&_not_cancel'] = ((df_comb.is_auto_renew == 1) == (df_comb.is_cancel == 0)).astype(np.int8)\n", "df_comb['autorenew_&_not_cancel'].unique()"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "a0196458-c60c-42ea-ac6e-99fc2cfa39ee", "_uuid": "0902951cf4d35a6bf0d6e918f0558dbf3e39fbb8"}, "source": ["## Feature 8 : ***notAutorenew_&_cancel***\n", "Binary feature to predict possible churning if \n", "* **auto_renew** = 0 and\n", "* **is_cancel** = 1"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "274a1b05-de59-4060-8fc6-4565b0668ee4", "_uuid": "93158526bd841a75a403ba62db33776331dfe024"}, "source": ["df_comb['notAutorenew_&_cancel'] = ((df_comb.is_auto_renew == 0) == (df_comb.is_cancel == 1)).astype(np.int8)\n", "df_comb['notAutorenew_&_cancel'].unique()"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "7712232d-5987-41d5-86c1-efacfaf739f7", "_uuid": "d4fc7afe31588850db840040a523fe5fb120f467"}, "source": ["## Feature 9 : *long_time_user*\n", "\n", "A binary feature to check whether user has been registered for more than a year. This can prompt the company to offer the user some discount."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "c63e855d-2b5c-4ddb-9fd9-2b17bec203ef", "_uuid": "bf8f42a7b78d9b84f6cf71019fbd9c00670c4430"}, "source": ["df_comb['long_time_user'] = (((df_comb['registration_duration'] / 365).astype(int)) > 1).astype(int)"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "6a6d6066-9d4a-4524-9769-4e76dad49fee", "_uuid": "5553a28d33a0d3eed5ead13407e92452e6595a2a"}, "source": ["## Important Note:\n", "I have noticed that columns of type **datetime64[ns]** consume a lot of memory. So after having extracted features from these columns they can be dropped to reduce memory!!"]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "17750d89-77cd-4ea3-a343-ca8fea86d9a9", "_uuid": "ebf52c9bdbc513724343e151cdde159838550326"}, "source": ["datetime_cols = list(df_comb.select_dtypes(include=['datetime64[ns]']).columns)"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "21b08c94-50ce-4389-921f-2dbe7958091d", "_uuid": "edd77c1237ceffd5c6261f035cecd6ed374f28be"}, "source": ["Dropping columns of type **datetime64[ns]**."]}, {"cell_type": "code", "outputs": [], "execution_count": null, "metadata": {"collapsed": true, "_cell_guid": "d7b1f23c-dfca-4836-ba6d-97fd47055f0d", "_uuid": "27d2a41909d2b057b38e00c93badd9c89d3859cf"}, "source": ["df_comb = df_comb.drop([datetime_cols], 1)"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "3f91b499-4fb5-4a1c-9c5b-aaa1fc73cfa9", "_uuid": "004532241957da2afd013254b9e2ac686de61e9b"}, "source": ["You can also include duration periods in terms of months and years to get insight analysis.\n", "\n", "## Share your ideas as well!!\n", "\n", "## Do upvote if you find it useful!!"]}], "metadata": {"language_info": {"pygments_lexer": "ipython3", "mimetype": "text/x-python", "codemirror_mode": {"name": "ipython", "version": 3}, "name": "python", "nbconvert_exporter": "python", "version": "3.6.3", "file_extension": ".py"}, "kernelspec": {"name": "python3", "language": "python", "display_name": "Python 3"}}}