{"metadata": {"language_info": {"version": "3.6.3", "pygments_lexer": "ipython3", "nbconvert_exporter": "python", "codemirror_mode": {"version": 3, "name": "ipython"}, "file_extension": ".py", "name": "python", "mimetype": "text/x-python"}, "kernelspec": {"display_name": "Python 3", "name": "python3", "language": "python"}}, "nbformat_minor": 1, "cells": [{"metadata": {"_uuid": "a4694b29be0ddda3f6667596312f758239d0fb5b", "_cell_guid": "06cb3429-1683-4dff-becf-d872a26454a8", "collapsed": true}, "cell_type": "markdown", "source": ["\n", "#Introduction#\n", "\n", "In this kernel are presented data exploration and preparation for [WSDM - KKBox's Churn Prediction Challenge](http://https://www.kaggle.com/c/kkbox-churn-prediction-challenge/) the goal of which is to build a predictive model to classify users between 2 classes: churn (1) and non-churn (0) based their activity and payment history along all their lifetime on service.\n", "\n", "![Model](https://i.imgur.com/ep0WSb0.png)\n"]}, {"metadata": {"_uuid": "9da65aec2d98b9961dd96925fbaf3b8fe31003f6", "_cell_guid": "0f55cadb-a699-4c3a-b5ce-71745d8aaab5"}, "cell_type": "markdown", "source": ["As input data for builing model we have:\n", "* members_v2.csv - dataset with users of KKbox with registration date from April 2014 till April 2017 (6.7M users)\n", "* train.csv  - users marked by churn attribute in March 2017\n", "* transactions.csv - payment information about users from Jan 2015 till Feb 2017\n", "* user_logs.csv - user behavior (% and number of listened songs)\n", "\n", "All datasets will be merged on 'msno' - user identification and possible that not all user from members.csv will be presented in other sets and vice versa.\n", "\n", "I will use *Python 3.6 with Pandas.*\n", "\n", "#Quick Data Exploration#\n", "\n", "##train.csv##\n", "\n", "First of all let make superficial look on data sets and begin from *train.csv.*\n", "Churn rate in Feb 2017 was 6.4% against 93.6% regular users.\n"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "b54183afc6a2e5a8e7fe110faa8a557cf179e55c", "_cell_guid": "e0c6a44b-8555-4dc1-b38d-c492c1139bc8", "collapsed": true, "_kg_hide-output": true}, "cell_type": "code", "source": ["import pandas as pd\n", "import matplotlib.pyplot as plt\n", "plt.style.use('seaborn-darkgrid')\n", "from datetime import datetime\n"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "9598f4c088e57889ba5b119684c185b0fbd02593", "_cell_guid": "a7184b3b-60a3-48bc-91ed-c4792d3e75a0"}, "cell_type": "code", "source": ["train=pd.read_csv('../input/train.csv')\n", "gr=train.groupby('is_churn').count()\n", "gr.plot.pie(subplots=True,autopct='%.2f',figsize=(5, 5))"]}, {"metadata": {"_uuid": "d38f4378a6d8b9efaa95f1c93bbc86667278e484", "_cell_guid": "cca36620-d105-4654-af12-9f80798968f6"}, "cell_type": "markdown", "source": ["Check dataset for data missing and outlier and found that data is clear."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "310f55b645b5a14d0e382e416d1c3328cfd87d5d", "_cell_guid": "b6ca3fb7-ea47-4fcc-9574-0a74e1fcd630"}, "cell_type": "code", "source": ["train.isnull().sum()"]}, {"metadata": {"_uuid": "2c24106ca29d23772b29b7ff53e5c1a23a73bfaf", "_cell_guid": "43159492-5a2b-411e-8573-0263450a9b2b"}, "cell_type": "markdown", "source": ["##Members.csv##\n"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "e759ef13ed0597cdd82e430b80d19924310546dd", "_cell_guid": "fa737b3a-f876-4d07-906a-032de7f06bba"}, "cell_type": "code", "source": ["members=pd.read_csv('../input/members_v2.csv')\n", "members.head()"]}, {"metadata": {"_uuid": "ee8d7daffca625e55bd26ab5c37c99aee1105ece", "_cell_guid": "16ec9d36-bd30-4059-b6fc-89d9900ec983"}, "cell_type": "markdown", "source": ["Check *members.csv* for missing data:"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "6f560297f90f4e17e3cf8d48950c04d238ca8f07", "_cell_guid": "48ee023e-b632-4e1c-8e50-447233ce853c"}, "cell_type": "code", "source": ["members.isnull().sum()"]}, {"metadata": {"_uuid": "7bfa98a97fe653912dec88c3714766b1283712be", "_cell_guid": "582bbe79-b23a-4437-b2b5-b55bd0e320f5"}, "cell_type": "markdown", "source": ["Looks like *gender* field has a large number of missing data. I prefer to exclude that field from input data for the predictive model."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "ee7437de71f04b352f71f07e2413411d208b92f1", "_cell_guid": "56ed071f-960c-44e7-b67a-722c6eef0688"}, "cell_type": "code", "source": ["members=members.fillna(\"NA\")\n", "members.groupby('gender').count().plot.bar(y='msno',figsize=(16, 5))"]}, {"metadata": {"_uuid": "c83110155ea9b198b1739f6320c9880f0688dc72", "_cell_guid": "8f215c91-38ec-41da-948c-65d382bf363d"}, "cell_type": "markdown", "source": ["Brief look to other column shows that '*bd'* (birthday field) include nubmer of outliers and more then half of not relevant data (age<0).\n", "'bd' filed will be excludet from predictive model as a 'gender'. \n"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "ebddd2af78b1b547c397c1ef06a19206be1b80ce", "_cell_guid": "d8f8318e-3bc9-4b80-b5b3-b72743e75dd5"}, "cell_type": "code", "source": ["members.hist(figsize=(16, 10))"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "ea5f6ff782c64b37943c06d412564f585ebd5527", "_cell_guid": "6c38b6b5-ccea-4bd4-ba59-19e5468462a8"}, "cell_type": "code", "source": ["bin=list(range(-3200,1980, 100))\n", "group=members.groupby(pd.cut(members.bd, bin)).count()\n", "group.plot.bar(y='msno',figsize=(16, 5))"]}, {"metadata": {"_uuid": "3880ccbe16938ce49d319155886d813aac40f003", "_cell_guid": "3827e914-a868-4cce-bebd-11fd3ee3c6a6"}, "cell_type": "markdown", "source": ["After data  munging *merge.csv* will look like:"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "d645b17a2310688850931a585f676d2cefa9120b", "_cell_guid": "44126172-cc30-47ce-a3dc-8ada4319e61f", "collapsed": true, "_kg_hide-output": true}, "cell_type": "code", "source": ["columns = ['gender','bd']\n", "members=members.drop(columns, axis=1)"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "07d4423fc9196c1ee9bbc96d02c3b632df929243", "_cell_guid": "9bd32cd8-d0fb-4fc9-9905-e9f59f566071", "collapsed": true, "_kg_hide-output": true}, "cell_type": "code", "source": ["members['registration_init_time']=pd.to_datetime(members['registration_init_time'],format='%Y%m%d')"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "aa94566f9f077ccc2ccc6d650d6391e813c572cc", "_cell_guid": "5d7952f2-1437-445a-b897-056aeb7df6b5"}, "cell_type": "code", "source": ["members.head()"]}, {"metadata": {"_uuid": "60c6787ea664f4efe50fe57a5771072e63e991e7", "_cell_guid": "3c715804-28d4-475d-a996-711dcadbc8e6"}, "cell_type": "markdown", "source": ["##Transactions.csv##\n", "\n", "Will check *transactions.csv* for missing data, outliers and not-meaningful attributes."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "7db758de3bb7e3d32c7eb73fdd48e8d67c42538c", "_cell_guid": "24a05d56-6c25-4b0a-80b3-d1b5293048a5"}, "cell_type": "code", "source": ["transactions = pd.read_csv('../input/transactions.csv')\n", "transactions.head()"]}, {"metadata": {"_kg_hide-input": true, "_uuid": "369936ad5154b180d931e6c5c1027666c38b58c6", "_cell_guid": "991178a4-40f5-4043-ab2e-cc4aeb1afd58", "collapsed": true}, "cell_type": "markdown", "source": ["Checking for missing data shown that all field filled."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "5f63468d765bd53f24c8d77005d8fe838c3aba3c", "_cell_guid": "9c3d381a-fa08-42b4-be4c-3e830ecc6277"}, "cell_type": "code", "source": ["transactions.isnull().sum()"]}, {"metadata": {"_uuid": "08fc9b5925fb74a982301315321ea5f4178b7d08", "_cell_guid": "a6da61d3-8c38-43a1-bb25-89a49d848802", "collapsed": true}, "cell_type": "markdown", "source": ["Converting date filds into timestamp"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "480a52033513fe32b9b63cc75af812580813dc96", "_cell_guid": "425abc01-0be3-4ea1-ab98-6431541a85ac", "collapsed": true, "_kg_hide-output": false}, "cell_type": "code", "source": ["transactions['transaction_date']=pd.to_datetime(transactions['transaction_date'],format='%Y%m%d') \n", "transactions['membership_expire_date']=pd.to_datetime(transactions['membership_expire_date'],format='%Y%m%d')"]}, {"metadata": {"_uuid": "49e260357948a3cb5d9f828fa20a346963ea5543", "_cell_guid": "17b82674-9f95-4d69-b75d-ad5da22d50a3"}, "cell_type": "markdown", "source": ["Numerical data looks clean:  without outliers."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "81c8837a970db0e166c3192235218dbe6806e76b", "_cell_guid": "937e7406-a773-464d-9757-b4f4baf2d13e"}, "cell_type": "code", "source": ["transactions.describe().transpose()"]}, {"metadata": {"_uuid": "11235718bce0c57716f45694c2b74ec16960b49e", "_cell_guid": "915d02cc-56d1-4bc8-9ffc-2f570e70dbaa"}, "cell_type": "markdown", "source": ["As we can see from plots most of the transactions are auto-renewal subscriptions, a number of canceled transaction is not big."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "cc85259752d89c989cd296f4a00c9ec1e9f9a581", "_cell_guid": "31fde0d2-eb38-4ca3-b312-72dce02659a1"}, "cell_type": "code", "source": ["transactions.hist(column=['is_cancel', 'is_auto_renew'],figsize=(16, 5))\n"]}, {"metadata": {"_uuid": "5be8035e6818dc2d915b983520835cf7579b81d3", "_cell_guid": "5b6d64b2-2cf0-40f2-a59c-34e3a1a0a05e"}, "cell_type": "markdown", "source": ["Most of users use payment method encoded as \"40\" and payment plan days \"30 days\" is absolute leader."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "63eca91e274179d281c81b80734413b3634ae3b7", "_cell_guid": "f68696d2-06cf-4b6d-a1c7-b8bd4bbccdd3"}, "cell_type": "code", "source": ["transactions.hist(column=['payment_method_id','payment_plan_days'],figsize=(16, 5),bins=50)"]}, {"metadata": {"_uuid": "6b6c8aaf25898f12ecc36b46cc061e9a2f7fcc2a", "_cell_guid": "e29e45b2-6098-47bb-9400-ae5c4daa0882"}, "cell_type": "markdown", "source": ["Plots for *\"plan list price\"* and *\"actual amount paid\"* are identical with median=149."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "7782d30a6222068c0549e7e2926030d1ecf17920", "_cell_guid": "ef9615c4-3eaf-49ac-9ed6-537bbbb40993"}, "cell_type": "code", "source": ["transactions.hist(column=['actual_amount_paid','plan_list_price'],figsize=(16, 5),bins=50)"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "4ab6eff5c6b08949d157c773dcf5ebe6dac08d72", "_cell_guid": "a2c77bad-c1e8-4ac7-92b5-2981580bd1cb", "_kg_hide-output": true}, "cell_type": "code", "source": ["transactions[\"actual_amount_paid\"].median()"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "ecbe2b1ac0ec0f5b059ddd205ae015bbfd4c5d6c", "_cell_guid": "edfc9153-6004-4641-aa3b-ddae7e7089c4", "_kg_hide-output": true}, "cell_type": "code", "source": ["transactions[\"plan_list_price\"].median()"]}, {"metadata": {"_uuid": "5c53fd8cec91cbdef7dda2013cdfc96db76e1c6b", "_cell_guid": "9c454aca-b461-427c-b584-5a0527a72ae6"}, "cell_type": "markdown", "source": ["##User_logs.csv##\n", "\n", "User_logs.csv size is 6.65 GB in archive format. For this trainnig kernal I will read a part of it - 30M rows."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "9c61c3597a4e27ac720bb24cf4a0fdb8f0142f3a", "_cell_guid": "0c70cc67-95d3-4403-bd9e-6fcf89f76808", "collapsed": true, "_kg_hide-output": true}, "cell_type": "code", "source": ["logs = pd.read_csv('../input/user_logs.csv', nrows=30000000)"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "fe735a32e83f3c0e6aa61e967d646c6ace7cbb82", "_cell_guid": "9db1c4ae-be94-4454-859a-db9b91343602"}, "cell_type": "code", "source": ["logs.head()"]}, {"metadata": {"_uuid": "86b0fce734e331fba3e6b665b34dd95e4f5c1770", "_cell_guid": "2f06446f-5381-47c5-9d77-c2d2092f3873"}, "cell_type": "markdown", "source": ["Checking data for missing valuas  and converting  date format into timestamp."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "3106c6369842e99eff36e16943a66f63de393f43", "_cell_guid": "2e30ca4f-588d-4d5d-810c-1a49cc30a375"}, "cell_type": "code", "source": ["logs.isnull().sum()"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "b47de5ef480b5b92410f63e45d19d1e6ed2b90b2", "_cell_guid": "9c0f025e-ae25-4b53-9ef1-040f900c0123", "collapsed": true, "_kg_hide-output": true}, "cell_type": "code", "source": ["logs['date']=pd.to_datetime(logs['date'],format='%Y%m%d') "]}, {"metadata": {"_uuid": "3fcb5c4b0faafa9724a4d22694e95687b9805818", "_cell_guid": "e56abd05-800c-4a5e-8862-d083964beeb7"}, "cell_type": "markdown", "source": ["According to plot below most songs are listened to the end every day."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "5d5758522e6906174e800896a30a3b26d0c9c3f8", "_cell_guid": "66b3c3af-31c9-4281-9e0e-8548383d7f39", "collapsed": true, "_kg_hide-output": true}, "cell_type": "code", "source": ["gr=logs.groupby(pd.Grouper(key='date', freq='D')).mean()"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "76c27268662db9d3eb38385dac73774dd03e30e6", "_cell_guid": "dab5a47d-168e-4a46-be51-7b1dfbb5bf52"}, "cell_type": "code", "source": ["gr.plot.line(y=['num_25','num_50','num_75','num_985','num_100'],figsize=(16, 5))"]}, {"metadata": {"_uuid": "d524f0e00f80aaae8b5f2528a8ed1ef0e17df646", "_cell_guid": "e45e2df0-3f4e-4a01-98d2-68126af56c70"}, "cell_type": "markdown", "source": ["Average number of unique songs played per day by users."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "998031416376ac2c15c9bb241f30d886fc462dcb", "_cell_guid": "48400b38-201d-4331-a43f-c4836a9e907d"}, "cell_type": "code", "source": ["gr.plot.line(y=['num_unq'],figsize=(16, 5))"]}, {"metadata": {"_uuid": "cd0f381c0e67cd102bd0e9f692ecac8d79913c3e", "_cell_guid": "aa818a9f-9eeb-4f78-b9e8-6a4167589187"}, "cell_type": "markdown", "source": ["#Hypothesis generation#\n", "\n", "First useful information dataset will be merged from *transaction.csv, train.csv *and* members.csv.* with 11.7M rows with information about transactions of ~700K users"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "255111b42db772df8d90b3bcabe88d1309c46b47", "_cell_guid": "5deb0904-c242-4684-a14c-a5b74c99ec8e"}, "cell_type": "code", "source": ["train_total=pd.merge(transactions,pd.merge(members,train,on='msno',how='inner'),on='msno',how='inner')\n", "train_total.head()"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "100f42bc1087204535fd5db53d03791b3bed6cf8", "_cell_guid": "9b4c8260-665a-46d5-911b-d888f1d9b823", "_kg_hide-output": true}, "cell_type": "code", "source": ["train_total.info()"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": false, "_uuid": "5ae09a17f147e051eba65e784308a624c90fd77c", "_cell_guid": "7b5caa60-c67c-4500-8a7f-5796dc9b1455", "collapsed": true, "_kg_hide-output": false}, "cell_type": "code", "source": ["gr=train_total.groupby(['transaction_date','is_churn']).count()"]}, {"metadata": {}, "cell_type": "markdown", "source": ["Below you can see on the plot number of transactions per user (chart once again shows the popularity of 30-day auto-renewal subscriptions)"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "f98afddf997eddd2d0178a62fe4acf0767486f39", "_cell_guid": "f9f214aa-416a-407e-98b1-53ad04ba799c"}, "cell_type": "code", "source": ["gr2=gr.unstack()\n", "gr2.plot.line(y='msno',figsize=(16, 5))"]}, {"metadata": {}, "cell_type": "markdown", "source": ["If look on users who churn on March 2017 we can see some transactions in February 2017..."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "f3e8ee5582b6c3d7c8f71a6f440bd61cb6b01f79", "_cell_guid": "d0cfccd8-2a78-4b60-ba64-a051f32247ea"}, "cell_type": "code", "source": ["gr2.plot.line(y='msno',figsize=(20, 8),ylim=(0,600),xlim=(pd.Timestamp('2015-10-01'), pd.Timestamp('2017-03-28')))"]}, {"metadata": {}, "cell_type": "markdown", "source": ["Transactions in February from churned user was \"canceled transactions\" only."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_kg_hide-output": true}, "cell_type": "code", "source": ["df1=train_total[train_total['is_cancel'] ==1]\n", "\n"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_kg_hide-output": true}, "cell_type": "code", "source": ["gr3=df1.groupby(['transaction_date','is_churn']).sum()\n"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true}, "cell_type": "code", "source": ["gr4=gr3.unstack()\n", "gr4.plot.line(y='is_cancel',figsize=(16, 5),ylim=(0,60),xlim=(pd.Timestamp('2015-10-01'), pd.Timestamp('2017-03-28')))"]}, {"metadata": {"_uuid": "2928ff1e28ee59b14107f5e877050b8837fa03b2", "_cell_guid": "5bf3e771-64e2-4fbd-8289-c9a6feb390eb"}, "cell_type": "markdown", "source": ["Second useful dataset will be merged from *user_logs.csv, train.csv *and* members.csv.* ~14.3M rows."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "688b04bae44a24ec1185a9d3daa9da5e1f9118eb", "_cell_guid": "9b7051e9-2cca-4605-8bc3-7e617b77d33a"}, "cell_type": "code", "source": ["logs_total=pd.merge(logs,pd.merge(members,train,on='msno',how='inner'),on='msno',how='inner')\n", "logs_total.head()"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "aba336b9ff6c79e8178dd74ccbcbe190252bffc0", "_cell_guid": "233b18ca-1ad4-427d-9cb7-18f0874a4787", "_kg_hide-output": true}, "cell_type": "code", "source": ["logs_total.info()"]}, {"metadata": {}, "cell_type": "markdown", "source": ["Number of unique songs played 100% by users depending on churn attribute."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true, "_uuid": "f20fed175cd7495d50447df1a02c4cb1dd71d94f", "_cell_guid": "ee1fee02-c6b8-4423-9e07-2952e0ee7e43", "_kg_hide-output": true}, "cell_type": "code", "source": ["group1=logs_total.groupby(['date','is_churn']).sum()"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true}, "cell_type": "code", "source": ["group2=group1.unstack()\n", "group2.plot.line(y='num_100',figsize=(16, 5))"]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true}, "cell_type": "code", "source": ["group2.plot.line(y=['num_100'],figsize=(16, 5),ylim=(0,15000),xlim=(pd.Timestamp('2015-10-01'), pd.Timestamp('2017-03-28')))"]}, {"metadata": {}, "cell_type": "markdown", "source": ["Unfortunately, I lost data about churned user behavior when decided to limit file importation. And now I see lack of data for data-driven hypothesis generation."]}, {"outputs": [], "execution_count": null, "metadata": {"_kg_hide-input": true}, "cell_type": "code", "source": ["group2.plot.line(y=['num_50'],figsize=(16, 5),ylim=(0,2000),xlim=(pd.Timestamp('2015-10-01'), pd.Timestamp('2017-03-28')))"]}, {"metadata": {}, "cell_type": "markdown", "source": ["Only, from logic and experience, I can guess that *user_logs.csv* will be more useful for predict user involment or satisfaction in service and *transactions.csv* can show financial pattern for usage and possible user revival after some churn period (seasonality etc.)"]}, {"metadata": {}, "cell_type": "markdown", "source": ["> I am new to Python and real big data analysis, and I will be thankful for advice, remarks or constructive criticism."]}], "nbformat": 4}