{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# American Express - Default Prediction \n## Predict If A Customer Will Default in the Future ...\nThe objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\n<img style=\"float: center;\" src=\"https://img.freepik.com/free-vector/brain-with-digital-circuit-programmer-with-laptop-machine-learning-artificial-intelligence-digital-brain-artificial-thinking-process-concept-vector-isolated-illustration_335657-2246.jpg?w=2000\" width = '550'>\n<a href='https://www.freepik.com/vectors/machine-learning'>Machine learning vector created by vectorjuice - www.freepik.com</a>\n\n#### Data Description\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\n* D_* = Delinquency variables\n* S_* = Spend variables\n* P_* = Payment variables\n* B_* = Balance variables\n* R_* = Risk variables\n\nWith the following features being categorical:\n\n**['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']**\n\nYour task is to predict, for each customer_ID, the probability of a future payment default (target = 1).\n\nNote that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.\n\n**Files**\n* train_data.csv - training data with multiple statement dates per customer_ID\n* train_labels.csv - target label for each customer_ID\n* test_data.csv - corresponding test data; your objective is to predict the target label for each customer_ID\n* sample_submission.csv - a sample submission file in the correct format\n\n---\n\n## My Strategy, or How I Will Aproach this Competition...\nWe have data from many Customers and there is many points of information by for each of the customers, the target labels are only one per customer id so aggregation will be requiered, from here there is quie a lot of possibilities, this is what I will folow in this Notebook...\n\n#### Loading the Datasets\nThe datasets is massive so I will rely on other Kaggles optimized datasets stored in a feather format to make my life easier in this competition.\n\n#### Quick EDA\nThe typical analysis that I always like to complete to undertstand the dataset better...\n* Information of the datasets, size and others.\n* Simple visualization of the first few records.\n* Data statistical analalysis using describe.\n* Visualization of the number of NaNs.\n* Understanding the amount of unique records.\n\n#### Exploring the Target Variable\nNothing in particular dataset seems to be quite inbalanced so I will get back to this part later...\n\n#### Structuring the Datasets\nHere is where everything happens, because we have time-base data o multiple points per customer we are trying to aggregate the information in certain way that's practical:\n* Statistical aggregation for numeric features\n* Only keep the last know record for analysis\n* Statictical aggregation for categorical features\n\n#### Feature Engineering\nAt this point the only thing that I can consider some type of feature will be the aggregation of the datasets, as I mentioned in the previous point\n* Statistical aggregation\n* Only keep the last know record for analysis\n\n#### Label Encoding\nBecause there is quite a lot of categorical variables and this is a NN model I will use the following encoding technique:\n* OneHot encoder, only train in the train dataset and applyed on test\n\n#### Fill NaNs**\nAt this point just to get started, I will fill everything with ceros, probably not a good idea.\n* Fill NaNs with 0\n\n#### Model Development and Training\nI'm going to go first with an NN in the last few competitions the NN models have been working quite well also we have so much data.\n* Simple NN tested, layer after later.\n* I also tested a more complex NN, that I learned from Ambross with Skip conections.\n\n#### Predictions and Submission\nNo much details here, just the simple average of all the predictions across multiple folds.\n* Average predictions across 5 folds\n\n---\n\n## Updates\n#### 05/28/2022\n* Build the initial model using Neuronal Nets and simple agg strategy (Last data point).\n* Evaluated the model and uploaded for Ranking.\n\n#### 05/29/2022\n* Improve model architecture.\n* Really dive deep into Feature Engineering (Not much here, memory is a big challenge)\n\n#### 05/30/2022\n* ...\n\n---\n\n## Resources, Inspiration\nI have taken Ideas or learned quite a lot from the Notebooks below, please check also if you like my work.\n\n* https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training/notebook\n* ...\n* ...\n* ...","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 1.0 Loading Model Libraries...","metadata":{}},{"cell_type":"code","source":"%%time\n# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-23T20:37:40.495166Z","iopub.execute_input":"2022-07-23T20:37:40.495508Z","iopub.status.idle":"2022-07-23T20:37:40.530191Z","shell.execute_reply.started":"2022-07-23T20:37:40.495436Z","shell.execute_reply":"2022-07-23T20:37:40.529405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport datetime # ...","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:37:40.531791Z","iopub.execute_input":"2022-07-23T20:37:40.532123Z","iopub.status.idle":"2022-07-23T20:37:40.536530Z","shell.execute_reply.started":"2022-07-23T20:37:40.532091Z","shell.execute_reply":"2022-07-23T20:37:40.535769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 2.0 Setting the Notebook Parameters and Default Configuration...","metadata":{}},{"cell_type":"code","source":"%%time\n# I like to disable my Notebook Warnings.\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:37:40.538000Z","iopub.execute_input":"2022-07-23T20:37:40.538959Z","iopub.status.idle":"2022-07-23T20:37:40.546942Z","shell.execute_reply.started":"2022-07-23T20:37:40.538922Z","shell.execute_reply":"2022-07-23T20:37:40.546043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Notebook Configuration...\n\n# Amount of data we want to load into the Model...\nDATA_ROWS = None\n# Dataframe, the amount of rows and cols to visualize...\nNROWS = 50\nNCOLS = 15\n# Main data location path...\nBASE_PATH = '...'","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:37:40.548523Z","iopub.execute_input":"2022-07-23T20:37:40.549428Z","iopub.status.idle":"2022-07-23T20:37:40.558112Z","shell.execute_reply.started":"2022-07-23T20:37:40.549393Z","shell.execute_reply":"2022-07-23T20:37:40.557106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Configure notebook display settings to only use 2 decimal places, tables look nicer.\npd.options.display.float_format = '{:,.5f}'.format\npd.set_option('display.max_columns', NCOLS) \npd.set_option('display.max_rows', NROWS)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:37:40.560337Z","iopub.execute_input":"2022-07-23T20:37:40.560792Z","iopub.status.idle":"2022-07-23T20:37:40.566558Z","shell.execute_reply.started":"2022-07-23T20:37:40.560756Z","shell.execute_reply":"2022-07-23T20:37:40.565839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 3.0 Loading the Dataset Information (Using Feather)...","metadata":{}},{"cell_type":"code","source":"%%time\n# Load the CSV information into a Pandas DataFrame...\ntrn_data = pd.read_feather('../input/parquet-files-amexdefault-prediction/train_data.ftr')\ntrn_lbls = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv').set_index('customer_ID')\n\ntst_data = pd.read_feather('../input/parquet-files-amexdefault-prediction/test_data.ftr')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:37:40.567894Z","iopub.execute_input":"2022-07-23T20:37:40.568648Z","iopub.status.idle":"2022-07-23T20:38:44.231637Z","shell.execute_reply.started":"2022-07-23T20:37:40.568606Z","shell.execute_reply":"2022-07-23T20:38:44.230777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsub = pd.read_csv('/kaggle/input/amex-default-prediction/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:38:44.233042Z","iopub.execute_input":"2022-07-23T20:38:44.233648Z","iopub.status.idle":"2022-07-23T20:38:46.162462Z","shell.execute_reply.started":"2022-07-23T20:38:44.233602Z","shell.execute_reply":"2022-07-23T20:38:46.161644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 4.0 Exploring the Dataset, Quick EDA...","metadata":{}},{"cell_type":"code","source":"%%time\n# Explore the shape of the DataFrame...\ntrn_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:38:46.163936Z","iopub.execute_input":"2022-07-23T20:38:46.164361Z","iopub.status.idle":"2022-07-23T20:38:46.172672Z","shell.execute_reply.started":"2022-07-23T20:38:46.164321Z","shell.execute_reply":"2022-07-23T20:38:46.171783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display simple information of the variables in the dataset...\ntrn_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:38:46.174171Z","iopub.execute_input":"2022-07-23T20:38:46.174572Z","iopub.status.idle":"2022-07-23T20:38:46.205167Z","shell.execute_reply.started":"2022-07-23T20:38:46.174532Z","shell.execute_reply":"2022-07-23T20:38:46.204250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display the first few rows of the DataFrame...\ntrn_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:38:46.206514Z","iopub.execute_input":"2022-07-23T20:38:46.206844Z","iopub.status.idle":"2022-07-23T20:38:46.229719Z","shell.execute_reply.started":"2022-07-23T20:38:46.206811Z","shell.execute_reply":"2022-07-23T20:38:46.228868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display the Min Date...\ntrn_data['S_2'].min()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:38:46.232316Z","iopub.execute_input":"2022-07-23T20:38:46.232943Z","iopub.status.idle":"2022-07-23T20:38:46.980632Z","shell.execute_reply.started":"2022-07-23T20:38:46.232916Z","shell.execute_reply":"2022-07-23T20:38:46.979830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display the Max Date...\ntrn_data['S_2'].max()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:38:46.982091Z","iopub.execute_input":"2022-07-23T20:38:46.982722Z","iopub.status.idle":"2022-07-23T20:38:47.728859Z","shell.execute_reply.started":"2022-07-23T20:38:46.982685Z","shell.execute_reply":"2022-07-23T20:38:47.728072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Generate a simple statistical summary of the DataFrame, Only Numerical...\ntrn_data.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:38:47.730223Z","iopub.execute_input":"2022-07-23T20:38:47.730723Z","iopub.status.idle":"2022-07-23T20:41:14.488258Z","shell.execute_reply.started":"2022-07-23T20:38:47.730683Z","shell.execute_reply":"2022-07-23T20:41:14.487283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Calculates the total number of missing values...\ntrn_data.isnull().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:14.489810Z","iopub.execute_input":"2022-07-23T20:41:14.490367Z","iopub.status.idle":"2022-07-23T20:41:20.017045Z","shell.execute_reply.started":"2022-07-23T20:41:14.490329Z","shell.execute_reply":"2022-07-23T20:41:20.016256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display the number of missing values by variable...\ntrn_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:20.018408Z","iopub.execute_input":"2022-07-23T20:41:20.019000Z","iopub.status.idle":"2022-07-23T20:41:25.349389Z","shell.execute_reply.started":"2022-07-23T20:41:20.018958Z","shell.execute_reply":"2022-07-23T20:41:25.348512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display the number of unique values for each variable...\ntrn_data.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:25.350902Z","iopub.execute_input":"2022-07-23T20:41:25.351519Z","iopub.status.idle":"2022-07-23T20:41:40.511173Z","shell.execute_reply.started":"2022-07-23T20:41:25.351482Z","shell.execute_reply":"2022-07-23T20:41:40.510144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display the number of unique values for each variable, sorted by quantity...\ntrn_data.nunique().sort_values(ascending = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:40.512481Z","iopub.execute_input":"2022-07-23T20:41:40.512905Z","iopub.status.idle":"2022-07-23T20:41:56.024688Z","shell.execute_reply.started":"2022-07-23T20:41:40.512868Z","shell.execute_reply":"2022-07-23T20:41:56.023761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 5.0 Understanding the Target Variable...","metadata":{}},{"cell_type":"code","source":"%%time\n# Explore the shape of the DataFrame...\ntrn_lbls.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:56.026145Z","iopub.execute_input":"2022-07-23T20:41:56.026528Z","iopub.status.idle":"2022-07-23T20:41:56.033370Z","shell.execute_reply.started":"2022-07-23T20:41:56.026494Z","shell.execute_reply":"2022-07-23T20:41:56.032598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display simple information of the variables in the dataset...\ntrn_lbls.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:56.034696Z","iopub.execute_input":"2022-07-23T20:41:56.035270Z","iopub.status.idle":"2022-07-23T20:41:56.049773Z","shell.execute_reply.started":"2022-07-23T20:41:56.035236Z","shell.execute_reply":"2022-07-23T20:41:56.048835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Check how well balanced is the dataset\ntrn_lbls['target'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:56.051267Z","iopub.execute_input":"2022-07-23T20:41:56.051952Z","iopub.status.idle":"2022-07-23T20:41:56.063186Z","shell.execute_reply.started":"2022-07-23T20:41:56.051909Z","shell.execute_reply":"2022-07-23T20:41:56.062270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Check some statistics on the target variable\ntrn_lbls['target'].describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:56.064693Z","iopub.execute_input":"2022-07-23T20:41:56.065352Z","iopub.status.idle":"2022-07-23T20:41:56.085967Z","shell.execute_reply.started":"2022-07-23T20:41:56.065316Z","shell.execute_reply":"2022-07-23T20:41:56.085256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 6.0 Structuring Data for the Model (Aggreations and More)","metadata":{}},{"cell_type":"markdown","source":"## 6.1 Training Dataset...","metadata":{}},{"cell_type":"code","source":"%%time\n# We have 458913 customers. and we have 458913 train labels...","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:56.087271Z","iopub.execute_input":"2022-07-23T20:41:56.087795Z","iopub.status.idle":"2022-07-23T20:41:56.093002Z","shell.execute_reply.started":"2022-07-23T20:41:56.087762Z","shell.execute_reply":"2022-07-23T20:41:56.091822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Calculates the amount of information by costumer or records available...\ntrn_num_statements = trn_data.groupby('customer_ID').size().sort_index()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:56.094443Z","iopub.execute_input":"2022-07-23T20:41:56.095008Z","iopub.status.idle":"2022-07-23T20:41:57.443090Z","shell.execute_reply.started":"2022-07-23T20:41:56.094973Z","shell.execute_reply":"2022-07-23T20:41:57.442188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Review some of the information created...\ntrn_num_statements","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:57.444279Z","iopub.execute_input":"2022-07-23T20:41:57.444951Z","iopub.status.idle":"2022-07-23T20:41:57.523155Z","shell.execute_reply.started":"2022-07-23T20:41:57.444912Z","shell.execute_reply":"2022-07-23T20:41:57.522291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Create a new dataset based on aggregated information\ntrn_agg_data = (trn_data\n                .groupby('customer_ID')\n                .tail(1)\n                .set_index('customer_ID', drop=True)\n                .sort_index()\n                .drop(['S_2'], axis='columns'))\n\n# Merge the labels from the labels dataframe\ntrn_agg_data['target'] = trn_lbls.target\ntrn_agg_data['num_statements'] = trn_num_statements\n\ntrn_agg_data.reset_index(inplace = True, drop = True) # forget the customer_IDs","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:41:57.524455Z","iopub.execute_input":"2022-07-23T20:41:57.524953Z","iopub.status.idle":"2022-07-23T20:42:00.571890Z","shell.execute_reply.started":"2022-07-23T20:41:57.524917Z","shell.execute_reply":"2022-07-23T20:42:00.570913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrn_agg_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:00.573252Z","iopub.execute_input":"2022-07-23T20:42:00.573703Z","iopub.status.idle":"2022-07-23T20:42:00.593505Z","shell.execute_reply.started":"2022-07-23T20:42:00.573666Z","shell.execute_reply":"2022-07-23T20:42:00.592675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 6.2 Test Dataset...","metadata":{}},{"cell_type":"code","source":"%%time\n# Calculates the amount of information by costumer or records available...\ntst_num_statements = tst_data.groupby('customer_ID').size().sort_index()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:00.597722Z","iopub.execute_input":"2022-07-23T20:42:00.598495Z","iopub.status.idle":"2022-07-23T20:42:03.422728Z","shell.execute_reply.started":"2022-07-23T20:42:00.598441Z","shell.execute_reply":"2022-07-23T20:42:03.421770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Create a new dataset based on aggregated information\ntst_agg_data = (tst_data\n                .groupby('customer_ID')\n                .tail(1)\n                .set_index('customer_ID', drop=True)\n                .sort_index()\n                .drop(['S_2'], axis='columns'))\n\n# Merge the labels from the labels dataframe\ntst_agg_data['num_statements'] = tst_num_statements\n\ntst_agg_data.reset_index(inplace = True, drop = True) # forget the customer_IDs","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:03.424294Z","iopub.execute_input":"2022-07-23T20:42:03.424672Z","iopub.status.idle":"2022-07-23T20:42:09.270903Z","shell.execute_reply.started":"2022-07-23T20:42:03.424636Z","shell.execute_reply":"2022-07-23T20:42:09.269919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntst_agg_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:09.272346Z","iopub.execute_input":"2022-07-23T20:42:09.272711Z","iopub.status.idle":"2022-07-23T20:42:09.292650Z","shell.execute_reply.started":"2022-07-23T20:42:09.272676Z","shell.execute_reply":"2022-07-23T20:42:09.291838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 7.0 Label / One-Hot Encoding the Categorical Variables...","metadata":{}},{"cell_type":"markdown","source":"## 7.1 One Hot Encoding Configuration...","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.preprocessing import StandardScaler, QuantileTransformer, OneHotEncoder, OrdinalEncoder","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:09.294228Z","iopub.execute_input":"2022-07-23T20:42:09.294566Z","iopub.status.idle":"2022-07-23T20:42:10.036824Z","shell.execute_reply.started":"2022-07-23T20:42:09.294534Z","shell.execute_reply":"2022-07-23T20:42:10.035815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# One-hot Encoding Configuration\ncat_features = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\n#trn_agg_data[cat_features] = trn_agg_data[cat_features].astype(object)\ntrn_not_cat_features = [f for f in trn_agg_data.columns if f not in cat_features]\ntst_not_cat_features = [f for f in tst_agg_data.columns if f not in cat_features]","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:10.040725Z","iopub.execute_input":"2022-07-23T20:42:10.041203Z","iopub.status.idle":"2022-07-23T20:42:10.054550Z","shell.execute_reply.started":"2022-07-23T20:42:10.041166Z","shell.execute_reply":"2022-07-23T20:42:10.053543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrn_agg_data[cat_features].head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:10.059153Z","iopub.execute_input":"2022-07-23T20:42:10.061380Z","iopub.status.idle":"2022-07-23T20:42:10.096451Z","shell.execute_reply.started":"2022-07-23T20:42:10.061337Z","shell.execute_reply":"2022-07-23T20:42:10.095639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n#encoder = OneHotEncoder(drop = 'first', sparse = False, dtype = np.float32, handle_unknown = 'ignore')\nencoder = OrdinalEncoder()\ntrn_encoded_features = encoder.fit_transform(trn_agg_data[cat_features])\n#feat_names = list(encoder.get_feature_names())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:10.101017Z","iopub.execute_input":"2022-07-23T20:42:10.103509Z","iopub.status.idle":"2022-07-23T20:42:10.934200Z","shell.execute_reply.started":"2022-07-23T20:42:10.103466Z","shell.execute_reply":"2022-07-23T20:42:10.932950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.2 Train Dataset One Hot Encoding...","metadata":{}},{"cell_type":"code","source":"%%time\n# One-hot Encoding\ntrn_encoded_features = pd.DataFrame(trn_encoded_features)\n#trn_encoded_features.columns = feat_names","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:10.935676Z","iopub.execute_input":"2022-07-23T20:42:10.936248Z","iopub.status.idle":"2022-07-23T20:42:10.943127Z","shell.execute_reply.started":"2022-07-23T20:42:10.936196Z","shell.execute_reply":"2022-07-23T20:42:10.941919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrn_agg_data = pd.concat([trn_agg_data[trn_not_cat_features], trn_encoded_features], axis = 1)\ntrn_agg_data.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:10.944559Z","iopub.execute_input":"2022-07-23T20:42:10.945451Z","iopub.status.idle":"2022-07-23T20:42:11.563033Z","shell.execute_reply.started":"2022-07-23T20:42:10.945352Z","shell.execute_reply":"2022-07-23T20:42:11.562089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 7.3 Test Dataset One-Hot Encoding...","metadata":{}},{"cell_type":"code","source":"%%time\ntst_agg_data[cat_features].head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:11.564664Z","iopub.execute_input":"2022-07-23T20:42:11.565153Z","iopub.status.idle":"2022-07-23T20:42:11.593198Z","shell.execute_reply.started":"2022-07-23T20:42:11.565094Z","shell.execute_reply":"2022-07-23T20:42:11.592206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# One-hot Encoding\ntst_encoded_features = encoder.transform(tst_agg_data[cat_features])\ntst_encoded_features = pd.DataFrame(tst_encoded_features)\n#tst_encoded_features.columns = feat_names","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:11.594927Z","iopub.execute_input":"2022-07-23T20:42:11.595507Z","iopub.status.idle":"2022-07-23T20:42:12.831499Z","shell.execute_reply.started":"2022-07-23T20:42:11.595470Z","shell.execute_reply":"2022-07-23T20:42:12.830626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntst_agg_data = pd.concat([tst_agg_data[tst_not_cat_features], tst_encoded_features], axis = 1)\ntst_agg_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:12.832843Z","iopub.execute_input":"2022-07-23T20:42:12.833418Z","iopub.status.idle":"2022-07-23T20:42:14.025416Z","shell.execute_reply.started":"2022-07-23T20:42:12.833375Z","shell.execute_reply":"2022-07-23T20:42:14.024383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 8.0 Pre-Processing the Data, Fill NaNs for model functionality...","metadata":{}},{"cell_type":"code","source":"%%time\n# Impute missing values\ntrn_agg_data.fillna(value = 0, inplace = True)\ntst_agg_data.fillna(value = 0, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:14.026747Z","iopub.execute_input":"2022-07-23T20:42:14.027259Z","iopub.status.idle":"2022-07-23T20:42:15.558238Z","shell.execute_reply.started":"2022-07-23T20:42:14.027204Z","shell.execute_reply":"2022-07-23T20:42:15.557388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 9.0 Feature Selection for Baseline Model...","metadata":{}},{"cell_type":"code","source":"%%time\nfeatures = [f for f in trn_agg_data.columns if f != 'target' and f != 'customer_ID']","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:15.559575Z","iopub.execute_input":"2022-07-23T20:42:15.560093Z","iopub.status.idle":"2022-07-23T20:42:15.566025Z","shell.execute_reply.started":"2022-07-23T20:42:15.560056Z","shell.execute_reply":"2022-07-23T20:42:15.565277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 10.0 NN Development","metadata":{}},{"cell_type":"code","source":"%%time\n# Release some memory by deleting the original DataFrames...\nimport gc\ndel trn_data, tst_data\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:15.567466Z","iopub.execute_input":"2022-07-23T20:42:15.568007Z","iopub.status.idle":"2022-07-23T20:42:15.714961Z","shell.execute_reply.started":"2022-07-23T20:42:15.567973Z","shell.execute_reply":"2022-07-23T20:42:15.713902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 10.1 Loading Specific Model Libraries...","metadata":{}},{"cell_type":"code","source":"%%time\nimport tensorflow as tf\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau, LearningRateScheduler, EarlyStopping\nfrom tensorflow.keras.layers import Dense, Input, InputLayer, Add, BatchNormalization, Dropout, Concatenate\nfrom tensorflow.keras.utils import plot_model\nfrom sklearn.metrics import log_loss\n\nfrom sklearn.preprocessing import StandardScaler, RobustScaler, MinMaxScaler\nimport random","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:15.716290Z","iopub.execute_input":"2022-07-23T20:42:15.716671Z","iopub.status.idle":"2022-07-23T20:42:20.678670Z","shell.execute_reply.started":"2022-07-23T20:42:15.716631Z","shell.execute_reply":"2022-07-23T20:42:20.677673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 10.2 Amex Metric, Function...","metadata":{}},{"cell_type":"code","source":"%%time\n# From https://www.kaggle.com/code/inversion/amex-competition-metric-python\n\ndef amex_metric(y_true, y_pred, return_components=False) -> float:\n    \"\"\"Amex metric for ndarrays\"\"\"\n    \n    def top_four_percent_captured(df) -> float:\n        \"\"\"Corresponds to the recall for a threshold of 4 %\"\"\"\n        \n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n    \n    \n    def weighted_gini(df) -> float:\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    \n    def normalized_weighted_gini(df) -> float:\n        \"\"\"Corresponds to 2 * AUC - 1\"\"\"\n        \n        df2 = pd.DataFrame({'target': df.target, 'prediction': df.target})\n        df2.sort_values('prediction', ascending=False, inplace=True)\n        return weighted_gini(df) / weighted_gini(df2)\n\n    \n    df = pd.DataFrame({'target': y_true.ravel(), 'prediction': y_pred.ravel()})\n    df.sort_values('prediction', ascending=False, inplace=True)\n    g = normalized_weighted_gini(df)\n    d = top_four_percent_captured(df)\n\n    if return_components: return g, d, 0.5 * (g + d)\n    return 0.5 * (g + d)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:20.680047Z","iopub.execute_input":"2022-07-23T20:42:20.681207Z","iopub.status.idle":"2022-07-23T20:42:20.694948Z","shell.execute_reply.started":"2022-07-23T20:42:20.681170Z","shell.execute_reply":"2022-07-23T20:42:20.692637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 10.3 Defining the NN Model Architecture...","metadata":{}},{"cell_type":"markdown","source":"## 10.3.1 Architecture 01, Simple NN","metadata":{}},{"cell_type":"code","source":"%%time\ndef nn_model():\n    '''\n    '''\n    regularization = 4e-4\n    activation_func = 'swish'\n    inputs = Input(shape = (len(features)))\n    \n    x = Dense(256, \n              #use_bias  = True, \n              kernel_regularizer = tf.keras.regularizers.l2(regularization), \n              activation = activation_func)(inputs)\n    \n    x = BatchNormalization()(x)\n    \n    x = Dense(64, \n              #use_bias  = True, \n              kernel_regularizer = tf.keras.regularizers.l2(regularization), \n              activation = activation_func)(x)\n    \n    x = BatchNormalization()(x)\n    \n    x = Dense(64, \n          #use_bias  = True, \n          kernel_regularizer = tf.keras.regularizers.l2(regularization), \n          activation = activation_func)(x)\n    \n    x = BatchNormalization()(x)\n\n    x = Dense(32, \n              #use_bias  = True, \n              kernel_regularizer = tf.keras.regularizers.l2(regularization), \n              activation = activation_func)(x)\n    \n    x = BatchNormalization()(x)\n\n    x = Dense(1, \n              #use_bias  = True, \n              #kernel_regularizer = tf.keras.regularizers.l2(regularization),\n              activation = 'sigmoid')(x)\n    \n    model = Model(inputs, x)\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:20.697655Z","iopub.execute_input":"2022-07-23T20:42:20.698840Z","iopub.status.idle":"2022-07-23T20:42:20.713800Z","shell.execute_reply.started":"2022-07-23T20:42:20.698810Z","shell.execute_reply":"2022-07-23T20:42:20.712942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 10.3.2 Architecture 02, Concatenated NN","metadata":{}},{"cell_type":"code","source":"%%time\ndef nn_model():\n    regularization = 4e-4\n    activation_func = 'swish'\n    inputs = Input(shape = (len(features)))\n\n    x0 = Dense(256,\n               kernel_regularizer = tf.keras.regularizers.l2(regularization), \n               activation = activation_func)(inputs)\n    x1 = Dense(128,\n               kernel_regularizer = tf.keras.regularizers.l2(regularization),\n               activation = activation_func)(x0)\n    x1 = Dense(64,\n               kernel_regularizer = tf.keras.regularizers.l2(regularization),\n               activation = activation_func)(x1)\n    x1 = Dense(32,\n           kernel_regularizer = tf.keras.regularizers.l2(regularization),\n           activation = activation_func)(x1)\n    \n    x1 = Concatenate()([x1, x0])\n    x1 = Dropout(0.1)(x1)\n    \n    x1 = Dense(16, kernel_regularizer=tf.keras.regularizers.l2(regularization),activation=activation_func,)(x1)\n    \n    x1 = Dense(1, \n              #kernel_regularizer=tf.keras.regularizers.l2(regularization),\n              activation='sigmoid')(x1)\n    \n    model = Model(inputs, x1)\n    \n    return model\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:20.715361Z","iopub.execute_input":"2022-07-23T20:42:20.715814Z","iopub.status.idle":"2022-07-23T20:42:20.728635Z","shell.execute_reply.started":"2022-07-23T20:42:20.715777Z","shell.execute_reply":"2022-07-23T20:42:20.727865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 10.4 Visualizing the Model Structure...","metadata":{}},{"cell_type":"code","source":"%%time\narchitecture = nn_model()\narchitecture.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:20.730244Z","iopub.execute_input":"2022-07-23T20:42:20.730599Z","iopub.status.idle":"2022-07-23T20:42:23.677481Z","shell.execute_reply.started":"2022-07-23T20:42:20.730564Z","shell.execute_reply":"2022-07-23T20:42:23.676673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nplot_model(nn_model(), show_layer_names = False, show_shapes = True, dpi = 60)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:23.678560Z","iopub.execute_input":"2022-07-23T20:42:23.679051Z","iopub.status.idle":"2022-07-23T20:42:24.842024Z","shell.execute_reply.started":"2022-07-23T20:42:23.679015Z","shell.execute_reply":"2022-07-23T20:42:24.841060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 10.5 Defining Model Training Parameters...","metadata":{}},{"cell_type":"code","source":"%%time\n# Defining model parameters...\nBATCH_SIZE         = 2048\nEPOCHS             = 192 \nEPOCHS_COSINEDECAY = 192 \nDIAGRAMS           = True\nUSE_PLATEAU        = False\nINFERENCE          = False\nVERBOSE            = 0 \nTARGET             = 'target'","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:24.843809Z","iopub.execute_input":"2022-07-23T20:42:24.844926Z","iopub.status.idle":"2022-07-23T20:42:24.851383Z","shell.execute_reply.started":"2022-07-23T20:42:24.844875Z","shell.execute_reply":"2022-07-23T20:42:24.850581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 10.6 Defining the Model Training Configuration...","metadata":{}},{"cell_type":"code","source":" %%time\n# Defining model training function...\ndef fit_model(X_train, y_train, X_val, y_val, run = 0):\n    '''\n    '''\n    lr_start = 0.01\n    start_time = datetime.datetime.now()\n    \n    scaler = StandardScaler()\n    X_train = scaler.fit_transform(X_train)\n\n    epochs = EPOCHS    \n    lr = ReduceLROnPlateau(monitor = 'val_loss', factor = 0.7, patience = 4, verbose = VERBOSE)\n    es = EarlyStopping(monitor = 'val_loss',patience = 12, verbose = 1, mode = 'min', restore_best_weights = True)\n    tm = tf.keras.callbacks.TerminateOnNaN()\n    callbacks = [lr, es, tm]\n    \n    # Cosine Learning Rate Decay\n    if USE_PLATEAU == False:\n        epochs = EPOCHS_COSINEDECAY\n        lr_end = 0.0002\n\n        def cosine_decay(epoch):\n            if epochs > 1:\n                w = (1 + math.cos(epoch / (epochs - 1) * math.pi)) / 2\n            else:\n                w = 1\n            return w * lr_start + (1 - w) * lr_end\n        \n        lr = LearningRateScheduler(cosine_decay, verbose = 0)\n        callbacks = [lr, tm]\n    \n    # Model Initialization...\n    model = nn_model()\n    optimizer_func = tf.keras.optimizers.Adam(learning_rate = lr_start)\n    loss_func = tf.keras.losses.BinaryCrossentropy()\n    model.compile(optimizer = optimizer_func, loss = loss_func)\n    \n    \n    X_val = scaler.transform(X_val)\n    validation_data = (X_val, y_val)\n    \n    history = model.fit(X_train, \n                        y_train, \n                        validation_data = validation_data, \n                        epochs          = epochs,\n                        verbose         = VERBOSE,\n                        batch_size      = BATCH_SIZE,\n                        shuffle         = True,\n                        callbacks       = callbacks\n                       )\n    \n    history_list.append(history.history)\n    \n    print(f'Training Loss: {history_list[-1][\"loss\"][-1]:.5f}, Validation Loss: {history_list[-1][\"val_loss\"][-1]:.5f}')\n    callbacks, es, lr, tm, history = None, None, None, None, None\n    \n    \n    y_val_pred = model.predict(X_val, batch_size = BATCH_SIZE, verbose = VERBOSE).ravel()\n    amex_score = amex_metric(y_val.values, y_val_pred, return_components = False)\n    \n    print(f'Fold {run}.{fold} | {str(datetime.datetime.now() - start_time)[-12:-7]}'\n          f'| Amex Score: {amex_score:.5f}')\n    \n    print('')\n    \n    score_list.append(amex_score)\n    \n    tst_data_scaled = scaler.transform(tst_agg_data[features])\n    tst_pred = model.predict(tst_data_scaled)\n    predictions.append(tst_pred)\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:24.853141Z","iopub.execute_input":"2022-07-23T20:42:24.854172Z","iopub.status.idle":"2022-07-23T20:42:24.869084Z","shell.execute_reply.started":"2022-07-23T20:42:24.854130Z","shell.execute_reply":"2022-07-23T20:42:24.868151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 10.7 Creating a Model Training Loop and Cross Validating in 5 Folds... ","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score, roc_curve\nimport math\n\n# Create empty lists to store NN information...\nhistory_list = []\nscore_list   = []\npredictions  = []\n\n# Define kfolds for training purposes...\nkf = KFold(n_splits = 5)\n\nfor fold, (trn_idx, val_idx) in enumerate(kf.split(trn_agg_data)):\n    X_train, X_val = trn_agg_data.iloc[trn_idx][features], trn_agg_data.iloc[val_idx][features]\n    y_train, y_val = trn_agg_data.iloc[trn_idx][TARGET], trn_agg_data.iloc[val_idx][TARGET]\n    \n    fit_model(X_train, y_train, X_val, y_val)\n    \nprint(f'OOF AUC: {np.mean(score_list):.5f}')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T20:42:24.870523Z","iopub.execute_input":"2022-07-23T20:42:24.870991Z","iopub.status.idle":"2022-07-23T21:01:09.464937Z","shell.execute_reply.started":"2022-07-23T20:42:24.870955Z","shell.execute_reply":"2022-07-23T21:01:09.464046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 11.0 Model Prediction and Submissions","metadata":{}},{"cell_type":"code","source":"%%time\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T21:01:09.466373Z","iopub.execute_input":"2022-07-23T21:01:09.466890Z","iopub.status.idle":"2022-07-23T21:01:09.482841Z","shell.execute_reply.started":"2022-07-23T21:01:09.466852Z","shell.execute_reply":"2022-07-23T21:01:09.481867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsub['prediction'] = np.array(predictions).mean(axis = 0)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T21:01:09.484223Z","iopub.execute_input":"2022-07-23T21:01:09.485062Z","iopub.status.idle":"2022-07-23T21:01:09.499159Z","shell.execute_reply.started":"2022-07-23T21:01:09.485022Z","shell.execute_reply":"2022-07-23T21:01:09.498357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsub.to_csv('my_submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T21:01:09.500543Z","iopub.execute_input":"2022-07-23T21:01:09.500902Z","iopub.status.idle":"2022-07-23T21:01:14.717877Z","shell.execute_reply.started":"2022-07-23T21:01:09.500867Z","shell.execute_reply":"2022-07-23T21:01:14.716924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T21:01:14.719177Z","iopub.execute_input":"2022-07-23T21:01:14.719698Z","iopub.status.idle":"2022-07-23T21:01:14.732248Z","shell.execute_reply.started":"2022-07-23T21:01:14.719658Z","shell.execute_reply":"2022-07-23T21:01:14.730753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}}]}