{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## kaggle kernels pull willkoehrsen/introduction-to-manual-feature-engineering","metadata":{}},{"cell_type":"markdown","source":"# Introduction: Manual Feature Engineering\n\nIf you are new to this competition, I highly suggest checking out [this notebook](https://www.kaggle.com/willkoehrsen/start-here-a-gentle-introduction/) to get started.\n\nIn this notebook, we will explore making features by hand for the Home Credit Default Risk competition. In an earlier notebook, we used only the application data in order to build a model. The best model we made from this data achieved a score on the leaderboard around 0.74. In order to better thsi score, we will have to include more information from the other dataframes. Here, we will look at using information from the bureau and bureau_balance data. The definiions of these data files are:\n\n- bureau : information about clinet's previous loans with other financial institutions reported to Home Credit. Each previous loan has its own row.\n- bureau_balance : monthly information about the previous loans. Each month has iths own row.\n\nManual feature engineering can be a tedious process (which is why we use automated feature engineering with featuretools!) and often relies on domain expertise. Since I have limited domain knowledge of loans and what makes a person likely to default, I will instead concentrate of getting as much info as possible into the final training dataframe. The idea is that the model will then pick up on which features are important rather than us having to decide that. Basically, our approach is to make as many features as possible and then give them all to the model to use! Later, we can perform feature reduction using the feature importances from the model or other techniques such as PCA.\n\nThe process of manual feature engineering will involve plenty of Pandas code, a little patience, and a lot of great parctice manipulation data. Even though automated feature engineering tools are starting to be made availavel, feature engineering will still have to be done using plenty of data wrangling for a little while longer.","metadata":{}},{"cell_type":"code","source":"# pandas and numpy for data manipulation\nimport pandas as pd\nimport numpy as np\n\n# matplotlib and saeborn for plotting\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nplt.style.use('fivethirtyeight')\n\n# Suppress warnings from pandas\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:00:51.168959Z","iopub.execute_input":"2022-07-29T05:00:51.169417Z","iopub.status.idle":"2022-07-29T05:00:51.821140Z","shell.execute_reply.started":"2022-07-29T05:00:51.169329Z","shell.execute_reply":"2022-07-29T05:00:51.819962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Example: Counts of a client's previous loans\n\nTo illustrate the general process of manual feature engineering, we will first simply get the count of a client's previous loans at other financial institutions. This requires a number of Pandas operations we will make heavy use of throughout the notebook:\n\n- groupby : group a dataframe by a column. In this case we will group by the unique client, the SK_ID_CURR column\n- agg : perform a calculation on the grouped data such as taking  the mean of columns. We can either call the function directly (grouped_df.mean()) or use the agg function together with a list of transforms (grouped_df.agg([mean, max, min, sum]))\n- merge : match the aggregated statistics to the appropriate client. We need to merge the original training data with the calculated stats on the SK_ID_CURR column which will insert NaN in any cell for which client does not have the corresponding statistic\n\nWe also use the (rename) function quite a bit specifying the columns to be renamed as a dictionary. This is useful in order to keep track of the new variables we create.\n\nThis might seem like a lot, which is why we'll eventually write a function to do this process for us. Let's take a look at implementing this by hand first.","metadata":{}},{"cell_type":"code","source":"# Read in bureau\nbureau = pd.read_csv('../input/home-credit-default-risk/bureau.csv')\nbureau.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:00:51.823324Z","iopub.execute_input":"2022-07-29T05:00:51.824038Z","iopub.status.idle":"2022-07-29T05:00:56.894667Z","shell.execute_reply.started":"2022-07-29T05:00:51.823997Z","shell.execute_reply":"2022-07-29T05:00:56.893525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Groupby the client id (SK_ID_CURR), count the number of previous loans, and rename the column\nprevious_loan_counts = bureau.groupby('SK_ID_CURR', as_index=False)['SK_ID_BUREAU'].count().rename(columns = {'SK_ID_BUREAU': 'previous_loan_counts'})\nprevious_loan_counts.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:00:56.896043Z","iopub.execute_input":"2022-07-29T05:00:56.896393Z","iopub.status.idle":"2022-07-29T05:00:57.084282Z","shell.execute_reply.started":"2022-07-29T05:00:56.896349Z","shell.execute_reply":"2022-07-29T05:00:57.083198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Join to the training dataframe\ntrain = pd.read_csv('../input/home-credit-default-risk/application_train.csv')\ntrain = train.merge(previous_loan_counts, on = 'SK_ID_CURR', how = 'left')\n\n# Fill the missing values with 0 \ntrain['previous_loan_counts'] = train['previous_loan_counts'].fillna(0)\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:00:57.086982Z","iopub.execute_input":"2022-07-29T05:00:57.087303Z","iopub.status.idle":"2022-07-29T05:01:05.361193Z","shell.execute_reply.started":"2022-07-29T05:00:57.087274Z","shell.execute_reply":"2022-07-29T05:01:05.359936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Scroll all the way to the right to see the new column.","metadata":{}},{"cell_type":"markdown","source":"## Assessing Usefulness of New Variable with r value\n\nTo determine if the new variable is useful, we can calculate the Pearsom Correlation Coefficient (r-value) between this variable and the target. This measures the strength of a linear relationship between two variables and ranges from -1 (perfectly negatively linear) to +1 (perfectly positively linear). The r-value is not best measure of the \"usefulness\" of a new variable, but it can give a first approximation of whether a variable will be helpful to a machin learning model. The larger the r-value of a variable with respect to the target, the more a change in this variable is likely to affect the value of the target. Therefore, we look for the variables with the greatest absolute value r-value relative to the target.\n\nWe can also visually inspect a relationship with the target using the Kernel Density Estimate (KDE) plot.","metadata":{}},{"cell_type":"markdown","source":"### Kernel Density Estimate Plots\n\nThe kernel density estimate plot shows the distribution of a single variable (think of it as a smoothed histogram). To see the different in distributions dependent on the value of a categorical variable, we cna color the distributions differently according to the category. For example, we can show the kernel density estimate of the previous_loan_counts colored by whether the Target = 1 or 0. Thr resulting KDE will show any significant differences in the distribution of the vairable between people who did not repay their loan(TARGET == 1) and the people who did(TARGET == 0). This can serve as an indicator of whether a variable will be 'relevant' to a machine learning model.\n\nWe will put this plotting functionality in a function to re-use for any variable.","metadata":{}},{"cell_type":"code","source":"# Plots the disribution of a variable colored by value of the target\ndef kde_target(var_name, df):\n    \n    # Calculate the correlation coefficient between the new variable and the target\n    corr = df['TARGET'].corr(df[var_name])\n    \n    # Calculate medians for repaid vs not repaid\n    avg_repaid = df.loc[df['TARGET'] == 0, var_name].median()\n    avg_not_repaid = df.loc[df['TARGET'] == 1, var_name].median()\n    \n    plt.figure(figsize = (12, 6))\n    \n    # Plot the distribution for target == 0 and target == 1\n    sns.kdeplot(df.loc[df['TARGET'] == 0, var_name], label = 'TARGET == 0')\n    sns.kdeplot(df.loc[df['TARGET'] == 1, var_name], label = 'TARGET == 1')\n    \n    # label the plot\n    plt.xlabel(var_name); plt.ylabel('Density'); plt.title('%s Distribution' % var_name)\n    plt.legend();\n    \n    # print out the correlation\n    print('The correlation between %s and the TARGET is %0.4f' % (var_name, corr))\n    # Print out average values\n    print('Median value for loan that was not repaid = %0.4f' % avg_not_repaid)\n    print('Median value for loan that was repaid =     %0.4f' % avg_repaid)\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:05.362817Z","iopub.execute_input":"2022-07-29T05:01:05.363256Z","iopub.status.idle":"2022-07-29T05:01:05.374813Z","shell.execute_reply.started":"2022-07-29T05:01:05.363212Z","shell.execute_reply":"2022-07-29T05:01:05.373952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can test this function using the EXT_SOURCE_3 variable which we [found to be one of the most important variables](https://www.kaggle.com/willkoehrsen/start-here-a-gentle-introduction) according to a Random Forest and Gradient Boosting Machine.","metadata":{}},{"cell_type":"code","source":"kde_target('EXT_SOURCE_3', train)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:05.376072Z","iopub.execute_input":"2022-07-29T05:01:05.377002Z","iopub.status.idle":"2022-07-29T05:01:06.701717Z","shell.execute_reply.started":"2022-07-29T05:01:05.376967Z","shell.execute_reply":"2022-07-29T05:01:06.700502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now for the new variable we just made, the number of previous loans at other insitutions.","metadata":{}},{"cell_type":"code","source":"kde_target('previous_loan_counts', train)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:06.703607Z","iopub.execute_input":"2022-07-29T05:01:06.703980Z","iopub.status.idle":"2022-07-29T05:01:08.566013Z","shell.execute_reply.started":"2022-07-29T05:01:06.703946Z","shell.execute_reply":"2022-07-29T05:01:08.564413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this it's difficult to tell if this variable will be important. The correlation coefficient is extremely weak and there is almost no noticeable difference in the distiribution.\n\nLet's move on to make a few more variables from the bureau dataframe. We will take the mean, min, and max of every numeric column in the bureau dataframe.","metadata":{}},{"cell_type":"markdown","source":"## Aggregating Numeric Columns\n\nTo account for the numeric information in the bureau dataframe, we can compute statistics for all the numeric columns. To do so, we groupby the client id, agg the grouped dataframe, and merge the result back into the training data. The agg function will only calculate the values for the numeric columns where the operation is considered valid. We will stick to using 'mean', 'max', 'min', 'sum' but any function can be passed in here. We can even write our own function and use it an agg call.","metadata":{}},{"cell_type":"code","source":"# Groupby the client id, calculate aggregation statistics\nbureau_agg = bureau.drop(columns = ['SK_ID_BUREAU']).groupby('SK_ID_CURR', as_index = False).\\\n            agg(['count', 'mean', 'max', 'min', 'sum']).reset_index()\nbureau_agg.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:08.568384Z","iopub.execute_input":"2022-07-29T05:01:08.568859Z","iopub.status.idle":"2022-07-29T05:01:13.561006Z","shell.execute_reply.started":"2022-07-29T05:01:08.568815Z","shell.execute_reply":"2022-07-29T05:01:13.559505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We need to create new names for each of these columns. The following code makes nwe names by appending the stat to the name. Here we have to deal with the fact that the dataframe has a multi-level index. I find these confusing and hard to work with, so I try to reduce to a single level index as quickly as possible.","metadata":{}},{"cell_type":"code","source":"# List of column names\ncolumns = ['SK_ID_CURR']\n\n# Iterate through the variables names\nfor var in bureau_agg.columns.levels[0]:\n    # Skip the id name\n    if var != 'SK_ID_CURR':\n        \n        # Iterate through the stat names\n        for stat in bureau_agg.columns.levels[1][:-1]:\n            # Make a new column name for the variable and stat\n            columns.append('bureau_%s_%s' % (var, stat))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:13.563147Z","iopub.execute_input":"2022-07-29T05:01:13.564340Z","iopub.status.idle":"2022-07-29T05:01:13.573969Z","shell.execute_reply.started":"2022-07-29T05:01:13.564292Z","shell.execute_reply":"2022-07-29T05:01:13.572821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assign the list of columns names as the dataframe column names\nbureau_agg.columns = columns\nbureau_agg.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:13.580580Z","iopub.execute_input":"2022-07-29T05:01:13.581617Z","iopub.status.idle":"2022-07-29T05:01:13.611774Z","shell.execute_reply.started":"2022-07-29T05:01:13.581578Z","shell.execute_reply":"2022-07-29T05:01:13.610497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we simply merge with the training data as we did before.","metadata":{}},{"cell_type":"code","source":"# Merge with the training data\ntrain = train.merge(bureau_agg, on = 'SK_ID_CURR', how= 'left')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:13.613175Z","iopub.execute_input":"2022-07-29T05:01:13.614149Z","iopub.status.idle":"2022-07-29T05:01:15.753414Z","shell.execute_reply.started":"2022-07-29T05:01:13.614103Z","shell.execute_reply":"2022-07-29T05:01:15.752205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Correlations of Aggregated Values with Target\n\nWe can calculate the correlation of all new values with the target. Again, we can use these as an approximation of the variables which may be important for modeling.","metadata":{}},{"cell_type":"code","source":"# List of new correlations\nnew_corrs = []\n\n# Iterate through the columns\nfor col in columns:\n    # Calculate correlation with the target\n    corr = train['TARGET'].corr(train[col])\n    \n    # Append the list as a tuple\n    new_corrs.append((col, corr))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:15.755414Z","iopub.execute_input":"2022-07-29T05:01:15.755910Z","iopub.status.idle":"2022-07-29T05:01:16.139409Z","shell.execute_reply.started":"2022-07-29T05:01:15.755863Z","shell.execute_reply":"2022-07-29T05:01:16.138310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the code below, we sort the correlations by the magnitude (absolute value) using the sorted Python function. We also make use of an anonymous lambda function, another important Python operation that is good to know.","metadata":{}},{"cell_type":"code","source":"# Sort the correlation by the absolute value.\n# Make sure tod reverse to put the largest values at the front of list\nnew_corrs = sorted(new_corrs, key = lambda x: abs(x[1]), reverse = True)\nnew_corrs[:15]","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:16.140994Z","iopub.execute_input":"2022-07-29T05:01:16.141339Z","iopub.status.idle":"2022-07-29T05:01:16.151068Z","shell.execute_reply.started":"2022-07-29T05:01:16.141301Z","shell.execute_reply":"2022-07-29T05:01:16.149998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"None of the new variables have a significant correlation with the TARGET. We can look at the KDE plot of  the highest correlated variable, bureau_DAYS_CREDIT_mean, with the target in terms of absolute magnitude correlation.","metadata":{}},{"cell_type":"code","source":"kde_target('bureau_DAYS_CREDIT_mean', train)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:16.152420Z","iopub.execute_input":"2022-07-29T05:01:16.152789Z","iopub.status.idle":"2022-07-29T05:01:17.585416Z","shell.execute_reply.started":"2022-07-29T05:01:16.152756Z","shell.execute_reply":"2022-07-29T05:01:17.583921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The definition of this column is: \"How many days befor current application did lient apply for Credit Bureau credit\". My interpretation is this is the number of days that the previous loan was applied for before the application for a loan at Home Credit. Therefore, a larger negative number indicates the loan was futher before the current loan application. We see an extremely  weak positive relationship between the average of this variable and the target meaning that clients who applied for loans futher in the past potentially are more likely to repay loans at Home Credit. With a carrelation this weak though, it is just as likely to be noise as a signal.\n\n**The Multiple Comparisions Problem**\n\nWhen we have lots of variables, we expect some of them to be correlated just by pure chance, a [problem known as multiple comparisons](https://towardsdatascience.com/the-multiple-comparisons-problem-e5573e8b9578). We can make hundreds of features, and some will turn out to be corelated with the target simply because of random noise in the data. Then, when our model trains, it may overfit to these variables because it thinks they have a relationshop with the target in the training set, but this does not necessarily generalize to the test set. There are many considerations that we have to take into account when making features!","metadata":{}},{"cell_type":"markdown","source":"## Function for Numeric Aggregation\n\nLet's encapsulate all of the previous work into a function. This will allow us to compute aggregate stats for numeric columns across any dataframe. We will re-use this function when we want  to apply the same operations for other dataframe.","metadata":{}},{"cell_type":"code","source":"'''Aggregates the numeric values in a dataframe. This can\nbe used to create for each instance of the grouping variable.\n\nParameters\n------------\n    df (dataframe):\n        the dataframe to calculate the statistics on\n    group_var (string):\n        the variable by witch to group df\n    df_name (string):\n        the variable used to rename the columns\n        \nReturn\n-----------\n    agg (dataframe):\n        a dataframe with the statistics aggregated for\n        all numeric columns. Each instance of the grouping variable will have\n        the statistics (mean, min, max, sum; currently supported) calculated.\n        The columns are also renamed to keep track of features created.\n        \n'''","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:17.587510Z","iopub.execute_input":"2022-07-29T05:01:17.587965Z","iopub.status.idle":"2022-07-29T05:01:17.597378Z","shell.execute_reply.started":"2022-07-29T05:01:17.587917Z","shell.execute_reply":"2022-07-29T05:01:17.596144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def agg_numeric(df, group_var, df_name):\n    # Remove id variable other than grouping variable\n    for col in df:\n        if col != group_var and 'SK_ID' in col:\n            df = df.drop(columns = col)\n            \n    group_ids = df[group_var]\n    numeric_df = df.select_dtypes('number')\n    numeric_df[group_var] = group_ids\n    \n    # Group by the specified variable and calculate the statistics\n    agg = numeric_df.groupby(group_var).agg(['count', 'mean', 'max', 'min', 'sum']).reset_index()\n    \n    # Need to create new column names\n    columns = [group_var]\n    \n    # Iterate through the vairables names\n    for var in agg.columns.levels[0]:\n        # Skip the grouping variable\n        if var != group_var:\n            # Iterate through the stat names\n            for stat in agg.columns.levels[1][:-1]:\n                # Make a new column names for the variable and stat\n                columns.append('%s_%s_%s' % (df_name, var, stat))\n                \n    agg.columns = columns\n    return agg","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:17.599014Z","iopub.execute_input":"2022-07-29T05:01:17.600248Z","iopub.status.idle":"2022-07-29T05:01:17.610764Z","shell.execute_reply.started":"2022-07-29T05:01:17.600204Z","shell.execute_reply":"2022-07-29T05:01:17.609747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bureau_agg_new = agg_numeric(bureau.drop(columns = ['SK_ID_BUREAU']), group_var = 'SK_ID_CURR', df_name = 'bureau')\nbureau_agg_new.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:17.612379Z","iopub.execute_input":"2022-07-29T05:01:17.613400Z","iopub.status.idle":"2022-07-29T05:01:20.362814Z","shell.execute_reply.started":"2022-07-29T05:01:17.613354Z","shell.execute_reply":"2022-07-29T05:01:20.361666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To make sure the function worked as intended, we should compare with the aggregated dataframe we constructed by hand.","metadata":{}},{"cell_type":"code","source":"bureau_agg.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:20.364387Z","iopub.execute_input":"2022-07-29T05:01:20.365203Z","iopub.status.idle":"2022-07-29T05:01:20.394315Z","shell.execute_reply.started":"2022-07-29T05:01:20.365156Z","shell.execute_reply":"2022-07-29T05:01:20.393230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we go through and inspect the value, we do find that they are equivalent. We will be able to reuse this function for calculating numeric stats for other dataframes. Using functions allows for consistent results and decreases the amount of work we have to do in the future!","metadata":{}},{"cell_type":"markdown","source":"### Correlation Function\n\nBefore we move on, we can also make the code to calculate coreelations with the target into a function.","metadata":{}},{"cell_type":"code","source":"# Function to calculate correlations with the target for a dataframe\ndef target_corrs(df):\n    # List of correlations\n    corrs = []\n    \n    # Iterate through the colums\n    for col in df.columns:\n        print(col)\n        # Skip the target column\n        if col != 'TARGET':\n            # calculate correlation with the target\n            corr = df['TARGET'].corr(df[col])\n            \n            # Append t he list as a tuple\n            corrs.append((col, corr))\n    \n    # Sort by absolute magnitude of correlatins\n    corrs = sorted(corrs, key = lambda x: abs(x[1]), reverse = True)\n    \n    return corrs","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:20.395978Z","iopub.execute_input":"2022-07-29T05:01:20.396304Z","iopub.status.idle":"2022-07-29T05:01:20.403185Z","shell.execute_reply.started":"2022-07-29T05:01:20.396274Z","shell.execute_reply":"2022-07-29T05:01:20.402336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Categorical Variables\n\nNow we move from the numeric columns  to the categorical columns. These are discreate string variables, so we cannot just calculate statistics such as mean and max which only work with numeric variables. Instead, we will rely on calculating value counts of each category within each categorical vairable. As an example, if we have the following dataframe:\n\n![image.png](attachment:00a199cd-07fe-4668-87e7-e2aa6b641820.png)\n\nwe will use this information counting the number of loans in each category for each client.\n\n![image.png](attachment:b72a8cd4-8a98-401f-8a8d-db4ae5a887cd.png)\n\nThen we can normalize these value count by the total number of occurences of that categorical variable for that observation (meaning that the normalized counts must sum to 1.0 for each observation).\n\n![image.png](attachment:877acc94-b90a-44ea-824e-df9e8b42f774.png)\n\nHopefuuly, encoding the categorical variables this way will allow us to capture  the information they contain. If anyone has a better idea for this process, please let me know in the comments! We will now go through this process step-by-step. At the end, we will wrap up all the code into one function to be re-used for mnay dataframe.","metadata":{},"attachments":{"00a199cd-07fe-4668-87e7-e2aa6b641820.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAKoAAAFwCAYAAAAopOB/AAAcBUlEQVR4nO3df3hU9aHn8fctnZE6qc2AzfgjUZNomGwhVEgJzIUGRAQNaDEIen2gXsVe4LHio/CoyXqF9RK7C+4CdtEtkfVifRCKcAWjkR9KDIaCCdaAm4APiRQUJ1Uml3I0TIzZP/ITmAwzIT/OFz6vv2DOmXO+c/jMOWfmzPnwD42NjY2I2NwPensAIpFQUMUICqoYQUEVI/yw5Q+fHT5Mnz59enMsF6zvvvuOH/7wh+ee8SLX0NDAdddeG3Ja69ZzOBy4Y2N7bFAXk6+PH6d/v369PQzbC9TWdjhNh34xgoIqRlBQxQgKqhhBQRUjKKhiBAVVjKCgihE6dbmkdn8h69/ZzeffQFzqBCbdmkG8C6CawmeKifvtDIa0XDv4+z5e/8O7xN41l7HXhF5e9VsLKY6by4z0WKrfWsjqsuYJl15NypAhjB02hDjXuUZlcXR3IZu37KeGS7l66ASmjB9EbJ8OxgVQu5fV/wG33zeEWKopfGY1u1um9UskI/0mRmXE42qZ9/nNVLcM7YoUhoy4lbEDu+4iSfvt0PPqsSxwuRy9sO5zizqoVmk+ObsSmf9ADvF9gwQqNrF0kZ8H824nkXqsj/1YDc0z1x3k9WfXwH0LOwwpQP03+/DXtf350hFLmTPCDVjU7C8m//FiRv3rXEbFdbSEGoqfX0Dx1TOZ8dsJuB0Wnxe9zoJnPmfugglnj6tFg4X/69ZRYH18KSP/5xyG9wfqa9i7YQWL/vYgeRMTm+b9OIVJ/347AwCsKgpX5ZBft5iZ6ed8F0Wk/XbocV//mRUvw4OPjcKO1yejPvR//lkhY8ePJTHWgaOvi7gbp/DIfYNw1585Zw3FK1dQM+kRslOifJf2jcHlcuFyxZGYkU3OvETeXrmdmg5mr9+/nTWxs5l/5xDiY5uel3LbbObeuJ3tpVZUq3a6XE3rjk1k1LQpJO3Y17oXBScxLdPjBpF99wQOfHSA6NbQWRY1H21n0xubKNx5kNrT3nT11OwvpvCNTWx6by81rQOqYe97B6mt3cf2lmmh3gjWQba/s5eaY3t5+41N7P2ynqO7NlH8Wft/1KPsfWM3R+vBOri9aZ7dhWx6YxPb959x6bOhloM7m6d9VNMl2yfqoMb2H0LZzmKq/97yiIO4lERi22exoWkPt/fGHGaP6HA3GLnrhjP2+/1Ud3ApuLrydUb9YhBnvh0S7zjPvV2dRSDc9GB9+Oldpp6D65eQX+1mSOZwkurfZ9mK4uY3bj3VG55m6Z4gSSPGMvyn1eQv2tT85qqluuBlVhZ+w4DM4Xi/LSbnfxdz1mZ0xpGUGEfMT+JISk4kzuXA82OLNSWVtEb14G7yrUvxOKC+Zi+FL+dTzEB+OcILe5ax+L2W3UgNxSuW8X59EsMzhxBz4AWWbKjmrP1YlKIOatyY+cz5WQ2Fz85hTs5i8t9q/w4G+Ia9a5eyrGgQwwd3QUgBcOK8NEjwzEM3ALX4D2cQ3yW/+fiGowf3sW//Pvbt383rK9fg/NVwElun11C1v3n6R9vJf7mMWzMH0TUH/jC+fJ/1B0Yy884hxMfGkTLm19zZfw3b99cDDhLvWMjC+8eSEucibuAoRrkPcLTllOazQdxxd0bT8267nduPHaDqzF2cI5bElHguvTSeAQMHEf9jcAwcytiS/VTWA9Szr3Q7o9K8rTuDQOoE7smIJzYuhbHT7yT+9e3sq286ur19zQxmjkkhLjaejLsfIOOT7ew7z91qJz5MOYjPyGZ2RjbU1XBw1yaWLqrmwYXZJDoA9uG/Mo/8vH0se7mYxN+O4vzjGoTvO5rmIjZuNwdr4fxXdJKa6mqqXVCzfzU16UvJGdl+oQGOHqomiEV18W7c9y6K/rSmM07UsnfAoHYvz0HCNYkU1lpALHxzlD8XF7O/rJqahlMEjrrJbp3X2W5BbtxXBrDq4NzvrhSGjFtJ8f+bwqD/Usn+/RMY/k9tr3XQdfFtszoSiE8spPYEWLVHqSlazcLytsmBw5eSHdE6OxZ1UOubPho2vbP6xpEyZib3+Wey96/ZJCYDZHDrmBRi+ydyT/nTLN0Qz8I7E886LEfl6wP82T+AKf1DTXSQkJjB6opqslMST5tSU/QCha4ZzEiPJXYwcOYe+e8Bapyedg/EMWT87YzqDwyGnD/spXp8fLs96gB+ecftpADW1RbzSw+QfeOQ7t+jAnx3+l+DDeB0Ahyl8Nn18JsHmTk+FkefWoqfW9klq0wcPIoVxZVYffazL2M497Sb5v9785ukaTQEv3ES0/xz5gF3zGZ+ZlcdTZtEeei32PvyXNZ81G4/bu2jrHQo8WeNy0HKlEe4tXolmw52/gyl/ut9vL5yDe67x5LYwTyxGVMYvnslaz5qO/uy/rqdNX9yMHCAC4gj6cYAb79/sO3EvqGG4s2biLtxQOhPudfdyj2pm1i/M/SJsSv9drJr1lB4Hq8tYokpZO98m+KW08C6gxS/U8/AxFjgG6yD8SRdF4ujD1B7gAMVnVhHHyfObyxOtn/suuFM+KyQFe8cZGzG6Vt/9+59rR9u6w8Ws/0HA0mMhdjrh1C/cy9H233zU/hKMUfPczNFuUd1kTFjLkeXz2fOnzx4fnQKvz+GsQ/OJePHoeaPY9T0W1n231ZQHPbrpdPtXjqTKUsBYolPH86EXy0m2xtmv+VIJPvJ+9j0hxymvBTLIGcN/v6/5J4nf82Q5nElTszlnvXLmD+zBs+1bgL+AIm35TJnZEdfxjgYdNuvKXx8E/vSZzAoxGv75V3DyfnT+4zKHdsFpzdN2l57k4xH8pk/chBT5n3O0t89wrtuN6cC4P2n2WRfAZDCqMe2k/foQtxu4KcDSRnQiRXHDuSm5CXkPbqX2x/OYcJ1AHEMGlnPCyUj+fUVp8+ecXUNry9YSE2fUwQavMyYk920Da4Yy8zxq3lhfg64L2kda/x5niH9Q8t9/Z9/8UV0v/Cvs7AanJF/Qfx1MYv/ZVnbF+rN7slbT3ZKZIuo3bmYmUvPWgJ567NpXURDPdb3DjocVkM9Vl0QZ8vpSw/oyl/4n3bq1V5DPVZd139hX/NeHv/umMP8dm/o2p2LWcmDzB8ZS71Vj6ODdYabFkqgtparr7oq5LTO38jT1xXduVn/UcxfP6rTqwOIHTmf9SPPMVMfB65wt371cdj26kskHK4OtnofBx1N6hSrmn0f7WXz6/FMWtLxDixcEKMJ6bnojjMJ7XsHxA5i5n9PIa7v6ZNcqVOY0jMfIVspqBLaj+MZNDD0JEf/xA4/2HYX/XpKjKCgihEUVDGCgipGUFDFCAqqGKH166n6+nq+Pn68N8dywfruu++0bSPQ0BDyd5xAu6D+5LKf4Hbb8SYE8wUCtdq2EQgEVJImhlNQxQgKqhhBQRUjKKhiBAVVjKCgihEUVDFC54N6opKNy3JZtadnukIuBlWbcyn4tLdHYU+d+IV/EP/OP/LiNvD8qJxAb5V6XYCCVjlWh0UbF7dOBNXCj495C7wceW0jG7t+TBe3UwHKt+3g0AknnqGj8V3b7t6kQCVFOys4Xh9Dwggf6Vc2TbMObKXC5cNzZAelx5wkjxlHmvMIJe+W4sdD+k0+EloXY+EvK6Hk8Emcl6fi+0cvbgP+H7xOHPrdpI309vCtXRePD97YSiDJh+9nTsqfX8LWY80TaopYsmQHwRQfo4e5qcrPY92BIABBfylrXy2gsl86vusDbFywkGfXlNMvw0cqJeTmlzQXbwSpfC2PF6vcpI/xkfyfW8hdVtJDRW/nRx+mbCY9eyqZSR48A8YxOctJ6acBIEh54St4pt3PuAEe3PHpTJ05moo1Rfibn9dvWCbjBnjwpI1j9FVB0rOy8MZ58E6ayMRtRzkCcKyItf4s5t2VToLbg3fSLO6PWcuOqt57vZHSXag2c0m7w7Cz9c8WgWPJXHNbu8KzKxNILSsnAHgA+rQvQ4vBGeqW+hMBSj/ZQt5TBa0PBf1V/CKzq0bffRRUk5zxQeuU+/SuvohkPsCie71dNaIeo0O/Edx4hwZ5891Kgs2P+N9/k6KbvcSHfd4ZEr1MLSqhrYTbT8lr6yg34CRVe1RDeG6exbT85TwyD9yXBAhclkXO3DScEHn1uDONaU8c4bnH57DR7YZAgJjx83jM3Y0D7yKtJWn6FXr36dJtG7Swvnfh6nvuWcMvpqlsLepTh24Ubjtpj2oaZ5TldB0uxqwvGHWOKkZQUMUICqoYQUEVIyioYgQFVYygoIoRFFQxgoIqRmi9MlVXVxe2pEo6T9s2MnV1Hd/X1BrUvn376lp/N9HvKCKjNj8xnoIqRlBQxQgKqhhBQRUjKKhiBAVVjKCgihHU5mcjavPrmNr8bERtfh1Tm5/dqM0vJLX52Yza/ELThymbUZtfaCqgsBm1+YWmoJpEbX5ib2rz0x7VEGrzU5tft1ObX2TU5nchUZufiH0pqGIEBVWMoKCKERRUMYKCKkZQUMUICqoYQSVpPUDbNjIqSetlujwdGZWkifEUVDGCgipGUFDFCAqqGEFBFSMoqGIEBVWMoJI0G1FJWsdUkmYjKknrmErS7EYlaSGpJM1mVJIWmj5M2YxK0kLTff02o5K00BRUk6gkTexNJWnaoxpCJWkqSet2KkmLjErSLiQqSROxLwVVjKCgihEUVDGCgipGUFDFCAqqGEFBFSOoJK0HaNtGRiVpvUyXpyOjkjQxnoIqRlBQxQgKqhhBQRUjKKhiBAVVjKCgihFUknYx+LSA3M1NLRPBslXMmFfQVEhBEMsKhnumbagk7WLwvUV5cyCdP5/MoqucJAB8tYvlL8GsxzOx+42oKkmzGetYKSW7jnDS0Y/UkZl43QB+SrcFSPAep3SPRfL4cXhdHc3bvJzDJewo88Plgxl9ebsV1Pmp/AT6XQYlb33Il1/AmxuOkzpiMulX9vCLjYJK0mwkeGAdeflVuIeNxpdykjcfX0LJVwABqja/xKpCC8/1yXgcZ8+75amWecHa8yJzVx/FM8xHWmwFa1/9sG0l3/opLfMTdHhITvIQc5mH5OuT8dj8H1S3S9uGn6I1FYz+l6eb92xZPPxvyRxt6UapSmPysnF4m+fduuZLsh59mPTLmuad9c9HmP9BFb47XJRsDjB97izSLwfis5hmHWLj4TNW53ST5L2GmA8gNS3tQjz0S/cIEChLJa3d4dcZ5yWp9W+XnDFvOVsW59Jad/atn6r00UCQwF+SSWt3uHf1uwLODKphFFTbcIL7VBTzZ/LAM9M5u+6sEjhjOQ3RLNee9D2qbcTjvbmIN99vaTwNUv5/Z/JKeaivj5Lw3lVESVlb65R/5zrWlQdap7Utx6J099bQq+zjwHHyJCe78FV0F+1RbcNJ2tQcjixbyJx33LhPBeDGh8lJC9UO1TLvfOZscOOmqTRt3tymM820qfM49LuW5cTgmzodX6iyXvdgbrkhj4UPlTL50afJSgoxj02oJK0HRL1t6yysH7hwRdJgFrSwCD2vHYvQwlFJmmn6RlGEFqY0zbQitHB0jipGUFDFCAqqGEFBFSMoqGIEBVWMoKCKERRUMYKCKkZQm18P0LaNjNr8epl+RxEZtfmJ8RRUMYKCKkZQUMUICqoYQUEVIyioYgQFVYzQqaBah0tYtyyX3Kee5cVtVVjnfor0pgugzS/6oH5VxPLnD5F079MseuphfF+9wJJt/nM/T3rPmW1+j41ua/P7/S5MKA6N+i5Uq+oQ7uzJpF/uBJyk3TyRgpcqCdzssX1/kQnU5hda1HtU17D7mTWibYsEjx7i+A0JCmkXUJtfx87vvv6aIpavCDJtiY0rNoyhNr9wOh/UukrWLdtB8hNPNm8sOT9q8wunc0Gtq2Lj4ldgxtNMTjKlMMbu1OYXTie+nvJTtOIlTk7KYeoAhbTrqM0vnKj3qIGdr7DkvXJ4727WtT46ncWbp4Z4d0vk1OYXjtr8eoDa/CKjNj/TqM3vLLrWL0ZQUMUICqoYQUEVIyioYgQFVYygoIoRFFQxgkrSeoC2bWRUktbLdHk6MipJE+MpqGIEBVWMoKCKERRUMYKCKkZQUMUICqoY4TxL0nJZvqFSJWl2d1GWpFVtZMHzh0i652kWLchhYt+NLHgj1C2OYhsXY0la8LLBzH4inqTmkrSkYT76vXSEwB1Jtq+FMYFK0kKLeo/qvDypOaRN/H8p4bhXJWldQSVpHev07dKVr01i/quQMOZhnnzIxs0FxlBJWjidDqr37s1svhsCZa+Ql7eVeQvG4enKkV10VJIWTvRBrbMIOlw4+zT91T30Fia+s4rKr8bhuTz8UyUclaSFE/U5qn/nEnL/VEnrlxo1lXz4WTxuVU+eJ5WkhRP1HtVz8yym5S9nzkPgcQcJBBKY/K+zCNnlJVFQSVo451GSFsSqc+Lq200ju4CoJC0y3VSSppB2G5WknUXX+sUICqoYQUEVIyioYgQFVYygoIoRFFQxgoIqRlBQxQhq8+sB2raRUZtfL1ObX2TU5ifGU1DFCAqqGEFBFSMoqGIEBVWMoKCKERRUMUKn7pmyDpdQ8B8FfPwVeG6czgN3eiO/x0d6SBUFT1WQ+kwWNr65NGJq87tgBbH+YmFGqeS5qc3PZjps6KvzU76nlENfBYm51odvqKftKBaopGhnBcfrY0gYmUl6XNsN0sFAOVvfO8RJh4f0m3wkGHroU5ufjXTY5hesYt1/XUJJMBnfeB+eT19sO4qdKOXFJTvgZ6O5ZaSbiv/xLFuPtSzxAzZuC5A80kfqj8pZvngr/g7WbXuNzY4fDzRGo2LNxMaJEyc2zn5uS+NfT0X11ItOZNv2y8YtTy9o3PJF2yOn/BWNh/7WvHFPtdvI/h2Neb/b0Xi8sbGx8eDaxun/58PGky3Tvmv5Q0Xj2omrGyvaLX/HorzGHX/r7KvofuG2k9r8bCN8m591eBc7dn5MSZUfvg1w5KfTmibckMnDRcuZ+9BG0ob6GDxyNJk3tBzf2zcAOqFPD7yMbqI2P9sI0+Z3uIAFG2D2b2aR5XbCV0U8+1LLRA/pMxeR3xAkcKycrfmzeeXOfKan9dS4e4ba/GwjTJtfnUVlfDJJ7qbPBoFPK6homevTAla954c+Ttzx6YwYloD/xIX333+ozc82wrX5ZfJk4ULm5LpxA56fe0lteVZCKp4NC5mzzY2bIMErJzLvN24w92NTSGrz6wFd1uYXruUvmgZAm1Kbn2k6avML1/IXTQOggXStX4ygoIoRFFQxgoIqRlBQxQgKqhhBQRUjKKhiBJWk9QBt28ioJK2XqSQtMipJE+MpqGIEBVWMoKCKERRUMYKCKkZQUMUICqoYQSVpFyyVpKkkzQgqSVNJWjdSSVpoKkmzEZWkhRFJQVUoKkmLnErSIqOSNCOoJC0claTZhkrSwlFJmm2oJC0claTZhkrSwlFJWg9QSVpkVJJmGpWknUXX+sUICqoYQUEVIyioYgQFVYygoIoRFFQxgoIqRlBJWg/Qto2MStJ6mUrSIqOSNDGegipGUFDFCAqqGEFBFSMoqGIEBVWMoKCKEc4vqFUFPPvUKkoDXTQa6R6fFpC7ualZJVi2ihnzCjjS9Dcsy4x2qvMIqp+tr+3g+Ak/VkPXDUi6wfcW5c2BdP58MoseG00CwFe7WP77XZiwn+n0zX3+d16kdMQDTNyzsSvHc9ELXZLmp3RbgATvcUr3WCSPH4fXFaZQjaZq0B1lfrh8MKPbF4PU+an8BPpdBiVvfciXX8CbG46TOmIy6VdiW53bo9YUsWqXj/vHqMSnK3VYkkaAqs0vsarQwnN9Mh7H2fNueaplXrD2vMjc1UfxDPORFlvB2lc/bFvJt35Ky/wEHR6SkzzEXOYh+fpkPDa/hbUTQQ1Q8scteGeoa6pr+SlaU8HomVNJj3fjGZDFw/82kStaOmmq0pg8cxzpaUm4nX6K1nxJ1ty2eWf9cwxrP6gC/JRsDjB99lTS4z0kpGUxbWLC2atzuknyXkNMzDWkpqWRYPOmm6gP/daetRQkP8CiC6HG2FbCl6SdXngWIFBWzpbFuRS0PPStn6r00UCQwF+SSWt3uHf1uwIOd+PQe0CUQa2i4PclBBKOkFva9Ejg8BH4X7mcvG8RWTd0/QAvHmFK0kLK5IFnpuM96/FK4IzlNESzXHuK8tCfxK9eeIHFOTnkPJFDzhOzmOxNZfJvchinPex5ClOSdpYkvHcVUVLWVobm37mOdeWB1mlty7Eo3b019Cr7OHCcPMnJLnwV3SXqQ7/T5aJd8XZT/aSrrYZSOitcSVpH885nzgY3bgIELsti3tymj/1pU+dx6Hcty4nBN3U6vlD/zYJ7MLfckMfCh0qZ/OjTZNl4Z3MeJWkSqS4rSQslaGERet6gZTXtRCJfc6/qppI06TbRFJ45O57X6bL5d05R0LV+MYKCKkZQUMUICqoYQUEVIyioYgQFVYygoIoRFFQxgtr8eoC2bWTU5tfL9DuKyKjNT4ynoIoRFFQxgoIqRlBQxQgKqhhBQRUjKKhiBLX52UjV5lwKPu3tUdjTedzc19Lm109tfl0kaJVjfd/bo7AntfnZzakA5dt2cOiEE8/Q0fiubXcnaaCSop0VHK+PIWGEj/Qrm6ZZB7ZS4fLhObKD0mNOkseMI815hJJ3S/HjIf0mHwmti7Hwl5VQcvgkzstT8f2jF7cBnQxq87OZD97YSiDJh+9nTsqfX8LWY80TaopYsmQHwRQfo4e5qcrPY92BphaVoL+Uta8WUNkvHd/1ATYuWMiza8rpl+EjlRJy80to6lQJUvlaHi9WuUkf4yP5P7eQu6zEiH5UtfnZTHr2VDKTPHgGjGNylpPSTwNAkPLCV/BMu59xAzy449OZOnM0FWuKaCnu6Tcsk3EDPHjSxjH6qiDpWVl44zx4J01k4rajTQ3Tx4pY689i3l3pJLg9eCfN4v6YtewI1aJiM2rzs5lL2h2G22qSLALHkrnmtnadJ1cmkFpWTgCadhh92vehxOB0hFj4iQCln2wh76nWDkCC/ip+kdlVo+8+avMzyRkftE65ib6uJ/MBFt17dgeg3UUZ1KY2v6zWvx+nZNkf4d6HyYzv4pFJO268Q4MsebeSEXd7cQL+99+k6OZbmAZY53p6i0QvU58rofRXXtJdAH5KXisiZvxU0tznenLvUpufITw3z2Ja/nIemQfuS5ra+3LmpuEkiqA605j2xBGee3wOG91uCASIGT+Px2weUlCbX4/o0m0btLC+d+Hqe76LsV/Tn9r8LiRh2vuiW4xZTX+61i9GUFDFCAqqGEFBFSMoqGIEBVWMoKCKERRUMYJK0nqAtm1kVJLWy3R5OjIqSRPjKahiBAVVjKCgihEUVDGCgipGUFDFCAqqGEElaTaikrSOqSTNRlSS1jGVpNmNStJCUkmazagkLTSVpNmMStJCU0mazagkLTSVpJlEJWmRUkla71BJmkrSDKGSNJWkdTuVpEVGJWkXEpWkidiXgipGUFDFCAqqGEFBFSMoqGIEBVWMoKCKEVSS1gO0bSMTriSt9RKqiJ3p0C9GUFDFCAqqGEFBFSMoqGIEBVWMoKCKERRUMYKCKkb4/3Qz6cG/h4nlAAAAAElFTkSuQmCC"},"b72a8cd4-8a98-401f-8a8d-db4ae5a887cd.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAagAAACXCAYAAABA6c8MAAAe1klEQVR4nO3dfXRU9b3v8XebmzkpG3EG6ATqRJ2hTcI9JC5irgk5BMSogPIghIIcD+iBeK/QFqwNq224lsQu8KwF7QHsAW/Fh4t68QFEwQhKqUJYPFiMhdBrkpaMSBRmljC5lA3jzIneP5KQhDyQTCbOjnxe/8Dsmdn7m+/89v7u328/feurr776ihY+/ewzHHY7En1nzp5l0MCBsQ7jG0d5jZxyF13KZ2TOnD1LksvVZvq3YxCLiIjIFalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJf2Xnny57thONr99iE8vgHP4BCZPzMJlAHjZ+esynD+ZS0bTJVV/r2DL7/+I/YeLybu+/fl53yqhzLmYuZl2vG+VsPGDxjf6XUdyRgZ5t2TgNK4UlUntoZ1sf+cYfvpx3c0TmDE+DXtcB3EB1JWz8XWY8kAGdrzs/PVGDjW9N9BNVuZt5Ga5MJo++8R2vE2hDUkmY9RE8kb0gWvHju+kpCqFZXe5CX+4kQWvOileMQEXYUwTDCM+1hE2qA9jhuMxEmIdSMe8b5VQlbKMCcNiHcnleuO39LLz11WkPDoBdxTn2mcETcLxBvFxUZxnN9t4uOFH5Yq/aqttWYxFIW8R96DMwxsoejeevPlFFC0pJH9YDeuWb2vccIcxj/gw65sCrWbL45tg8sIOixNA+EIFvmDz//uNmk/hI4UUPjSF7H5eNvx8DWX+zqLyU/bEEjZ+6mbyTwopfOR+xrCP4l/vbD+uJvUmvjNmUxSYR/ox+v7ChmXfl4vxl3Usf9Pb/NkjyUx+pOH9hVNHUPdmERsOm1hevUnFhTAA8elTKF6UiwvgzEHWPXmQupgG18Lxbdz/ZnWso+hU+EJF23ZkBb3yW4Yxj5iEozrPvqP6zfvZdjzKM+1WG6/j4JPrOHimCx9ttS2LrWjkLeIe1Kcf7yRv/GbcdoB4nCNn8LBRiy0Mrcu8n7Kn1uGfXMSC5G7u1SX0xzAMwMCdlU9R4jaKntpNytI8nO18PHxsN5vsC1gzPa0xBIPkuxawOLyE3YdzKcjs+qJthoFhALjJnTWDqmUVeCe5G/cgbfQ3jIYelZFG/r0TWLKrCjMzgyt28LrIPF3OwUO1mPEOUnNySbYD+Cl/tw5XcoDyw+fx3J5HstHRZxvn88khyj70waA0cge3WEDQT9VH4LgGDr5djv8U7HgjQErWFDKGdDUeABP/hwc5+ImJbVAq2aOSG3ur4D+8jVpX8/xavvYf3k3d92/miw/34j1n4Bo1hgxnPJwuZ9ueKvAH2PbGSVIb/8aoqKumbH8lgbCBKyubjCFNMw7jP3aQ8uMBQgNcZLfsqZt+yt8/SO05G44RueQOaxFMqI6Kd/fiPWcjcWQuWdd3EGh9HdUHDlJ5JoRxfTbZI53N7aSDmMzq3Rwkm7zktq/N6t1UGdk4a8soP01zXGY1u7vwW0bqi7oKdu/xYsYnkjG2abQEOm8DLX7ni43fCzfOZ4CbMWPSLn220zy10vHyOp6Hn/I3anFNzWjcdrR4bVaz+6N+ZDs/bV5XRrsxaGize6vAH9jGtk9TyRuX3PN1vKM23m5bMKl+dwflp/zw9jYCydlMyXRC0E/F4XK8V8xVL+atcXvkGfkFH+zxYg5wkT0qA2dC9PIWcQ/KPiiDD/aV4f1705R4nMlu7C1rUL2fsieKKR9ZxIJR7ZWUbroxm7wvj+HtYPfQW7mF3P+W1qYb7J66koLMHjSroEmgs/dD4c7f76Zw9RZWPefFkTmG7O+b7Pifazh0BqAOb+lz/O9d53EO8+C0tf3s7seaPtvQy13yf2pxZmYzwl7Flpc+aF5I0Ed5uY+wzYnH7aT/tU48w9ztDqF2HE+Y6s2r2OB1kDE2G094L2vWldHUya37eCPec83zafm67uOdPPf0Di4kjyF7eJiyotWUnQEMJ+7rHTDIhbvxb4wKfxlrVu8l9P1sxmT2p2rtcrYcDwNhvK8tY/X7ITyj8sj+rpcNl0YCvGxbuQHvd7PJG5tC6K0lbDrW3I84+OZu6tzZZA+Pp+LJNew+3e6CKVu3hr1hD9ljM+j/0WqWv+Zt6I20ismB97lVbKlumH/YX065v3lZLV+H/eVsfmoLxxwZjMl0UPt0ccP3uvBbRu4g2/9Yh2dUNqnfqWDd6t2Nv/OV2sBONr20l3ByNmlGBWtWr2D9zgukjMog8dSm5pGJy/NUtZ5VTXlq5bLlndtB0bpDjT3G1vNweDewanN14zzq8D7vbdGzbPE66KP8tafZcsxBxtgMHJ88RXHj94yhblyDwHG9G/f1TqLSHNtr4x22BRvO6z04r+2P0+3GPdSAsJctj63mUNhD9u3ZOI9vaJHHjvRC3hq3R0/tvEDK2GxSL5ZR9B9l1EUxbxEXKOe4JSz8Rz87H1/IwqKVbHirHH+rnuUFyl9ezZo9aWTfFIXiBIANW78QoXaHVurwncjCFZXbYF2gtrqCimMVVBw7xJanNmG7J7vF+LufmmON73+4mw3PfcDEsWlR6j352ftqFaMfyCfjOjvO5Aks/NVEnE2r6sdpTH0gj4wRbuzxfva+6mfCj5o/WzDHYPMBL+Dn4I4Asx/MJ+M6J64RE8ifmNR2cfF23Mku+vVzkTIiDdc13Yjn9F42V42mYHoGLruT5HH3M33QJnYf68pgkJe0u2eT1TjPKff4qfKacI2LNLcTBntIG3HZDk/EwlT8YRPOGfeTl+zEfl0Ws3+5kOSEEBCPe2oJJfPySHYaOEfkkuuoovYMUOfHG5dB1nAnht1N3sJ1zB7RHFDGPfnk3ujEmZzHlPHxlP+t7Z5T+NhuNg2azv3jknHaXWT981IWJscTahNTBvkPjKbq1b10OordyDFuAhMa/5Yp01M4WF3bhd+yJzKYMT0Xt9NJ8rgpTLSVU3WGLrQBL8m3TiHjOifusWPI/tDF6PwsXE4XWeMn4vxbLXWNedpx/VwKmvJ073yy/rKbistHqy5f3l0LKR7vhHDbXGdML2B01Wb2trvjcHlCb2PCXY3LnjqDlIOV1ALGdWl4BoPTnUbaMPuVjwF1RZs23llbiMc+LAVXv364ktNIu86AeDf5JSUNuTKcpI3JxVFV2/mwbm/l7eM0pt6b1TjPKUw5VUWNGb289eAkiXhcWfksyMqHoJ/qA9tYvdzLgyX5uOMBKvANXcGGFRWsea4M909y2x2W654QfNnRewZ25yGq66DnCzqP3+vFa4D/2Eb8maspGt1ypgFqj3sJYeItO4TjvuXkd3f4skN11H2YQlqLoZl4Z3KL4mi77LMV7F5dws6mSRd9eDPGNLx3xEPaoOZPGw4nnIxiPNV1lKektUh3PEnXu9lZZ0IXDtHaWuweOexOAsHeOsphUnfKjWt8i9/oGhdpTRvwC7UcLCvj2Ade/PVfEKh1kA9gz2DKzatZsXgv7pvSyMjMY8xI56WVrWX8tg4OBJt1tbivn9hiBTVwjTCAurYxDUki5cMKutSE45rbgTHAgffj3j5CZGv9/6a//dyV24Ct1W6wwT+0s6qYdbX492yk5GjztMCJfuQHodWeXzvLcyY3rB11bXLtJCmlnIpzwJWGOr/d4i80+uP4OPw1HnNrp31eoS2YnxykbP8xDn3sh4sBTg7K73wRvZW3Vu3CgWNoAPPy36wHIi5Qrc4qSXCSPK6AB3wFlH+Sj3sYQBYTxyVjH+Rm9tFlrH7NRcl0d8/2QM5UcdCXwoxB7b0ZT5I7i40feclPbn2ukX/PenYac5mbacd+E3B5D+zvAfy2xBYTnGSMn0LuIOAmKPp9Od7xrhZFIoUxU6eQDJjXmSw5XEX+yGgdf4oHe6gbn8/lgUdnk9xmejVw2Xy+7M58uxjPf7Z+GaoHW7SG5aLIZqODnnctOx/fDP/9QQrG24mPq6PsN081vheP+64lrLsrjOn3cvCNlaysW0LRuG7sAcXZINRJ/i6LKWQnOnvpX6cotYGUqQtYMvYKuY0D+3928n79Zbn+TztE8+y73tTVtvDJTpa/Dg/OL2CCPR7OlLHyuSvMu4/mLcIhPpPy5xaz6cMW/W+zgg8O34yrTfuKJ3nGw0z0PsW26sj3ScJnKtjy1CYc9+Z1eKqrPWsG2YeeYtOHzZ1d85PdbHo1nhEpBuDEMzLAjr3VXIq83k/Z9m04R6a0v89/40RmD9/G5n3td6CNzCnk+zexswd/W2sukm8rY8e+poGeMBXPL2x17KOZm+TpZRxs8Tv4D2xhy7G6S+81z8ek/E+7219knA3bBZPz3Y3HnUz+vh3NZ1YGqyl7O8yIhjNniP+2nZpPG980K/jgYBdTEAf2c+ejuAdrxz0izI691Zfm6f/DCla86wcuYFa78Nxobzgdtq6Kqo+aPnSIDW9UECYew5nMmMwUagLdO0PK7h5B+O0yqhvPTuX0blYs340fOykZl8W0bwdltyU3nFkZZ8Nfe7Khndb7+eDQofZm31anv2UvuEIb6Cr79zMI7yuntsWZvzufL6P28kZwfTJ5rZZXwcYfb6Ii3DiPlrn2l7FjXx7J10PDjlYNJ5ua47EP6GpzjP+2nYAZ5f5UqzZ+hbYA2GwXGnomAEGT6us8uBvHv+v+VkXVlZbXR/MWYQ/KIGvuYmrXLmHhq4kkfucLfL7+5D24mKx2x72d5M6ZyJrH1lH2q8XkdnEH9NDqAmasBrDjysxmwj0ryU/tpJ8S7yb/lw+w7fdFzHjaTprNj2/QGGb/8n4yGuNyT1rK7M1rWFLgJ/EGBwFfAPddS1k4uqMVKp60u+5n58+3UZE5l7R2/rYxP8ym6NW95HZwdmH3xJM2fQmf/scKHv6DA0coAOkLKBzR3r5U02eX8vAbDhwECFwzgcU/avhb0qY/jPc3TfPpR9b02WR93M5s7CO4bdgqVjxSzpRFRUy4savxpDGj8FNW/9vD/NHh4IsApP7zAvIbhwXcoxeSuK6Igi02+g+/n/xMunR8hWFZzH59BYsfyabg13PJiELX1DmugBnPrWdJEThsAQLfnUzhfCfgJPdnu1nxSAkOB/DdESSnNH3JTUZgPUuKXsPxnS/44ts3s3hRN68EGpJHwfSNrC8qAsc/EAg4mfyzgoZ2cnlM10xgyY8aTvKxj5xC3v6V3P9jcA3OI39UVteW1+lv2QviO28DXTYkj4LxG1m/pCFPTfNxXd7s21le2oOFpMXTTq77MaHw4Yb3cJP7UCLriwrYktCftPvyyehaa8SdNZvA44t5OKuA5XOiNFJyeRvvpC2AnRHjkln1+MOUT15M0fhcluxawcPFDhyAMz2ZlM6X1mfz9q0eP7AwaGLW27p+YeCZMlb+jzVcvj84e8Vm8tuOU7Wrbt9KCla3mQMrNuc3D3XVhzG/jKfDsOrDmMEQtq5c/BYl3X6YWdDEjDM6/htaCpuYtP/ZLl/k14N4oraMCHQrr2ETs95oe4FkfRgz2MEFrh19p1s6uYA2KvOPTDQfsBetNhA2w8R3odF3vLzYXXje43x2py10Z/vQahHWzFt7DyzseYGSLtPTNnuH8ho55S66lM/I6Im6IiLSp6hAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJbW5F184HCZQZ5mHf3+j1NfXK7e9QHmNnHIXXcpnZOrr233UQNsCde2Aa3E4dKuj3hAI1Cm3vUB5jZxyF13KZ2QCgfaLuob4RETEklSgRETEklSgRETEklSgRETEklSgRETEklSgRETEklSgRETEklSgRETEknpeoM5VsnXNUp55PxCFcKQV5Ta6zlVSuv5xlj66lLWvVWLGOp4+xeTkgVdY++hSlj66lq3/V9mLmppSHn/0GQ5rNW+jzZ0kui6Eb98LPPkHSPzOUQLB6AUlym301bD1sacJzy1mWSrUvrue4hdh+X2p2GIdWh9Q80Yx6//fLAqXziQxVEPp+mK2JqxkmifWkfV1Pna99B5nzw3EbP9uP1e1HvSgTHzkUFg8j1tviF5AAspt9IX+/B5bM+czM93AZjPwjJ9OzvvvcVQdgS4I4fivCyi6N5PEBGCAh5xRA6ms1S5/T/nefpLDo+Yz6XuxjsSaelCgHKSPTsWIXixyiXIbbbUnjpDjSWoxxcPwW0qpqY1ZSH2IDccPPDiaupr1Pg4fOEuqyxHTqPo8/x6eOZDDvHGJsY7EsnSShFwVQhdrGHjtZSU/Ljax9F2VvDJ5MpPvKeHIzYu4W8N7PRBg/wvvkDr3DlSeOqYCJVcFW5yD8xdDsQ6jj0tl5vbtbH99Obd+tpbH3/bFOqA+y3z/ZUqHzdcxvCvowUkSIn1H4g05HPcF4NL+agDfiXQcY2MZVd8RMkPYjMYxvjgHmRMmsevZSgLjE9FAX3fVUPq7/QSSTrL0cMOUwImT8O9LOf/Acu7+QWyjsxIVKLkqGKmZJD32DkdHzyHdgFDVLt68eCuFQ2MdWV/gY8/KVQRmL2dmSkOR8v3lTxwfeqeOk0bEwz3r13P3pddn2b/mBbhvEWNdMQzLglSg5OowIJN5D9ZQ8tOFvJxoEPhyOPMWz9T4f5ckcsfCWTyzZiELScTxRYDAjdNYVpCuU/QjZDOMFrkLYYsDDKPhX7nkW1999dVXLSfoiZC9R7ntHd3Na6vhqqtct9tkyCQUpw1pR7SOR6ajvKkHJVcdFacesBnqNcnXRmfxiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJbW5k8SpU6dJSEiIVTzfaMFgULntBcpr5JS76FI+IxMMBhk6dEib6W3uJJGQkKBbdfQS3QaldyivkVPuokv5jEwgUNfudA3xiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJalAiYiIJfXoke/mif2Uvl7Kkc/7kzR2FnNu92BEKzKBc5VsffZ5AqMKmXeLI9bR9H3nKil9cSv7PztP4sg5zJ+eqvbaZSFOHnidrW8dwdd/GHfe9y+Mdenh7z2mdbxTkfegPt/D2ieO47lvGcsfXUTO5+tZ9QdfFEO7moXw7XuGkt/uxxc8ii8Y63i+CWrY+tjTmP+0iGWPFjHJeJPiFysJxTqsPsJ8/xlKDiQybekyls1P55PVv2GPP9ZR9WVax7si4gJl1hzHkT+JzME2sBmk3z4J2weVBKIZ3VXLxEcOhcXzuPWGWMfyzRD683tszZzPzHQDm83AM346Oe+/x1Ez1pH1BT72bw8wZ+5YkhJs2AZnMj3f4LUDNbEOrA/TOt4VERco45Z5PDSquUsaqj3O2R8koU5qNDhIH63hp2iqPXGEHE9Siykeht9SSk1tzELqO8yTHD+XStLg5knGsJsYUnlSO6QR0zreFdE5ScK/h7XrQsy63ROV2YlEW+hiDQOvvWxzEBebWPqciyaB7w1svfOp06vka9DzZhas5JU17zHsF/PIHBCFiER6gS3OwfmLOuIUkbh44s+HdLxOvnY9K1DBGraufB7m/pJpHp3RI9aVeEMOx30tB6QC+E6k49BO1ZU5hjDs3El8LSvUGR8fOR0aopJe1YMC5WPPuqc5P7mImSkqTmJtRmomSbveuXRSRKhqF29evJX0obGNq2/wcNNtH1G6t+ksXR97St/hjptT0ZovvSni66AC+55n1btH4d17eeXS1Dms3D6T1KiEJhJFAzKZ92ANJT9dyMuJBoEvhzNv8UwSYx1XH+GZWsjYNUsp2JlIYjBA//FF/Cxd5Ul6V5tHvuuJkL1Hue0d3c1ryAxhM7RxhQjaZMgkFGdg0wkm7dI6HpmO8tajO0mI9EUqTj1gMzSsJ18bnSwqIiKWpAIlIiKWpAIlIiKWpAIlIiKWpAIlIiKWpAIlIiKWpAIlIiKWpAIlIiKW1OZOEqdOnSYhISFW8XyjBYNB5bYXKK+RU+6iS/mMTDAYZOjQIW2mt7mTREJCgm7V0Ut0G5TeobxGTrmLLuUzMoFAXbvTNcQnIiKWpAIlIiKWpAIlIiKWpAIlIiKWpAIlIiKWpAIlIiKWpAIlIiKWpAIlIiKW1KNHvpsn9lP6eilHPofEkXOYPz0VI1qRCZyrZOuzzxMYVci8WxyxjqbvCxyl9KVS9n92nsR/nMasH2aSGBfroPqKECcPvM7Wt47g6z+MO+/7F8a69PD3ntD288oi70HVbKX4ieN4Zi9jeXERkxK2UvxGTRRDu5qF8O17hpLf7scXPIovGOt4vgFClbyy/B1s4xexvHgZ0xLfY+mLRwnFOq4+wnz/GUoOJDJt6TKWzU/nk9W/YY8/1lH1Ydp+dknEBSo04CYW/GIWmU4bxBl4bslhYOVJAtGM7qpl4iOHwuJ53HpDrGP5hqg9zke507nDY0CcjaRxk7jz1Uq0SegKH/u3B5gzdyxJCTZsgzOZnm/w2gFlL1LafnZNxAXKNtiDZ3BzF9/35/2cTU1CA1HR4CB9tLr7UeW5m2VTPc2vT53ko9tdJMUuor7DPMnxc6kkDW6eZAy7iSHaoEZM28+u6fFJEpUvTWby5MmUHM1k0UTPlb8gEmvBSl75zTukT87RTkBXXDQJfG9g642nTq+KCm0/O9fjZpZ673a2b9/O8rGnWbtiF75oRCXSa3zsWfc8Z2cWMk3bg66Jiyf+fEjH63qBtp+di7xABU1C9c0vHTffySTbYSo/j0JUIr3Cx+Hfr+VPmYt46JbEWAfTdziGMOzcSXwtK9QZHx85HeqBRkrbzy6JuED59q1i6auVzXtV/kr+9LELx4DoBCYSXSEqX1rLezc8xKIxKk7d4+Gm2z6idG/T/r2PPaXvcMfNqehE88ho+9k1EV8HlXj7Q8zasJaFP4ZER4hAIIlpv3qIdLVYsaK/vs6KF48SYCF7ftc0MYfCZ3/J2MGdfVEAPFMLGbtmKQU7E0kMBug/voifaWWPmLafXdPmke/dfyJkCDNow9BTjq9IT9vsHcpr5Lqdu5BJKM7Apguc26XtZ2Q6yluP7iTRQMkVuWrYDA3rRZW2n53RyaIiImJJKlAiImJJKlAiImJJKlAiImJJKlAiImJJKlAiImJJKlAiImJJKlAiImJJKlAiImJJbW51dOrUaRISdGlzbwgGg8ptL1BeI6fcRZfyGZlgMMjQoUPaTG9zq6OEhATd16yX6J5xvUN5jZxyF13KZ2QCgbp2p2uIT0RELEkFSkRELEkFSkRELEkFSkRELEkFSkRELEkFSkRELEkFSkRELEkFSkRELKnNhboRqSnl8Wd93PHIPDIdUZmjAJyrZOuzzxMYVci8W5TYHjtXSemLW9n/2XkSR85h/vRUjFjH1GeYnDxQyta3juDrn8St0+dwxw+UvZ4wT+yn9PVSjnyO2mMHotCD8rHrpfc4e86HWd/zuQlACN++Zyj57X58waP4grGO55ughq2PPY35T4tY9mgRk4w3KX6xklCsw+ojAnvXsvavHuY8upxlC3II/K9V7DoV66j6sJqtFD9xHM/sZSwvLmJSwlaK36iJdVSW0+MC5Xv7SQ6Pms+k70UjHGlg4iOHwuJ53HpDrGP5Zgj9+T22Zs5nZrqBzWbgGT+dnPff46gZ68j6ApPjxx1MuysThw1sA9K5Y4qNw38NxDqwPis04CYW/GIWmU4bxBl4bslhYOVJlNHWelag/Ht45kAO88YlRikcaeAgfbS6+9FUe+IIOZ6kFlM8DL+llJramIXUhxhk/utD5Axueh3i5PGzpLo07Bwp22APnsG2S699f97P2dQklNHWelCgAux/4R1S596BypNYXehiDQOvvazkx8Umlr7Ot3cta4OzuNMT60j6vsqXJjN58mRKjmayaKISermIC5T5/suUDpvPNOVU+gBbnIPzF3XEqadCVa+wds8wiv41Uz38KEi9dzvbt29n+djTrF2xC1+sA7KYCM/iq6H0d/sJJJ1k6eGGKYETJ+Hfl3L+geXc/YPoBSgSDYk35HDcF4BL/f0AvhPpOMbGMqq+JVSzlcc3wpxHp+HRI496JmgSijewNfbiHTffyaS3n6Hy8ztIHNz5V68mERYoD/esX8/dl16fZf+aF+C+RYx1RSkykSgyUjNJeuwdjo6eQ7oBoapdvHnxVgqHxjqyPsK/h7VPn2faz+eQquLUY759q1j1+SyW35uKDcBfyZ8+dnHngFhHZi0RXwdlMwyaD/GFGvYEjOY9AhFLGZDJvAdrKPnpQl5ONAh8OZx5i2fq+GmXBNj/wir2HIU9973SPPm+lWy/NzV2YfVhibc/xKwNa1n4Y0h0hAgEkpj2q4dIt135u1eTNo981xMhe49y2zu6m9eQGcJmaEsAapPR1v18hjCDNoyrvFfaUd50qyO56qg4iXWoOHVGBUpERCxJBUpERCxJBUpERCxJBUpERCxJBUpERCxJBUpERCxJBUpERCxJBUpERCypzZ0kTp06TUKCrhzrDcFgULntBcpr5JS76FI+IxMMBhk6dEib6W0KlIiIiBVoiE9ERCxJBUpERCxJBUpERCxJBUpERCzp/wOsmqaA7+irYgAAAABJRU5ErkJggg=="},"877acc94-b90a-44ea-824e-df9e8b42f774.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAwUAAACfCAYAAACocHXBAAAgAElEQVR4nO3dfXhTZb7v/7fWrqmmgw0wCWAr03QGipuWsWYEqlAU8RFQqAd1KzgCzkb2PuJW+I2WnxvQA851wNnA7CPuPYAOjEdRERSqWEQpxSKIOBSUgrYV2hHTLaR2uiSk1p4/2tIH+pCElDTN53VdXtJkZeWb732v+17ftVZWLqitra1FREREREQi1oWhDkBEREREREJLRYGIiIiISIS7qOEfXx09SlRUVChj6bZ++OEHLrrooo4XFL8or4FT7oJL+QyM8tY9qV3Dm9qv+6mpqeHn/ft3uNyZVo+OjsYaF9epQUWqEydP0qtnz1CH0e0or4FT7oJL+QyM8tY9qV3Dm9qv+3FXVPi0nC4fEhERERGJcCoKREREREQinIoCEREREZEIp6JARERERCTCqSgQEREREYlwKgpERERERCKcigIRERERkQinokBEREREJMJFzZ8/fz7A3//+dy6OifHpRRUHt/CXtS+xaVsuh/77Yvr2j6eHAVDClqc3UTlkCH0bVvX3A6z/j1c4Hj8Mx6Wtr6/k7QVs+u5XDOkXQ8nbC1j+Si65O3LJ3fslx0+D3dYXi9FRVCZluzfx0guv8faOj/nS7IHDYSfmwjbiAqjYx5pXjtP/V32JoYQtTy/npR31733wGJ7aXvSN74HRsOzi/2Jj/fMfHzmO5ycJOGwd5+zUqVNccvHFHS7XaYq2sOCjKEb90kr1p2v4pz8eJ230L+hBNab5I4bRRX7JuqYa0xuF4eMPKYYiryVvL2D3haP4RZf7XRf/2tK33JWw5endRGX8Auu5Bxh+PCbVFxhE+XDoxOe+6GcfrzZNfjQMOmzVZmNZiHVG3jpLuIyNfuS0KwhJu3albaCl7th+XTnf54HPY3MX4fF46PHTn3a4nN9d1Ny7kqwPohk9LYusObPJTCrmuYVvUQJANeZ+F2ZNQxRHWP/MyzBuJqMvb3ud1d8fwOVp/Pclw6cx+9HZzJ4xnmGXlLDyd8vIK28vqnLy/jiHNX9LZNz/nM3sR+9nJDuZ//SW1uNqUGPiOmE2RIG5/xKuvX923XvfOwLLZ8+xcHNJ47L7BzDu0brnZ94+mIrNWazca9Ll1Zgc+L4agOjU8cx/eATxACc+4rnnP8K337k7D4re4v7NR0IdRbuqvz9wdj/qCjqlLasx95tUB3Wd4ePI5vt5qyjIK/Wrj1fw0fPP8dEJHxZtNpaFVqfkrbOEydgYVjkNlS60DbTULduvC+e78/kxNocZv88UfJW/CMs1j5N+WRRRFxlY+iaTknAxsT2sxESd4NCrh7GMu4b+l5STt+IPlFw3hym/6tHuOk98/iqHY8dzzeUxdf+Ou5NRv+iBYViwxl/BiCvK+c8Xy0gZ6cDSyuurD77F4orbePruX2GNMTAMC71++WsGfvefbKu6hrR+f28SV5MXnjpG/l/hqvT+xFAXe4+Jo/jFpQZGjJX+A6yUvPgl1hsHYj11jPzNHlLvG0Jfw8Cw2LkivoaXP6zmmqv60t6JDH+Ompjf7CNv28ccLC7ngp/1p1cMQDn7PvgbUcaX5H1wmB8THPQy2lq2fj3HdvP+9n18+a1B34uPs7HczqQresGpY3y8/+/YbSfIy85j/xfHMU+VUx03kL6xvsYDYFL+aR7b9hyk5NsL6H1Zr/qzMlC+9y0OX9i4vqZ/l+/dxt8u6cnXu3LYXVCG2TOBvpYo+GYfb+XsZv+xb4lxu7mg/jO2x+e8Vhwh7/0P2f95GWacnb6xDSuupvxgPnm79vP5cRNr0zNSZjn7PtzGxwUllEf3pX/PuidOfP4q5dYbqfksh90FJVRc0pf4S9sItKaCI/nb+fDTzyk7ZcXe19LYT9qIyTyyjbwTdhy9zv7bPLKNgx47NYfeZ/snXzbGZR5hmw9t6X/uTnDo1XLibqzh4Lu7OVhcgaVfw1lBaL8PNGnnQ/Wv+/4A297dzcHjXuyX288s226emmn7/dpeRzn73jxMVHLf+rGjyd/mEbYVnML+w6HGbeVyKwZ1fTZn936OfRuDu/IC4hN7nfs23lYfb7UvmBz5IJu8fV9y/HsP5aetDOxnAU85Bz7KY3fLz9lsLOvkvNWPRxdbv+bDd3dz8LiJtW9fLBd1Ut4aPkXYj42dlNN2+nGH8bTIXcxXTceYo3j7OLDXlLH7ve3s+6Ll9h+Mdm2nb/69jN3529lXUEL5hb3p36QRK47ksX1n/bhdn6eGbWBg/Dfkb9/H5y3HGV/b73xvEx21X3vxhLL9Osx3+/2uIVdHvrPQN74H3x/cRs6ug5T9YG9+BUab83cLAcy3mEfY9uEJ7A1t0vTvNtuljbG5WZO2MVf7Ek/BKeyVH7PlwwqsyVEc9idX7ei0MwVxvdL4ZGceJX9veCQa24BE4qKbLFRTTt4f57PvyiweGm7z9y3O9vNhjP7xICVtHLYpKVzPiF+nEN3i8cTbFzPd2VoZ4SOPibu9573V7T/vp+oj61nyYglW50iG/cLknf9/GbtPAFRQkv0if95ahS3Jgc04e9ltTzUsW3c2Z87/LcPmHMbguMOsf+WTJp/Jxb59LqoNG45EG7GX2nAkJWJrJU1tx1PNkdeXsLLESlrGMBzVO1j2XB4NJ3MqvlpDSWXjepr+XfHVFl5c9Q7fDxjJsEHV5GUtJe8EYLGReLkVesWTWP8Zg6I8j2VLd+D9xTBGOmM5vHwh64uqgWpK3pjH0j1eHMNHM+xnJaw8c8arhLcWr6TkZ8MYnTEQ79tzePlg4/HyjzZvoyJxGMMGRXPg+WVs+6bVNybvuWXsqHYwLCON2ENLWfhGSd1R92YxWSl5cQnrj9Stv7p8H/vKG9+r6d/V5ft4/U/rOWhNY6TTStmq+XWv86EtA/cRm96vwDF8GMkXH+C5pdvq27mjPrCFl1/ZQfWAYaRYDrBs6SJWbPmegcPTsB9/ufEMXMs8HV7BkoY8NdPi/SrfIeu53fVHcpuvw1qykiWvH6lfRwUla0uaHPFt8rfHxb43VrH+oJW0jDSsx/7E/PrXWfomEt8LrJcnkni5rd0dW5+11sfb7AsGtssd2C6NxZaYSGJfC1SXsP6ppeyudjDshmHYilY2yWNbOiFv9ePRn7Z8z8CMYSSfyiPr/+RR0Vl5o7uMjZ2U03b6cYfxtMxd+T5ef2ULh61pDEuqYNP/WsTi1w5idQ4jmY+Y/+JugndsuJ2+eWI3y55YQ1lcGiMzBkLu/z5zVt7cu5JlO2FgxmiGxR1m6bPbzow7HH+HLQcvYfDwNOLdm8had6DVM51tt18Itol226+DeELafu3lu6N+9zIv51YzcHgKloJlLPv9Ct4xBzLMaaf8/y7kra/q199y/n5+Sf383VJg8+2Zbb5pWzT83Wa7tDI2t9DmXO1LPG+sYs3nsSQm2bD4k6sg8bsosF03h5n/UM6WZ2YyM2sxK9/eR3mzXvY9+9YtZVluCsOGBKEgAMDAuMSLt9XLNipwHR1KfFCu8f6esiMHOHDwAAcO7mb9n17GuGMYiWeeL6f4YP3zn25j5YufcEtGSqtnL/xXzo7XDnPtbzJJuywO24Cbmflvt2Br6K5fpXD7b0aTNjiRuOhydrxWzs3/3Ljs9MkWXt9VApTz0Ttu7nkwk7TLbMQPvpnMWxLOfrvoOBIHxHPJJfEMHJxC/FkFZDvxfLOD1w9fy/SJacTH2Rhw3f1M7PUy2w76cqFJCSm33cPQ+nWOv6OcwyUm/DSelEQb9HaQMrhFkRmwag689zK2O+9n9AAbcZcN5Z4nZjIgxgtEk3j7AhZMHc0AmwXb4BGMsB6m7ARQUU5JVBpDB9mwxCUyeuZz3DO4MaC0OzIZ8XMbtgGjGX9TNPu+PLtarT64jZd7TeT+6wZgi4tn6D/OZeaAaLxnxZRG5m+u5fBrO2j3Crl61utu5ub6zzJ+4kA+OlLmQ1ueizTunDiCRJuNAdeN5xZjH4dP4EMfKGHAqPGkXWYjMWMkwz6N59rMocTb4hl60y3Yviyjoj5P71w+hekNebp7GkM/28aBljNXy/e7dSbzb7JB9dm5Tps4nWsPv86OVou1lgm9nptvrX/v2+9k4EeFlAGWy1Jw9AZbYgopSXFnHXAIyFl9vL2+EE1c0kDiL7mE+AEppFxmgehEMhcsqMuVxUbKyBFYD5e1f4lLZ+XtqxRuv3to/TrHM/74YYrNTspbpIyN55LTNvpxh/E0y139qpwjGD3Ahm3waEb0rSbt5psZYLMx4NZbuOWDMv7mQ4v5pJ2+Sa+hzPz3LDKvjCcuLpERI1LY8lXdO5cf+4jLUtNIjLNgu/IeFj0+mjN7GvZbGH9rCvG2eNIm3smIN45wdtncdvuFbJtoo/06jCeU7Qdt57vDfjeA0benEW9LZETGMPZddi13Do3HdtlQbrnVxuGyCurm73dI/M30xvn7waEcfP/AWYVNZ823rbdLK2Nzay9tba72JR7r9WROHErK4Pi6fUufchU8Pn7dralo4odm8tDQTPCUc2TXWyxdWMKDCzJJjAY4gKvvIlYuOsCyF/NI/J8jOPfSwAs/tvWchTjbbo5UwLm/URXlJSWUWKD84BrKnUvJurbpSt2UFZXgxaQkbzfWexeSOSA40x5UUPHpQFL6ND4SbRvQpCAxWix7gG1LF7Cl4aFTLkrSRtY9t99BSq/GpS1WG5QGMZ4jFewbmNIk3dEkXJ7IlgoTiOtwzUaTUtQaZ8Pt6ayr1k0qjicSf1OTNvppPCkNk/z3ZXyUl8fBT0oorzmNu8xKJkBcGuOvWsqiWTtIHJJCmnM0I6+0nRnMm8bf1vcQzYoyEi+/pckEYCF+sAWoODumPgkM/PQAPnXhqMZ+YOlhpeSrzr7i32j+74bPXtlxHzCaHXKw8JNWNhWzoozy3DUsKGh8zH30EjI90KzabuX9bAPqto6Ks3JtI2HgPg5UAk36b6subPIJLbFYv6o+j9+haKV/dtAXzGMfkZd/kN1flcMpN6W9Mtt/i87KW7N+YcXa143Zss2CpruMjR05h5y21Y/b3U5bvme9qKaPXYIRrCmupXb6JlTjPriDbXv3ceSb78HtgvSrAEi8dibGc3N4ZHMKaVcPZsSIESQ2jOkXnjVitaLt9qs4FKJtoo32O3seaRlPCNsP2s63v/3u4p+0UiyZVBwv5/0XF/Bpwzxb46b0p5lnjdGdNt+ey/zQ6lztw5h/Ycvs+JKr4PG7KKg2TbBY6oKKsTHguun8xjWdfccySUwCGMot1w0grlci9xTMY+kb8SyYmHhuH+LEYT5yDeTOXq09GU1C4lDWHCohc0Bis2fKc1ewxTKFKc444oYALc80/N1NuWFv8oCNtJvGM6IXMASy/msfJTfFN5l8BjLy9vEMAMzLTObsPUzmlWlBmgejIc7rx/Ij+M2T9zDgrMePAC3W86M/6/Uxnh+a/+mtASNY1woEkWHQxhmmMrY88zr89kGm3xRHdFQFec/+qf65aBJvncNzt1Zjlpfw0ZuLWVwxh6zr/Kg6owzwtpO/FjF54+jUDb1TBKkPDLz9IeZkdJDbKIj7oZ3na1rk+oc4wua2EL72hWNbWLgRHpw2nZvjouFEHotf7GDd3SJvGhvPSVeLp0E7fdP89M8sPZjGI//4CPdYouHIeu5sOHBgS2PK/OeY4qmg7NA21vzry4xbcQ8pPr9xB+3X1baJrhaPr4LS7waSOXMOIzqaesNtvu1q8TTh5+VDJvtenMXLnzY56mEe4JO9VxF/VqNFM+DOR7il5E+8dSTwY2/VJw6w/k8vY717NIltLBM39E6G7f4TL3/aeBrFPLaNl1+LZvBAC2DDcaWbd3YcaTztVFNO3qa3sF05sPVj2z+/hXsGvcXrO1s/NWNxjiez/GW2nMNnay6eAdfn8c7OhpNI1RxYO7PZteyNEhkwMY+PmrRD+a71rD9Ycea5xvWY7Pt4W+tvGWVgfG9S5W88iQPI3PlO4x2hPEfIe7eawYl1mYy+MI7iv9U/aR7gk498TEEUxFVWBfFIbRyJg6t5Z8eRM+ssf28Riz4oB77HPBKP4+dxREcBFYc5fKhhod2sfPMA1URjsQ1gpHMgxW7/rsSMSxxM9bt5HKm/qxbfbGPRwm2UE8fAtBYx7XyHvOsH1N31JMqgvKy0rp/WlPPJ7t2+vWG7bdkJOugDvor7RRrVO/dR1uSOZVvW5lHWshNcPoDRzd7vAGv+5WUOVNevo2muy/N4Z+doBlwOdTsAxZQ2dMeDn+Brd4y+MA63GeTzBs36eAd9ATCM7+uONgJ4TI5c5iCx/jqBii8Pc7ij9+sWeesuY2MIchqk7bRTtNM3q0031svjsVnq+nrJoYbvflRT8vYa8r4BYuKIv3Ioaf3L+d6v4bnt9utq20T78XRhQel3deNj3qdlZx6pPrKFNbllZ+0jnMt8axwvo7S+/5QXfISPM27zsdnPz9TemB9qfp4psDB0yizKls9h5mt27BefxuWKZfSDsxja6nXMNkZMvoVlTz1H3r/N6rjaq7d76XTuXAoQR7xzGDffsZjM5HaOx0cnkvnEb3jrv7K4c1UcKUY5rl4jueeJ+0mrjytx7FzueX0Zc6aXY+9vxe1yk3jrXGZe21YnjSbl1vvZ8ru3OOCc0spRCBsj/8cwsl7bwYi5o4NwiVQ0KRPn8Lf/s4hH3rNi9boh9SFmD26tfmxYdi6PvGnFihv3T29m1j/XfZaUiY9Q8mzDei5h6MR7GNral1HiBnN90hIWPbqP8Q9ncfPPfY0nhTtn/42lv3+E961WTrsh+R8fIrP+9GritTOxP5fF9PUGsYPuJ9OJb9fvJQ3lno2LmPXoMKY/PYW0IJyCsV03nTtfXMGcLLAabtw/G8fsaTbAxojHtrHo0QVYrcDPBjNgYMOLEklzr2BO1htYLz7N6QuvYtbDbZWkbegzmukT17AiKwusP8HttjHusel1/aRlTD+9mTn/XPdF+bgrxzM6fzH3/wvE9x5N5vChvr1fu23ZCaLb7wM+6zOa6TetYcWcujw1rCe+Zbdv5f1SHpxNSjSt5PoSbp79SN1zJDJihp0VWdNZHxNLyr2ZpPnWG0kceg/uZ2bxyNDpLJwcpDOCLft4O30B4hh83QCWPPMI+8bNIuumEczZuohH5luxArbUAQxs/926Sd66y9gYgpwGazvtDO31TeftOH6/iKwdVn5SE83gwY6GFxF/hY3X/3fdazh1GttNjzA6DvD59pDttV8X2yba3Ua7sCD1O9t107n5xRU8kgXWi0/jrklmysz4s4+qBzjfEpfG+Ot3sPj+mXC5ndETh+LbjNtybPZ9/+CsfZJmY37oXVBbW1sL8Levv8Ya50cV5zExawwsFh8/yok8Fv/TsrOqsHsWvU7m2ed5W1WxczHTl561Bha9ntl4qrimGvPHaNoMq6Ya0+PFaLgE6jw4cfIkvXr68U1oj4kZZWn7MzRVbWLS+rLNLvU6F+3EE7T3CIBfea02MWssWFrevaumGtND6/24rdf4pZq6FHXW+gPjd59sR7D6QLVZTbQPnb7t92sn153snPPpT1/wZ3xo9hbdIG8aGwN2PuPxt13biq3aNCHGUnc218fX+KXN9gvdNtG68xtPl5sfaqoxq6N9GB+75nzbqvMcj7uigsv69etwucCLAvFZMDcwaaS8Bk65Cy7lMzDKW/ekdg1var/ux9eiIEx+dFtERERERDqLigIRERERkQinokBEREREJMKpKBARERERiXAqCkREREREIpyKAhERERGRCKeiQEREREQkwqkoEBERERGJcGd+vOyro0eJimrlZwPlnP3www9cdNFFoQ6j21FeA6fcBZfyGRjlrXtSu4Y3tV/3U1NTw8/79+9wuTOtfmmPS7Fa9YvGncHtrlBuO4HyGjjlLriUz8Aob92T2jW8qf26H7e7wqfldPmQiIiIiEiEU1EgIiIiIhLhVBSIiIiIiEQ4FQUiIiIiIhFORYGIiIiISIRTUSAiIiIiEuFUFIiIiIiIRDgVBSIiIiIiES7wn6yrLGTDC2txD5/N1KutQQxJlNsgqywk+6UN5H9dhf3KyUybmIwl1DGFDZPSXdlseHs/Luw475nGhCuUvaAozuaZF1yMeXQqTm3mHfKW5bNxfTb7v40lacx93DcyAaONZc2j+WRvrFs2IeMuJt/gaLLNu9m7cgkbjjZ5wdXTWDjO0anxS2u8lO7aWDe+xCZx4733kRHfVqtK12NSuGktG/aUUtXbyeQHJpDco7Xlisl+chX5LR9uut19kc3cNU2XSGKCxsbzLoCiwItr5194/j2wX1yA2xP8oCKXcht8xWx4ahXVU+YzLxnKPljB/Jdg4b3Jbe5QSKPiN+ez4ru7mD13EnZvMdkr5rMhZjETtP90jlxsfWU7Jyt7YtaEOpYwULmX1U/lM+jf5jHPZlLwyiKerZnNE9fZz17221yW//EYYx6fx6Qe1RS8MZ8l781m3g0Ny7oofjOZCa9MZFDDa6JU6IaCuWc1C3YNYt7cedirClj3+2fJ/f+eIMMW6sjEF8VvzmfVqcnMfzIZynJZ8dRa+F+TSY5puaSDMY9nMarJI0Ub55B9aeMev/t4AbHXzuDha3ueeSxam+V5F8DlQyYu0pk9fyqj+gc/oMim3Aab96/b2eCcxqRUC4ZhwXHTRNL3bKfADHVk4cCL9YqHyLrbiT0G6OEgfXhPCsvcoQ4s7LnefZ69w6cxtl+oIwkPrl0bcN83jYx4A8Ow4rxjApaN+RS3sqxZXIQ1cyzO3gYYFlJvGIvxSSFneq1Zxclf9SHBYsHS8N9ZOzHS+Vzkb3IzeUoGCTEGRm8nEzMtvLGrtVaVLsdbwPb1v2ba3alYDAOLYwwT0/eyvY3J1Wi6vdUcIv/TUUwY1qQocH+D3W5vXMZi0YG7EAigKLCSeq0uv+gcym2wlR3dT7ojockjDgZdnU1xWchCCiMG1l86sDaMzDUu9u46SXK8zueek/JcVu9KZ2prR7mlFSalxVXN+12PJIbYCyn99uylLVdPZcbwxmW9ZUWc/GUCZx6pdFN6cRWF761l9Qtr2fqJC2+nxi+tMkspqkwmoXfjQ5akIfQpLEWHHcJAWRH7hztoNrsmO8kuLu3wpcUfrMM9fgzJZ/b6vVSdLMb79V6yX1jN6k25FFd2RtDSEX3RWLo176liel7aosyKCk0s4auQV8eNY9wdC9h/1cPcpkuHzoGb/L/kkDxlDCoJfOXFrOxDz2bXKhu+bcfluSx/zstdNzTttFacybFYksdy122/xrtzLnPf1NHp8+6UibtfT5odYtAeSfiorqbYGtv8IGbUTzp+XeVecvLSuWtk05avxkiYSkJsT5yTJjCm9zFWzX6evSoMzjttgtKtGVFWqk7pOOC5SWbSpk1s2riQUV8v55l3XaEOKGyZe9aRnTRN38nwU3R0Fd5qP1/kKeTVZdtJenwqzqYFRV8nEyaOwRlvxWJL5rZ/mkHS+r0UBjNg6VhUNNFVXp2lCVcXgvWU/+1X/N5avJk30nwItJB8wwRuG5mM3WIlYfhkZtxRyta/6pzR+aaiQLo1e/90ilxNBxY3rqOpWFu9Q4K05DWbDPlRVpw3j8X4a6FO7wekmOz/yMe9ZxVzn5zL3CeXsO6zQ6z797lkfxHq2LoyK32Sqih1Nd39cOH6zN72duwpZsPitTDlCSY4WlyZ7DVp2q2JiaWn+yRV+p7R+WXtQ1JlKc2a9YSLQzarLqENB/0cpBe7ms0FblcRqdZ2Li+tzGdz7ijGDj+7hZvNNUBsbCzfuDXTnG8qCqRbsyQ7Sdiac+aLxd7DW9l8ahSpfUMbV3hwkbt4Lq8ebhysXZ99TFFfTdqBcXDHihUszsoi6/Essh6fwYTkQUz4bRZjdOagXY6UURx6NxdX/Z2aXDs2k3ODs/6aZC+uw4W4ztytzUXuc6uoGpfFpIFnf1XRW/gGs/4jl4bzXWbBx3x4QyqD1KnPMwdDrj9E9o6GlnCRm53DmKt0Z7iwYBmEs18OOQ2Tq6eQrW95GZVaf2FkZSkFR5tX2sXvrYO7W54lgLPmmhoXe3d9w6gUDYznW+C/UyASDno4mfpgMQv+dSbr7BbcPw5i6qxJup7bJ3bGzLyL1ctmMhM71tNu3D+fwLzpqZq0A2Q0u6OGFyMKsFjq/i9tc0xg9jXLmftQDvafmbh73EbWrPp+6C0kZ+ESeHwNk68A9861LPmgAD64m1fPrGAyizdNIhkwUu8i6+izzP3t5rp1RY9ixqPpKnRDwHH7bDKWzWX6Fjt2j5vYm7J4LFWjS3iw4HxgGsVPz2LmOjuW72DQlIeZVH/AzbVnNXP3jGJlVkbdfNtwlmBZa1ta87nG8p2bhDueYIZqgvPugtra2loAt7sCqzUu1PF0S8pt5/A3r17Ti2HRhAMB9EmviTdKO69t0TYeGP/z5sX0GEG6hWgw1yVNaXwJb363n8fEGxOkW4gGc11yhq9tqjMFEjFUEJwDQ4O0dAXB3IlXQdBlaHwJb8HciVdBEFL6ToGIiIiISIRTUSAiIiIiEuFUFIiIiIiIRDgVBSIiIiIiEU5FgYiIiIhIhFNRICIiIiIS4VQUiIiIiIhEOBUFIiIiIiIRTkWBiIiIiEiEu6C2trYW4Pjxb4iJ0c87dgaPx6PcdgLlNXDKXXApn4FR3rontWt4U/t1Px6Ph759+3S43EUN/4iJicFqjevUoCKV212h3HYC5TVwyl1wKZ+BUd66J7VreFP7dT9ud4VPy+nyIRERERGRCKeiQEREREQkwqkoEBERERGJcCoKREREREQinIoCEREREZEIp6JARERERCTCqSgQEREREYlwKgpERERERCLcRR0vcjbzaD7ZG7PZ/20sCRl3MfkGB5ZgRxbJKgvZ8MJa3MNnM/Vqa6ijCX+VhWS/tIH8r6uwXzmZaROT1V995qV010Y2vL0fV2wSN957H1j19P0AABTzSURBVBnxRqiDCn/axgPjY968ZflsXF83RyWNuY/7RiZg+PG8nC8aX8KbSeGmtWzYU0pVbyeTH5hAco9QxyTnwv8zBd/msvyPRTjuncfCJx8m/dsVLHnP1QmhRSIvrp2rWfCHfFyeAlyeUMfTHRSz4alVmNc8zLwnsxhr2cz8lwrxhjqsMGHuWc2CXXYmzJ3HvGmpHFv6LLnloY4qnGkbD4wfeavcy+qn8rFnzmPekzNI/Wo5z37g8v15OW80voS34jfns8pM5+En55E1LpbNT62lUGNaWPO7KDCLi7BmjsXZ2wDDQuoNYzE+KcTdGdFFHBMX6cyeP5VR/UMdS/fg/et2NjinMSnVgmFYcNw0kfQ92ykwQx1ZOHCRv8nN5CkZJMQYGL2dTMy08Mau4lAHFsa0jQfG97y5dm3Afd80MuINDMOK844JWDbmU+zj83K+aHwJa94Ctq//NdPuTsViGFgcY5iYvpftmlzDmt9FgeXqqcwY3nja1ltWxMlfJqAT4MFgJfVaXdoSTGVH95PuSGjyiINBV2dTXBaykMKHWUpRZTIJvRsfsiQNoU9hqQ4CBEzbeGB8zZtJaXEVyfFNZqQeSQyxF1L6rS/Py3mj8SW8lRWxf7iDZrNrspPs4tKQhSTn7ty+aFyey/LnvNx1gyNI4YgEl/dUMT0vbbErERWaWMLOKRN3v57NC37dmkC6NC9mZR96Nruu2WiyzXf0vJw3Gl/CW3U1xdbY5oV61E9CFY0ESeCboKeQV5dtJ+nxqTj1xRLpoowoK1Wn9A2CgERFE13l1fcvJKxER1fhrQ78eTlPNL6EtwvBekrt190EVhR4itmweC1MeYIJDt0pQLoue/90ilxNT0a7cR1NxapCtmPWPiRVluJqOuqfcHHIZtXlL9JFWemTVEVps07rwvWZvX6b7+h5OW80voS3fg7Si13NLvVyu4pItepi8nAWQFHgIve5VVSNy2LSQBUE0rVZkp0kbM0588Vi7+GtbD41itS+oY0rPDgYcv0hsnc03JnFRW52DmOuStbtG6UL8eI6XHjmjkSOlFEcejcXV03d364dm8m5wUmy4dvzcr5ofAlrlkE4++WQ0zC5egrZ+paXUan20MYl58Tv3ylw71zLkg8K4IO7efXMo5NZvGkSyUENTSQIejiZ+mAxC/51JuvsFtw/DmLqrElo2PKN4/bZZCyby/QtduweN7E3ZfFYqqZs6UK8heQsXAKPr2HyFYBjArOvWc7ch3Kw/8zE3eM2smalNu5odvS8nDcaX8KZBecD0yh+ehYz19mxfAeDpjzMJB1wC2sX1NbW1gK43RVYrXGhjqdbUm47h7959ZpeDIsmHAigT3pNvFEWDH0hs1XaxgPTuXnzYnoMLDGBPi+B0vgS3vxuP4+JN8aiwroL87VNA/pFY5FwpILgHBga8CXcdLTDr4Kgy9D4Et5UEHQbugGYiIiIiEiEU1EgIiIiIhLhVBSIiIiIiEQ4FQUiIiIiIhFORYGIiIiISIRTUSAiIiIiEuFUFIiIiIiIRDgVBSIiIiIiEe7MLxofP/4NMTH6JZfO4PF4lNtOoLwGTrkLLuUzMMpb96R2DW9qv+7H4/HQt2+fDpc784vGMTExnfhz85HN758MF58or4FT7oJL+QyM8tY9qV3Dm9qv+3G7K3xaTpcPiYiIiIhEOBUFIiIiIiIRTkWBiIiIiEiEU1EgIiIiIhLhVBSIiIiIiEQ4FQUiIiIiIhFORYGIiIiISIRTUSAiIiIiEuEu6niRs5lH88nemM3+b8F+5WSmTUzGEuzIIlllIRteWIt7+GymXm0NdTThz11A9ivZ5H9dhf0fJnDX/3Bijwp1UOHCS+mujWx4ez+u2CRuvPc+MuKNUAcV1jR+BsZbls/G9dns/zaWpDH3cd/IBNrqiR3lWG3QVfg3vpifZ7N2Uz6lVXaG3DqBO4Y37QNN15VA+rjJ3HaFWrVzmRRuWsuGPaVU9XYy+YEJJPcIdUxyTmrrnTzprvVJ0Ru1sx9bU/ux63Rt7Q9VtUXZi2pnbyzy7bURyufc1p6u/SZvVe38eatqV/x+bO2i3JOdGle48ymvpw/VrntscW1OUVVt7Q+na4+9v7h22p/3157u/PC6NF/7ZNXuFbXTnt1ee+zU6drT//1x7ZrHFtVud3VycGFI42dgfM7bdx/Xrnhwce320tO1p0+frP34z7NrF73/TevLdpRjtUGn64zx5fSRdbVZz+bUFn13uvZ01bHa7c9Oq12xu6rxPfMW12b9+ePab06drj19cn/tusdm176hZg2Ir+1XtHF27eyX99dWnT5dW1WUU7v4sTW1h051cnASEF/b1O/Lh7w9hvDQ43fhtBkQZcFxdTo9C0txd0bFEnFMXKQze/5URvUPdSzdRFkRh0ZMZIzDAlEGCdeN5cbXCikOdVxhwUX+JjeTp2SQEGNg9HYyMdPCG7uUvUBp/AyMa9cG3PdNIyPewDCsOO+YgGVjfqvbcUc5Vht0Ff6MLyYF7x1i1D1jcPQwMCwJZEyfR3pvL14Aitm+0c5ddzuxxxgY1lTuePwhBl1snt+PFEm8BWxf/2um3Z2KxTCwOMYwMX0v2wuU83Dmd1Fg9Hbg6N14ws7113xOJiegi1yCwUrqtTqNHVSO25h3u6Px7+OlHLohnoTQRRQ+zFKKKpNJ6N34kCVpCH20AxUwjZ+BMCktriI5vkmWeiQxxF5I6bdnL91RjtUGXYRf40spxbsGkWApJnfTWta+lE2+uyepDmvd5UPfllLYLwl7xV62vrSWtW9spTDKQXJfzaadpqyI/cMdzeZSR7KT7OLSkIUk5y7gLxoXvjKOcePGsaDAycO3ODp+gUioeQp59dkcUselq/DyxSkTd7+ezXeWdGuCoND46Q8vZmUfeja7VtmADr4X1FGO1QYh5s/44j7JMevHrF2xF8uVYxmbYaXoj/N59Qtv/fMuDn29gdUb3SSMmciN/+Al53dLyG+laJQgqa6m2BrbfC6N+kmoopEgCXiKT757E5s2bWJhxjcsX7QVVzCjEgk6F7nPreXkpNlM0Pzvm6hooqsaTs9LMGn89E90dBXeav9e01GO1QYh5s/4YhjEFscyasoknPFWrPHpTJ6ZzofvFWDWr8tansptvxlDss2CfeBtPPSAwboPdaljp7kQrKc0P3Q3/hcFHhNvTeOf1qtuZKyxl0JV5NJludj7X8v52PkwM662hzqY8GHtQ1JlKa6mo/4JF4dsVp1pCZTGzwBY6ZNURWmzjujC9Zkda2t3Oukox2qDrsGf8cXSk57WBOy9mjwW25M+lWbdTqnVTkL/Ptib3LjIcqmV4lPaZe00/RykF7uaXerldhWRatWFeOHM76LAtXMJc18rbKwOywv5+Kv41gdnkZDzUvjKcrb3n8HDI1UQ+MfBkOsPkb2j4Riqi9zsHMZcldzmrSClfRo/A+NIGcWhd3Nx1e/Mu3ZsJucGJ8kGgBfX4UJcnvrnOsix2qCr6Gh8MSktKK07E4ADZ2YROR82LGtS+P5mvFcl111+ZB3CqH7byf28/kuuNS5yt+Qy6QqdFu40lkE4++WQ0/DFYk8hW9/yMipV82w4u6C2trYWwO2uwGqN8+ElLvauXM7zfwW71YvbncCEx2cwRvcub5PvuW1U+Mo4NvRbwxMjVXW3xae8fvEqUx5d2+KLa+nMfuEJMnq38ZoI4M/2nr9sLqtL7dg9bmJvms1j4xwqClrQ+BkYf8ZG187lzF1Tiv1nJu4etzF71m04YgBvAWunL4HH1zD5Cug4x2qDzhaU8eX4Vhb8di+jVj1Bhg3wlLJ15TNs+MqK1eOCX83g4elOzuyCVhbw6h+eZ3u1Fct/u+g5LktjVYB8bj9PIa8+vYTt2LF8B4OmPMxUnY3vknxt0wCKggZeTI+BJSbQECNHIEWBdEx5DZzfufOaeKMsGPrRt1Zp/AxM5+ato2XVBp2lU8eXGi9ejLaX9XrxGoaKgXPgd/t5TLwxFuW8C/O1TQP6ReM6GkxFIoahAT+4NH4Gxp+8dbSs2qDL8Gd8iepgh18FwfmngqDb0A0GRUREREQinIoCEREREZEIp6JARERERCTCqSgQEREREYlwKgpERERERCKcigIRERERkQinokBEREREJMKpKBARERERiXBnftH4+PFviInRL7l0Bo/Ho9x2AuU1cMpdcCmfgVHeuie1a3hT+3U/Ho+Hvn37dLjcmV80jomJ8fPn5sVXfv9kuPhEeQ2cchdcymdglLfuSe0a3tR+3Y/bXeHTcrp8SEREREQkwqkoEBERERGJcCoKREREREQinIoCEREREZEIp6JARERERCTCqSgQEREREYlwKgpERERERCKcigIRERERkQh3UceLtKM4m2decDHm0ak4rUGKSKCykA0vrMU9fDZTr1Ziz1llIdkvbSD/6yrsV05m2sRkLKGOKWyYlO7KZsPb+3HFJjBq4mTG/FLZOxfm0XyyN2az/1vUH/3gLctn4/ps9n8bS9KY+7hvZAJGG8t2lGO1QVfhpXTXxvrxJYkb772PjPi2WlW6HpPCTWvZsKeUqt5OJj8wgeQe7Sz9eTZrN+VTWmVnyK0TuGN40224aV9IIH3cZG67Qlvl+XYOZwpcbH1lOycrXZg1wQsosnlx7VzNgj/k4/IU4PKEOp7uoJgNT63CvOZh5j2ZxVjLZua/VIg31GGFCfeO5Sz/wsHkJxcy76F03P+5hK3HQx1VGCvewPw/FuG4Zx4L52cxNmYD898sDnVUXV/lXlY/lY89cx7znpxB6lfLefYDV+vLdpRjtUGXYe5ZzYJddibMnce8aakcW/osueWhjkp8VfzmfFaZ6Tz85DyyxsWy+am1FLax3+L94lUWbTG48aF5zHt8AvZdC1i9xzzzvHvncp7/wsFdc+cx77fpmKvns0Gb5XkXcFHgevd59g6fxth+wQwn0pm4SGf2/KmM6h/qWLoH71+3s8E5jUmpFgzDguOmiaTv2U6B2fFrxaSoyMqEW51YDTB6pDJmvMHeL9yhDixseXsM4aHH78JpMyDKguPqdHoWlqKMts+1awPu+6aREW9gGFacd0zAsjGf1vYZOsqx2qCrcJG/yc3kKRkkxBgYvZ1MzLTwxi7tCYYFbwHb1/+aaXenYjEMLI4xTEzfy/ZWJ1eTgvcOMeqeMTh6GBiWBDKmzyO9t7f+AF0x2zfauetuJ/YYA8Oayh2PP8SgizVRn2+BFQXluazelc7U6+xBDifSWUm9Vqexg6ns6H7SHQlNHnEw6OpsistCFlIYseB8YAbpvRv+9lJadJLkeF3SFiijtwNH78YT5q6/5nMyOQFltD0mpcVVzftdjySG2Asp/fbspTvKsdqgizBLKapMJqF340OWpCH0UYEWHsqK2D/cQbPZNdlJdnFpKwuXUrxrEAmWYnI3rWXtS9nku3uS6rDWXT70bSmF/ZKwV+xl60trWfvGVgqjHCT31d7Q+RZAUeAm/y85JE8Zg0oC6eq8p4rpeWmLgSUqNLGEO9eO5Sz33MWNjlBHEv4KXxnHuHHjWFDg5OFblND2eTEr+9Cz2bXKRofbcUc5VhuE2CkTd7+ezYsx3fokfFRXU2yNbX4QM+onrS/rPskx68esXbEXy5VjGZthpeiP83n1i/oLed0uDn29gdUb3SSMmciN/+Al53dLyG+l6JfO5fcmaO5ZR3bSNCZoDJUwYERZqTqlbxCcK+/hV1mem0TWA06dyQqC5Ls3sWnTJhZmfMPyRVtp4+p4qRcdXYW32r/XdJRjtUGIRUUTXeXV97vC1YVgPeVj+xkGscWxjJoyCWe8FWt8OpNnpvPhewWYAFHRWMtTue03Y0i2WbAPvI2HHjBY96EuJTvf/CwKisn+j3zce1Yx98m5zH1yCes+O8S6f59L9hedE6DIubD3T6fI1fRktBvX0VSs7dwhQZrzFm/gmTUwec4EHDGhjibMeUy8TW7MYL3qRsYaeynUEbF2WOmTVEWpq+nuhwvXZ/bWt+OOcqw26BqsfUiqLKVZs55wcchm1YGHcNDPQXqxq9mlXm5XEanWVi7Es/SkpzUBe68mj8X2pE+lWVdUWO0k9O+DvcmNpyyXWinWAb3zzs+iwMEdK1awOCuLrMezyHp8BhOSBzHht1mM0ZkD6YIsyU4Stuac+WKx9/BWNp8aRWrf0MYVNspzWb6qigm/m0SyCoJz5tq5hLmvNbn7VXkhH38VryK1A46UURx6NxdX/c68a8dmcm5wkmwAeHEdLjxzt7aOcqw26CocDLn+ENk7Gs7RuMjNzmHMVclt3mpWuhDLIJz9cshpmFw9hWx9y8uo1PoLyytLKTja8EVhB87MInI+bGhrk8L3N+O9Krnu8jHrEEb1207u5/XL17jI3ZLLpCu0Y3m++f07BYbF0uy+skYUYLHU/V+kq+nhZOqDxSz415mss1tw/ziIqbMm6fswPnGT/5cl5BZA7r2vNj5872I23Z0curDCmP2GGdy1cjkz/wXsVi9udwIT/m0GqdoLap9jArOvWc7ch3Kw/8zE3eM2smal1s1F3kJyFi6Bx9cw+YqOc6w26Doct88mY9lcpm+xY/e4ib0pi8fUEGHCgvOBaRQ/PYuZ6+xYvoNBUx5mUv0BN9ee1czdM4qVWRnYAcdNMxiycgEzs61YPS741QwevtveuK5776L4D3OYWW3F8t8ueo5TXwiFC2pra2sB3O4KrNa4UMfTLSm3ncPfvHpNL4ZFgwyoTwab//n0YnoMLBF+9qVz89bRsmqDzuJ3u3pNvFE6uNhV+N1+HhNvjMW3Mzw1XrwYbbe114vXMHS2KMh8bdNz+0VjkTCigkC6Du2MBsafvHW0rNqgyzB83KGUrsnXggAgqoMdfhUEIaUbgImIiIiIRDgVBSIiIiIiEU5FgYiIiIhIhFNRICIiIiIS4VQUiIiIiIhEOBUFIiIiIiIRTkWBiIiIiEiEU1EgIiIiIhLhVBSIiIiIiES4C2pra2sBjh//hpgY/bxjZ/B4PMptJ1BeA6fcBZfyGRjlrXtSu4Y3tV/34/F46Nu3T4fLnSkKREREREQkMunyIRERERGRCKeiQEREREQkwqkoEBERERGJcCoKREREREQinIoCEREREZEIp6JARERERCTCqSgQEREREYlw/w86AFDONFB0BQAAAABJRU5ErkJggg=="}}},{"cell_type":"markdown","source":"First we one-hot encode a dataframe with noly the categorical columns(dtype == 'object')","metadata":{}},{"cell_type":"code","source":"categorical = pd.get_dummies(bureau.select_dtypes('object'))\ncategorical['SK_ID_CURR'] = bureau['SK_ID_CURR']\ncategorical.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:20.404643Z","iopub.execute_input":"2022-07-29T05:01:20.405065Z","iopub.status.idle":"2022-07-29T05:01:21.394165Z","shell.execute_reply.started":"2022-07-29T05:01:20.405033Z","shell.execute_reply":"2022-07-29T05:01:21.393139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_grouped = categorical.groupby('SK_ID_CURR').agg(['sum', 'mean'])\ncategorical_grouped.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:21.395300Z","iopub.execute_input":"2022-07-29T05:01:21.395616Z","iopub.status.idle":"2022-07-29T05:01:23.299337Z","shell.execute_reply.started":"2022-07-29T05:01:21.395587Z","shell.execute_reply":"2022-07-29T05:01:23.298224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The sum columns represent the count of that category for the associated client and the mean represent the normalized count. One-hot encoding makes the process of calculating these figures very easy!\n\nWe can use a similar function as befroe to rename the columns. Again, we have to deal with the multi-level index for the columns. We iterate through the first level (level 0) which is the name of the categorical variable appended with the value of the category (from one-hot encoding). Then we iterate stats we calculated for each client. We will rename the column with the level 0 name appende with the stat. As an example, the column with CREDIT_ACTIVE_Active as level 0 and sum as level 1 will become CREDDIT_ACTIVE_Active_count.","metadata":{}},{"cell_type":"code","source":"categorical_grouped.columns.levels[0][:10]","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:23.301049Z","iopub.execute_input":"2022-07-29T05:01:23.301594Z","iopub.status.idle":"2022-07-29T05:01:23.309772Z","shell.execute_reply.started":"2022-07-29T05:01:23.301551Z","shell.execute_reply":"2022-07-29T05:01:23.308719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_grouped.columns.levels[1]","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:23.311421Z","iopub.execute_input":"2022-07-29T05:01:23.311843Z","iopub.status.idle":"2022-07-29T05:01:23.324884Z","shell.execute_reply.started":"2022-07-29T05:01:23.311809Z","shell.execute_reply":"2022-07-29T05:01:23.324153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"group_var = 'SK_ID_CURR'\n\n# Need to create new column names\ncolumns = []\n\n# Iterate through the variable names\nfor var in categorical_grouped.columns.levels[0]:\n    # Skip the grouping variable\n    if var != group_var:\n        # Iterate through the stat names\n        for stat in ['count', 'count_norm']:\n            columns.append('%s_%s' % (var, stat))\n\n# Rename the columns\ncategorical_grouped.columns = columns\n\ncategorical_grouped.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:23.326142Z","iopub.execute_input":"2022-07-29T05:01:23.326918Z","iopub.status.idle":"2022-07-29T05:01:23.365452Z","shell.execute_reply.started":"2022-07-29T05:01:23.326882Z","shell.execute_reply":"2022-07-29T05:01:23.364467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The sum column records the counts and the mean column records the normalized count.\n\nWe can merge this dataframe into the training data.","metadata":{}},{"cell_type":"code","source":"train = train.merge(categorical_grouped, left_on = 'SK_ID_CURR', right_index = True, how = 'left')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:23.366890Z","iopub.execute_input":"2022-07-29T05:01:23.367258Z","iopub.status.idle":"2022-07-29T05:01:26.076587Z","shell.execute_reply.started":"2022-07-29T05:01:23.367226Z","shell.execute_reply":"2022-07-29T05:01:26.075496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:26.077767Z","iopub.execute_input":"2022-07-29T05:01:26.078074Z","iopub.status.idle":"2022-07-29T05:01:26.085252Z","shell.execute_reply.started":"2022-07-29T05:01:26.078045Z","shell.execute_reply":"2022-07-29T05:01:26.084213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.iloc[:10, 123:]","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:26.087145Z","iopub.execute_input":"2022-07-29T05:01:26.087611Z","iopub.status.idle":"2022-07-29T05:01:26.143760Z","shell.execute_reply.started":"2022-07-29T05:01:26.087565Z","shell.execute_reply":"2022-07-29T05:01:26.142619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Function to Handle Categorical Variables\n\nTo make the code more efficient, we can now write a function to handle the categorical variables for us. This will take the same form as the agg_numeric function in that it accepts a dataframe and a grouping variable. Then it will calculate the counts and normalized counts of each category for all categorical variables in the dataframe.","metadata":{}},{"cell_type":"code","source":"'''Computes counts and normalized counts for each observation\nof 'group_var' of each unique category in every categorical variable\n\nParameter\n----------\ndf : dataframe\n    The dataframe to calculate the value counts for\n    \ngroup_var : string\n    The variable by which to group the dataframe. For each unique\n    value of this variable, the final dataframe will have one row\n    \ndf_name : string\n    Variable added to the front of column names to keep track of columns\n    \n    \nReturn\n----------\ncategorical : dataframe\n    A dataframe with counts and normalized counts of each unique category\n    in every categorical variable\n    with one row for every unique value of the 'group_var.'\n\n'''","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:26.152691Z","iopub.execute_input":"2022-07-29T05:01:26.153278Z","iopub.status.idle":"2022-07-29T05:01:26.161647Z","shell.execute_reply.started":"2022-07-29T05:01:26.153228Z","shell.execute_reply":"2022-07-29T05:01:26.160508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def count_categorical(df, group_var, df_name):\n    # Select the categorical columns\n    categorical = pd.get_dummies(df.select_dtypes('object'))\n    \n    # Make sure to put the identifying id on the column\n    categorical[group_var] = df[group_var]\n    \n    # Groupby the group var and calculate the sum and mean\n    categorical = categorical.groupby(group_var).agg(['sum', 'mean'])\n    \n    column_names = []\n    \n    # Iterate through the columns in level 0\n    for var in categorical.columns.levels[0]:\n        # Iterate through the stats in level 1\n        for stat in ['count', 'count_norm']:\n            # Make a new column name\n            column_names.append('%s_%s_%s' % (df_name, var, stat))\n            \n    categorical.columns = column_names\n    \n    return categorical","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:26.163200Z","iopub.execute_input":"2022-07-29T05:01:26.163555Z","iopub.status.idle":"2022-07-29T05:01:26.172411Z","shell.execute_reply.started":"2022-07-29T05:01:26.163514Z","shell.execute_reply":"2022-07-29T05:01:26.171302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bureau_counts = count_categorical(bureau, group_var = 'SK_ID_CURR', df_name = 'bureau')\nbureau_counts.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:26.174115Z","iopub.execute_input":"2022-07-29T05:01:26.174781Z","iopub.status.idle":"2022-07-29T05:01:29.205886Z","shell.execute_reply.started":"2022-07-29T05:01:26.174735Z","shell.execute_reply":"2022-07-29T05:01:29.204812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Applying Operations to another dataframe\n\nWe will now turn to the bureau balance dataframe. This dataframe has monthly information about ech client's previous loans with other financial institutions. Instead of grouping this dataframe by the SK_ID_CURR which is the client id, we will first group the dataframe by the SK_ID_BUREAU which is the id of the previous loan. This will give us one row of the dataframe for each loan. Then, we can groupby the SK_ID_CURR and calculate the aggregations across the loans of each client. The final result will be a dataframe with one row fore each client, with stats calculated for therir loans.","metadata":{}},{"cell_type":"code","source":"# Read in bureau balance\nbureau_balance = pd.read_csv('../input/home-credit-default-risk/bureau_balance.csv')\nbureau_balance.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:29.207350Z","iopub.execute_input":"2022-07-29T05:01:29.207864Z","iopub.status.idle":"2022-07-29T05:01:41.401134Z","shell.execute_reply.started":"2022-07-29T05:01:29.207829Z","shell.execute_reply":"2022-07-29T05:01:41.398822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First, we can calculate the value counts of each status for each loan. Fortunately, we already have a function that does this for us!","metadata":{}},{"cell_type":"code","source":"# Counts of each of status for each previous loan\nbureau_balance_counts = count_categorical(bureau_balance, group_var = 'SK_ID_BUREAU', df_name = 'bureau_balance')\nbureau_balance_counts.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:41.402672Z","iopub.execute_input":"2022-07-29T05:01:41.403329Z","iopub.status.idle":"2022-07-29T05:01:54.202801Z","shell.execute_reply.started":"2022-07-29T05:01:41.403284Z","shell.execute_reply":"2022-07-29T05:01:54.201535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can handle the one numeric column. The MONTHS_BALANCE column has the \"months of balance relative to application data.\" This might not necessarily be that important as a numeric variable, and in future work we might want to consider this as a time vairable. For now, we can just calculate the same aggregation statistics as previously.","metadata":{}},{"cell_type":"code","source":"# Calculate value count statistics for each 'SK_ID_CURR'\nbureau_balance_agg = agg_numeric(bureau_balance, group_var = 'SK_ID_BUREAU', df_name = 'bureau_balance')\nbureau_balance_agg","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:54.204799Z","iopub.execute_input":"2022-07-29T05:01:54.205252Z","iopub.status.idle":"2022-07-29T05:01:56.649472Z","shell.execute_reply.started":"2022-07-29T05:01:54.205195Z","shell.execute_reply":"2022-07-29T05:01:56.648163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above dataframe have the calculations done on each loan. Now we need  to aggregate these for each client. We can do this by merging the dataframes together first and then since all the variables are numeric, we just need to aggregate the statistics again, this time grouping by the SK_ID_CURR.","metadata":{}},{"cell_type":"code","source":"# Dataframe grouped by the loan\nbureau_by_loan = bureau_balance_agg.merge(bureau_balance_counts, right_index = True,\n                                          left_on = 'SK_ID_BUREAU', how = 'outer')\n\ndisplay(bureau_by_loan.head())\n\n# Merger to include the SK_ID_CURR\nbureau_by_loan = bureau_by_loan.merge(bureau[['SK_ID_BUREAU', 'SK_ID_CURR']],\n                                     on = 'SK_ID_BUREAU', how = 'left')\n\nbureau_by_loan.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:56.651063Z","iopub.execute_input":"2022-07-29T05:01:56.651399Z","iopub.status.idle":"2022-07-29T05:01:58.218142Z","shell.execute_reply.started":"2022-07-29T05:01:56.651369Z","shell.execute_reply":"2022-07-29T05:01:58.216970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bureau_balance_by_client = agg_numeric(bureau_by_loan.drop(columns = ['SK_ID_BUREAU']), \n                                      group_var = 'SK_ID_CURR', df_name = 'client')\nbureau_balance_by_client.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:58.219621Z","iopub.execute_input":"2022-07-29T05:01:58.219940Z","iopub.status.idle":"2022-07-29T05:01:59.951593Z","shell.execute_reply.started":"2022-07-29T05:01:58.219911Z","shell.execute_reply":"2022-07-29T05:01:59.950508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To recap, for the bureau_balance dataframe we:\n\n1. Calculated numeric stats grouping by each loan\n2. Made value counts of each categorical variable grouping by loan\n3. Merged the stats and the value counts on the loans\n4. Calculated numeric stats for the resulting dataframe grouping by the client id\n\nThe final resulting dataframe has one row for each client, with statistics calculated for all of their loans with monthly balance information.\n\nSome of these variables are a little confusing, so let's try to explain a few:\n\n- client_bureau_balance_MONTHS_BALANCE_mean_mean : For each loan calculate the mean value of MONTHS_BALANCE. Then for each client, calculate the mean of this value for all of their loan.\n- client_bureau_balance_STATUS_X_count_norm_sum : For each loan, calculate the number of occurences of STATUS == X divided by the number of total STATUS values for the loan. Then, for each client, add up the values for each loan.","metadata":{}},{"cell_type":"markdown","source":"we will hold off on calculating the correlations until we have all the variable together in one dataframe","metadata":{}},{"cell_type":"markdown","source":"# Putting the Functions Together\n\nWe now have all the pieces in place to take the information from the previous loans at other institutions and the monthly payment information about these loans and put them into the main training dataframe. Let's do a reset of all the variables and then use the functions we built to do this from the ground up. This demonstrate the benefit of using functions for repeatable workflows!","metadata":{}},{"cell_type":"code","source":"# Free up memory by deleting old objects\nimport gc\ngc.enable()\ndel train, bureau, bureau_balance, bureau_agg, bureau_agg_new, bureau_balance_agg,\\\nbureau_balance_counts, bureau_by_loan, bureau_balance_by_client, bureau_counts\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:01:59.953396Z","iopub.execute_input":"2022-07-29T05:01:59.953770Z","iopub.status.idle":"2022-07-29T05:02:00.145840Z","shell.execute_reply.started":"2022-07-29T05:01:59.953738Z","shell.execute_reply":"2022-07-29T05:02:00.144343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read in new copies of all the data dataframe\ntrain = pd.read_csv('../input/home-credit-default-risk/application_train.csv')\nbureau = pd.read_csv('../input/home-credit-default-risk/bureau.csv')\nbureau_balance = pd.read_csv('../input/home-credit-default-risk/bureau_balance.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:00.147978Z","iopub.execute_input":"2022-07-29T05:02:00.148417Z","iopub.status.idle":"2022-07-29T05:02:17.731025Z","shell.execute_reply.started":"2022-07-29T05:02:00.148374Z","shell.execute_reply":"2022-07-29T05:02:17.729580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Counts of Bureau Dataframe","metadata":{}},{"cell_type":"code","source":"bureau_counts = count_categorical(bureau, group_var = 'SK_ID_CURR', df_name = 'bureau')\nbureau_counts.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:17.732557Z","iopub.execute_input":"2022-07-29T05:02:17.733008Z","iopub.status.idle":"2022-07-29T05:02:20.414996Z","shell.execute_reply.started":"2022-07-29T05:02:17.732961Z","shell.execute_reply":"2022-07-29T05:02:20.413672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Aggergated Stats of Bureau Dataframe","metadata":{}},{"cell_type":"code","source":"bureau_agg = agg_numeric(bureau.drop(columns = ['SK_ID_BUREAU']), group_var = 'SK_ID_CURR',\n                        df_name = 'bureau')\nbureau_agg.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:20.416505Z","iopub.execute_input":"2022-07-29T05:02:20.416927Z","iopub.status.idle":"2022-07-29T05:02:22.884633Z","shell.execute_reply.started":"2022-07-29T05:02:20.416894Z","shell.execute_reply":"2022-07-29T05:02:22.883524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Value counts of Bureau Balance dataframe by loan","metadata":{}},{"cell_type":"code","source":"bureau_balance_counts = count_categorical(bureau_balance, group_var = 'SK_ID_BUREAU',\n                                         df_name = 'bureau_balance')\nbureau_balance_counts.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:22.885988Z","iopub.execute_input":"2022-07-29T05:02:22.886323Z","iopub.status.idle":"2022-07-29T05:02:35.391231Z","shell.execute_reply.started":"2022-07-29T05:02:22.886281Z","shell.execute_reply":"2022-07-29T05:02:35.390081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Aggregated stats of Bureau Balance dataframe by loan","metadata":{}},{"cell_type":"code","source":"bureau_balance_agg = agg_numeric(bureau_balance, group_var = 'SK_ID_BUREAU',\n                                df_name = 'bureau_balance')\nbureau_balance_agg.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:35.392761Z","iopub.execute_input":"2022-07-29T05:02:35.393755Z","iopub.status.idle":"2022-07-29T05:02:37.823140Z","shell.execute_reply.started":"2022-07-29T05:02:35.393709Z","shell.execute_reply":"2022-07-29T05:02:37.822044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Aggregated Stats of Bureau Balance by Client","metadata":{}},{"cell_type":"code","source":"# Dataframe grouped by the loan\nbureau_by_loan = bureau_balance_agg.merge(bureau_balance_counts,\n                                         right_index = True,\n                                         left_on = 'SK_ID_BUREAU', how = 'outer')\n# Merge to include the Sk_ID_CURR\nbureau_by_loan = bureau[['SK_ID_BUREAU', 'SK_ID_CURR']].merge(bureau_by_loan,\n                                                             on = 'SK_ID_BUREAU',\n                                                             how = 'left')\n# Aggregate the stats for each client\nbureau_balance_by_client = agg_numeric(bureau_by_loan.drop(columns = ['SK_ID_BUREAU']), \n                                      group_var = 'SK_ID_CURR', df_name = 'client')","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:37.824707Z","iopub.execute_input":"2022-07-29T05:02:37.825056Z","iopub.status.idle":"2022-07-29T05:02:44.130609Z","shell.execute_reply.started":"2022-07-29T05:02:37.825026Z","shell.execute_reply":"2022-07-29T05:02:44.129607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Insert Computed Feautres into Training Data","metadata":{}},{"cell_type":"code","source":"original_features = list(train.columns)\nprint('Original Number of Features: ', len(original_features))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:44.131921Z","iopub.execute_input":"2022-07-29T05:02:44.132254Z","iopub.status.idle":"2022-07-29T05:02:44.138098Z","shell.execute_reply.started":"2022-07-29T05:02:44.132224Z","shell.execute_reply":"2022-07-29T05:02:44.137264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge with the value counts of bureau\ntrain = train.merge(bureau_counts, on = 'SK_ID_CURR', how = 'left')\n\n# Merge with the stats of bureau\ntrain = train.merge(bureau_agg, on = 'SK_ID_CURR', how = 'left')\n\n# Merge with the monthly information grouped by client\ntrain = train.merge(bureau_balance_by_client, on = 'SK_ID_CURR', how = 'left')","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:44.139764Z","iopub.execute_input":"2022-07-29T05:02:44.140382Z","iopub.status.idle":"2022-07-29T05:02:55.580023Z","shell.execute_reply.started":"2022-07-29T05:02:44.140349Z","shell.execute_reply":"2022-07-29T05:02:55.578885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_features = list(train.columns)\nprint('Number of features using previous loans from other insititutions data: ', len(new_features))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:55.581249Z","iopub.execute_input":"2022-07-29T05:02:55.581578Z","iopub.status.idle":"2022-07-29T05:02:55.587712Z","shell.execute_reply.started":"2022-07-29T05:02:55.581548Z","shell.execute_reply":"2022-07-29T05:02:55.586420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineering Outcomes\n\nAfter all that work, now we want to take a look at the variables we have created. We can look at the percentage of missing values, the correlations of variablese withe the target, and alos the correlation of variables with other variables. The correlation of vairables with the other variables. The correlations between variables can show if we have collinear variables, that is, variables that are highly correlated with one another. Often, we want to remove one in a pair of collinear variables because having both variables would be redundant. We can also use the percentage of missing values to remove features with a substantial majority of values that are not present. **Feature selection** will be an important focus going forward, because reducing the number of features can help the model learn during training and also generalize better to the testing data. The \"curse of dimensionality\" is the name given to the issues caused by having too many features (too high of a dimension). As the number of variables increase, the number of datapoints needed to learn the relationship between these variables and the target value increases exponentially.\n\nFeautre selection is the process of removing variables to help our model to learn and generalize better to the testing set. The objective is to remove useless/redundant variables while preserving those that are useful. There are a number of tools we can use for this process, but in this notebook we will stick to removing columns with a high percentage of missing values and avriables that have a high correlation with one another. Later we can look at using the feature imporatnces returned from models such as the Gradient Boosting Machine or Random Forest to perform feature selection.","metadata":{}},{"cell_type":"markdown","source":"## Missing Values\n\nAn important consideration is the missing values in the dataframe. Columns with woo many missing values might have to be dropped.","metadata":{}},{"cell_type":"code","source":"# Function to claculate missing values by column# Funct\ndef missing_values_table(df):\n    # Total missing values\n    mis_val = df.isnull().sum()\n    \n    # Percentage of missing values\n    mis_val_percent = 100 * df.isnull().sum() / len(df)\n    \n    # Make a table with the results\n    mis_val_table = pd.concat([mis_val, mis_val_percent], axis = 1)\n    \n    # Rename the columns\n    mis_val_table_ren_columns = mis_val_table.rename(\n    columns = {0: 'Missing Values', 1: '% of Total Values'})\n    \n    # Sort the table by percentage of missing descending\n    mis_val_table_ren_columns = mis_val_table_ren_columns[\n        mis_val_table_ren_columns.iloc[:, 1] != 0].sort_values(\n    '% of Total Values', ascending = False).round(1)\n    \n    # Print some summary information\n    print('Your selected dataframe has ' + str(df.shape[1]) + ' columns.\\n'\n         'There are ' + str(mis_val_table_ren_columns.shape[0]) + \n         ' columns that have missing values.')\n    \n    # Return the dataframe with missing information\n    return mis_val_table_ren_columns","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:55.589669Z","iopub.execute_input":"2022-07-29T05:02:55.590052Z","iopub.status.idle":"2022-07-29T05:02:55.600324Z","shell.execute_reply.started":"2022-07-29T05:02:55.590021Z","shell.execute_reply":"2022-07-29T05:02:55.599437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_train = missing_values_table(train)\nmissing_train.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:55.602042Z","iopub.execute_input":"2022-07-29T05:02:55.602359Z","iopub.status.idle":"2022-07-29T05:02:57.007808Z","shell.execute_reply.started":"2022-07-29T05:02:55.602329Z","shell.execute_reply":"2022-07-29T05:02:57.006738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see there are a number of columns with a high percentage of missing values. There is no well-established threshold for removing missing values, and the best course of action depends on the problem. Here, to reduce the number of features, we will remove any columns in either the training or the testing data that have greater than 90% missing values.","metadata":{}},{"cell_type":"code","source":"missing_train_vars = list(missing_train.index[missing_train['% of Total Values'] > 90 ])\nlen(missing_train_vars)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:57.009320Z","iopub.execute_input":"2022-07-29T05:02:57.009688Z","iopub.status.idle":"2022-07-29T05:02:57.019319Z","shell.execute_reply.started":"2022-07-29T05:02:57.009655Z","shell.execute_reply":"2022-07-29T05:02:57.018125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Before we remove the missing values, we will find the missing value percentages in the testing data. We'll then remove any columns with greater than 90% missing values in either the training or testing data. Let's now read in the testing data, perform the same operations, and look at the missing values in the testing data. We already have calculated all the counts and aggreagation statistics, so we only need to merge the testing data with the appropriate data.","metadata":{}},{"cell_type":"markdown","source":"## Calculate Information for Testing Data","metadata":{}},{"cell_type":"code","source":"# Read in the test dataframe\ntest = pd.read_csv('../input/home-credit-default-risk/application_test.csv')\n\n# Merge with the value of counts of bureau\ntest = test.merge(bureau_counts, on = 'SK_ID_CURR', how = 'left')\n\n# Merge with the stats of bureau\ntest = test.merge(bureau_agg, on = 'SK_ID_CURR', how = 'left')\n\n# Merge with the value counts of bureau_balance\ntest = test.merge(bureau_balance_by_client, on = 'SK_ID_CURR', how = 'left')","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:57.020784Z","iopub.execute_input":"2022-07-29T05:02:57.021104Z","iopub.status.idle":"2022-07-29T05:02:59.961728Z","shell.execute_reply.started":"2022-07-29T05:02:57.021074Z","shell.execute_reply":"2022-07-29T05:02:59.960629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Shape of Testing Data: ', test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:59.963151Z","iopub.execute_input":"2022-07-29T05:02:59.963520Z","iopub.status.idle":"2022-07-29T05:02:59.969590Z","shell.execute_reply.started":"2022-07-29T05:02:59.963481Z","shell.execute_reply":"2022-07-29T05:02:59.967943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We need to align the testing and training dataframe, which means matching up the columns so they have  the exact same columns. This shouldn't be an issue here, but when we one-hot encode variables, we need to align the dataframe to make sure they have the same columns.","metadata":{}},{"cell_type":"code","source":"train_labels = train['TARGET']\n\n# Align the dataframe, this will remove the 'TARGET' column\ntrain, test = train.align(test, join = 'inner', axis = 1)\n\ntrain['TARGET'] = train_labels","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:02:59.971248Z","iopub.execute_input":"2022-07-29T05:02:59.971912Z","iopub.status.idle":"2022-07-29T05:03:00.992642Z","shell.execute_reply.started":"2022-07-29T05:02:59.971864Z","shell.execute_reply":"2022-07-29T05:03:00.991520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Training Data Shape: ', train.shape)\nprint('Testing Data Shape: ', test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:03:00.994223Z","iopub.execute_input":"2022-07-29T05:03:00.994805Z","iopub.status.idle":"2022-07-29T05:03:01.002334Z","shell.execute_reply.started":"2022-07-29T05:03:00.994757Z","shell.execute_reply":"2022-07-29T05:03:01.001103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dataframes now have the same columns (with the exception of the TARGET column in the training data). This means we can use them in a machine learning model which needs to see the same columns in both the training and testing dataframes.\n\nLet's now look at the percentage of missing values in the testing data so we cna figure out the columns that should be dropped.","metadata":{}},{"cell_type":"code","source":"missing_test = missing_values_table(test)\nmissing_test.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:03:01.004217Z","iopub.execute_input":"2022-07-29T05:03:01.004625Z","iopub.status.idle":"2022-07-29T05:03:01.249801Z","shell.execute_reply.started":"2022-07-29T05:03:01.004593Z","shell.execute_reply":"2022-07-29T05:03:01.248543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_test_vars = list(missing_test.index[missing_test['% of Total Values'] > 90])\nlen(missing_test_vars)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:03:01.251226Z","iopub.execute_input":"2022-07-29T05:03:01.252163Z","iopub.status.idle":"2022-07-29T05:03:01.260440Z","shell.execute_reply.started":"2022-07-29T05:03:01.252127Z","shell.execute_reply":"2022-07-29T05:03:01.259407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_columns = list(set(missing_test_vars + missing_train_vars))\nprint('There are %d columns with more that 90%% missing in either the training or testing data.' \n     % len(missing_columns))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:03:01.261944Z","iopub.execute_input":"2022-07-29T05:03:01.262423Z","iopub.status.idle":"2022-07-29T05:03:01.268958Z","shell.execute_reply.started":"2022-07-29T05:03:01.262390Z","shell.execute_reply":"2022-07-29T05:03:01.268048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop the missing columns\ntrain = train.drop(columns = missing_columns)\ntest = test.drop(columns = missing_columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:03:01.270238Z","iopub.execute_input":"2022-07-29T05:03:01.271013Z","iopub.status.idle":"2022-07-29T05:03:01.689835Z","shell.execute_reply.started":"2022-07-29T05:03:01.270942Z","shell.execute_reply":"2022-07-29T05:03:01.688655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We ended up removing no columns in this round because there are no columns with more than 90% missing values. We might have to apply another feature selection method to reduce the dimensionality.","metadata":{}},{"cell_type":"markdown","source":"At this point we will save bouth the training and testing data. I encourage anyone to try different percentages for dropping the missing columns and compare the outcomes.","metadata":{}},{"cell_type":"code","source":"train.to_csv('train_bureau_raw.csv', index = False)\ntest.to_csv('test_bureau_raw.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:03:01.691608Z","iopub.execute_input":"2022-07-29T05:03:01.692323Z","iopub.status.idle":"2022-07-29T05:04:23.586850Z","shell.execute_reply.started":"2022-07-29T05:03:01.692277Z","shell.execute_reply":"2022-07-29T05:04:23.585654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Correlation\n\nFirst let's look at the correlations of the variables with the target. We can see in any of the variables we created have a greater correlation than those already present in the training data (from application).","metadata":{}},{"cell_type":"code","source":"# Calculate all correlations in dataframe\ncorrs = train.corr()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:04:23.588306Z","iopub.execute_input":"2022-07-29T05:04:23.588807Z","iopub.status.idle":"2022-07-29T05:05:53.199910Z","shell.execute_reply.started":"2022-07-29T05:04:23.588772Z","shell.execute_reply":"2022-07-29T05:05:53.198887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corrs = corrs.sort_values('TARGET', ascending = False)\n\n# Ten most positibe correlations\npd.DataFrame(corrs['TARGET'].head(10))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:05:53.201340Z","iopub.execute_input":"2022-07-29T05:05:53.201685Z","iopub.status.idle":"2022-07-29T05:05:53.216234Z","shell.execute_reply.started":"2022-07-29T05:05:53.201654Z","shell.execute_reply":"2022-07-29T05:05:53.215090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Ten most negative correlations\npd.DataFrame(corrs['TARGET'].dropna().tail(10))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:05:53.217789Z","iopub.execute_input":"2022-07-29T05:05:53.218108Z","iopub.status.idle":"2022-07-29T05:05:53.230411Z","shell.execute_reply.started":"2022-07-29T05:05:53.218079Z","shell.execute_reply":"2022-07-29T05:05:53.229613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The highest correlated variable with the target (other than the TARGET which of course has a correlation of 1), is a variable we created. However, just because the variable is correlated does not mean that it wil lbe useful, and we have to remember that if we generate hundreds of new variable, some are going to be correlated with the target simply beacuse of random noise.\n\nViewing the correlation skepically, it does apear that several of the newly created  variables may be useful. To assess the \"usefulness\" of variables, we will look at the feature importances returned by the model. For curiousity's sake (and because we already wrote the function) we can make a kde plot of two for the newly created variables.","metadata":{}},{"cell_type":"code","source":"kde_target(var_name = 'client_bureau_balance_MONTHS_BALANCE_count_mean', df = train)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:05:53.231778Z","iopub.execute_input":"2022-07-29T05:05:53.232353Z","iopub.status.idle":"2022-07-29T05:05:53.951406Z","shell.execute_reply.started":"2022-07-29T05:05:53.232320Z","shell.execute_reply":"2022-07-29T05:05:53.950307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This variable represent the average number of monthly records per loan for each client. For example, if a client had three previous loans with 3, 4 and 5 records in the monthly data, the value of this variable for them would be 4. Based on the distribution, clients with a greater number of a verage monthly records per loan were likely to repay their loans with Home Credit. Let's not read too much into this value, but it could indicate that clients who have had more previous credit history are generally more likely to repay a loan.","metadata":{}},{"cell_type":"code","source":"kde_target(var_name = 'bureau_CREDIT_ACTIVE_Active_count_norm', df = train)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:10:01.325247Z","iopub.execute_input":"2022-07-29T05:10:01.325775Z","iopub.status.idle":"2022-07-29T05:10:02.574269Z","shell.execute_reply.started":"2022-07-29T05:10:01.325739Z","shell.execute_reply":"2022-07-29T05:10:02.572703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well this distribution is all over the place. This variable represents the number of previous loans with a CREDIT_ACTIVE value of ACTIVE divided by the total number of previous loans for a client. The correlation here is so weak that I do not think we should draw any conclusions!","metadata":{}},{"cell_type":"markdown","source":"### Collinear Variables\n\nWe can calculate not only the correlations of t he variables with the target, but also the correlation of each variable with every other variable. This will allow us to see if there are highly collinear variables that should perhaps be removed from the data.\n\nLet's look for any variable that have a greater than 0.8 correlation with other variables.","metadata":{}},{"cell_type":"code","source":"# Set the threshold\nthreshold = 0.8\n\n# Empty dictionary to hold correlated variables\nabove_threshold_vars = {}\n\n# For each column, record the variables that are above the threshold\nfor col in corrs:\n    above_threshold_vars[col] = list(corrs.index[corrs[col] > threshold])","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:26:17.085996Z","iopub.execute_input":"2022-07-29T05:26:17.086478Z","iopub.status.idle":"2022-07-29T05:26:17.151863Z","shell.execute_reply.started":"2022-07-29T05:26:17.086428Z","shell.execute_reply":"2022-07-29T05:26:17.150729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For each of these pairs of highly correlated variables, we only want to remove one of the variables. The following code creates a set of variables to remove by only adding one of each pair","metadata":{}},{"cell_type":"code","source":"# Track columns to remove and columns already examined\ncols_to_remove = []\ncols_seen = []\ncols_to_remove_pair = []\n\n# Iterate through columns and correlated columns\nfor key, value in above_threshold_vars.items():\n    # Keep track of columns already examined\n    cols_seen.append(key)\n    for x in value:\n        if x == key:\n            next\n        else:\n            # Only want to remove one in a pair\n            if x not in cols_seen:\n                cols_to_remove.append(x)\n                cols_to_remove_pair.append(key)\n                \ncols_to_remove = list(set(cols_to_remove))\nprint('Number of columns to remove: ', len(cols_to_remove))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:25:52.042091Z","iopub.execute_input":"2022-07-29T05:25:52.042529Z","iopub.status.idle":"2022-07-29T05:25:52.056148Z","shell.execute_reply.started":"2022-07-29T05:25:52.042495Z","shell.execute_reply":"2022-07-29T05:25:52.055158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can remove these columns from both the training and the testing dataset. We will have to compare performance after removing thses variables with performance keeping these variables (the raw csv file we saved earilier).","metadata":{}},{"cell_type":"code","source":"train_corrs_removed = train.drop(columns = cols_to_remove)\ntest_corrs_removed = test.drop(columns = cols_to_remove)\n\nprint('Training Corrs Removed Shape: ', train_corrs_removed.shape)\nprint('Testing Corrs Removed Shape: ', test_corrs_removed.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T05:50:12.541162Z","iopub.execute_input":"2022-07-29T05:50:12.541777Z","iopub.status.idle":"2022-07-29T05:50:12.806762Z","shell.execute_reply.started":"2022-07-29T05:50:12.541721Z","shell.execute_reply":"2022-07-29T05:50:12.804640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_corrs_removed.to_csv('train_bureau_corrs_removed.csv', index = False)\ntest_corrs_removed.to_csv('test_bureau_corrs_removed.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T06:01:36.574012Z","iopub.execute_input":"2022-07-29T06:01:36.574445Z","iopub.status.idle":"2022-07-29T06:02:23.049641Z","shell.execute_reply.started":"2022-07-29T06:01:36.574407Z","shell.execute_reply":"2022-07-29T06:02:23.048478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling\n\nTo actually test the performance of these new datasets, we will try using them for machine learning! Here we will use a function I developed in another notebook to compare the features (the raw version with the highlt correlated variables removed). We can run this kind of like an experiment, and the control will be the performance of just the application data in this function when shbmitted to the competition. I've already recorded  that performance, so we can list out our control and our two test conditions:\n\n**For all datasets, use the model shown below (with the exact hyperparameters).**\n\n- control: only the data in the application files.\n- test one: the data in the application files with all of the data recorded from the bureau and bureau_balance files\n- test two: the data in the application files with all of the data recorded from the bureau and bureau_balacne files with highly correlated variables removed.","metadata":{}},{"cell_type":"code","source":"import lightgbm as lgb\n\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.preprocessing import LabelEncoder\n\nimport gc\n\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2022-07-29T06:20:18.105774Z","iopub.execute_input":"2022-07-29T06:20:18.106570Z","iopub.status.idle":"2022-07-29T06:20:19.747944Z","shell.execute_reply.started":"2022-07-29T06:20:18.106531Z","shell.execute_reply":"2022-07-29T06:20:19.746771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''Train and test a light gradient boosting model using cross validation\n\nParameters\n-----------\n    features (pd.DataFrame):\n        dataframe of training features to use\n        for training a model. Must include the TARGET column.\n    test_features (pd.DataFrame):\n        dataframe of testing features to use\n        for making predictions with the model.\n    encoding (str, default = 'ohe'):\n        method for encoding categorical variables. Either 'ohe' for \n        one-hot encoding or 'le' for integer label encoding\n    n_folds (int, default = 5):\n        number of folds to use for cross validation\n\nReturn\n----------\n    submission (pd.DataFrame):\n        dataframe with 'SK_ID_CURR' and 'TARGET' probabilities\n        predicted by the model.\n    feature_importances (pd.DataFrame):\n        dataframe with the feature importances from the model.\n    valid_metrics (pd.DataFrame):\n        dataframe with training and validation metrics (ROC AUC) \n        for each fold and overall.\n        \n'''","metadata":{"execution":{"iopub.status.busy":"2022-07-29T06:26:24.122158Z","iopub.execute_input":"2022-07-29T06:26:24.122605Z","iopub.status.idle":"2022-07-29T06:26:24.131205Z","shell.execute_reply.started":"2022-07-29T06:26:24.122570Z","shell.execute_reply":"2022-07-29T06:26:24.129959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def model(features, test_features, encoding = 'ohe', n_folds = 5):\n    # Extract the ids\n    train_ids = features['SK_ID_CURR']\n    test_ids = test_features['SK_ID_CURR']\n    \n    # Extract the labels for training\n    labels = features['TARGET']\n    \n    # Remove the ids and target\n    features = features.drop(columns = ['SK_ID_CURR', 'TARGET'])\n    test_features = test_features.drop(columns = ['SK_ID_CURR'])\n    \n    \n    # One Hot Encoding\n    if encoding == 'ohe':\n        features = pd.get_dummies(features)\n        test_features = pd.get_dummies(test_features)\n        \n        # Align the dataframes by the columns\n        features, test_features = features.align(test_features, join = 'inner', axis = 1)\n        \n        # No categorical indices to record\n        cat_indices = 'auto'\n    \n    # Integer label encoding\n    elif encoding == 'le':\n        \n        # Create a label encoder\n        label_encoder = LabelEncoder()\n        \n        # List for storing categorical indices\n        cat_indices = []\n        \n        # Iterate through each column\n        for i, col in enumerate(features):\n            if features[col].dtype == 'object':\n                # Map the categorical features to integers\n                features[col] = label_encoder.fit_transform(np.array(features[col].astype(str)).reshape((-1,)))\n                test_features[col] = label_encoder.transform(np.array(test_features[col].astype(str)).reshape((-1,)))\n\n                # Record the categorical indices\n                cat_indices.append(i)\n    \n    # Catch error if label encoding scheme is not valid\n    else:\n        raise ValueError(\"Encoding must be either 'ohe' or 'le'\")\n        \n    print('Training Data Shape: ', features.shape)\n    print('Testing Data Shape: ', test_features.shape)\n    \n    # Extract feature names\n    feature_names = list(features.columns)\n    \n    # Convert to np arrays\n    features = np.array(features)\n    test_features = np.array(test_features)\n    \n    # Create the kfold object\n    k_fold = KFold(n_splits = n_folds, shuffle = True, random_state = 50)\n    \n    # Empty array for feature importances\n    feature_importance_values = np.zeros(len(feature_names))\n    \n    # Empty array for test predictions\n    test_predictions = np.zeros(test_features.shape[0])\n    \n    # Empty array for out of fold validation predictions\n    out_of_fold = np.zeros(features.shape[0])\n    \n    # Lists for recording validation and training scores\n    valid_scores = []\n    train_scores = []\n    \n    # Iterate through each fold\n    for train_indices, valid_indices in k_fold.split(features):\n        \n        # Training data for the fold\n        train_features, train_labels = features[train_indices], labels[train_indices]\n        # Validation data for the fold\n        valid_features, valid_labels = features[valid_indices], labels[valid_indices]\n        \n        # Create the model\n        model = lgb.LGBMClassifier(n_estimators=10000, objective = 'binary', \n                                   class_weight = 'balanced', learning_rate = 0.05, \n                                   reg_alpha = 0.1, reg_lambda = 0.1, \n                                   subsample = 0.8, n_jobs = -1, random_state = 50)\n        \n        # Train the model\n        model.fit(train_features, train_labels, eval_metric = 'auc',\n                  eval_set = [(valid_features, valid_labels), (train_features, train_labels)],\n                  eval_names = ['valid', 'train'], categorical_feature = cat_indices,\n                  early_stopping_rounds = 100, verbose = 200)\n        \n        # Record the best iteration\n        best_iteration = model.best_iteration_\n        \n        # Record the feature importances\n        feature_importance_values += model.feature_importances_ / k_fold.n_splits\n        \n        # Make predictions\n        test_predictions += model.predict_proba(test_features, num_iteration = best_iteration)[:, 1] / k_fold.n_splits\n        \n        # Record the out of fold predictions\n        out_of_fold[valid_indices] = model.predict_proba(valid_features, num_iteration = best_iteration)[:, 1]\n        \n        # Record the best score\n        valid_score = model.best_score_['valid']['auc']\n        train_score = model.best_score_['train']['auc']\n        \n        valid_scores.append(valid_score)\n        train_scores.append(train_score)\n        \n        # Clean up memory\n        gc.enable()\n        del model, train_features, valid_features\n        gc.collect()\n        \n    # Make the submission dataframe\n    submission = pd.DataFrame({'SK_ID_CURR': test_ids, 'TARGET': test_predictions})\n    \n    # Make the feature importance dataframe\n    feature_importances = pd.DataFrame({'feature': feature_names, 'importance': feature_importance_values})\n    \n    # Overall validation score\n    valid_auc = roc_auc_score(labels, out_of_fold)\n    \n    # Add the overall scores to the metrics\n    valid_scores.append(valid_auc)\n    train_scores.append(np.mean(train_scores))\n    \n    # Needed for creating dataframe of validation scores\n    fold_names = list(range(n_folds))\n    fold_names.append('overall')\n    \n    # Dataframe of validation scores\n    metrics = pd.DataFrame({'fold': fold_names,\n                            'train': train_scores,\n                            'valid': valid_scores}) \n    \n    return submission, feature_importances, metrics","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:42:33.678221Z","iopub.execute_input":"2022-07-29T07:42:33.678626Z","iopub.status.idle":"2022-07-29T07:42:33.708394Z","shell.execute_reply.started":"2022-07-29T07:42:33.678583Z","shell.execute_reply":"2022-07-29T07:42:33.707292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''Plot importances returned by a model. This can work any measure of\nfeature importance provided that higher importance is better.\n\nArgs:\n    df (dataframe):\n        feature importances. Must have the features in a column called\n        'features' and the importances in a column called 'importance'\n        \nReturns:\n    shows a plot of  the 15 most importance features\n    \n    df (dataframe) :\n        features importances sorted by importance (highest to lowest)\n        with a column for normalized importance\n'''","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:28:59.242658Z","iopub.execute_input":"2022-07-29T07:28:59.243075Z","iopub.status.idle":"2022-07-29T07:28:59.250691Z","shell.execute_reply.started":"2022-07-29T07:28:59.243042Z","shell.execute_reply":"2022-07-29T07:28:59.249548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_feature_importance(df):\n    # Sort features according to importance\n    df = df.sort_values('importance', ascending = False).reset_index()\n    \n    # Normalize the feature importances to add up to one\n    df['importance_normalized'] = df['importance'] / df['importance'].sum()\n    \n    # Make a horizontal bar chart of feature importances\n    plt.figure(figsize = (10, 6))\n    ax = plt.subplot()\n    \n    # Need to reverse the index to plot most important on top\n    ax.barh(list(reversed(list(df.index[:15]))),\n           df['importance_normalized'].head(15),\n           align = 'center', edgecolor = 'k')\n    \n    # Set the yticks and labels\n    ax.set_yticks(list(reversed(list(df.index[:15]))))\n    ax.set_yticklabels(df['feature'].head(15))\n    \n    # Plot labeling\n    plt.xlabel('Normalized Importance'); plt.title('Feature Importance')\n    plt.show()\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:20:13.938534Z","iopub.execute_input":"2022-07-29T07:20:13.939029Z","iopub.status.idle":"2022-07-29T07:20:13.949843Z","shell.execute_reply.started":"2022-07-29T07:20:13.938959Z","shell.execute_reply":"2022-07-29T07:20:13.948424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Control\n\nThe first step in any experiment is establishing a control. For this we will use the function defined above (that implements a Gradient Boosting Machine model) and the single main data source (application)","metadata":{}},{"cell_type":"code","source":"train_control = pd.read_csv('../input/home-credit-default-risk/application_train.csv')\ntest_control = pd.read_csv('../input/home-credit-default-risk/application_test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:22:21.276214Z","iopub.execute_input":"2022-07-29T07:22:21.276679Z","iopub.status.idle":"2022-07-29T07:22:28.588644Z","shell.execute_reply.started":"2022-07-29T07:22:21.276638Z","shell.execute_reply":"2022-07-29T07:22:28.587175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Fortunately, once we have taken the time to write a function, using it is simple (if there's central theme iin this notebook, it's use functions to make things simpler and reproducible!). The function above returns a submission dataframe we can upload to competition, a fi dataframe of feature importances, and a metrics dataframe with validation and test performance.","metadata":{}},{"cell_type":"code","source":"submission, fi, metrics = model(train_control, test_control)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:42:38.191645Z","iopub.execute_input":"2022-07-29T07:42:38.192260Z","iopub.status.idle":"2022-07-29T07:45:46.439943Z","shell.execute_reply.started":"2022-07-29T07:42:38.192224Z","shell.execute_reply":"2022-07-29T07:45:46.438860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metrics","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:45:46.442274Z","iopub.execute_input":"2022-07-29T07:45:46.442754Z","iopub.status.idle":"2022-07-29T07:45:46.455933Z","shell.execute_reply.started":"2022-07-29T07:45:46.442707Z","shell.execute_reply":"2022-07-29T07:45:46.454726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The control slightly overfits because the training score is higher than the validation score. We can address this in later notebooks when we look at regularization (we already perform some regularization in this model by using reg_lambda and reg_alpha as well as early stopping).\n\nWe can visualize the feature importance with another function, plot_feature_importances. The feature importances may be useful when it's time for feature selection.","metadata":{}},{"cell_type":"code","source":"fi_sorted = plot_feature_importance(fi)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:46:09.971885Z","iopub.execute_input":"2022-07-29T07:46:09.972357Z","iopub.status.idle":"2022-07-29T07:46:10.299434Z","shell.execute_reply.started":"2022-07-29T07:46:09.972317Z","shell.execute_reply":"2022-07-29T07:46:10.298163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('control.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:46:13.768007Z","iopub.execute_input":"2022-07-29T07:46:13.769117Z","iopub.status.idle":"2022-07-29T07:46:13.949208Z","shell.execute_reply.started":"2022-07-29T07:46:13.769074Z","shell.execute_reply":"2022-07-29T07:46:13.947862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The control scores 0.745 when submitted to the competition.**","metadata":{}},{"cell_type":"markdown","source":"### Test One\n\nLet's conduct the first test. We will just need to pass in the data to the function, which does most of the work for us.","metadata":{}},{"cell_type":"code","source":"submission_raw, fi_raw, metrics_raw = model(train, test)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:46:16.949351Z","iopub.execute_input":"2022-07-29T07:46:16.949785Z","iopub.status.idle":"2022-07-29T07:53:06.069699Z","shell.execute_reply.started":"2022-07-29T07:46:16.949748Z","shell.execute_reply":"2022-07-29T07:53:06.068456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metrics_raw","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:53:06.072162Z","iopub.execute_input":"2022-07-29T07:53:06.072545Z","iopub.status.idle":"2022-07-29T07:53:06.085443Z","shell.execute_reply.started":"2022-07-29T07:53:06.072510Z","shell.execute_reply":"2022-07-29T07:53:06.084019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on these numbers, the engineered features perform better than control case. However, we will have to submit the predictions to the leaderboard before we can say if this better validation performance transfers to the testing data.","metadata":{}},{"cell_type":"code","source":"fi_raw_sorted = plot_feature_importance(fi_raw)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:53:06.086917Z","iopub.execute_input":"2022-07-29T07:53:06.087282Z","iopub.status.idle":"2022-07-29T07:53:06.377532Z","shell.execute_reply.started":"2022-07-29T07:53:06.087249Z","shell.execute_reply":"2022-07-29T07:53:06.376520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Examining the feature importances, it looks as if a few of the feature we constructed are among the most important. Let's find the percentage of the top 100 most important features that we made in this notebook. Howerver, rather than just compare to the original features, we need to compare to the one-hot encoded original features. These are already recorded for us in fi (from the original data).","metadata":{}},{"cell_type":"code","source":"top_100 = list(fi_raw_sorted['feature'])[:100]\nnew_features = [x for x in top_100 if x not in list(fi['feature'])]\n\nprint('%% of Top 100 Features created from the bureau data = %d.00' % len(new_features))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:53:06.379730Z","iopub.execute_input":"2022-07-29T07:53:06.380131Z","iopub.status.idle":"2022-07-29T07:53:06.391776Z","shell.execute_reply.started":"2022-07-29T07:53:06.380100Z","shell.execute_reply":"2022-07-29T07:53:06.390682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Over half of top 100 features were made by us! That should give us confidence that all the hard work we did was worthwhile.","metadata":{}},{"cell_type":"code","source":"submission_raw.to_csv('test_one.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:53:06.549779Z","iopub.execute_input":"2022-07-29T07:53:06.550319Z","iopub.status.idle":"2022-07-29T07:53:06.731137Z","shell.execute_reply.started":"2022-07-29T07:53:06.550284Z","shell.execute_reply":"2022-07-29T07:53:06.729945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Test one scores 0.759 when submitted to the competition**","metadata":{}},{"cell_type":"markdown","source":"### Test Two\n\nThat was easy, so lets do another run! Same as befor but with the highly collinear variables removed.","metadata":{}},{"cell_type":"code","source":"submission_corrs, fi_corrs, metrics_corrs = model(train_corrs_removed, test_corrs_removed)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:54:10.815695Z","iopub.execute_input":"2022-07-29T07:54:10.816191Z","iopub.status.idle":"2022-07-29T07:58:22.286732Z","shell.execute_reply.started":"2022-07-29T07:54:10.816149Z","shell.execute_reply":"2022-07-29T07:58:22.285531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metrics_corrs","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:58:22.288896Z","iopub.execute_input":"2022-07-29T07:58:22.289658Z","iopub.status.idle":"2022-07-29T07:58:22.301594Z","shell.execute_reply.started":"2022-07-29T07:58:22.289614Z","shell.execute_reply":"2022-07-29T07:58:22.300357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These results are better than the control, but slightly lower than the raw features.","metadata":{}},{"cell_type":"code","source":"fi_corrs_sorted = plot_feature_importance(fi_corrs)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:58:22.303275Z","iopub.execute_input":"2022-07-29T07:58:22.303940Z","iopub.status.idle":"2022-07-29T07:58:22.606945Z","shell.execute_reply.started":"2022-07-29T07:58:22.303906Z","shell.execute_reply":"2022-07-29T07:58:22.605701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_corrs.to_csv('test_two.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:58:22.765069Z","iopub.execute_input":"2022-07-29T07:58:22.765780Z","iopub.status.idle":"2022-07-29T07:58:22.948350Z","shell.execute_reply.started":"2022-07-29T07:58:22.765732Z","shell.execute_reply":"2022-07-29T07:58:22.947188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Test Two scores 0.753 when submitted to the competition**","metadata":{}},{"cell_type":"markdown","source":"# Result\n\nAfter all that work, we can say that including the extra information did imporve performance! The model is definitely not optimized to our data, but we still had a noticeable improvement over the original dataset when using the calculated features. Let's officially summarize the performances:\n\n![image.png](attachment:483f960b-1fac-4d78-a1e4-c582ecb6de5d.png)\n\n(Note that these scores may change from run to run of the notebook. I have not observed that the general ordering changes however.)\n\nAll of our hand work translates to a small improvement of 0.014 ROC AUC over the original testing data. Removing the highly collinear variables slightly decreases performance so we will want to consider a differnet method for feature selection.\nMoreover, we can say that some of the features we built are among the most important as judged by the model.\n\nIn a competition such as this, even an improvement of this size is enough to move us up 100s of spots on the leaderboard. By making numerous small imporvements such as in  this notebook, we can gradually achieve better and better performance.\nI encourage other to use the results here to make their own improvements, and I will continue to documnet the steps I take to help others.","metadata":{},"attachments":{"483f960b-1fac-4d78-a1e4-c582ecb6de5d.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAVoAAAB8CAYAAAAo03jpAAAgAElEQVR4nO29f1hU17Xw/2kJc80daxhjGZNCDdAiNIIVp0GpBlKjkgjEH6k/bopJlOSqt1fSKzYtJBdpL7T3xeQbbF7J2ypJIHkRG0BBEsSQgFoUIVjAG0b7nSFmiHFIwiF0ToSZmvP+MQMMCDgIo5jsz/Pw6Oy9zzp7r73O2uusc2b2NxRFURAIBAKB27jlg/Pn8fT0vNH9+Ephs9mETkeB0NfNgZinkbHZbNw1Y8aQdbd4enqi8fK6zl36avNZR4fQ6SgQ+ro5EPM0MlJn57B137yO/RAIBIKvJcLRCgQCgZsRjlYgEAjcjHC0AoFA4GaEoxUIBAI3IxytQCAQuBnhaAUCgcDN3HyO1iYj2250JyYeNqEU17lsQ+52j2ibLCNmQjCYW1xr1knDK1mUmgaXh7P+2Wj8xrtXw9JK4ebt5Hsn8EJGND7X6aw2WcaKCrV6/L4V0/pmGrnvDV3n90Ai63WuvxjeeSKLhOeaiU7JImGOetR9kU/nsvNQKxBI7H+sI6xPRO+8D5znzvpcst5qJfxfUokOsJfZ2huoPFhJ7cUv4J+/w5wF0SwK92H0vRmatnez2Hu8k1k/TWZVUP88yKfz2XnoHF4LNpJ433AW0Ur5b3Op9Y0l8TF/zu5OILMxmuSsBKex9g6ugdw/lNI6dz2pD17Fsm0yshVUk9R4egCGQjY/nY/3hhfIeHC8rbOV8udyqf0C/nn+Rrbf7yTfUE7a/60dZDfOYw7DXirTeryc8nfP0A5437WABbGLCBm37yA4zjlknR+x/76esGs512UbcrcVVGqGvwRlGvJ2UvoBEBBL0r+E9dvekHM6tG3LhmOUl7/DmQ5gqh8LIuNYNGvsCnI5opU/a6a50UzPmE85FnxYuCmBhDVh183JQicnX3qURx8t4Zxb5Pdgbmym+bx0zRK8Zj3A5vhHWRR8LW5N5kxNCc2NzTQ3FtLQIg+s/ayZ5sZBUVq3mebGZuTL9o+2c4Wkbslgz+FWu318dJLczKdIf+PcuEV3Pn5+WBqbya/TO8mUOVtfSHOjBT+/kSzChtzYTPNnMuDFrKWbWR+/iJCh1HVZts/HF1fveWftbh599FFKDI6C7y5ky4YE1uncYJ0fNFN5opnmxmZqyxtoHdTn5sZmzAOidOcxA7Rz7A+JbH8hn/ckgB6aD2aTlpTFsfbx7y6X7DZiunaz7sdQwqOPPsru2uG/ecXfz1B70K6f5qIGzjqb8TBzOti2249nkfh0Fvmn7Z3uaSkhe0cSWcfHriAXI9peFvHYs6sIdC6ytVF7uAHzpCAW3R+Imk6a3zxKq4cf9y71p/3dSvRdPoTd781H1Q2Y0RIWGY5Pn5HLtJ8+yckPZZjiw7x7wvBWA7TTcPAkbVOCmBcg0XDak7CHfGj7yIp1Sicyaj7qlf0TDa3HmpFuDWLRfYHwYS3HTpthehgLnaIq+WIDJ2vbkFHjEz6PsOlqQObcCH1sr3+Lho8BznL0YCXcv4jAcQjT/B5MJfVBgHMUPpxMfvAqErcttEceFxsoOdiGOngeM6UGGjzCiNN5w9/baKhvoK0L1NNDmKvzw8sD5I9NyMhIneBHAyW1baiD72Wu9T2OGmTUAfcOvyr//QwN74LX0mjmHi6n/PRZ1unCRhGJtvLWH/M557WQxIxEFnoDyNT+KZHMfX/iLV0mcXeNWV1wVwgL74LW6jPo14YQ4gnYjJw5DMyOZt5dADY6z53kZIuE1VODX1gYIdMHj0Tmo49k6JKQ8MMboPMcx2r0SGgJmz34xMPIvNjAW6ftF+DZ6hIqWcSiKW202ayoO2UcRnxNNjeklhuP0UoY0UttlB8+RvMHcfiNQq9ybSFZ1Z0Ers0g7eFAPAHbByVkJuWSdXAhYU+MZs6Hw4/oZ1OJBjhXyMPJ+cxcmcj2BU62d7mTcydOov/Miur2IObND8TLw9HHi800NLQi2UD93XnMm+ONWj5HZfVZANpPv0XJpEX2a2Hw+N5voBIvopfOpfxwOQ0t6wjTjWJEf6+l8IVjdAauI+M/VxE4CehupeS57eS+UMLCOUPc/YyCUTpaKxZZpm+x8FCjnuSD36RcMl86Cd/L4IHPS8jKOcq8X2XhhY3mhlxyT4Rx5lw7NhlMZ9rIPd47mHaO/SGZrGoVgbO1YM4ltzCa5N8lEPatTlrzcsm/K5CTnec417mOjIcm28vmJ3Lvfd6YG3LJPRHIydNf8MWnbbRdhJMN4dD6EV90t9HWmUvDv+0m+T5v2o9nkfzCMVSBIWgxk5tXSHRKJglzbP1yajv5ottK24ed5B5fT2ZGHLSfo7UToA1jkw9hkdeubJfpaiU3Lx+/wJPknztH59oM4vxayXo6k+YpYcwNgOa8XPIjt7Pz38OhvYHcvFrWBccRRv+xhd1fwIdttFNC66+ySBjC8Dobj1IJxN3zMLMop/JwLWfWhhH+LRf7amig5APweyzO4WQB1ISv3UHqfInJ37IB45Fy8SMkyg9eqeTM2YcJmeWJ7f3TlAAh80PwxkZrUSrb/28ngfeFoDWUkJvzFuv+O41VAc5yeud7HUEPheH99wb2JGVQ3ulN4Gwv3jnu3NsRZNLOuVZ75NPWehqfOfeCQ/fhT93LokD6bI7vhuA7yUxuXj4Ln8ogcYFqRJu7MmFxjobSVpgdTdximbOHcznW2ErcXa4m7To5c6oSiObhBwL7xud51wMkZPhhtmpGPx3XQvc5Cn+TTH67DyEzNEjnc3nj7QRSdkTjYygk/el82mctYlGAhfL0XAofziAr2sJHrW0ASK3nOB0wj7ghxtdwvBKII3zlLDhcSXntGdbpwl1ePDrPnKQSiF4ZbXeyAJP8eOCJDPzae9B8Obahj9LRFpLxaGH/x7UZvPFwIN6R61j3znZyy3Lp/KiEzvmJPKxTA72hvsSsn75A3F3QXp3Jlj/k89bpRfh9q5JXqzuJTsknYY4n/L2W7MczKa2NI+x+x6EfaHhg9xtkeAOco/mKPk1m0aYMFk1vpSRpO7kfzyTzxe34fXaMzH/NovazTrCZqXzlGJ0PJpO/IQxPZGr/96NkHjpJ3Jy5/XK2ZrBouo3mvM2kHbRHDXEPJrKqJYGsE0NE826m9fYH2P1GBt6A7bNWfvJUBuu+F4j3JGjWPExa0UnOxoczc4hjpeB1ZMeH4HmxkoyfZw8TqXbSXFsLxDHnB17426LhcDkN768nPNxFE71sn+WZXoMu1m/5EDJrfG+h/WYtxI9cKvWtrJvlh/5/jgIhLJjlDch4Bq4j9Xl/Qr6rhvZK5C3ZvNXYNsjRDqS9tpTyTli0PZPN4Wps+nySn+lNElmHl7kymsSVzSS8UMuix1JZFQgDcku2ZrvNzdnM7pRFeNNOZfoWsl+p5CfhDzgaDW1zV0Sq55p5qxNC1obgfZfMwrsg91gzrQ/5ufx8xGoF0DB5wLR64h0YwpXxoXtoP/EG+edCSMhKJfo7wAclbE/aw7HGRcT93cg5fFi1Yj3rZquJi2zG2PXPWL0CWf/YIkqGio57+ayZkyeAh+YQdLs/tqVQfriBM+vDXQ8YLlsB0EwZaPee3oGEjIOCRp06SNixgO/0ftQ4/ufpR9xPo8lPL6EEP9bvdNwC9zGPIIfxeN/hD9TS/LGEfLmNTqA8fR3lzs2dfwVn/ryrDHQmvtMBNGjucPwLcLsW+5mArk7aOoE3M1j35oAT0XmFHE98/WYCtdisIyrD7YSH918Enl4aqH2DV4vz+eh8s308WLFeHvrYmX6+9shlui8zgYauIZ6G9xrofC9sZ5sx2tSEAJX1Z1gf7no0APQZqlu5ax7Rs3PJLm/mXEwPZ97phDnrCJkOoEYz5QtqS3dS9JGZ5nOOvNqXI+daOzubgXBmfc8+Ws+AWYRQ6MiBXptMoN/mZvo65tAb35nA6TY6u3obuWZz55reohMf7vUw03wGVN8D3h59+gBurEF3ftYAwJ7Eh9njVO79dxmvwHtZND2Twt8+SqGXDyFzFhK9PM4lG+xsOUktED7Fhv6MEdu3QoDK0QUMDqxjjFyHY5SO1hv/WSFDRHU29GdOgpcXXp2tnDwz+LZGwiIDauBL+2T7TlGDhwqAVSmvEucs1EMN4/noyUOFCmBlMq8+5Bz/qVAjYx6/M7mN1rfSSXvFk1Xbt/DELC2mt9aRtm9sMtubjtkXohO5ZJxwqnj3KA1rw1l4e2+BhKUbmARgo/1iO+AFHsB0P6KB8uPNtEcu6lsY2g6n8dSfTKzL2GOP9sYFb0Lmh0DjSZoPW6nshLAIx2IkN5D/H5mcjEzgl08tJOnLBnb/PGuYJ+D9eH7TC2hH+jtwOyB19tvDNcoE+m3uUg99yZN/XMOQbc28V94JdFLyv9Mo6atopaT+HHF39Sv3C9kCvSHOZ2aMACoV4IXPTD84UcnJxocJnO1IHnQ3sOdnGZTPT2TPtsHB0fhj1/VMNj+/hXm3O1Wo1OAZzuasfNZdbMXY+h7lr+STeVqyvxkyotR23qu2z0htXsaAuak83sC6cKdxdVn658LWTvvHgJf9s5fPTPyopbKumYeDQhzpFRsNOevIeDOcxP+z3el6GD2jdLRnOXqwBH3fZx/mPRSGxlBC/sFOwp/azcOdL7D9lTc4ds92p5xdOW8cnIlnqIpzBwsBP+Z8zxuvb4URzjEqj59knt881K0lvJBeiObf97B9PHOhXjMJmw/H3jnGyQh/5v2zkZIXMii8PZE920JcFGJEf7qN78wZv1eWRoOtuxUIx9tHg6fczOmTY5XYTnNNAxBN8qv9iX65fg+P/r6cky2dLFzgRUjEQrxOlLP79zYWzfkOaukMJYdaIXA9YQEAYSx6LJDyV7JJ3nHO3qbrLOUHm2HOZhaOc67Fe9YCwsgmP68VCCM82GFk3TISoPn2d/BW22g/ftIlh+gTuAgvCinZV4g2xg/Lqbf6j3NRpvFcA213DHoTxmsm8yK9OHYwnzcCNSxU1/LGG9jv0G5niBTY0NjOnqGyE/weyyQzpjd4caTJyps593AggX5hrAvMJ/9PGWS2RzNzisxHtYXU4kX0/Fl4AV6R64g+kEHJbxMxP+Roc7qSSryIu3+e250sgM+sB/Ajn8rjiwh5yB/5xKukv3SWBzJ2E/7/J/NUDqxKeYq4WfNoDy6koUWNp6r/+PbWZs4FziPQ2ynnf7GZ2tPA0mRe7XugJ9Pwp0fJOHyS5s8WsvD2EO6N9KL2zd1kWBcx5w41nS0llHwAgY+F2dMvdy1i3dISMg6mkdgeR/T3vJA/Pk3l2+D1UPTAheEaGKWjbaA8r8Hp8zqCHvKheV8+5wLXk7nAG7/udcQdSOPVg7WEPNEbPa7iAb8zZO+opB1vwjds4YG7ABay5b8kdr+YzfaEbMCbsIeSeTTSHmWMH14sfDID6Y8vkJ2UQDbgPSeO5PiFeDHCKyOOY+fFruOdlnxy079A/WIqi6aPY9dcxE+XQHjVHrITa8mftY4HfgB8MAaBfQYaxkynlUM9M4RFlFNZ/R7tCxbhPX8LKf+mIvf1SgrPAHjhs2A9v9zY/9DGLyaNF7zyyd1XQmGevU1ITCLJaxeOf/5vegjhc6DhNDB/IXN75+L2EBbFhLD7jTQS3vBm0b8sJJzaq0e0sx4mZUMbmTn5ZNYHEr1hLuGcsx93FZleuodYN+ss+a9k8MWk3aR+11myF+EbUkiwZrIn8ykKAe/568l40hWb68WGvqmSTvyIm+V8h+hHyEI/yHuL987EETjLj1VPpUJeNvkHHe+xTg8j7qn1rOt9APqtMBL+vwy89+6mpK9NOOv/awtxQddn1wTPwDgSn5LZ+0oGW4oALx8WbthOXCB4fvcJEs5l8Ub6Fgp7+5YUZ3+7JHAh2++rJPNgFsl4kR8f0vdAr72llgYges5MpwBIzczQRXC4kmNN7Sy8z5vwJ1LY7JlL/tuF9kXOy4eF8b8koW/xUhP2RBYZ3nvYXVpC7gkAb8LjM9jyUOCYH+d+o+2jjxT3/Wp6J8eeSyDrxDoy3lhF4GUbNg/PITttk2VQq8fl+fSI2GRkRnrxeex81tHB7VOnukW2Tbbh6c7Oj3juq8/RtczjuOrLJmP7puMLBKNhBNu8ZplOx7vb5vq4bEO2eaKeNMY2QzCe8zSsHV9j30ZxZuwmOvJkXMt1JnV28p077xyybpQR7RgZzpABT/V1uiH3VN+QW//x4kY5Wfu5r6656zaPw3fg2hbrEWzzmmU6HX/dtOLhifpqC4IrbdzMsHbs9r554oqJjvd15mZHqyZkeSqpSzX9byoIBALB1ww3O1pPvAJCrkuiXSAQCCYqN9+vdwkEAsFNhnC0AoFA4GaEoxUIBAI3IxytQCAQuBnhaAUCgcDNCEcrEAgEbkY4WoFAIHAzt9hsNqROV797LXCFy5cvC52OAqGvmwMxTyNjsw3/85m33DblNjQa8ZWC8USSOoVOR4HQ182BmKeRkaThFyGROhAIBAI3IxytQCAQuBnhaAUCgcDNCEcrEAgEbkY4WoFAIHAzwtEKBAKBmxGOViAQCNyMcLQCgUDgZka3w0K3mfojxRw5ZcLCZHznRLJsaQS+47AhklWWseGJWq26euMhkE7lsLPUQMT6dJZ9f+z9uT7I6EvzKD5lwjJNR/zjKwiaMkzTyxJNB/MoOG2GabNZtmo5ET6DdNVtojr/NSqmrSE91t+pQqJ+z06KzzsV3bNxUJubASumEwcofrMR8+QAljzyMyIH6wAAI2XP7qVmcPGAMTvL8iUiNp5lP3Ay5C49Za8XU3PBgnZOPBtXBt3Ue81dV1zVnVRPzvPFGAYVB8QmseEezSCZ9eRkFqMdcH3fRHbd0SEpLnGpRSnYFqPExMQoG7clK8nJm5X4mBglZluB0nLJNRHD06FU/T5GiYkpUFquVUJ1hhITE6MU6Mfal7Hjqk4NB5KUpPxGxdLTo1gMFUrmttxhdNmjNL66UcksNyiWfyhKj+kvSva2bKXu8/4WFn2RkrEtQ9m7O1mJyR+sxRalICZXqbNYFEvv35jnbPxwVV+W2mxl43NVyoeXepSeT+qU3G0ZSpV56LY9zmO1WJTG1zYrGdUd/ec8lqkkv1qnXLzUo/R0NCoF25KUIkNvrUEp2pakFDRalJ4ei2Ioz1SSXmtResY2zJse1+ZpNLrrGTBHFsuHSsX/2qoU6Ae37lEac5KVpG2Dr++JZdcj6cfl1IHxzWzyzmqI3L6HPTvTSU/fTXbyMjRn88g+bOxrJ39cz5GiYoqLjlD/sdxbiv7tYoqL6jHJJmpKiykurcHkqDafOkTdBYAWqoqOoJdBPnuE4qJi6s9LNL1dzJGz8gjyb0KsTVQV/oiNa0NRq1So/RezMqKeqqahxiRhNgbwo7n+qD1A5RNBlK4M48f9LTosU/nZf/2aFXdPvvJw2ULHD6fjq1aj7v1z23bO7sJMTalE/PpIfCepUE3TsXKVmqITxiFbq5zHermFmtNRrJjXGyUZqTqgZc1aHdpJKlSaUJb/ajPBt9p1b/1rFcW6jawOVaNSqfFfupKIU1UMOTWCAYxOd6r+OVKrUZvrqfhyJYtnqgbJfI08VTxrdIMOv4ns2kVHa6T+gBH8V7DyXm1fqXp+POnp6Wz8gQYrYD66k81PppH3Tj31NQWkPbmZnUfNgBXzeznkvFzArtQ0XjtcTM4ff8eW1GKMgGzWY5QATBhOm7BYwWquJ+flHAr+kEFKVg71ZusI8m9C2gw0zvfH16nIP0hHmdE0RGMNAXdbqGt2jFVuoq5+Gf539LfwnRuJ73BG1iVhutWC/u08cl7O48h7ZqzjNIzrhmzC0BWE77T+InXAbKbrTUhXOdT4bgFS3GKCeq/fT03o7wxA21nPkdfzyCs6gt7Dn6A77De4becbifAfMDME31OGsW08B/TV5Np1J1N/uIYfLZ/PgKRBt56C1yB+VRBXhBA3kV276GitIAF3Th2oBNT4hoYS+n0NKmsTFXuqkeZuJfPFdNJ3prN1rkT1ngqa+kY/mSXb9rD7xT2kr9TA2SoajeAfm8SauwGWsPG3G9A5ncQ6ZyOFpaX8ep7JBfk3ETYbRs3kgbkrj38aprEK/6UbCTieQGxsLLFrX0JatQbdcPncK9CgC5qMOiiGNct+hPV4CikHh44EJyyXZKTB9ueK9XbVU3EsgjX3Oh0pmWm5UEzOAQnfxStZcreViqd3UvOpvdp6ycjU2wZlFT3G2P+vCdesO2MFBdYVLBsQzVoxluYhr15D6JBBxM1j16N76+Ayw68YXRJtEhDkiz3m1eIbBEhtSF29jYLxvQNAhW9AMGBkhF8WA+DHuiBULsu/ifgmaC5ZXVyBjRT/pgDV+n2UlpZSeiCV2Scy2H/WxRXmDh0rVi5G56NB7R3Esn/dREBhPfoxdP+64+GJp8VVffVjfDsP66olDHg84uGJpj2UZY8tJshbjXbmMjY/rqLgL/aLVOWhwXLpZly9bzzXpjuZmtIqImIjBgYexjL2frqC+HuGeQx5E9m1i47WF/8HgRNVNDnlBWkrIyU2lvX79PYLAcDZeVwex566W/715k5/IozmAbe9ktlAqEZzZdtPTei/HUHEDIfBeWiJjA4l77SLq7dVRna2/UmTmSp1YLmZco6a6QR0mTA7j+MzMy3emuHfBuiq4VB1FDHzB7XQaPGdMR2tU/Ckvk2D0eEgtDMiMJgHzAzm86FoXL6D+PpyTbozVlDAGpYMWA0lag7kYDYUk/FsCinPprD3Haj4Uwo5pxzybyK7dtHRqtEt3UAQ9ex6OoVdrxdTXJTD737zEk3oiI8MAk0wEfdpoCiPghMmTE37KdgHLIggdNpVT+DAQMt7JobU07jIn0Cog9HdWUFF71OCbj1HSqxEhTpy4F0mms476m5VozEYMPQpxor+dD26aUM45SGw6otIfLGa3my23FTHX+4PJfimel/Jn9k/aaGsLydvprqsgsVzHXc8yJiaBtqO8e0CWDsomgXQzCbqziqq33e0vmymurya1T+wt1QH6fA9UtH3AMd69giHLkUResdgQYLBXE131nY9+nZn72iPZqMGR7No0P3rPrLSkkn+lf1vzQKIXJfMmh/a7f5msmvX36P1X0F69lRe25NH8b4cADShK0j6z58ReQeAhoh/3cEmWwYvZWxhP6BdsIHMzZForvq4QsP85fGE/k8eOTssTP5jOoMfMI4sn6ueYeKhRvf4Roy/TWRLgRb15xC8fiurHQZpPpVDyqko9iRHolXriN9uZtcvEijQalFJEqq5G9h0v3bkUzhQha4h+fxzpDx5CO23ZSTPKDb9x2DDnvj4P5REZFYKCeVatN0Sk5cmsy3UEZZ+XENOSj1Re39NpDf90WzWUKNUo3tkDcbnt7PFpkH9iZmpsU6ypujY8ISRtF9soUCrRvoymA2Jq3FN219zRtSdFX15Bjs9ksl9JMje3hHNZgzx6qtKrcY5YztZBf80WU3vq/Y3k11/o6NDUkb9q+lWGZn+AY+6fqy4W/4YGfUv0XfLWCcNNKrhsMoyqF1rO8TRyN2qCfcKzKj1ZZWxeqhRjccDKqsVq0o1rD6tshXVRDW068xo5+n66W5i2PVI+hndN8N6UalHXjWuVj9W3C3/euOikwX7Kn/t3HhjHBdU17rQDCVreCdrP5VwstfK9dPdxLdr8VsHAoFA4GaEoxUIBAI3IxytQCAQuBnhaAUCgcDNCEcrEAgEbkY4WoFAIHAzwtEKBAKBmxGOViAQCNzMLd3d3UhS543ux1cKodPRIfR1cyDmaWS6u7uHrbtl0qRJo/v6o+CqjPorpV9zhL5uDsQ8jcxIi5BIHQgEAoGbEY5WIBAI3IxwtAKBQOBmhKMVCAQCNyMcrUAgELgZ4WgFAoHAzQhHKxAIBG5GOFqBQCBwM65tZSPVk/N8MYYhKyPY+NtlV+406gpWGdkGnpOG3/9J/ls1ZW9W0PgpMG02S6IXEznTtd1fJz4y+tI8ik+ZsEzTEf/4CoKG25b5skTTwTwKTpth2myWrVpOhI/TViFSE2X7yqi5YGGy/xJ+ti4SX8f2HsbSFPaeGixwDPN2w7BiOnGA4jcbMU8OYMkjPyPSZ6jtUoyUPbuXmsHF92wkPbZ3xM6yfImIjWfZD+zbBEmncthZOtjaA1jxHxvQfVVMz5106Sl7vZiaCxa0c+LZuDJo6K2nhvErAbFJbLjHrugrbHfGCpISdPROg/x+GXmlNZgsWmY/uILl833Hb5uj8aSjQ1KuSkedsveZZCX5mWQl+d/ilZiYGGXjNsfnZw4phqtLGFpsdYYSExOjFOiHrr9YnanEx8QoMfGbleRnkpWkJ2KUmJh4JbnQoPRc4zmvBy7pVFEUw4EkJSm/UbH09CgWQ4WSuS1Xabk0VMsepfHVjUpmuUGx/ENRekx/UbK3ZSt1nzuqP69Tsp/IVCoMFqXnHz1KR+1eJWl3nWLpPfySRbFYnP5O5yqbf1+ldIx9qOOCq/qy1GYrG5+rUj681KP0fFKn5G7LUKrMQ7ftsQwcc+Nrm5WM6v4RdxzLVJJfrVMuXupRejoalYJtSUqRoe/ggfoyVSiZWwuUlolsdNcB1+bJoBRtS1IKGi1KT49FMZRnKkmvtQxzvfYM1LPlQ6Xif21VCvS9rTuUqt9nKBUfOLfpl9RzrkBJ2par1Jl7lJ5LF5W6V5OUHeUXxz7Qa2Qk/bjmaJ2FDeMcLRfqlIrCIqWosEKpu2BxrlEuNlYphwbXXahTcp/bqsTExCg7dhcpFXrLQIGf/0XJiolRYrYV9Duff1xUKn4fo8TE7FAqLthlFBUWKRX6DqWjscL+/8ZB7qOjRakqKVKKCg8pVfrr41pc0mlPo7I3vkBpcSoyFG5VsmstQzS+qFSkZihVn/SXtOQ7zcHnHyqNJufjWpSCmIGy+7EodbuTnIz5xuOaDV5UKp4ZqANLTZay9YALy2pP1OsAABQaSURBVPzndUr2NmdHaVCKtuUqjU4q6PnEoLRcGEr39gXR2Ul/XXFlnnpO71Xi8wdYtVK0NVupG1q1AzEUKUkDAgCDUrR174B5cjqT0pgTo+T+j3NRo7J3a9E1B35jZST9jEuO1nx0J5ufTKOgpp76mgLSnkzkpfdkwIrxzztISNlLzQUL0t/sdXnvW8FiRm+UADAZ6jF9bh0gU2qq4QiwbPUygnp3uPTQErl0BVBP/d8k6DKS83IOh/6UwfY/FlP2cg67UhJ56ZQMgPXsfrav387ew/XUv1fG3qREUkqN4zHksdNmoHG+P75ORf5BOsqMpiEaawi420Jds9n+UW6irn4Z/nc4qqf4EurTf3MmN9VR9+BA2X0YKyjoimHxzAl5gzU8sglDVxC+0/qL1AGzma43IV3lUOO7BUhxiwnqHfKnJvR3BqDtrOfI63nkFR1B7+FP0B1D3OB21VNx7EesmCdyBq7Qdr6RCP8BVk3wPWUY2652pEz94Rp+tHx+X1oAqwXJaMX0Xhk5L+dQdtSIPJIIlQpPox7Tp2MYgJsYu6O1NlGxpxopNpXdO9NJ35nFr+83U3agBjMyJqMefBaz5vF4Njydxe70rcyeYoPvLyNpdTAAS55I78vJ9PGlDYCptw00ftWtkwGouWDuK5Pujmf3i7vZ88et6JAoe68FGTPV+Xnof7iJ9BfTSU/fzY7HNTT9sYqmgT79xmCzYdRMHpi78vinYRqr8F+6kYDjCcTGxhK79iWkVWvQDcrnSkd/R2xsLGsLIP4R3RB5MbsxR/w0kpvObVySke6cOrDfrlhvVz0VxyJYc6/TkZKZlgvF5ByQ8F28kiV3W6l4eic1Q1ygxncLsK5a1u+kBSNivWS84pplmOcvAzBWUGBdwTLnAMCmwvdxXyZP1bFm+WI0pr0kZtc7nK2KgLuXUV1ejfkygIyxqIAjE9SwXXsYNhJdEm0SUJrGqlLnCgkJDUE/Xoz2+H5S1u5HMyMU3YJlrFg+ZGp8SCyXrOCc3r7cA0Copl+jwQGOBPgdvgQD9V0yViSk9wBeYkvsS04SpyN1AU6R0Q3hm6C5ZGXQ6IbBSPFvClD9+z5Kn1XDZTPVWRnsn5rOaifD1Nz7a0rvBfn8EbKTcmDXBkKd97s3VpBnXUHGzfUEzI6HJ54WmdGukca387Cuyhj40M/DE017KMsyHFGu9zI2P24g+S9GIh5yatlVw6F3IojJct1ev+6oPDRXXrNXRaamtIqI2KyBwYE6iMUrg/o+RjyyCenZ16j/VEfkNFDfs4HU7td4afNODCodUevXsPHuKtS3jtNgxpGxO1oPTzwBfprKvlXBThWedqUt2MoeXTym8waMtWXsff131Helkv2kbkSxGp8g/Kmh6rSen/0w1DFtVvSn6wF/dN/Xgm0kCSrQAEFb2Z0YwVTnnk2E6+ZOfyKMZiRA6yiSzAZCNZFXtv3UhP7bEWya4ei4h5bI6FBiTxtZPTPI/vYGatQO21bPWExMZCxN5zcQOrNXiN2YowYb882CZjoBXVWYraDtvYY/M9PiPcwTbbA7yuqoKx2lRovvjMn9cgD1bRqM5we6cePbBbA24yZ7M+PGop0RgcE8wKoxnw9lKLPuw1hBAWuuDAAuW5FtKtR9wcJkJk++SEdvoNRtQ/vjDaTeu8FebW0i53UtURPQwMeeOtAE86MFwNtV1JitILdQlLqWtS+eQMJEWVIssc8cwqwJJnSBjmCASSq7c3Zg0Ndj6hok138J8Q9qkIpS2PLfORQXFZP3YhoZ+4xo7ltD1FWt34fQB/3hRAVVegtYzdT8n82s/UUZxstjHvXYUQeju7OCiiZH1qlbz5ESK1GhDgPtMtF03lF3qxqNwYChL0FlX3B00+xRvVVfROKL1fQlU2Q9Tad1aJxTCw5jXnLTeg1/Zv+khbKjvaM0U11WweK5QY5FWMbUZBqQw7M7yiVXOkrNbKLurKL6fUfry2aqy6tZ/YNB0Wx1FDHzJ+BVO4FRB+nwPVJBr1lbzx7h0KUoQh3PE6ztevTtzgtabwAQceWC2V7Nzmf2o+/9Pe32emouRDHbMU3m4zvZvk/fd5djPlqGKTZiQi6MY49o0RC5OZOO7J3sSlzPLkA7dwWpj9vzgIu3bKLl+QLSNu4HQLtgA8nL7RGq6p4VxIe2kLcnDcukPaQv1TrJVaPbnE2mNptdB4rJOQ6gJeKRVDb8VOdCjlFF0PIkkiwvsXdHAvsBzYxINv1qOUGu5Izcjhrd4xsx/jaRLQVa1J9D8PqtrHYYpPlUDimnotiTHIlWrSN+u5ldv0igQKtFJUmo5m5g0/12falC15B8/jlSnjyEVqtCklREJSSxuPdh2c0ezTrwfyiJyKwUEsq1aLslJi9NZluoIyz9uIaclHqi9v6aSG+Gj2YBUKN7ZA3G57ezxaZB/YmZqbFOshDR7DUzRceGJ4yk/WILBVo10pfBbEhc7YhvrejLM9jpkUzuI46UwHDRLMAdi9m0OoddW7eAVo0k+bLiV5v65kR7/ybi9+ziqSTQIEHQBrYmaIcQdOP5RkeHpIzbr6YPuoUdQLeM1XP4LyZcXbQV1ZCCXeCKWxD3Mupfou+WsU5Su5TVssoyqIdvOyY93SBGrS+rjNXj2m1poCwrVpVqYr7kPsEY7TyNpy2OKOuyFSuq8bGHMTCSfsYhonVCpR4+YnLRkQwvegxHe6hQT4godhhGoRuVeuSY9GZzsteEamy2NFCWcLLuYjxtcURZHhN/DsVvHQgEAoGbEY5WIBAI3IxwtAKBQOBmhKMVCAQCNyMcrUAgELgZ4WgFAoHAzQhHKxAIBG5GOFqBQCBwM7d0d3cjSZ03uh9fKYROR4fQ182BmKeR6e7uHrbulkmTJo3u64+CqzLqr5R+zRH6ujkQ8zQyIy1CInUgEAgEbkY4WoFAIHAzwtEKBAKBmxGOViAQCNyMcLQCgUDgZoSjFQgEAjcjHK1AIBC4GeFoBQKBwM24tpWNVE/O88UYhqyMYONvl13bJnZWGdkGnpOu3P/JWJrC3lNDHxYQm8SGe66+PePER0ZfmkfxKROWaTriH19B0JRhml6WaDqYR8FpM0ybzbJVy4nwcdrAQ2qibF8ZNRcsaO9ewZqf6tA667RLT9nrxY76ZaxYHoHvddpDbfywYjpxgOI3GzFPDmDJIz8j0meoTUyMlD27l5rBxfdsJD2211KdZfkSERvPsh/0bxN0hf3NWEFSgiubggoG2NqceDauHGZL+GH8ivP1fbV5kM/XUHagjMZPGflcN5jR7xkmmWg6L6GdGYr21rGdXDq5i/WZNcTvLGX1zOFaWTH/VY9Z40vojK+WmRsP7mDvpXh2PBsEbdVk/yYP/iueoCscoJWm17dToU0meYc/nh/XkPNCDqr/3IRuCmDVsz+9As2WraTP8MR0dBcpr6vYvd6+2zBWI/v/+xCajVtJvRPMJ7JJe1lF1mbdhDTK4ZBP5ZB2IpjUlFS0liYKfv8c1b907Ho7AH8W/yqZKKcSw4HtlN3Wbz/S8V28ZIxia8pyNJf0HEjfQfGWTFb4A0iY3p9MVMJWIqb1HuF5U+nqxmGk+Dd7sa3fQWoQtL2bzY7XIf2RoCv39dKEsuZXwU4FHdTs3ol0W6+mrzIPxmJ27Law5peprL7dhvHwLnYcXEPmQxNw7+KODkkZDR3VGUpMTIxSoB9YbrlQp1QUFilFhRVK3QWLc41ysbFKOTS47kKdkvvcViUmJkbZsbtIqdBblKFpUQpiYpSY31cpHYqiKEqP8uExu6wWi11+y5EipaiwTvnQ3kOlsaRIKSprdLS3KBfrK5SiwiKl6EidcnG404wjLum0p1HZG1+gtDgVGQq3Ktm1Q3XwolKRmqFUfdJf0pLvNAeGQ8qOAwan9i1KQUy/bEtttrLjyEWneovyYaNB6ehxZTTuxzUbvKhUPDNQB5aaLGXrgHEPw+d1Sva2AqWlb7wGpWhbrtLoNP6eTwxKS5/dGpSirXsH1Atcm6ee03uV+PwBVq0Ubc1W6ly57gxFSlLfdd577PDz0POJQTF84lT5SZWSMeD468tI+hmXHK356E42P5lGQU099TUFpD2ZyEvvyYAV4593kJCyl5oLFqS/2evy3reCxYzeKAFgMtRj+tzq4tlUTP5ST87LeTRdAKwGarJyyHk5j3ojIBuo+WMOxV0qNJipfn4zCTvyOPJePTX700jYvJPq9vEY9RhpM9A43x9fpyL/IB1lRtMQjTUE3G2hrtls/yg3UVe/DP87eg9cRqrzKv6xiZb7ffpkm4w1BPtOxni0jLzX8yg70cHUUH80E33rUGdkE4auIHyn9RepA2YzXW9CusqhxncLkOIWE9Q73k9N6O8MQNtZz5HX88grOoLew5+gOxyxktWCZLRieq+MnJdzKDtqRHbHmL6CtJ1vJMJ/gFUTfE8ZxrarHSlTf7iGHy2f35+euco8qKb54z+t34jNf62hI8h3QqZ3xu5orU1U7KlGik1l98500ndm8ev7zZQdqMGMjMmoB5/FrHk8ng1PZ7E7fSuzp9jg+8tIWm2/bVjyRPqocq6a7+sIRaLRaLY7LDRoNEb0bRK0GakBooL8sTZVsPddCd3PM9mdnk5m2lZ0UjV7y5tw1a27DZsNo2bywNtRj38aprEK/6UbCTieQGxsLLFrX0JatcaeNhhMt579z1UQGhvhkC3RcUFDXe4u6tWhxDwYheZvu9jxZ+ON18FouCQj3Tl14EXkivV21VNxLII19zodKZlpuVBMzgEJ38UrWXK3lYqnd1LzqaPepsL3cV8mT9WxZvliNKa9JGbXC2frAtZLRqbeNijJ4jF02wEYKyiwrmDZTKfV38V50O+LJTY2lrQmHVsfmIBpA8bD0XZJtElAaRqrYmOJjV3L794G/iohoSHox4vRtu0nZW0s63+eQfH7FjTTxpjtuiMAnT80GU0Yz+sx/nAN8Q9qqPkfA/oPGpGIJGiGCrmzDQkIvkvrOM6XYED6WLrxF803QXPJ6qKzM1L8mwJU6/dRWlpK6YFUZp/IYP/ZwUebqd6dR8fqJEeuEUCF6lYjk+/bwOq5vmg0vkSs30zE8QqabrgSRoGHJ54WV/XVj/HtPKyrlgx8WOvhiaY9lGWPLSbIW4125jI2P66i4C9Ge706iMUrlxE5U4ta40vEI5tYceEI9Z8OdQaBMyoPDZZLo50lmZrSKiL6ggMHLs5D0NpSSktLSY+8yK6MI5jHOAZ3MHZH6+GJJ8BPU9m3b5/T33KCAO2Crez5cy67d6ay8R4V9a//jpTcsUYH/gTP18CJeiqaavDXBRPhHwEn6qnSN8EPZxOgAb7pCTBg4nvGdN5x5E5/IozmAbe9ktlAqGaIyP5TE/pvRxAxw2GGHloio0PJO210amSm/o+7qNNtZdM9WqdyNVOnafCdphlYppWQL43fcNyOZjoBXSbMztfwZ2ZavDXDP6TqquFQdRQx8we10GjxnTEdrVPwpL5Ng7HXTi5bkQf8tOhkJk++SEfXmEfxlUc7IwKDeYBVYz4fima4t2nAHs2yhiWDg9GrzUO3jPVyf61m7hJiVPXoJ+CCOHZHqwnmRwuAt6uoMVtBbqEodS1rXzyBhImypFhinzmEWRNM6AIdwQCTVHbn7MCgr8c0SiP2D4oCqYyytzX8OMgftX8wEVIZZW+Dvy4ALaAJiiBSA8WvFVDTZqJpXwH7gYj5oTc+j6MORndnBRW9YWW3niMlVqJCHU6yy0TTeUfdrWo0BgOGvtXJiv50Pbo+52lFv28XVTM2sfVeZydrx3/uCgzvnMDsMEr5/WoOWXUETbui6QTGn9k/aaHsaG+8Yqa6rILFc3ufZsuYmkwDFnDj2wWwdsmVrx5qZhN1ZxXV7ztaXzZTXV7N6h84WrZXs/OZ/eh7L/L2emouRDF7Yt6VTijUQTp8j/TfLVnPHuHQpShCHc8TrO169O3Oq6U9mo0aHM3CVefBfHwnKX/W99/ltOup+8BnZKd+gxj9611XoCFycyYd2TvZlbieXYB27gpSH49EAyzesomW5wtI27gfAO2CDSQvt792pLpnBfGhLeTtScMyaQ/pS690EsOh+n4oyyimjCiC/ABVELofQs1f7Y4XgGkRbH52E7bMl/jd5v2AlojHM9l67w13s4Aa3eMbMf42kS0FWtSfQ/D6rax2GKT5VA4pp6LYkxyJVq0jfruZXb9IoECrRSVJqOZuYNP9Dn397QAZrzchsYXqF3vlR5D08q+JnAb4L2NT6EukJR5Cc5uMGR2bElfjurYnBv4PJRGZlUJCuRZtt8TkpclsC3WEpR/XkJNST9Rex+tevdFs1lDxrhrdI2swPr+dLTYN6k/MTI11knXHYjatzmHX1i2gVSNJvqz41aZre1f868YUHRueMJL2iy0UaNVIXwazoc/WrOjLM9jpkUzuI0H29o5oNmMo5V5lHrT3b2LNnl1s+TloNVZ7/X9uInQCPuT9RkeHpIzbr6ZbZWTUqIcaaLeM1fPKLyZcL6yyDGr1le/yuYFR/xJ9t4x1kmt9G+s4rN2gmmBfVBi1vqwyVo9xsiWrFatKNaw+rbIV1ZAG/fVjtPM0nrobWZYVuVuF+gbb9Uj6GYeI1gmVevh8mYuOxF2o1BP4dfNR6Gas45hoTvaaUI2jLY3gZO2nEk72WhlP3Y0s68Y72ashfutAIBAI3IxwtAKBQOBmhKMVCAQCNyMcrUAgELgZ4WgFAoHAzQhHKxAIBG5GOFqBQCBwM8LRCgQCgZsRjlYgEAjczC3d3d1IUueN7sdXCqHT0SH0dXMg5mlkuru7h637hqIoynXsi0AgEHztEKkDgUAgcDPC0QoEAoGbEY5WIBAI3Mz/A2S0yjf9r2bkAAAAAElFTkSuQmCC"}}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}