{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-31T23:58:27.703913Z","iopub.execute_input":"2022-07-31T23:58:27.704321Z","iopub.status.idle":"2022-07-31T23:58:27.715809Z","shell.execute_reply.started":"2022-07-31T23:58:27.704291Z","shell.execute_reply":"2022-07-31T23:58:27.714374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploring the Data","metadata":{}},{"cell_type":"code","source":"# Import libraries necessary for this project\nimport numpy as np\nimport pandas as pd\nfrom time import time\nfrom IPython.display import display # Allows the use of display() for DataFrames\n\n\n# Pretty display for notebooks\n%matplotlib inline\n\n# Load the Census dataset\ndata = pd.read_csv(\"../input/udacity-mlcharity-competition/census.csv\")\n\n# Display the first five records\ndisplay(data.head(n=5))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:27.718783Z","iopub.execute_input":"2022-07-31T23:58:27.719324Z","iopub.status.idle":"2022-07-31T23:58:27.879557Z","shell.execute_reply.started":"2022-07-31T23:58:27.719290Z","shell.execute_reply":"2022-07-31T23:58:27.878087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Implementation: Data Exploration\nA cursory investigation of the dataset will determine how many individuals fit into either group, and will tell us about the percentage of these individuals making more than \\$50,000. In the code cell below, you will need to compute the following:\n- The total number of records, `'n_records'`\n- The number of individuals making more than \\$50,000 annually, `'n_greater_50k'`.\n- The number of individuals making at most \\$50,000 annually, `'n_at_most_50k'`.\n- The percentage of individuals making more than \\$50,000 annually, `'greater_percent'`.\n","metadata":{}},{"cell_type":"code","source":"data.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:27.882263Z","iopub.execute_input":"2022-07-31T23:58:27.882614Z","iopub.status.idle":"2022-07-31T23:58:27.890484Z","shell.execute_reply.started":"2022-07-31T23:58:27.882584Z","shell.execute_reply":"2022-07-31T23:58:27.889308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.income.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:27.891919Z","iopub.execute_input":"2022-07-31T23:58:27.892355Z","iopub.status.idle":"2022-07-31T23:58:27.912172Z","shell.execute_reply.started":"2022-07-31T23:58:27.892312Z","shell.execute_reply":"2022-07-31T23:58:27.911001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Total number of records\nn_records = data.shape[0]\n\n# Number of records where individual's income is more than $50,000\nn_greater_50k = data[data.income == '>50K']['income'].count()\n\n# Number of records where individual's income is at most $50,000\nn_at_most_50k = data[data.income == '<=50K']['income'].count()\n\n# Percentage of individuals whose income is more than $50,000\ngreater_percent = n_greater_50k*100/n_records\n\n# Print the results\nprint(\"Total number of records: {}\".format(n_records))\nprint(\"Individuals making more than $50,000: {}\".format(n_greater_50k))\nprint(\"Individuals making at most $50,000: {}\".format(n_at_most_50k))\nprint(\"Percentage of individuals making more than $50,000: {:.2f}%\".format(greater_percent))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:27.914321Z","iopub.execute_input":"2022-07-31T23:58:27.914669Z","iopub.status.idle":"2022-07-31T23:58:27.955893Z","shell.execute_reply.started":"2022-07-31T23:58:27.914639Z","shell.execute_reply":"2022-07-31T23:58:27.954300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"** Featureset Exploration **\n\n* **age**: continuous. \n* **workclass**: Private, Self-emp-not-inc, Self-emp-inc, Federal-gov, Local-gov, State-gov, Without-pay, Never-worked. \n* **education**: Bachelors, Some-college, 11th, HS-grad, Prof-school, Assoc-acdm, Assoc-voc, 9th, 7th-8th, 12th, Masters, 1st-4th, 10th, Doctorate, 5th-6th, Preschool. \n* **education-num**: continuous. \n* **marital-status**: Married-civ-spouse, Divorced, Never-married, Separated, Widowed, Married-spouse-absent, Married-AF-spouse. \n* **occupation**: Tech-support, Craft-repair, Other-service, Sales, Exec-managerial, Prof-specialty, Handlers-cleaners, Machine-op-inspct, Adm-clerical, Farming-fishing, Transport-moving, Priv-house-serv, Protective-serv, Armed-Forces. \n* **relationship**: Wife, Own-child, Husband, Not-in-family, Other-relative, Unmarried. \n* **race**: Black, White, Asian-Pac-Islander, Amer-Indian-Eskimo, Other. \n* **sex**: Female, Male. \n* **capital-gain**: continuous. \n* **capital-loss**: continuous. \n* **hours-per-week**: continuous. \n* **native-country**: United-States, Cambodia, England, Puerto-Rico, Canada, Germany, Outlying-US(Guam-USVI-etc), India, Japan, Greece, South, China, Cuba, Iran, Honduras, Philippines, Italy, Poland, Jamaica, Vietnam, Mexico, Portugal, Ireland, France, Dominican-Republic, Laos, Ecuador, Taiwan, Haiti, Columbia, Hungary, Guatemala, Nicaragua, Scotland, Thailand, Yugoslavia, El-Salvador, Trinadad&Tobago, Peru, Hong, Holand-Netherlands.","metadata":{}},{"cell_type":"markdown","source":"----\n## Preparing the Data\nBefore data can be used as input for machine learning algorithms, it often must be cleaned, formatted, and restructured — this is typically known as **preprocessing**. Fortunately, for this dataset, there are no invalid or missing entries we must deal with, however, there are some qualities about certain features that must be adjusted. This preprocessing can help tremendously with the outcome and predictive power of nearly all learning algorithms.","metadata":{}},{"cell_type":"markdown","source":"### Transforming Skewed Continuous Features\nA dataset may sometimes contain at least one feature whose values tend to lie near a single number, but will also have a non-trivial number of vastly larger or smaller values than that single number.  Algorithms can be sensitive to such distributions of values and can underperform if the range is not properly normalized. With the census dataset two features fit this description: '`capital-gain'` and `'capital-loss'`. \n","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\", category = UserWarning, module = \"matplotlib\")\nfrom IPython import get_ipython\nget_ipython().run_line_magic('matplotlib', 'inline')\nimport matplotlib.pyplot as pl\nimport matplotlib.patches as mpatches\n\n\ndef distribution(data, transformed = False):\n    \"\"\"\n    Visualization code for displaying skewed distributions of features\n    \"\"\"\n    \n    # Create figure\n    fig = pl.figure(figsize = (11,5));\n\n    # Skewed feature plotting\n    for i, feature in enumerate(['capital-gain','capital-loss']):\n        ax = fig.add_subplot(1, 2, i+1)\n        ax.hist(data[feature], bins = 25, color = '#00A0A0')\n        ax.set_title(\"'%s' Feature Distribution\"%(feature), fontsize = 14)\n        ax.set_xlabel(\"Value\")\n        ax.set_ylabel(\"Number of Records\")\n        ax.set_ylim((0, 2000))\n        ax.set_yticks([0, 500, 1000, 1500, 2000])\n        ax.set_yticklabels([0, 500, 1000, 1500, \">2000\"])\n\n    # Plot aesthetics\n    if transformed:\n        fig.suptitle(\"Log-transformed Distributions of Continuous Census Data Features\", \\\n            fontsize = 16, y = 1.03)\n    else:\n        fig.suptitle(\"Skewed Distributions of Continuous Census Data Features\", \\\n            fontsize = 16, y = 1.03)\n\n    fig.tight_layout()\n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:27.957992Z","iopub.execute_input":"2022-07-31T23:58:27.958512Z","iopub.status.idle":"2022-07-31T23:58:27.974557Z","shell.execute_reply.started":"2022-07-31T23:58:27.958465Z","shell.execute_reply":"2022-07-31T23:58:27.973242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into features and target label\nincome_raw = data['income']\nfeatures_raw = data.drop('income', axis = 1)\n\n# Visualize skewed continuous features of original data\ndistribution(data)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:27.977101Z","iopub.execute_input":"2022-07-31T23:58:27.977597Z","iopub.status.idle":"2022-07-31T23:58:28.546745Z","shell.execute_reply.started":"2022-07-31T23:58:27.977553Z","shell.execute_reply":"2022-07-31T23:58:28.545371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For highly-skewed feature distributions such as `'capital-gain'` and `'capital-loss'`, it is common practice to apply a <a href=\"https://en.wikipedia.org/wiki/Data_transformation_(statistics)\">logarithmic transformation</a> on the data so that the very large and very small values do not negatively affect the performance of a learning algorithm. Using a logarithmic transformation significantly reduces the range of values caused by outliers. Care must be taken when applying this transformation however: The logarithm of `0` is undefined, so we must translate the values by a small amount above `0` to apply the the logarithm successfully.\n","metadata":{}},{"cell_type":"code","source":"# Log-transform the skewed features\nskewed = ['capital-gain', 'capital-loss']\nfeatures_log_transformed = pd.DataFrame(data = features_raw)\nfeatures_log_transformed[skewed] = features_raw[skewed].apply(lambda x: np.log(x + 1))\n\n# Visualize the new log distributions\ndistribution(features_log_transformed, transformed = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:28.548169Z","iopub.execute_input":"2022-07-31T23:58:28.548491Z","iopub.status.idle":"2022-07-31T23:58:29.119216Z","shell.execute_reply.started":"2022-07-31T23:58:28.548461Z","shell.execute_reply":"2022-07-31T23:58:29.117755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Normalizing Numerical Features\nIn addition to performing transformations on features that are highly skewed, it is often good practice to perform some type of scaling on numerical features. Applying a scaling to the data does not change the shape of each feature's distribution (such as `'capital-gain'` or `'capital-loss'` above); however, normalization ensures that each feature is treated equally when applying supervised learners. Note that once scaling is applied, observing the data in its raw form will no longer have the same original meaning","metadata":{}},{"cell_type":"code","source":"# Import sklearn.preprocessing.StandardScaler\nfrom sklearn.preprocessing import MinMaxScaler\n\n# Initialize a scaler, then apply it to the features\nscaler = MinMaxScaler() # default=(0, 1)\nnumerical = ['age', 'education-num', 'capital-gain', 'capital-loss', 'hours-per-week']\n\nfeatures_log_minmax_transform = pd.DataFrame(data = features_log_transformed)\nfeatures_log_minmax_transform[numerical] = scaler.fit_transform(features_log_transformed[numerical])\n\n# Show an example of a record with scaling applied\ndisplay(features_log_minmax_transform.head(n = 5))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:29.120932Z","iopub.execute_input":"2022-07-31T23:58:29.121418Z","iopub.status.idle":"2022-07-31T23:58:29.155524Z","shell.execute_reply.started":"2022-07-31T23:58:29.121379Z","shell.execute_reply":"2022-07-31T23:58:29.154090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Preprocessing","metadata":{}},{"cell_type":"code","source":"# One-hot encode the 'features_log_minmax_transform' data using pandas.get_dummies()\nfeatures_final = pd.get_dummies(features_log_minmax_transform)\n\n# Encode the 'income_raw' data to numerical values\nincome = income_raw.replace({ \"<=50K\" :0 , \">50K\" : 1})\n\n# Print the number of features after one-hot encoding\nencoded = list(features_final.columns)\nprint(\"{} total features after one-hot encoding.\".format(len(encoded)))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:29.157489Z","iopub.execute_input":"2022-07-31T23:58:29.157888Z","iopub.status.idle":"2022-07-31T23:58:29.266511Z","shell.execute_reply.started":"2022-07-31T23:58:29.157854Z","shell.execute_reply":"2022-07-31T23:58:29.265570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(encoded)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:29.271699Z","iopub.execute_input":"2022-07-31T23:58:29.272131Z","iopub.status.idle":"2022-07-31T23:58:29.280624Z","shell.execute_reply.started":"2022-07-31T23:58:29.272096Z","shell.execute_reply":"2022-07-31T23:58:29.279033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Shuffle and Split Data","metadata":{}},{"cell_type":"code","source":"# Import train_test_split\nfrom sklearn.model_selection import train_test_split\n\n# Split the 'features' and 'income' data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(features_final, \n                                                    income, \n                                                    test_size = 0.2, \n                                                    random_state = 0)\n\n# Show the results of the split\nprint(\"Training set has {} samples.\".format(X_train.shape[0]))\nprint(\"Testing set has {} samples.\".format(X_test.shape[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:29.282515Z","iopub.execute_input":"2022-07-31T23:58:29.282907Z","iopub.status.idle":"2022-07-31T23:58:29.330114Z","shell.execute_reply.started":"2022-07-31T23:58:29.282867Z","shell.execute_reply":"2022-07-31T23:58:29.328910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Implementation - Creating a Training and Predicting Pipeline","metadata":{}},{"cell_type":"code","source":"# Import two metrics from sklearn - fbeta_score and accuracy_score\nfrom sklearn.metrics import fbeta_score, accuracy_score\ndef train_predict(learner, sample_size, X_train, y_train, X_test, y_test): \n    '''\n    inputs:\n       - learner: the learning algorithm to be trained and predicted on\n       - sample_size: the size of samples (number) to be drawn from training set\n       - X_train: features training set\n       - y_train: income training set\n       - X_test: features testing set\n       - y_test: income testing set\n    '''\n    \n    results = {}\n    \n    # Fit the learner to the training data using slicing with 'sample_size' using .fit(training_features[:], training_labels[:])\n    start = time() # Get start time\n    learner = learner.fit(X_train[:sample_size],y_train[:sample_size])\n    end = time() # Get end time\n    \n    # Calculate the training time\n    results['train_time'] = end  - start\n        \n    # Get the predictions on the test set(X_test),\n    #       then get predictions on the first 300 training samples(X_train) using .predict()\n    start = time() # Get start time\n    predictions_test = learner.predict(X_test)\n    predictions_train = learner.predict(X_train[:300])\n    end = time() # Get end time\n    \n    # Calculate the total prediction time\n    results['pred_time'] = end - start\n            \n    # Compute accuracy on the first 300 training samples which is y_train[:300]\n    results['acc_train'] = accuracy_score(y_train[:300],predictions_train)\n        \n    # Compute accuracy on test set using accuracy_score()\n    results['acc_test'] =accuracy_score(y_test,predictions_test)\n    \n    # Compute F-score on the the first 300 training samples using fbeta_score()\n    results['f_train'] = fbeta_score(y_train[:300], predictions_train, beta=0.5)\n        \n    # Compute F-score on the test set which is y_test\n    results['f_test'] = fbeta_score(y_test,predictions_test,beta = 0.5)\n       \n    # Success\n    print(\"{} trained on {} samples.\".format(learner.__class__.__name__, sample_size))\n        \n    # Return the results\n    return results","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:29.331648Z","iopub.execute_input":"2022-07-31T23:58:29.332042Z","iopub.status.idle":"2022-07-31T23:58:29.345794Z","shell.execute_reply.started":"2022-07-31T23:58:29.331996Z","shell.execute_reply":"2022-07-31T23:58:29.344493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Implementation: Initial Model Evaluation","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\ndef visualize_classification_performance(results):\n    \"\"\"\n    Visualization code to display results of various learners.\n    \n    inputs:\n      - results: a list of dictionaries of the statistic results from 'train_predict_evaluate()'\n    \"\"\"\n  \n    # Create figure\n    sns.set()\n    sns.set_style(\"whitegrid\")\n    fig, ax = plt.subplots(2, 3, figsize = (11,7))\n    # print(\"VERSION:\")\n    # print(matplotlib.__version__)\n    # Constants\n    bar_width = 0.3\n    colors = [\"#e55547\", \"#4e6e8e\", \"#2ecc71\"]\n    \n    # Super loop to plot four panels of data\n    for k, learner in enumerate(results.keys()):\n        for j, metric in enumerate(['train_time', 'acc_train', 'f_train', 'pred_time', 'acc_test', 'f_test']):\n            for i in np.arange(3):\n                \n                # Creative plot code\n                ax[j//3, j%3].bar(i+k*bar_width, results[learner][i][metric], width = bar_width, color = colors[k])\n                ax[j//3, j%3].set_xticks([0.45, 1.45, 2.45])\n                ax[j//3, j%3].set_xticklabels([\"1%\", \"10%\", \"100%\"])\n                ax[j//3, j%3].set_xlabel(\"Training Set Size\")\n                ax[j//3, j%3].set_xlim((-0.1, 3.0))\n    \n    # Add unique y-labels\n    ax[0, 0].set_ylabel(\"Time (in seconds)\")\n    ax[0, 1].set_ylabel(\"Accuracy Score\")\n    ax[0, 2].set_ylabel(\"F-score\")\n    ax[1, 0].set_ylabel(\"Time (in seconds)\")\n    ax[1, 1].set_ylabel(\"Accuracy Score\")\n    ax[1, 2].set_ylabel(\"F-score\")\n    \n    # Add titles\n    ax[0, 0].set_title(\"Model Training\")\n    ax[0, 1].set_title(\"Accuracy Score on Training Subset\")\n    ax[0, 2].set_title(\"F-score on Training Subset\")\n    ax[1, 0].set_title(\"Model Predicting\")\n    ax[1, 1].set_title(\"Accuracy Score on Testing Set\")\n    ax[1, 2].set_title(\"F-score on Testing Set\")\n    \n    # Add horizontal lines for naive predictors\n    ax[0, 1].axhline(y = 1, xmin = -0.1, xmax = 3.0, linewidth = 1, color = 'k', linestyle = 'dashed')\n    ax[1, 1].axhline(y = 1, xmin = -0.1, xmax = 3.0, linewidth = 1, color = 'k', linestyle = 'dashed')\n    ax[0, 2].axhline(y = 1, xmin = -0.1, xmax = 3.0, linewidth = 1, color = 'k', linestyle = 'dashed')\n    ax[1, 2].axhline(y = 1, xmin = -0.1, xmax = 3.0, linewidth = 1, color = 'k', linestyle = 'dashed')\n    \n    # Set y-limits for score panels\n    ax[0, 1].set_ylim((0, 1))\n    ax[0, 2].set_ylim((0, 1))\n    ax[1, 1].set_ylim((0, 1))\n    ax[1, 2].set_ylim((0, 1))\n\n    # Create patches for the legend\n    patches = []\n    for i, learner in enumerate(results.keys()):\n        patches.append(mpatches.Patch(color = colors[i], label = learner))\n    plt.legend(handles = patches, bbox_to_anchor = (-.80, 2.53), \\\n               loc = 'upper center', borderaxespad = 0., ncol = 3, fontsize = 'x-large')\n    \n    # Aesthetics\n    plt.suptitle(\"Performance Metrics for Three Supervised Learning Models\", fontsize = 16, y = 1.10)\n    plt.tight_layout(pad=1, w_pad=2, h_pad=5.0)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:29.348259Z","iopub.execute_input":"2022-07-31T23:58:29.348998Z","iopub.status.idle":"2022-07-31T23:58:29.376135Z","shell.execute_reply.started":"2022-07-31T23:58:29.348956Z","shell.execute_reply":"2022-07-31T23:58:29.374374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import the three supervised learning models from sklearn\n\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\n# Initialize the three models\nclf_A = LogisticRegression(random_state= 43)\nclf_B = DecisionTreeClassifier(random_state= 42)\nclf_C = KNeighborsClassifier()\n\n# Calculate the number of samples for 1%, 10%, and 100% of the training data\n# samples_100 is the entire training set i.e. len(y_train)\n# samples_10 is 10% of samples_100 (ensure to set the count of the values to be `int` and not `float`)\n# samples_1 is 1% of samples_100 (ensure to set the count of the values to be `int` and not `float`)\nsamples_100 = len(y_train)\nsamples_10 = int(samples_100*0.1)\nsamples_1 = int(samples_100 * 0.01)\n\n# Collect results on the learners\nresults = {}\nfor clf in [clf_A, clf_B, clf_C]:\n    clf_name = clf.__class__.__name__\n    results[clf_name] = {}\n    for i, samples in enumerate([samples_1, samples_10, samples_100]):\n        results[clf_name][i] = \\\n        train_predict(clf, samples, X_train, y_train, X_test, y_test)\n\n# Run metrics visualization for the three supervised learning models chosen\nvisualize_classification_performance(results)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:29.378056Z","iopub.execute_input":"2022-07-31T23:58:29.378559Z","iopub.status.idle":"2022-07-31T23:58:42.558622Z","shell.execute_reply.started":"2022-07-31T23:58:29.378514Z","shell.execute_reply":"2022-07-31T23:58:42.557382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Choosing the Best Model","metadata":{}},{"cell_type":"markdown","source":"**The best model is Decision Tree**\n- low time in trianing and prediction\n- accuracy 80% and close to other algorithms (LG,KNN)\n- high accuracy in training and f-score \n","metadata":{}},{"cell_type":"markdown","source":"### Implementation: Model Tuning","metadata":{}},{"cell_type":"code","source":"clf_B.get_params()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:42.560279Z","iopub.execute_input":"2022-07-31T23:58:42.560651Z","iopub.status.idle":"2022-07-31T23:58:42.569012Z","shell.execute_reply.started":"2022-07-31T23:58:42.560618Z","shell.execute_reply":"2022-07-31T23:58:42.567943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import 'GridSearchCV', 'make_scorer', and any other necessary libraries\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.metrics import make_scorer\n# Initialize the classifier\nclf = DecisionTreeClassifier(random_state=42)\n\n# Create the parameters list you wish to tune, using a dictionary if needed.\n# parameters = {'parameter_1': [value1, value2], 'parameter_2': [value1, value2]}\nparameters = parameters = {'criterion': ['gini', 'entropy'],\n                           'max_depth': list(range(5,15))\n                          }\n\n# Perform grid search on the classifier \ngrid_obj = GridSearchCV(clf,parameters)\n\n# Fit the grid search object to the training data and find the optimal parameters using fit()\ngrid_fit =  grid_obj.fit(X_train, y_train)\n\n# Get the estimator\nbest_clf = grid_fit.best_estimator_\n\n# Make predictions using the unoptimized and model\npredictions = (clf.fit(X_train, y_train)).predict(X_test)\nbest_predictions = best_clf.predict(X_test)\n\n# Report the before-and-afterscores\nprint(\"Unoptimized model\\n------\")\nprint(\"Accuracy score on testing data: {:.4f}\".format(accuracy_score(y_test, predictions)))\nprint(\"F-score on testing data: {:.4f}\".format(fbeta_score(y_test, predictions, beta = 0.5)))\nprint(\"\\nOptimized Model\\n------\")\nprint(\"Final accuracy score on the testing data: {:.4f}\".format(accuracy_score(y_test, best_predictions)))\nprint(\"Final F-score on the testing data: {:.4f}\".format(fbeta_score(y_test, best_predictions, beta = 0.5)))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:58:42.570680Z","iopub.execute_input":"2022-07-31T23:58:42.571042Z","iopub.status.idle":"2022-07-31T23:59:05.873515Z","shell.execute_reply.started":"2022-07-31T23:58:42.571011Z","shell.execute_reply":"2022-07-31T23:59:05.871805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Final Model Evaluation\n","metadata":{}},{"cell_type":"markdown","source":"#### Results:\n\n|     Metric     | Unoptimized Model | Optimized Model |\n| :------------: | :---------------: | :-------------: | \n| Accuracy Score |        0.8186     |       0.8547    |\n| F-score        |        0.6279     |       0.7217    |\n\n- The accuracy of optimiaed model is 85.26 % and F-score 72.28% \n- Optimized model is better than unoptimized model.\n- Optimized model is better than naive predictor.","metadata":{}},{"cell_type":"markdown","source":"## Feature Importance","metadata":{}},{"cell_type":"markdown","source":"**The most important five features from my observation**\n\n- **capital-gain** ==> this feature is more important which is specified that the person is enable to make income above 50K\n- **education num** ==> number of educatoin years is more imortant Because the higher the number of years of education, the more it affects the income\n- **Hours-per-week** ==> the higher the number of hours in work. that help individuals to increase income\n- **age** ==> Age is very important in the income process, as there are some positions in which you need a number of years of experience\n- **marital-status** when the person is married this means he has a more chance to gain more income","metadata":{}},{"cell_type":"markdown","source":"### Extracting Feature Importance\n","metadata":{}},{"cell_type":"code","source":"def feature_plot(importances, X_train, y_train):\n    \n    # Display the five most important features\n    indices = np.argsort(importances)[::-1]\n    columns = X_train.columns.values[indices[:11]]\n    values = importances[indices][:11]\n\n    sns.set()\n    sns.set_style(\"whitegrid\")\n\n    # Creat the plot\n    fig = plt.figure(figsize = (12,5))\n    plt.title(\"Normalized Weights for First Five Most Predictive Features\", fontsize = 16)\n    plt.bar(np.arange(11), values, width = 0.2, align=\"center\", label = \"Feature Weight\")\n    # plt.bar(np.arange(11) - 0.3, np.cumsum(values), width = 0.2, align = \"center\", color = '#00A0A0', \\\n    #       label = \"Cumulative Feature Weight\")\n    plt.xticks(np.arange(11), columns)\n    plt.xlim((-0.5, 4.5))\n    plt.ylabel(\"Weight\", fontsize = 12)\n    plt.xlabel(\"Feature\", fontsize = 12)\n    \n    plt.legend(loc = 'upper center')\n    plt.tight_layout()\n    plt.show() ","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:05.875427Z","iopub.execute_input":"2022-07-31T23:59:05.875791Z","iopub.status.idle":"2022-07-31T23:59:05.886913Z","shell.execute_reply.started":"2022-07-31T23:59:05.875759Z","shell.execute_reply":"2022-07-31T23:59:05.885607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import a supervised learning model that has 'feature_importances_'\nfrom sklearn.ensemble import RandomForestClassifier\n\n# Train the supervised model on the training set using .fit(X_train, y_train)\nmodel = RandomForestClassifier()\nmodel.fit(X_train,y_train)\n\n# Extract the feature importances using .feature_importances_ \nimportances = model.feature_importances_ \n\n# Plot\nfeature_plot(importances, X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:05.888575Z","iopub.execute_input":"2022-07-31T23:59:05.889466Z","iopub.status.idle":"2022-07-31T23:59:13.282411Z","shell.execute_reply.started":"2022-07-31T23:59:05.889427Z","shell.execute_reply":"2022-07-31T23:59:13.281486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**I choose all features correctly and the oreder of feature not the same in figure**","metadata":{}},{"cell_type":"markdown","source":"### Feature Selection\nHow does a model perform if we only use a subset of all the available features in the data? With less features required to train, the expectation is that training and prediction time is much lower — at the cost of performance metrics. From the visualization above, we see that the top five most important features contribute more than half of the importance of **all** features present in the data. This hints that we can attempt to *reduce the feature space* and simplify the information required for the model to learn. The code cell below will use the same optimized model you found earlier, and train it on the same training set *with only the top five important features*. ","metadata":{}},{"cell_type":"code","source":"# Import functionality for cloning a model\nfrom sklearn.base import clone\n\n# Reduce the feature space\nX_train_reduced = X_train[X_train.columns.values[(np.argsort(importances)[::-1])[:5]]]\nX_test_reduced = X_test[X_test.columns.values[(np.argsort(importances)[::-1])[:5]]]\n\n# Train on the \"best\" model found from grid search earlier\nclf = (clone(best_clf)).fit(X_train_reduced, y_train)\n\n# Make new predictions\nreduced_predictions = clf.predict(X_test_reduced)\n\n# Report scores from the final model using both versions of data\nprint(\"Final Model trained on full data\\n------\")\nprint(\"Accuracy on testing data: {:.4f}\".format(accuracy_score(y_test, best_predictions)))\nprint(\"F-score on testing data: {:.4f}\".format(fbeta_score(y_test, best_predictions, beta = 0.5)))\nprint(\"\\nFinal Model trained on reduced data\\n------\")\nprint(\"Accuracy on testing data: {:.4f}\".format(accuracy_score(y_test, reduced_predictions)))\nprint(\"F-score on testing data: {:.4f}\".format(fbeta_score(y_test, reduced_predictions, beta = 0.5)))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.284106Z","iopub.execute_input":"2022-07-31T23:59:13.285342Z","iopub.status.idle":"2022-07-31T23:59:13.352473Z","shell.execute_reply.started":"2022-07-31T23:59:13.285293Z","shell.execute_reply":"2022-07-31T23:59:13.350987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"when reducing data the accuracy score and f-score is reduced\nAnd If training time was a factor,I would consider using the reduced data as training set. Because the change is very small and the model still doing good.","metadata":{}},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"# read test data from Kaggle\ndata_test = pd.read_csv('../input/udacity-mlcharity-competition/test_census.csv')\ndata_test","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.353907Z","iopub.execute_input":"2022-07-31T23:59:13.354232Z","iopub.status.idle":"2022-07-31T23:59:13.499627Z","shell.execute_reply.started":"2022-07-31T23:59:13.354200Z","shell.execute_reply":"2022-07-31T23:59:13.498501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# read example submission\nexample_submission = pd.read_csv('../input/udacity-mlcharity-competition/example_submission.csv')\nexample_submission","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.501387Z","iopub.execute_input":"2022-07-31T23:59:13.502105Z","iopub.status.idle":"2022-07-31T23:59:13.524867Z","shell.execute_reply.started":"2022-07-31T23:59:13.502074Z","shell.execute_reply":"2022-07-31T23:59:13.523901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_test.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.526375Z","iopub.execute_input":"2022-07-31T23:59:13.526892Z","iopub.status.idle":"2022-07-31T23:59:13.581021Z","shell.execute_reply.started":"2022-07-31T23:59:13.526861Z","shell.execute_reply":"2022-07-31T23:59:13.579767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# sum of total nan values in each column\ndata_test.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.582700Z","iopub.execute_input":"2022-07-31T23:59:13.583065Z","iopub.status.idle":"2022-07-31T23:59:13.635762Z","shell.execute_reply.started":"2022-07-31T23:59:13.583033Z","shell.execute_reply":"2022-07-31T23:59:13.634106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop the Unnamed column \ndata_test.drop('Unnamed: 0', axis = 1 ,inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.638145Z","iopub.execute_input":"2022-07-31T23:59:13.638695Z","iopub.status.idle":"2022-07-31T23:59:13.650290Z","shell.execute_reply.started":"2022-07-31T23:59:13.638643Z","shell.execute_reply":"2022-07-31T23:59:13.648967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fill nan values with the mean of each column ('age','education-num','hours-per-week')\ndata_test['age'].fillna(data_test['age'].mean(),inplace=True)\ndata_test['education-num'].fillna(data_test['education-num'].mean(),inplace=True)\ndata_test['hours-per-week'].fillna(data_test['hours-per-week'].mean(),inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.652762Z","iopub.execute_input":"2022-07-31T23:59:13.653281Z","iopub.status.idle":"2022-07-31T23:59:13.667176Z","shell.execute_reply.started":"2022-07-31T23:59:13.653232Z","shell.execute_reply":"2022-07-31T23:59:13.666123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fill nan values of columns ('capital-gain','capital-loss') with median because this columns are skewed\ndata_test['capital-gain'].fillna(data_test['capital-gain'].median(),inplace=True)\ndata_test['capital-loss'].fillna(data_test['capital-loss'].median(),inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.668878Z","iopub.execute_input":"2022-07-31T23:59:13.669292Z","iopub.status.idle":"2022-07-31T23:59:13.681593Z","shell.execute_reply.started":"2022-07-31T23:59:13.669258Z","shell.execute_reply":"2022-07-31T23:59:13.680533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fill all categorical columns with the mode of each column in test data\ncategorical_columns = ['workclass', 'occupation','sex','race', 'marital-status', 'education_level',  'native-country', 'relationship']\nfor col in categorical_columns:\n    data_test[col].fillna(data_test[col].mode()[0],inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.683838Z","iopub.execute_input":"2022-07-31T23:59:13.684222Z","iopub.status.idle":"2022-07-31T23:59:13.790315Z","shell.execute_reply.started":"2022-07-31T23:59:13.684190Z","shell.execute_reply":"2022-07-31T23:59:13.788941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_test.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.797822Z","iopub.execute_input":"2022-07-31T23:59:13.798241Z","iopub.status.idle":"2022-07-31T23:59:13.850603Z","shell.execute_reply.started":"2022-07-31T23:59:13.798208Z","shell.execute_reply":"2022-07-31T23:59:13.849174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Log-transform the skewed features\nskewed = ['capital-gain', 'capital-loss']\ndata_test[skewed] = data_test[skewed].apply(lambda x: np.log(x + 1))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.853124Z","iopub.execute_input":"2022-07-31T23:59:13.853608Z","iopub.status.idle":"2022-07-31T23:59:13.866146Z","shell.execute_reply.started":"2022-07-31T23:59:13.853564Z","shell.execute_reply":"2022-07-31T23:59:13.864777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# normalize the numirical columns\ndata_test[numerical] = scaler.fit_transform(data_test[numerical])\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.867901Z","iopub.execute_input":"2022-07-31T23:59:13.868486Z","iopub.status.idle":"2022-07-31T23:59:13.880937Z","shell.execute_reply.started":"2022-07-31T23:59:13.868453Z","shell.execute_reply":"2022-07-31T23:59:13.879584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get_dummies\ndata_test = pd.get_dummies(data_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.883380Z","iopub.execute_input":"2022-07-31T23:59:13.883982Z","iopub.status.idle":"2022-07-31T23:59:13.972431Z","shell.execute_reply.started":"2022-07-31T23:59:13.883920Z","shell.execute_reply":"2022-07-31T23:59:13.971326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_test.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.974144Z","iopub.execute_input":"2022-07-31T23:59:13.975516Z","iopub.status.idle":"2022-07-31T23:59:13.985186Z","shell.execute_reply.started":"2022-07-31T23:59:13.975464Z","shell.execute_reply":"2022-07-31T23:59:13.983786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_predict = best_clf.predict(data_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:13.986733Z","iopub.execute_input":"2022-07-31T23:59:13.987796Z","iopub.status.idle":"2022-07-31T23:59:14.053920Z","shell.execute_reply.started":"2022-07-31T23:59:13.987754Z","shell.execute_reply":"2022-07-31T23:59:14.052316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example_submission['income'] = data_predict","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:14.055764Z","iopub.execute_input":"2022-07-31T23:59:14.056677Z","iopub.status.idle":"2022-07-31T23:59:14.062008Z","shell.execute_reply.started":"2022-07-31T23:59:14.056639Z","shell.execute_reply":"2022-07-31T23:59:14.060935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the submission file\nexample_submission.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T23:59:14.063810Z","iopub.execute_input":"2022-07-31T23:59:14.064502Z","iopub.status.idle":"2022-07-31T23:59:14.151096Z","shell.execute_reply.started":"2022-07-31T23:59:14.064460Z","shell.execute_reply":"2022-07-31T23:59:14.149554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}