{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Different Types Imputation Techniques : Tutorial","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:350%;text-align:center;display:fill;border-radius:5px;background-color:#8c1707;overflow:hidden;font-weight:500\">Different Types Imputation Techniques : Tutorial</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#d13621;overflow:hidden;font-weight:500\">What is Imputation?</div>","metadata":{}},{"cell_type":"markdown","source":"# 1. What is Imputation?","metadata":{}},{"cell_type":"markdown","source":"In statistics, imputation is the process of replacing missing data with substituted values. When substituting for a data point, it is known as \"unit imputation\"; when substituting for a component of a data point, it is known as \"item imputation\".\nImputation is a technique used for replacing the missing data with some substitute value to retain most of the data/information of the dataset. These techniques are used because removing the data from the dataset every time is not feasible and can lead to a reduction in the size of the dataset to a large extend, which not only raises concerns for biasing the dataset but also leads to incorrect analysis.","metadata":{}},{"cell_type":"markdown","source":"![](https://editor.analyticsvidhya.com/uploads/63685Imputation.JPG)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#d13621;overflow:hidden;font-weight:500\">Why Imputation is Important?</div>","metadata":{}},{"cell_type":"markdown","source":"# 2. Why Imputation is Important?","metadata":{}},{"cell_type":"markdown","source":"- **Incompatible with most of the Python libraries used in Machine Learning**:- Yes, you read it right. While using the libraries for ML(the most common is skLearn), they don’t have a provision to automatically handle these missing data and can lead to errors.\n- **Distortion in Dataset**:- A huge amount of missing data can cause distortions in the variable distribution i.e it can increase or decrease the value of a particular category in the dataset.\n- **Affects the Final Model**:- the missing data can cause a bias in the dataset and can lead to a faulty analysis by the model.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#d13621;overflow:hidden;font-weight:500\">Imputation Tecniques Discussed</div>","metadata":{}},{"cell_type":"markdown","source":"# 3. Imputation Tecniques Discussed","metadata":{}},{"cell_type":"markdown","source":"#### 1. Complete Case Analysis(CCA)\n#### 2. Arbitrary Value Imputation\n#### 3. Frequent Category Imputation\n#### 4. statistical Values imputation\n#### 5. Using a linear regression\n#### 6. Iterative imputation\n#### 7. Nearest neighbors imputation\n#### 8. Marking imputed values","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#d13621;overflow:hidden;font-weight:500\">Exploring Datasets</div>","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport os\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom scipy import stats\nimport warnings\nimport json\nimport random\nimport sys","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-09-18T15:01:02.465048Z","iopub.execute_input":"2022-09-18T15:01:02.466134Z","iopub.status.idle":"2022-09-18T15:01:03.564956Z","shell.execute_reply.started":"2022-09-18T15:01:02.465989Z","shell.execute_reply":"2022-09-18T15:01:03.563920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df=pd.read_csv(\"/kaggle/input/tabular-playground-series-jun-2022/data.csv\")\nsubmission = pd.read_csv(\"/kaggle/input/tabular-playground-series-jun-2022/sample_submission.csv\", index_col='row-col')\n\ndf=df[:1000]\n\nwarnings.filterwarnings('ignore')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-09-18T15:01:03.570982Z","iopub.execute_input":"2022-09-18T15:01:03.571365Z","iopub.status.idle":"2022-09-18T15:01:25.234895Z","shell.execute_reply.started":"2022-09-18T15:01:03.571329Z","shell.execute_reply":"2022-09-18T15:01:25.234121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Basic Exploration","metadata":{}},{"cell_type":"code","source":"print(\"Data Shape: There are {:,.0f} rows and {:,.0f} columns.\\nMissing values = {}, Duplicates = {}.\\n\".\n      format(df.shape[0], df.shape[1],df.isna().sum().sum(), df.duplicated().sum()))","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:25.236017Z","iopub.execute_input":"2022-09-18T15:01:25.236888Z","iopub.status.idle":"2022-09-18T15:01:25.263114Z","shell.execute_reply.started":"2022-09-18T15:01:25.236851Z","shell.execute_reply":"2022-09-18T15:01:25.262015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import missingno as msno\nmsno.matrix(df.sample(100))","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:25.264744Z","iopub.execute_input":"2022-09-18T15:01:25.265121Z","iopub.status.idle":"2022-09-18T15:01:25.705686Z","shell.execute_reply.started":"2022-09-18T15:01:25.265072Z","shell.execute_reply":"2022-09-18T15:01:25.704940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set(rc={'figure.figsize':(24,20)})\nfor i, column in enumerate(list(df.columns), 1):\n    plt.subplot(8,11,i)\n    p=sns.histplot(x=column,data=df.sample(1000),stat='count',kde=True,color='green')","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:25.708095Z","iopub.execute_input":"2022-09-18T15:01:25.708714Z","iopub.status.idle":"2022-09-18T15:01:39.473434Z","shell.execute_reply.started":"2022-09-18T15:01:25.708675Z","shell.execute_reply":"2022-09-18T15:01:39.472620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Very interesting co-relations!","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#d13621;overflow:hidden;font-weight:500\">Imputation Tecniques Discussed</div>","metadata":{}},{"cell_type":"markdown","source":"# 5. Imputation Techniques of Different Types of Variables.","metadata":{}},{"cell_type":"markdown","source":"![](https://editor.analyticsvidhya.com/uploads/30381Imputation%20Techniques%20types.JPG)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:5px;color:black;margin:0;font-size:185%;text-align:left;display:fill;border-radius:5px;background-color:#daf0ec;overflow:hidden;font-weight:500\">Estimators that handle NaN values</div>\n\nSome estimators are designed to handle NaN values without preprocessing. Below is the list of these estimators, classified by type (cluster, regressor, classifier, transform) :\n\n<p></br></p>\n<div style=\"padding:5px;color:black;margin:0;font-size:150%;text-align:left;display:fill;border-radius:5px;background-color:#daf0ec;overflow:hidden;font-weight:500\">Estimators that allow NaN values for type regressor:</div>\n\n- HistGradientBoostingRegressor\n\n<p></br></p>\n<div style=\"padding:5px;color:black;margin:0;font-size:150%;text-align:left;display:fill;border-radius:5px;background-color:#daf0ec;overflow:hidden;font-weight:500\">Estimators that allow NaN values for type classifier:</div>\n\n- HistGradientBoostingClassifier\n\n<p></br></p>\n<div style=\"padding:5px;color:black;margin:0;font-size:150%;text-align:left;display:fill;border-radius:5px;background-color:#daf0ec;overflow:hidden;font-weight:500\">Estimators that allow NaN values for type transformer :</div>\n\n- IterativeImputer\n- KNNImputer\n- MaxAbsScaler\n- MinMaxScaler\n- MissingIndicator\n- PowerTransformer\n- QuantileTransformer\n- RobustScaler\n- SimpleImputer\n- StandardScaler\n- VarianceThreshold","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#f7055e;overflow:hidden;font-weight:500\">Complete Case Analysis(CCA)</div>","metadata":{}},{"cell_type":"markdown","source":"# >> 1. Complete Case Analysis(CCA)","metadata":{}},{"cell_type":"markdown","source":"This is a quite straightforward method of handling the Missing Data, which directly removes the rows that have missing data i.e we consider only those rows where we have complete data i.e data is not missing. This method is also popularly known as “Listwise deletion”.\n\n## Assumptions:-\n- Data is Missing At Random(MAR).\n- Missing data is completely removed from the table.\n\n## Advantages:- \n- Easy to implement.\n- No Data manipulation required.\n","metadata":{}},{"cell_type":"markdown","source":"## Limitations:-\n- Deleted data can be informative.\n- Can lead to the deletion of a large part of the data.\n- Can create a bias in the dataset, if a large amount of a particular type of variable is deleted from it.\n- The production model will not know what to do with Missing data.\n\n## When to Use:-\n- Data is MAR(Missing At Random).\n- Good for Mixed, Numerical, and Categorical data.\n- Missing data is not more than 5% – 6% of the dataset.\n- Data doesn’t contain much information and will not bias the dataset.\n\n## Code:-","metadata":{}},{"cell_type":"code","source":"#shape of DF\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.474338Z","iopub.execute_input":"2022-09-18T15:01:39.474635Z","iopub.status.idle":"2022-09-18T15:01:39.481630Z","shell.execute_reply.started":"2022-09-18T15:01:39.474606Z","shell.execute_reply":"2022-09-18T15:01:39.480224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Finding the columns that have Null Values(Missing Data) \n## We are using a for loop for all the columns present in dataset with average null values greater than 0\nna_variables = [ var for var in df.columns if df[var].isnull().mean() > 0 ]","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.483283Z","iopub.execute_input":"2022-09-18T15:01:39.484087Z","iopub.status.idle":"2022-09-18T15:01:39.517833Z","shell.execute_reply.started":"2022-09-18T15:01:39.484051Z","shell.execute_reply":"2022-09-18T15:01:39.516996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_na = df[na_variables].isnull().mean()\n## Implementing the CCA techniques to remove Missing Data\ndata_cca = df.dropna(axis=0)\n## Verifying the final shape of the remaining dataset\ndata_cca.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.519160Z","iopub.execute_input":"2022-09-18T15:01:39.519653Z","iopub.status.idle":"2022-09-18T15:01:39.535759Z","shell.execute_reply.started":"2022-09-18T15:01:39.519618Z","shell.execute_reply":"2022-09-18T15:01:39.534772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## But we won't use it in this competition.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#f7055e;overflow:hidden;font-weight:500\">Arbitrary Value Imputation</div>","metadata":{}},{"cell_type":"markdown","source":"# >> 2. Arbitrary Value Imputation","metadata":{}},{"cell_type":"markdown","source":"This is an important technique used in Imputation as it can handle both the Numerical and Categorical variables. This technique states that we group the missing values in a column and assign them to a new value that is far away from the range of that column. Mostly we use values like 99999999 or -9999999 or “Missing” or “Not defined” for numerical & categorical variables.","metadata":{}},{"cell_type":"markdown","source":"## Assumptions:-\n- Data is not Missing At Random.\n- The missing data is imputed with an arbitrary value that is not part of the dataset or Mean/Median/Mode of data.\n\n## Advantages:-\n- Easy to implement.\n- We can use it in production.\n- It retains the importance of “missing values” if it exists.\n\n## Disadvantages:-\n- Can distort original variable distribution.\n- Arbitrary values can create outliers.\n- Extra caution required in selecting the Arbitrary value.","metadata":{}},{"cell_type":"markdown","source":"## When to Use:-\n- When data is not MAR(Missing At Random).\n- Suitable for All.\n\n## Code:-","metadata":{}},{"cell_type":"code","source":"train_df=df.copy()\nna_variables = [ var for var in train_df.columns if train_df[var].isnull().mean() > 0 ]","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.537236Z","iopub.execute_input":"2022-09-18T15:01:39.537838Z","iopub.status.idle":"2022-09-18T15:01:39.567711Z","shell.execute_reply.started":"2022-09-18T15:01:39.537803Z","shell.execute_reply":"2022-09-18T15:01:39.566251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Here nan represent Missing Data\n## Using Arbitary Imputation technique, we will Impute missing Gender with \"Missing\"  {You can use any other value also}\narb_impute = train_df['F_1_0'].fillna(random.choice(train_df['F_1_0'].unique()))\narb_impute.unique()\n#continue for each column","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-09-18T15:01:39.569314Z","iopub.execute_input":"2022-09-18T15:01:39.569590Z","iopub.status.idle":"2022-09-18T15:01:39.589788Z","shell.execute_reply.started":"2022-09-18T15:01:39.569536Z","shell.execute_reply":"2022-09-18T15:01:39.589128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## But we won't use it in this competition.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#f7055e;overflow:hidden;font-weight:500\">Frequent Category Imputation</div>","metadata":{}},{"cell_type":"markdown","source":"# >> 3. Frequent Category Imputation","metadata":{}},{"cell_type":"markdown","source":"This technique says to replace the missing value with the variable with the highest frequency or in simple words replacing the values with the Mode of that column. This technique is also referred to as Mode Imputation.","metadata":{}},{"cell_type":"markdown","source":"## Assumptions:-\n- Data is missing at random.\n- There is a high probability that the missing data looks like the majority of the data.\n\n## Advantages:-\n- Implementation is easy.\n- We can obtain a complete dataset in very little time.\n- We can use this technique in the production model.\n\n## Disadvantages:-\n- The higher the percentage of missing values, the higher will be the distortion.\n- May lead to over-representation of a particular category.\n- Can distort original variable distribution.","metadata":{}},{"cell_type":"markdown","source":"## When to Use:-\n- Data is Missing at Random(MAR)\n- Missing data is not more than 5% – 6% of the dataset.\n\n## Code:- ","metadata":{}},{"cell_type":"code","source":"train_df['F_1_0'].groupby(train_df['F_1_0']).count()","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.590655Z","iopub.execute_input":"2022-09-18T15:01:39.590905Z","iopub.status.idle":"2022-09-18T15:01:39.601464Z","shell.execute_reply.started":"2022-09-18T15:01:39.590882Z","shell.execute_reply":"2022-09-18T15:01:39.600552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['F_1_0'].mode()","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.602983Z","iopub.execute_input":"2022-09-18T15:01:39.603444Z","iopub.status.idle":"2022-09-18T15:01:39.612997Z","shell.execute_reply.started":"2022-09-18T15:01:39.603409Z","shell.execute_reply":"2022-09-18T15:01:39.611998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Using Frequent Category Imputer\nfrq_impute = train_df['F_1_0'].fillna('-0.665735')\nfrq_impute.unique()\n\n#continue for each column","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-09-18T15:01:39.614386Z","iopub.execute_input":"2022-09-18T15:01:39.615022Z","iopub.status.idle":"2022-09-18T15:01:39.628155Z","shell.execute_reply.started":"2022-09-18T15:01:39.614978Z","shell.execute_reply":"2022-09-18T15:01:39.627071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## But we won't use it in this competition.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#f7055e;overflow:hidden;font-weight:500\">Statistical Values imputation</div>","metadata":{}},{"cell_type":"markdown","source":"# >> 4. statistical Values imputation","metadata":{}},{"cell_type":"markdown","source":"This is done using statistical values like mean, median. However, none of these guarantees unbiased data, especially if there are many missing values.\n\nMean is most useful when the original data is not skewed, while the median is more robust, not sensitive to outliers, and thus used when data is skewed.\n\n In a normally distributed data, one can get all the values that are within 2 standard deviations from the mean. Next, fill in the missing values by generating random numbers between **(mean — 2 * std) & (mean + 2 * std)**","metadata":{}},{"cell_type":"code","source":"xdf=df.copy()\n\naverage=xdf.F_1_0.mean()\nstd=xdf.F_1_0.std()\ncount_nan_age=xdf.F_1_0.isnull().sum()\n\nrand = np.random.randint(average - 2*std, average + 2*std, size = count_nan_age)\n\nprint(\"Before\",xdf.F_1_0.isnull().sum())\n\nxdf[\"F_1_0\"][np.isnan(xdf[\"F_1_0\"])] = rand\nprint(\"After\",xdf.F_1_0.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.630869Z","iopub.execute_input":"2022-09-18T15:01:39.631212Z","iopub.status.idle":"2022-09-18T15:01:39.642758Z","shell.execute_reply.started":"2022-09-18T15:01:39.631187Z","shell.execute_reply":"2022-09-18T15:01:39.641695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#f7055e;overflow:hidden;font-weight:500\">Using a linear regression</div>","metadata":{}},{"cell_type":"markdown","source":"# >> 5. Using a linear regression. \n\nBased on the existing data, one can calculate the best fit line between two variables, from checking correlations.\n\nIt is worth mentioning that linear regression models are sensitive to outliers.","metadata":{}},{"cell_type":"code","source":"#seeing co-relation\nsns.set(rc={'figure.figsize':(8,8)})\nsns.scatterplot(data=xdf, x=\"F_4_8\", y=\"F_4_11\",legend=\"full\",sizes=(20, 200))","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.644341Z","iopub.execute_input":"2022-09-18T15:01:39.644993Z","iopub.status.idle":"2022-09-18T15:01:39.963194Z","shell.execute_reply.started":"2022-09-18T15:01:39.644955Z","shell.execute_reply":"2022-09-18T15:01:39.961276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Interesting corelation!","metadata":{}},{"cell_type":"code","source":"print(\"F_4_11\",xdf.F_4_11.isnull().sum())\nprint(\"F_4_8\",xdf.F_4_8.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.970095Z","iopub.execute_input":"2022-09-18T15:01:39.970764Z","iopub.status.idle":"2022-09-18T15:01:39.985084Z","shell.execute_reply.started":"2022-09-18T15:01:39.970704Z","shell.execute_reply":"2022-09-18T15:01:39.984160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Now, we can use liner regression to do it; for each column. We aso can use multiple features for each linear regression too!","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#f7055e;overflow:hidden;font-weight:500\">Iterative imputation with XGBoost</div>","metadata":{}},{"cell_type":"markdown","source":"# >> 6. Iterative imputation","metadata":{}},{"cell_type":"markdown","source":"Iterative imputation refers to a process where each feature is modeled as a function of the other features, e.g. a regression problem where missing values are predicted. Each feature is imputed sequentially, one after the other, allowing prior imputed values to be used as part of a model in predicting subsequent features.\n\nIt is iterative because this process is repeated multiple times, allowing ever improved estimates of missing values to be calculated as missing values across all features are estimated.\n\nThis approach may be generally referred to as fully conditional specification (FCS) or multivariate imputation by chained equations (MICE).","metadata":{}},{"cell_type":"code","source":"# loading modules\nfrom sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import IterativeImputer\nimport xgboost\n\ndata=df.copy()\n\n#setting up the imputer\n\nimp = IterativeImputer(\n    estimator=xgboost.XGBRegressor(\n        n_estimators=5,\n        random_state=1,\n        tree_method='gpu_hist',\n    ),\n    missing_values=np.nan,\n    max_iter=5,\n    initial_strategy='mean',\n    imputation_order='ascending',\n    verbose=2,\n    random_state=1\n)\n\ndata[:] = imp.fit_transform(data)\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:01:39.986693Z","iopub.execute_input":"2022-09-18T15:01:39.987405Z","iopub.status.idle":"2022-09-18T15:02:10.165435Z","shell.execute_reply.started":"2022-09-18T15:01:39.987365Z","shell.execute_reply":"2022-09-18T15:02:10.164544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- You can use any other model, too!","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#f7055e;overflow:hidden;font-weight:500\">Nearest neighbors imputation</div>","metadata":{}},{"cell_type":"markdown","source":"#  >> 7. Nearest neighbors imputation","metadata":{}},{"cell_type":"markdown","source":"The KNNImputer class provides imputation for filling in missing values using the k-Nearest Neighbors approach. By default, a euclidean distance metric that supports missing values, nan_euclidean_distances, is used to find the nearest neighbors. Each missing feature is imputed using values from n_neighbors nearest neighbors that have a value for the feature. The feature of the neighbors are averaged uniformly or weighted by distance to each neighbor. If a sample has more than one feature missing, then the neighbors for that sample can be different depending on the particular feature being imputed. When the number of available neighbors is less than n_neighbors and there are no defined distances to the training set, the training set average for that feature is used during imputation. If there is at least one neighbor with a defined distance, the weighted or unweighted average of the remaining neighbors will be used during imputation. If a feature is always missing in training, it is removed during transform. ","metadata":{}},{"cell_type":"code","source":"from sklearn.impute import KNNImputer\nnan = np.nan\nX = [[1, 2, nan], [3, 4, 3], [nan, 6, 5], [8, 8, 7]]\nimputer = KNNImputer(n_neighbors=2, weights=\"uniform\")\nimputed_x=imputer.fit_transform(X)\nprint(X,imputed_x)","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:02:10.167006Z","iopub.execute_input":"2022-09-18T15:02:10.167693Z","iopub.status.idle":"2022-09-18T15:02:10.177746Z","shell.execute_reply.started":"2022-09-18T15:02:10.167650Z","shell.execute_reply":"2022-09-18T15:02:10.176567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#f7055e;overflow:hidden;font-weight:500\">Marking imputed values</div>","metadata":{}},{"cell_type":"markdown","source":"# >> 8. Marking imputed values","metadata":{}},{"cell_type":"markdown","source":"The MissingIndicator transformer is useful to transform a dataset into corresponding binary matrix indicating the presence of missing values in the dataset. This transformation is useful in conjunction with imputation. When using imputation, preserving the information about which values had been missing can be informative. Note that both the SimpleImputer and IterativeImputer have the boolean parameter add_indicator (False by default) which when set to True provides a convenient way of stacking the output of the MissingIndicator transformer with the output of the imputer.\n\nNaN is usually used as the placeholder for missing values. However, it enforces the data type to be float. The parameter missing_values allows to specify other placeholder such as integer. In the following example, we will use -1 as missing values:","metadata":{}},{"cell_type":"code","source":"from sklearn.impute import MissingIndicator\nX = np.array([[-1, -1, 1, 3],\n              [4, -1, 0, -1],\n              [8, -1, 1, 0]])\nindicator = MissingIndicator(missing_values=-1)\nmask_missing_values_only = indicator.fit_transform(X)\nmask_missing_values_only","metadata":{"execution":{"iopub.status.busy":"2022-09-18T15:02:10.179294Z","iopub.execute_input":"2022-09-18T15:02:10.179687Z","iopub.status.idle":"2022-09-18T15:02:10.207673Z","shell.execute_reply.started":"2022-09-18T15:02:10.179650Z","shell.execute_reply":"2022-09-18T15:02:10.206786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- later we can use this to develop a method for imputation.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#d13621;overflow:hidden;font-weight:500\">Comments</div>","metadata":{}},{"cell_type":"markdown","source":"# 6. Comments\nOne can take the random approach where we fill in the missing value with a random value. Taking this approach one step further, one can first divide the dataset into two groups (strata), based on some characteristic, say gender, and then fill in the missing values for different genders separately, at random.\n\nIn sequential hot-deck imputation, the column containing missing values is sorted according to auxiliary variable(s) so that records that have similar auxiliaries occur sequentially. Next, each missing value is filled in with the value of the first following available record.\n\nWhat is more interesting is that 𝑘 nearest neighbour imputation, which classifies similar records and put them together, can also be utilized. A missing value is then filled out by finding first the 𝑘 records closest to the record with missing values. Next, a value is chosen from (or computed out of) the 𝑘 nearest neighbours. In the case of computing, statistical methods like mean (as discussed before) can be used.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#d13621;overflow:hidden;font-weight:500\">Other Related Notebooks</div>","metadata":{}},{"cell_type":"markdown","source":"# Other Related Notebooks\n- ### [✅ 07 Cross Validation Methods ➡️ Tutorial 📊](https://www.kaggle.com/code/azminetoushikwasi/07-cross-validation-methods-tutorial)\n- ### [➡️[Tutorial] 🛠 Feature Engineering ⚙📝](https://www.kaggle.com/code/azminetoushikwasi/tutorial-feature-engineering)\n- ### [📋 Bias-Variance Tradeoff ➡️ with NumPy & Seaborn](https://www.kaggle.com/code/azminetoushikwasi/bias-variance-tradeoff-with-numpy-seaborn)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:250%;text-align:center;display:fill;border-radius:5px;background-color:#d13621;overflow:hidden;font-weight:500\">References</div>","metadata":{}},{"cell_type":"markdown","source":"# References\n- [Defining, Analysing, and Implementing Imputation Techniques](https://www.analyticsvidhya.com/blog/2021/06/defining-analysing-and-implementing-imputation-techniques/)\n- [The Ultimate Guide to Data Cleaning](https://towardsdatascience.com/the-ultimate-guide-to-data-cleaning-3969843991d4#b498)\n- [Scikit Learn Documentation](https://scikit-learn.org/stable/modules/impute.html)\n- [TPS Jun 2022 IterativeImputer baseline](https://www.kaggle.com/code/hiro5299834/tps-jun-2022-iterativeimputer-baseline)","metadata":{}}]}