{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"A sample script to generate pre-imputed versions of the test and train sets. Imputation is performed on a per-product-code basis. This guarantees several useful properties:\n\n+ It corresponds to an intuition about lab testing being done in separate per-product sessions.\n+ There is no \"leakage\" between train and test set.\n+ If you use GroupKFold to cross-validate per-product-code, there will be no leakage between cross-validation folds.\n\nThus, there is no risk in pre-imputing the data. As a bonus, the smaller batch sizes mean that KNN imputation runs much faster. This gives you the option to impute as a preprocessing step, or as part of data loading, or as part of your cross-validation.\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom sklearn.impute import KNNImputer","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:53:08.378674Z","iopub.execute_input":"2022-08-04T17:53:08.379494Z","iopub.status.idle":"2022-08-04T17:53:09.706437Z","shell.execute_reply.started":"2022-08-04T17:53:08.379381Z","shell.execute_reply":"2022-08-04T17:53:09.705126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A standalone imputation function. Extra care has been taken to leave columns in exactly the same form as the original input.","metadata":{}},{"cell_type":"code","source":"def kimpute(X: pd.DataFrame, n=10, weights=\"uniform\"):\n    \"\"\"Impute missing values in TPS2208 data.\n    \n    Imputation is performed over separate \"per-product-code\" batches, and is designed to leave all non-imputed \n    data in the exact same format as before imputation.\"\"\"\n    def transform(X):\n        return pd.DataFrame(\n            KNNImputer(n_neighbors=n, weights=weights).fit_transform(X), index=X.index,\n            columns=X.columns)\n\n    cats = [\"product_code\", \"attribute_0\", \"attribute_1\", \"attribute_2\", \"attribute_3\"]\n    ints = [\"measurement_0\", \"measurement_1\", \"measurement_2\"]\n    right = pd.concat([transform(gdf.drop(columns=cats)) for g, gdf in X.groupby(\"product_code\")],\n                      axis=\"rows\")\n    right[ints] = right[ints].round().astype(int)\n    return pd.concat([X[cats], right], axis=\"columns\").reindex(columns=X.columns)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:53:09.708689Z","iopub.execute_input":"2022-08-04T17:53:09.709463Z","iopub.status.idle":"2022-08-04T17:53:09.719033Z","shell.execute_reply.started":"2022-08-04T17:53:09.709421Z","shell.execute_reply":"2022-08-04T17:53:09.717949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Separately impute train and test sets (to make extra-super-duper sure that there is no leakage) and write back out to CSV files. Probably not necessary, given the speed of ```kimpute```, but it's a nice proof of concept.","metadata":{}},{"cell_type":"code","source":"Xy_train = pd.read_csv(\"../input/tabular-playground-series-aug-2022/train.csv\", index_col='id')\nXt, yt = kimpute(Xy_train.drop(columns=[\"failure\"])), Xy_train.failure\nXt.assign(failure=yt).to_csv(\"train-imputed.csv\")\n\nXv = kimpute(pd.read_csv(\"../input/tabular-playground-series-aug-2022/test.csv\", index_col='id'))\nXv.to_csv(\"test-imputed.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:53:09.720732Z","iopub.execute_input":"2022-08-04T17:53:09.721862Z","iopub.status.idle":"2022-08-04T17:53:23.744411Z","shell.execute_reply.started":"2022-08-04T17:53:09.721820Z","shell.execute_reply":"2022-08-04T17:53:23.743198Z"},"trusted":true},"execution_count":null,"outputs":[]}]}