{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom sklearn import calibration as skc, linear_model as sklm\ntransform = lambda X: pd.DataFrame({\n    **pd.get_dummies(X.fillna(0)).to_dict(),\n    **{f\"{c}_NA\": X[c].isna() for c in X.columns if X[c].isna().any()}})\nXy_train = pd.read_csv(\"../input/tabular-playground-series-aug-2022/train.csv\", index_col='id')\nXt, yt = transform(Xy_train.drop(columns=[\"failure\"])), Xy_train.failure\nXv = transform(pd.read_csv(\"../input/tabular-playground-series-aug-2022/test.csv\", index_col='id'))\nXv = pd.DataFrame({col: Xv.get(col, 0) for col in Xt.columns})\ny_pred = skc.CalibratedClassifierCV(sklm.RidgeClassifier()).fit(Xt, yt).predict_proba(Xv)[:, 1]\npd.DataFrame({'id': Xv.index, 'failure': y_pred}).to_csv(\"submission.csv\", index=False)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-02T23:42:27.810319Z","iopub.execute_input":"2022-08-02T23:42:27.811166Z","iopub.status.idle":"2022-08-02T23:42:34.161197Z","shell.execute_reply.started":"2022-08-02T23:42:27.811045Z","shell.execute_reply":"2022-08-02T23:42:34.159698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explanation:\nThat's really all there is. 12 lines for (at the time of writing) position #106 out of 259. This is intended as a confirmation that an effective entry doesn't have to be complex. It is *not*, obviously an example of how to write clear code, or show your work, or really anything else worthwhile. In particular, the short variable names allow me to meet the self-imposed goal of keeping all the lines under 100 characters.\n\nWe transform the data by (a) imputing every missing value as '0'; (b) doing one-hot encoding as necessary; and (c) adding columns identifying which values are imputed. (This last step is what lets us impute '0' rather than computing a mean or some such.) Line 10 aligns the test set columns with the train set, adding (zero-filled) missing columns and dropping extra columns.\n\nActual training and prediction are done with off-the-shelf models. The \"RidgeClassifier\" is great at ignoring redundant data and, for whatever reason, linear classifiers beat out gradient boosting for this set. We need the \"CalibratedClassifierCV\" in order to get the prediction probabilities for the output.\n\nFeel free to take this as a starting point, clean it up for readability, and then proceed to the most important final step: Profit!!\n","metadata":{}}]}