{"cells":[{"metadata":{},"cell_type":"markdown","source":"# SIIM-ISIC : An Ensemble Beginner's Approach","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"> This notebook consists of a simple approach to solve SIIM ISIC Melanoma Classification Problem. <br>\n> Using ensemble of my main model submission and a simple XGBoost model to improve my final blend model's performance!<br>\n> NOTE: I have achieved nearly **94.68%** score using this strategy and which gave me a good intial push for the competition on the Leaderboard.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nimport xgboost as xgb\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.datasets import load_iris","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"train= pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ntest= pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.target.value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> It is clear that the data is highly imbalanced and skewed towards the 0 target class","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Handling missing values in Train and Test Datasets ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['sex'] = train['sex'].fillna('na')\ntrain['age_approx'] = train['age_approx'].fillna(0)\ntrain['anatom_site_general_challenge'] = train['anatom_site_general_challenge'].fillna('na')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test['sex'] = test['sex'].fillna('na')\ntest['age_approx'] = test['age_approx'].fillna(0)\ntest['anatom_site_general_challenge'] = test['anatom_site_general_challenge'].fillna('na')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Feature Engineering","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['sex'] = train['sex'].astype(\"category\").cat.codes +1\ntrain['anatom_site_general_challenge'] = train['anatom_site_general_challenge'].astype(\"category\").cat.codes +1\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test['sex'] = test['sex'].astype(\"category\").cat.codes +1\ntest['anatom_site_general_challenge'] = test['anatom_site_general_challenge'].astype(\"category\").cat.codes +1\ntest.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Data Manipulation for training and validation","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"x_train = train[['sex', 'age_approx','anatom_site_general_challenge']]\ny_train = train['target']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x_test = test[['sex', 'age_approx','anatom_site_general_challenge']]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_DMatrix = xgb.DMatrix(x_train, label= y_train)\ntest_DMatrix = xgb.DMatrix(x_test)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Building a simple XGBoost model for training \n(Hyperparameter tuning already done and model best values used)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"xgb_model = xgb.XGBClassifier(n_estimators=2000, \n                        max_depth=8, \n                        objective='multi:softprob',\n                        seed=0,  \n                        nthread=-1, \n                        learning_rate=0.15, \n                        num_class = 2, \n                        scale_pos_weight = (32542/584))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"xgb_model.fit(x_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"xgb_pred_result = xgb_model.predict_proba(x_test)[:,1]\nprint(xgb_pred_result)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"xgb_df = pd.DataFrame({\n        \"image_name\": test[\"image_name\"],\n        \"target\": xgb_pred_result\n    })\n\nxgb_df.to_csv('tuned_XGBClassifier_submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Loading older submission files","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"main_submission = pd.read_csv('../input/my-siim-isic-submissions/my_siim_isic_main_submission.csv')\nefficient_b7 = pd.read_csv('../input/my-siim-isic-submissions/EfficientNetB7_submission.csv')\nefficient_b7_blend_6 = pd.read_csv('../input/my-siim-isic-submissions/EfficientNetB7_submission_Blend_6.csv')\nmodel_blend_0_6 = pd.read_csv('../input/my-siim-isic-submissions/submission_models_blended_0-6.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"final_target =  main_submission.target *0.85 + efficient_b7.target *0.05 + xgb_df.target *0.10","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = pd.DataFrame({\n        \"image_name\": test[\"image_name\"],\n        \"target\": final_target\n    })\n\nresult.to_csv('final_submission_blend.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> That is a simple approach to ensemble blend my main submission with a simple XGBoost Model. \n> This has improved my final accuracy to some good extent and performed much better after hyperparameter tuning. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Read my other notebooks at:\nhttps://www.kaggle.com/blurredmachine/notebooks<br>\nCompetition Link:\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification\n\nI hope you like this approach and it might be useful for someone to get a good score in competition.<br>\nI am continuously working on this notebook to keep it updated with new features and easy approaches for beginners to understand the concepts easily.\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Consider upvoting if it was helpful! 😃","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}