{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Melanoma Classification : The Naive Approach\n## Strictly for Beginners\n\n![Imgur](https://i.imgur.com/7U9tlVA.jpg)\n\nThis notebook is ideal for beginners who want to start out with Computer vision. This kernel uses fastai library for computer vision processing. It is one of the easy library which can be used for cv. You can check out the official page of fastai to get a hang of it or to know more about it. But for the sake of your understanding, I have explained most of the stuffs used here. Thank me later.","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        os.path.join(dirname, filename)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**In this Competition, as you can see, there are meta data as well as images. I am trying to process both and attain results, so that we can use the combined result to build our classifier.**\n\n# Training the Metadata\n\n**I used XGBoost to train the meta data and predict the targets.**","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import warnings;\nwarnings.filterwarnings('ignore');\n\nimport numpy as np\nimport pandas as pd\n\nfrom sklearn.datasets import load_iris\nimport xgboost as xgb\nfrom sklearn.metrics import accuracy_score","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train= pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ntest= pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')\nsub   = pd.read_csv('../input/siim-isic-melanoma-classification/sample_submission.csv')\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.target.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['sex'] = train['sex'].fillna('na')\ntrain['age_approx'] = train['age_approx'].fillna(0)\ntrain['anatom_site_general_challenge'] = train['anatom_site_general_challenge'].fillna('na')\n\ntest['sex'] = test['sex'].fillna('na')\ntest['age_approx'] = test['age_approx'].fillna(0)\ntest['anatom_site_general_challenge'] = test['anatom_site_general_challenge'].fillna('na')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['sex'] = train['sex'].astype(\"category\").cat.codes +1\ntrain['anatom_site_general_challenge'] = train['anatom_site_general_challenge'].astype(\"category\").cat.codes +1\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test['sex'] = test['sex'].astype(\"category\").cat.codes +1\ntest['anatom_site_general_challenge'] = test['anatom_site_general_challenge'].astype(\"category\").cat.codes +1\ntest.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x_train = train[['sex', 'age_approx','anatom_site_general_challenge']]\ny_train = train['target']\n\nx_test = test[['sex', 'age_approx','anatom_site_general_challenge']]\n\ntrain_DMatrix = xgb.DMatrix(x_train, label= y_train)\ntest_DMatrix = xgb.DMatrix(x_test)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"param = {\n    'booster':'gbtree', \n    'eta': 0.3,\n    'num_class': 2,\n    'max_depth': 5\n}\nepochs = 100","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"clf = xgb.XGBClassifier(n_estimators=1000, \n                        max_depth=8, \n                        objective='multi:softprob',\n                        seed=0,  \n                        nthread=-1, \n                        learning_rate=0.15, \n                        num_class = 2, \n                        scale_pos_weight = (32542/584))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"clf.fit(x_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub[\"meta_target\"] = clf.predict_proba(x_test)[:,1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#sub.to_csv('submission.csv', index = False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now that we predict the target from the meta data, we can move on to image training. For this we are using the fastai library.\n\n# Training the Images\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"from fastai.imports import *\nfrom fastai import *\nfrom fastai.vision import *\nfrom torchvision.models import *","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Since the quantity of images are gigantic, it will take some time to train the whole dataset even with the GPU's. So I have used some images for training and the predictions are done with the aid of that.**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label1 = train[train['target']==1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label1.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label2 = train[train['target']==0].iloc[:584]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label2.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_small = pd.concat([label1,label2])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_small.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_small.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_small['image_name'] = train_small['image_name']+'.jpg'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test['image_name'] = test['image_name']+'.jpg'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_small.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_small.to_csv('train_jpg.csv', index = False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.to_csv('test_jpg.csv', index = False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tfms = get_transforms(flip_vert=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"path = \"/kaggle/\"\ndata = ImageDataBunch.from_csv(path, folder= 'input/siim-isic-melanoma-classification/jpeg/train', \n                              valid_pct = 0.2,\n                              csv_labels = 'working/train_jpg.csv',\n                              ds_tfms = tfms, \n                              fn_col = 'image_name',\n                              label_col = 'target',\n                              bs = 32,\n                              size=256).normalize(imagenet_stats);\ntest_data = ImageDataBunch.from_csv(path, folder= 'input/siim-isic-melanoma-classification/jpeg/test', \n                              valid_pct = 0.2,\n                              csv_labels = 'working/test_jpg.csv',\n                              ds_tfms = tfms, \n                              fn_col = 'image_name',\n                              bs = 32,\n                              size=256).normalize(imagenet_stats);\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.show_batch(rows=3,figsize=(8,8));","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn = cnn_learner(data, models.resnet34, metrics=error_rate)\ndevice = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")\nlearn.model.to(device)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#learn.model_dir = \"/kaggle/working\"\nlearn.lr_find()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.recorder.plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(8,slice(0.015));","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.freeze()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Resources\n\n[Fastai documentation](https://docs.fast.ai/tutorial.resources.html)\n\n\n[Fastai Course](https://course.fast.ai/)\n\n## Happy Coding!","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}