{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Note\n> **This is a notebook that is being created for a Portuguese version of a Data Science class for Awari School. Hence, there will be many comments in portuguese. If there's any question, please leave a comment**\n\n### Playlist em Vídeo Passo a Passo\n- [Link da Playlist](https://loom.com/share/folder/8f3d5415a9fb4d37b8d6626d30b000b3)\n- [Notebook auxiliar utilizado na Playlist](https://github.com/WittmannF/course/blob/master/day-4/assignment-3-cats-dogs-solved.ipynb)","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\n#import os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"## 1. Leitura e análise dos metadados em CSV \nsample_submission = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/sample_submission.csv')\ntest = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/test.csv')\ntrain = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/train.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.tail()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.describe(include='all')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.describe(include='all')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cat_cols = ['patient_id', 'sex', 'anatom_site_general_challenge', 'diagnosis', 'benign_malignant']\nprint('Contagens dos atributos categóricos do conjunto de treino')\nfor col in cat_cols:\n    print('Contagem de valores da coluna {}'.format(col))\n    print(train[col].value_counts().head(20))\n    print('='*80)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cat_cols = ['patient_id', 'sex', 'anatom_site_general_challenge']\nprint('Contagens dos atributos categóricos do conjunto de teste')\nfor col in cat_cols:\n    print('Contagem de valores da coluna {}'.format(col))\n    print(test[col].value_counts().head(20))\n    print('='*80)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission = sample_submission","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.read_csv('submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"## 2. Visualização de Imagens\nDATA_PATH = '../input/siim-isic-melanoma-classification/jpeg/'\nTRAIN_PATH = f'{DATA_PATH}train/'\nTEST_PATH = f'{DATA_PATH}test/'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import glob\n\nfilepaths = glob.glob(TRAIN_PATH+'/*.jpg')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nimport random\nfrom keras.preprocessing.image import load_img\n\nimg2diag = train[['image_name', 'benign_malignant']].set_index('image_name')['benign_malignant'].to_dict()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_path = random.choice(filepaths)\nimg_name = img_path.split('/')[-1].replace('.jpg', \"\")\nimg = load_img(img_path)\nimg_diagnostic = img2diag[img_name]\nimg_np = np.asarray(img)\nplt.imshow(img_np)\nplt.title(img_diagnostic)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"## 3 - Criação do modelo Baseline (Ponto de Partida)\n## 3.1 - Image data generator\n\n# TODO: Import the model and the preprocess_input function\nfrom keras.applications.resnet50 import preprocess_input\n\n# TODO: Import the ImageDataGenerator class\nfrom keras.preprocessing.image import ImageDataGenerator","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Shape in which all images are going to be reshaped\nTARGET_SHAPE = (224, 224, 3)\n\n# TODO: Initialize the data generator class \ndatagen = ImageDataGenerator(preprocessing_function=preprocess_input)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df_datagen = train[['image_name', 'benign_malignant']].copy()\ntrain_df_datagen['image_name'] = train_df_datagen['image_name']+'.jpg'\ntrain_df_datagen.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"N_BENIGN = 584\n\nfilter_benign = train_df_datagen['benign_malignant']=='benign'\nfilter_malignant = train_df_datagen['benign_malignant']=='malignant'\nsample_benign = train_df_datagen[filter_benign].sample(N_BENIGN, random_state=10)\n# Let's try to ignore the class balance test make before\n#sample_benign = train_df_datagen[filter_benign]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_val_sampled = pd.concat([sample_benign, train_df_datagen[filter_malignant]])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\ntrain_df, valid_df = train_test_split(train_val_sampled, \n                                      test_size=0.2, \n                                      random_state=1,\n                                      stratify=train_val_sampled['benign_malignant']\n                                     )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_gen = datagen.flow_from_dataframe(train_df,\n                           TRAIN_PATH,\n                           'image_name',\n                           'benign_malignant',\n                           target_size=TARGET_SHAPE[:2],\n                            class_mode='sparse'\n                           )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"valid_gen = datagen.flow_from_dataframe(valid_df,\n                           TRAIN_PATH,\n                           'image_name',\n                           'benign_malignant',\n                            target_size=TARGET_SHAPE[:2],\n                            class_mode='sparse',\n                            shuffle=False\n                           )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df_datagen = test[['image_name']]+'.jpg'\ntest_df_datagen.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_gen = datagen.flow_from_dataframe(test_df_datagen,\n                            TEST_PATH,\n                            'image_name',\n                            target_size=TARGET_SHAPE[:2],\n                            class_mode=None,\n                            shuffle=False\n                           )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"## 3.2 Criando modelo base\n# Leitura recomendada sobre ResNet e outros modelos: https://medium.com/analytics-vidhya/timeline-of-transfer-learning-models-db2a0be39b37 \nfrom keras.models import Sequential\nfrom keras.layers import Flatten, Dense, GlobalAveragePooling2D\nfrom keras.applications.resnet50 import ResNet50\n\n\nresnet_model = ResNet50(include_top=False, input_shape=TARGET_SHAPE, pooling='avg')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for layer in resnet_model.layers:\n    layer.trainable = False","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"base_model = Sequential([resnet_model,\n                         Dense(1024, activation='relu'),\n                         Dense(2, activation='softmax')\n                        ])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"## 3.3 Treinar modelo","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import tensorflow as tf","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from keras.optimizers import Adam\nbase_model.compile(optimizer=Adam(lr=1e-4), \n                   loss='sparse_categorical_crossentropy',\n                   metrics=['accuracy']\n                  )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"base_model.fit_generator(train_gen,\n                         validation_data=valid_gen,\n                         epochs=3\n                        )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_gen","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Predict in the test set\npred = base_model.predict(test_gen)\n# Get the malignant columns\npred = pred[:, 1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission_dict = {'image_name': test.image_name.values,\n              'target': pred}\n\nsubmission = pd.DataFrame(submission_dict)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.read_csv('submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Ideas and Next Steps\n- Include best practices on Keras from [here](https://github.com/WittmannF/course/blob/master/day-4/Best_Practices_Playground.ipynb) and [here](https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet)\n    - Augmix, LRFinder, [attention](https://www.kaggle.com/ibtesama/melanoma-classification-with-attention), [effnet](https://www.kaggle.com/andradaolteanu/melanoma-competiton-aug-resnet-effnet-lb-0-91), [another effnet](https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet), [include more data](), \n- [Include metafeatures](https://www.kaggle.com/titericz/simple-baseline)\n- Unfreeze layers\n- Create multiple feature extractor and try different TL models\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}