{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.10","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":20270,"databundleVersionId":1222630,"sourceType":"competition"},{"sourceId":643971,"sourceType":"datasetVersion","datasetId":319080}],"dockerImageVersionId":30097,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div align=\"center\">\n<font size=\"6\"> SIIM-ISIC Melanoma Classification  </font>  \n</div> \n\n\n<div align=\"center\">\n<font size=\"4\"> Identify melanoma in lesion images  </font>  \n</div> ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<img align=\"left\" src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/logo.png\" data-canonical-src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/logo.png\" width=\"280\" height=\"280\" />\n\nSkin cancer is the most prevalent type of cancer. **Melanoma**, specifically, is responsible for **75%** of skin cancer deaths, despite being the least common skin cancer. The American Cancer Society estimates over 100,000 new melanoma cases will be diagnosed in 2020. It's also expected that almost 7,000 people will die from the disease. As with other cancers, early and accurate detection—potentially aided by data science—can make treatment more effective.\n\nCurrently, dermatologists evaluate every one of a patient's moles to identify outlier lesions or “ugly ducklings” that are most likely to be melanoma. Existing AI approaches have not adequately considered this clinical frame of reference. Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma. If successful, classifiers would be more accurate and could better support dermatological clinic work.\n\nAs the leading healthcare organization for informatics in medical imaging, the [Society for Imaging Informatics in Medicine (SIIM)](https://siim.org/)'s mission is to advance medical imaging informatics through education, research, and innovation in a multi-disciplinary community. SIIM is joined by the [International Skin Imaging Collaboration (ISIC)](https://www.isic-archive.com/), an international effort to improve melanoma diagnosis. The ISIC Archive contains the largest publicly available collection of quality-controlled dermoscopic images of skin lesions.\n\nIn this competition, you’ll identify melanoma in images of skin lesions. In particular, you’ll use images within the same patient and determine which are likely to represent a melanoma. Using patient-level contextual information may help the development of image analysis tools, which could better support clinical dermatologists.\n\nMelanoma is a deadly disease, but if caught early, most melanomas can be cured with minor surgery. Image analysis tools that automate the diagnosis of melanoma will improve dermatologists' diagnostic accuracy. Better detection of melanoma has the opportunity to positively impact millions of people.","metadata":{}},{"cell_type":"markdown","source":"<img align=\"left\" src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/melanoma.png\" data-canonical-src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/melanoma.png\" width=\"1200\" height=\"450\" />","metadata":{"execution":{"iopub.status.busy":"2021-06-05T23:37:34.304398Z","iopub.execute_input":"2021-06-05T23:37:34.304863Z","iopub.status.idle":"2021-06-05T23:37:34.313401Z","shell.execute_reply.started":"2021-06-05T23:37:34.304767Z","shell.execute_reply":"2021-06-05T23:37:34.312036Z"}}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 1. Libraries </h1>","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"\n\nimport numpy as np \nimport pandas as pd \nimport os\nif False:\n    for dirname, _, filenames in os.walk('/kaggle/input'):\n        for filename in filenames:\n            print(os.path.join(dirname, filename))\n","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:28:44.835690Z","iopub.execute_input":"2023-11-17T14:28:44.836021Z","iopub.status.idle":"2023-11-17T14:28:44.846016Z","shell.execute_reply.started":"2023-11-17T14:28:44.835945Z","shell.execute_reply":"2023-11-17T14:28:44.845222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport re\nimport glob\nimport pathlib\nimport time\nimport math\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nimport cv2\n\nimport PIL\nfrom PIL import Image\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.utils import class_weight\n\nfrom collections import Counter\n\nfrom warnings import filterwarnings\nfilterwarnings('ignore')\n\nSEED=123\nnp.random.seed(SEED)","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:30:20.541494Z","iopub.execute_input":"2023-11-17T14:30:20.541848Z","iopub.status.idle":"2023-11-17T14:30:21.432396Z","shell.execute_reply.started":"2023-11-17T14:30:20.541814Z","shell.execute_reply":"2023-11-17T14:30:21.431611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.models import Model,Sequential\nfrom tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow.keras.layers import Dropout, BatchNormalization\nfrom tensorflow.keras.layers import (\n    Input, Dense, Conv2D, Flatten, Activation, \n    MaxPooling2D, AveragePooling2D, ZeroPadding2D, GlobalAveragePooling2D, GlobalMaxPooling2D, add\n)\n\nfrom tensorflow.python.keras.callbacks import EarlyStopping, ModelCheckpoint\nfrom tensorflow.keras.preprocessing import image\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nfrom tensorflow.keras.utils import plot_model\n","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:30:27.274164Z","iopub.execute_input":"2023-11-17T14:30:27.274555Z","iopub.status.idle":"2023-11-17T14:30:31.995813Z","shell.execute_reply.started":"2023-11-17T14:30:27.274519Z","shell.execute_reply":"2023-11-17T14:30:31.995020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 2. Configs </h1>","metadata":{}},{"cell_type":"code","source":"CFG = dict(\n        batch_size        =  8,     # 8; 16; 32; 64; bigger batch size => moemry allocation issue\n        epochs            =  2,   # 5; 10; 20;\n        verbose           =   1,    # 0; 1\n        workers           =   4,    # 1; 2; 3\n\n        optimizer         = 'adam', # 'SGD', 'RMSprop'\n\n        RANDOM_STATE      =  123,   \n    \n        # Path to save a model\n        path_model        = '../working/',\n\n        # Images sizes\n        img_size          = 224, \n        img_height        = 224, \n        img_width         = 224, \n\n        # Images augs\n        ROTATION          = 180.0,\n        ZOOM              =  10.0,\n        ZOOM_RANGE        =  [0.9,1.1],\n        HZOOM             =  10.0,\n        WZOOM             =  10.0,\n        HSHIFT            =  10.0,\n        WSHIFT            =  10.0,\n        SHEAR             =   5.0,\n        HFLIP             = True,\n        VFLIP             = True,\n\n        # Postprocessing\n        label_smooth_fac  =  0.00,  # 0.01; 0.05; 0.1; 0.2;    \n)","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:31:12.640881Z","iopub.execute_input":"2023-11-17T14:31:12.641270Z","iopub.status.idle":"2023-11-17T14:31:12.647459Z","shell.execute_reply.started":"2023-11-17T14:31:12.641213Z","shell.execute_reply":"2023-11-17T14:31:12.646628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 3. Paths </h1>","metadata":{"execution":{"iopub.status.busy":"2021-06-05T23:43:09.018882Z","iopub.execute_input":"2021-06-05T23:43:09.019279Z","iopub.status.idle":"2021-06-05T23:43:09.761092Z","shell.execute_reply.started":"2021-06-05T23:43:09.019246Z","shell.execute_reply":"2021-06-05T23:43:09.759959Z"}}},{"cell_type":"code","source":"BASEPATH = \"../input/siim-isic-melanoma-classification\"\ndf_train_full = pd.read_csv(os.path.join(BASEPATH, 'train.csv'))\ndf_test  = pd.read_csv(os.path.join(BASEPATH, 'test.csv'))\ndf_sub   = pd.read_csv(os.path.join(BASEPATH, 'sample_submission.csv'))","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:31:18.599801Z","iopub.execute_input":"2023-11-17T14:31:18.600166Z","iopub.status.idle":"2023-11-17T14:31:18.726427Z","shell.execute_reply.started":"2023-11-17T14:31:18.600135Z","shell.execute_reply":"2023-11-17T14:31:18.725636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_path = '../input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration/Train'\ntest_path  = '../input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration/Test'","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:43:49.890417Z","iopub.execute_input":"2023-11-17T14:43:49.890756Z","iopub.status.idle":"2023-11-17T14:43:49.894573Z","shell.execute_reply.started":"2023-11-17T14:43:49.890729Z","shell.execute_reply":"2023-11-17T14:43:49.893629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dir = pathlib.Path(train_path)\ntest_dir  = pathlib.Path(test_path)","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:43:55.939778Z","iopub.execute_input":"2023-11-17T14:43:55.940105Z","iopub.status.idle":"2023-11-17T14:43:55.944395Z","shell.execute_reply.started":"2023-11-17T14:43:55.940076Z","shell.execute_reply":"2023-11-17T14:43:55.943294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dir","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:43:58.157925Z","iopub.execute_input":"2023-11-17T14:43:58.158270Z","iopub.status.idle":"2023-11-17T14:43:58.163409Z","shell.execute_reply.started":"2023-11-17T14:43:58.158230Z","shell.execute_reply":"2023-11-17T14:43:58.162500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dir","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:44:00.875776Z","iopub.execute_input":"2023-11-17T14:44:00.876131Z","iopub.status.idle":"2023-11-17T14:44:00.881145Z","shell.execute_reply.started":"2023-11-17T14:44:00.876099Z","shell.execute_reply":"2023-11-17T14:44:00.880302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_full","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:44:27.628780Z","iopub.execute_input":"2023-11-17T14:44:27.629109Z","iopub.status.idle":"2023-11-17T14:44:27.656682Z","shell.execute_reply.started":"2023-11-17T14:44:27.629081Z","shell.execute_reply":"2023-11-17T14:44:27.655783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can load images by\n- `.flow_from_directory()` using information from subdirectories which has names from labels (we need to prepare data for that, 9 classes give 9 subdirs). Check [here](https://keras.io/api/preprocessing/image/#flowfromdataframe-method).\n- `.flow_from_dataframe()` using information about labels from dataframe. Check [here](https://keras.io/api/preprocessing/image/). ","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 4. Dataset </h1>\n<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 4.1 Description </h1>","metadata":{}},{"cell_type":"markdown","source":"This set consists of **2357** images of **malignant** and **benign** oncological diseases, which were formed from [The International Skin Imaging Collaboration (ISIC)](https://www.isic-archive.com/).    \n   - All images were sorted according to the classification taken with ISIC, and all subsets were divided into the same number of images, with the exception of melanomas and moles, whose images are slightly dominant.\n\nThe data set contains the following diseases:  \n- actinic keratosis\n- basal cell carcinoma\n- dermatofibroma\n- melanoma\n- nevus\n- pigmented benign keratosis\n- seborrheic keratosis\n- squamous cell carcinoma\n- vascular lesion","metadata":{}},{"cell_type":"code","source":"classes=[\n    'pigmented benign keratosis',\n    'melanoma',\n    'vascular lesion',\n    'actinic keratosis',\n    'squamous cell carcinoma',\n    'basal cell carcinoma',\n    'seborrheic keratosis',\n    'dermatofibroma',\n    'nevus'\n]","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:44:33.262732Z","iopub.execute_input":"2023-11-17T14:44:33.263068Z","iopub.status.idle":"2023-11-17T14:44:33.267556Z","shell.execute_reply.started":"2023-11-17T14:44:33.263039Z","shell.execute_reply":"2023-11-17T14:44:33.266458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 4.2 EDA </h1>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\n# Display the first few rows of the dataset\nprint(df_train_full.head())\n\n# Check the dimensions of the dataset\nprint(\"Dimensions of the dataset:\", df_train_full.shape)\n\n# Calculate basic statistics for numerical columns\nprint(\"Summary statistics for numerical columns:\")\nprint(df_train_full.describe())\n\n# Check for missing values\nmissing_values = df_train_full.isnull().sum()\nprint(\"Missing values in the dataset:\")\nprint(missing_values)","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:44:35.629007Z","iopub.execute_input":"2023-11-17T14:44:35.629580Z","iopub.status.idle":"2023-11-17T14:44:35.673651Z","shell.execute_reply.started":"2023-11-17T14:44:35.629536Z","shell.execute_reply":"2023-11-17T14:44:35.672764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n\n# Set the style for Seaborn plots\nsns.set(style=\"whitegrid\")\n\n# Visualize the distribution of 'sex' column\nplt.figure(figsize=(8, 6))\nsns.countplot(data=df_train_full, x='sex', palette='Set2')\nplt.title('Distribution of Sex')\nplt.xlabel('Sex')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:44:41.770044Z","iopub.execute_input":"2023-11-17T14:44:41.770406Z","iopub.status.idle":"2023-11-17T14:44:41.947341Z","shell.execute_reply.started":"2023-11-17T14:44:41.770371Z","shell.execute_reply":"2023-11-17T14:44:41.946525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize the distribution of 'anatom_site_general_challenge' column\nplt.figure(figsize=(12, 6))\nsns.countplot(data=df_train_full, x='anatom_site_general_challenge', palette='Set3')\nplt.title('Distribution of Anatom Site')\nplt.xlabel('Anatom Site')\nplt.ylabel('Count')\nplt.xticks(rotation=45)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:35:28.585084Z","iopub.execute_input":"2023-11-17T14:35:28.585476Z","iopub.status.idle":"2023-11-17T14:35:29.018660Z","shell.execute_reply.started":"2023-11-17T14:35:28.585442Z","shell.execute_reply":"2023-11-17T14:35:29.017802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Explore the relationship between 'sex' and 'target' (benign or malignant)\nplt.figure(figsize=(8, 6))\nsns.countplot(data=df_train_full, x='sex', hue='benign_malignant', palette='pastel')\nplt.title('Relationship between Sex and Benign/Malignant')\nplt.xlabel('Sex')\nplt.ylabel('Count')\nplt.legend(title='Benign/Malignant')\nplt.show()\n\n# Explore the relationship between 'age_approx' and 'target'\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=df_train_full, x='target', y='age_approx', palette='Set2')\nplt.title('Relationship between Age and Melanoma (Target)')\nplt.xlabel('Target (0: Benign, 1: Malignant)')\nplt.ylabel('Age')\nplt.show()\n\n# Explore the relationship between 'anatom_site_general_challenge' and 'target'\nplt.figure(figsize=(12, 6))\nsns.countplot(data=df_train_full, x='anatom_site_general_challenge', hue='benign_malignant', palette='Set3')\nplt.title('Relationship between Anatom Site and Benign/Malignant')\nplt.xlabel('Anatom Site')\nplt.ylabel('Count')\nplt.xticks(rotation=45)\nplt.legend(title='Benign/Malignant')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:44:45.912764Z","iopub.execute_input":"2023-11-17T14:44:45.913123Z","iopub.status.idle":"2023-11-17T14:44:46.614038Z","shell.execute_reply.started":"2023-11-17T14:44:45.913086Z","shell.execute_reply":"2023-11-17T14:44:46.613120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count number of images in each set.\nimg_count_train = len(list(train_dir.glob('*/*.jpg')))\nimg_count_test  = len(list(test_dir.glob('*/*.jpg')))\nprint('{} train images'.format(img_count_train))\nprint('{} test  images'.format(img_count_test))","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:44:49.302214Z","iopub.execute_input":"2023-11-17T14:44:49.302567Z","iopub.status.idle":"2023-11-17T14:44:50.208960Z","shell.execute_reply.started":"2023-11-17T14:44:49.302538Z","shell.execute_reply":"2023-11-17T14:44:50.207894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Define the data directory\ndata_directory = '/kaggle/input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration'\n\n# Define the ImageDataGenerator \ndatagen = ImageDataGenerator(\n    rescale=1.0/255.0,      # Rescale pixel values to the range [0, 1]\n    rotation_range=20,      # Randomly rotate images by up to 20 degrees\n    width_shift_range=0.2,  # Randomly shift images horizontally by up to 20% of the width\n    height_shift_range=0.2, # Randomly shift images vertically by up to 20% of the height\n    horizontal_flip=True,   # Randomly flip images horizontally\n    validation_split=0.2    # Split the data into training (80%) and validation (20%)\n)\n\n# Create the training data generator\ntrain_generator = datagen.flow_from_directory(\n    data_directory,\n    target_size=(224, 224), # Resize images to 224x224 pixels (adjust as needed)\n    batch_size=32,          # Batch size for training\n    class_mode='binary',    # Set class_mode to 'binary' or 'categorical' based on your problem\n    subset='training'       # Specify that this is the training set\n)\n\n# Create the validation data generator\nvalidation_generator = datagen.flow_from_directory(\n    data_directory,\n    target_size=(224, 224), # Resize images to 224x224 pixels (should match training)\n    batch_size=32,          # Batch size for validation\n    class_mode='binary',    \n    subset='validation'     \n)","metadata":{"execution":{"iopub.status.busy":"2023-11-17T14:45:18.201464Z","iopub.execute_input":"2023-11-17T14:45:18.201786Z","iopub.status.idle":"2023-11-17T14:45:18.421621Z","shell.execute_reply.started":"2023-11-17T14:45:18.201758Z","shell.execute_reply":"2023-11-17T14:45:18.420898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 5. Model </h1>","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.applications import MobileNetV2\nfrom tensorflow.keras.layers import Dense, GlobalAveragePooling2D\nfrom tensorflow.keras.models import Model\n\n# Define the data directory\ndata_directory = '/kaggle/input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration'\n\n# Create an ImageDataGenerator with data augmentation settings\ndatagen = ImageDataGenerator(\n    rescale=1.0/255.0,          # Rescale pixel values to [0, 1]\n    rotation_range=20,          # Randomly rotate images by up to 20 degrees\n    width_shift_range=0.2,      # Randomly shift images horizontally by up to 20% of the width\n    height_shift_range=0.2,     # Randomly shift images vertically by up to 20% of the height\n    horizontal_flip=True,       # Randomly flip images horizontally\n    zoom_range=0.2,             # Randomly zoom into or out of the image by 20%\n    brightness_range=[0.8, 1.2] # Randomly adjust brightness by a factor in this range\n)\n\n# Create a data generator from the directory with augmentation\ndata_generator = datagen.flow_from_directory(\n    data_directory,\n    target_size=(224, 224),     # Target size for resizing images\n    batch_size=32,              # Batch size\n    class_mode='binary',        # Set to 'binary' for binary classification, 'categorical' for multiclass\n    shuffle=True                 # Shuffle the data\n)\n\n# Load a pre-trained MobileNetV2 model without the top (classification) layer\nbase_model = MobileNetV2(weights='imagenet', include_top=False, input_shape=(224, 224, 3))\n\n# Add a custom top layer for binary classification\nx = base_model.output\nx = GlobalAveragePooling2D()(x)\nx = Dense(1024, activation='relu')(x)\npredictions = Dense(1, activation='sigmoid')(x)\n\n# Create the final model\nmodel = Model(inputs=base_model.input, outputs=predictions)\n\n# Compile the model\nmodel.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n\n# Train the model using the data generator\nmodel.fit(data_generator, epochs=2, steps_per_epoch=len(data_generator))\n","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:24:14.465058Z","iopub.execute_input":"2023-11-17T15:24:14.465429Z","iopub.status.idle":"2023-11-17T15:26:34.541160Z","shell.execute_reply.started":"2023-11-17T15:24:14.465399Z","shell.execute_reply":"2023-11-17T15:26:34.540301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the accuracy\naccuracy = model.evaluate(validation_generator)[1]","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:26:39.866654Z","iopub.execute_input":"2023-11-17T15:26:39.867003Z","iopub.status.idle":"2023-11-17T15:26:50.231209Z","shell.execute_reply.started":"2023-11-17T15:26:39.866973Z","shell.execute_reply":"2023-11-17T15:26:50.230330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Accuracy: {accuracy * 100:.2f}%')","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:01:30.127901Z","iopub.execute_input":"2023-11-17T15:01:30.128268Z","iopub.status.idle":"2023-11-17T15:01:30.133062Z","shell.execute_reply.started":"2023-11-17T15:01:30.128225Z","shell.execute_reply":"2023-11-17T15:01:30.132160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the file path where you want to save the model\nmodel_file_path = '/kaggle/working/.h5'\n\n# Save the trained model to the specified file path\nmodel.save(model_file_path)\n\n# Print a message to confirm that the model has been saved\nprint(f\"Model saved to {model_file_path}\")","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:28:33.723372Z","iopub.execute_input":"2023-11-17T15:28:33.723754Z","iopub.status.idle":"2023-11-17T15:28:58.622490Z","shell.execute_reply.started":"2023-11-17T15:28:33.723719Z","shell.execute_reply":"2023-11-17T15:28:58.621366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 6. Result </h1>","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.models import Model, Sequential\nfrom tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow.keras.layers import Dropout, BatchNormalization\nfrom tensorflow.keras.layers import (\n    Input, Dense, Conv2D, Flatten, Activation, \n    MaxPooling2D, AveragePooling2D, ZeroPadding2D, GlobalAveragePooling2D, GlobalMaxPooling2D, add\n)\n\nfrom tensorflow.python.keras.callbacks import EarlyStopping, ModelCheckpoint\nfrom tensorflow.keras.preprocessing import image\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nfrom tensorflow.keras.utils import plot_model\n\nfrom tensorflow.keras.applications import MobileNetV2\n","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:37:45.597211Z","iopub.execute_input":"2023-11-17T15:37:45.597625Z","iopub.status.idle":"2023-11-17T15:37:45.604784Z","shell.execute_reply.started":"2023-11-17T15:37:45.597591Z","shell.execute_reply":"2023-11-17T15:37:45.603741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pandas as pd\nfrom tensorflow.keras.models import load_model\n\n# Load the trained model\nmodel = load_model('/kaggle/working/.h5')","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:37:48.318209Z","iopub.execute_input":"2023-11-17T15:37:48.318612Z","iopub.status.idle":"2023-11-17T15:37:56.058056Z","shell.execute_reply.started":"2023-11-17T15:37:48.318578Z","shell.execute_reply":"2023-11-17T15:37:56.057139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pandas as pd\n\n# Define the directory containing test images\ntest_image_dir = '/kaggle/input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration/Test'\n\n# List to store predictions and corresponding image filenames\npredictions = []\nimage_filenames = []\n\n# Iterate through test images\nfor image_filename in os.listdir(test_image_dir):\n    if image_filename.endswith('.jpg'):\n        image_path = os.path.join(test_image_dir, image_filename)\n        # Load and preprocess the image\n        img = cv2.imread(image_path)\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)  # Convert to RGB\n        img = cv2.resize(img, (224, 224))  # Resize to match model input size\n        img = img / 255.0  # Normalize pixel values to [0, 1]\n        img = np.expand_dims(img, axis=0)  # Add batch dimension\n        \n        # Make a prediction\n        prediction = model.predict(img)\n        predictions.append(prediction[0][0])  \n        image_filenames.append(image_filename)  \n\nif len(predictions) == len(image_filenames):\n    # Create a DataFrame for predictions\n    df_predictions = pd.DataFrame({\n        'image_name': image_filenames,\n        'target': predictions\n    })\n\n    # Save the predictions to a CSV file\n    df_predictions.to_csv('predictions.csv', index=False)\nelse:\n    print(\"Error: Predictions and image filenames have different lengths.\")","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:37:59.080280Z","iopub.execute_input":"2023-11-17T15:37:59.080645Z","iopub.status.idle":"2023-11-17T15:37:59.101031Z","shell.execute_reply.started":"2023-11-17T15:37:59.080614Z","shell.execute_reply":"2023-11-17T15:37:59.100262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nfrom IPython.display import display, Image\n\n# Define the directory path where your images are located\nimage_dir = '/kaggle/input/skin-cancer9-classesisic'\n\n# List all image files in the directory\nimage_files = [f for f in os.listdir(image_dir) if f.endswith('.jpg')]\n\n# Display each image\nfor image_file in image_files:\n    image_path = os.path.join(image_dir, image_file)\n    display(Image(filename=image_path))","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:38:02.475095Z","iopub.execute_input":"2023-11-17T15:38:02.475496Z","iopub.status.idle":"2023-11-17T15:38:02.482024Z","shell.execute_reply.started":"2023-11-17T15:38:02.475460Z","shell.execute_reply":"2023-11-17T15:38:02.481169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Visualize some sample images from the training dataset\nsample_images, sample_labels = next(train_generator)  # Get a batch of images and labels\n\n# Define a function to display images with their labels\ndef plot_images(images, labels, num_images=5):\n    plt.figure(figsize=(12, 6))\n    for i in range(num_images):\n        plt.subplot(1, num_images, i + 1)\n        plt.imshow(images[i])\n        plt.title(f'Label: {labels[i]}')\n        plt.axis('off')\n\n# Display the sample images\nplot_images(sample_images, sample_labels, num_images=5)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-17T15:38:10.522544Z","iopub.execute_input":"2023-11-17T15:38:10.522885Z","iopub.status.idle":"2023-11-17T15:38:12.091052Z","shell.execute_reply.started":"2023-11-17T15:38:10.522856Z","shell.execute_reply":"2023-11-17T15:38:12.090070Z"},"trusted":true},"execution_count":null,"outputs":[]}]}