{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div align=\"center\">\n<font size=\"6\"> SIIM-ISIC Melanoma Classification  </font>  \n</div> \n\n\n<div align=\"center\">\n<font size=\"4\"> Identify melanoma in lesion images  </font>  \n</div> ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<img align=\"left\" src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/logo.png\" data-canonical-src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/logo.png\" width=\"280\" height=\"280\" />\n\nSkin cancer is the most prevalent type of cancer. **Melanoma**, specifically, is responsible for **75%** of skin cancer deaths, despite being the least common skin cancer. The American Cancer Society estimates over 100,000 new melanoma cases will be diagnosed in 2020. It's also expected that almost 7,000 people will die from the disease. As with other cancers, early and accurate detection—potentially aided by data science—can make treatment more effective.\n\nCurrently, dermatologists evaluate every one of a patient's moles to identify outlier lesions or “ugly ducklings” that are most likely to be melanoma. Existing AI approaches have not adequately considered this clinical frame of reference. Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma. If successful, classifiers would be more accurate and could better support dermatological clinic work.\n\nAs the leading healthcare organization for informatics in medical imaging, the [Society for Imaging Informatics in Medicine (SIIM)](https://siim.org/)'s mission is to advance medical imaging informatics through education, research, and innovation in a multi-disciplinary community. SIIM is joined by the [International Skin Imaging Collaboration (ISIC)](https://www.isic-archive.com/), an international effort to improve melanoma diagnosis. The ISIC Archive contains the largest publicly available collection of quality-controlled dermoscopic images of skin lesions.\n\nIn this competition, you’ll identify melanoma in images of skin lesions. In particular, you’ll use images within the same patient and determine which are likely to represent a melanoma. Using patient-level contextual information may help the development of image analysis tools, which could better support clinical dermatologists.\n\nMelanoma is a deadly disease, but if caught early, most melanomas can be cured with minor surgery. Image analysis tools that automate the diagnosis of melanoma will improve dermatologists' diagnostic accuracy. Better detection of melanoma has the opportunity to positively impact millions of people.","metadata":{}},{"cell_type":"markdown","source":"<img align=\"left\" src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/melanoma.png\" data-canonical-src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/melanoma.png\" width=\"1200\" height=\"450\" />","metadata":{"execution":{"iopub.status.busy":"2021-06-05T23:37:34.304398Z","iopub.execute_input":"2021-06-05T23:37:34.304863Z","iopub.status.idle":"2021-06-05T23:37:34.313401Z","shell.execute_reply.started":"2021-06-05T23:37:34.304767Z","shell.execute_reply":"2021-06-05T23:37:34.312036Z"}}},{"cell_type":"markdown","source":"<h2 style=color:Teal align=\"left\"> Table of Contents </h2>\n\n#### 1. ResNet\n#### 2. Libraries\n##### 2.1 Load Required Libraries\n##### 2.2 Load TensorFlow\n#### 3. Configs\n#### 4 Paths\n#### 5. Dataset\n##### 5.1 Description\n##### 5.2 EDA\n#### 6. Keras image data processing\n#### 7. Class weights\n#### 8. Model\n##### 8.1 Build model with ResNet50\n##### 8.2 Visualize model with ResNet50\n#### 9 Fit model\n#### 10. Visualize performance\n#### 11. Evaluate on test\n#### 12. Submit predictions\n#### References","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n\nif False:\n    for dirname, _, filenames in os.walk('/kaggle/input'):\n        for filename in filenames:\n            print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:11:55.757897Z","iopub.execute_input":"2023-09-30T17:11:55.758278Z","iopub.status.idle":"2023-09-30T17:11:55.766617Z","shell.execute_reply.started":"2023-09-30T17:11:55.758195Z","shell.execute_reply":"2023-09-30T17:11:55.765776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 1. ResNet </h1>","metadata":{}},{"cell_type":"markdown","source":"Deeper neural networks are more difficult to train. A residual learning framework is easy to train. The layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions are explicitly reformulated. It has been shown that residual networks are easier to optimize, and can gain accuracy from considerably increased depth. Residual nets (ResNets) are with a depth of up to 152 layers, i.e., x8 deeper than e.g. VGG nets but still having lower complexity. [Kaiming He et all. 2015]\n\nResNet50 stands for ResNet with 50 layers. See [architecture visualization](http://ethereon.github.io/netscope/#/gist/db945b393d40bfa26006).\n\n&nbsp;\n\n<div align=\"center\">\n<font size=\"4\"> Residual learning: a building block.  </font>  \n</div> \n\n<img align=\"left\" src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/A-cell-from-the-Residual-Network-architecture-The-identity-connection-helps-to-reduce.png\" data-canonical-src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/A-cell-from-the-Residual-Network-architecture-The-identity-connection-helps-to-reduce.png\" width=\"350\" height=\"350\" />\nThe degradation (of training accuracy) indicates that not all systems are similarly easy to optimize. In [Kaiming He et all. 2015] the degradation problem is adressed by introducing a deep residual learning framework. Instead of hoping each few stacked layers directly fit a desired underlying mapping, we explicitly let these layers fit a residual mapping. Formally, denoting the desired underlying mapping as H(x), we let the stacked nonlinear layers fit another mapping of F(x) := H(x)−x. The original mapping is recast into F(x)+x. We hypothesize that it is easier to optimize the residual mapping than to optimize the original, unreferenced mapping. To the extreme, if an identity mapping were optimal, it would be easier to push the residual to zero than to fit an identity mapping by a stack of nonlinear layers. \n\nThe formulation of F(x)+x can be realized by feedforward neural networks with ''shortcut connections'' (see scheme). Shortcut connections are those skipping one or\nmore layers. In our case, the shortcut connections simply perform identity mapping, and their outputs are added to the outputs of the stacked layers (see scheme). Identity shortcut connections add neither extra parameter nor computational complexity. The entire network can still be trained end-to-end by SGD with backpropagation, and can be easily implemented using common libraries\n\n&nbsp;\n&nbsp;\n\n\n<div align=\"center\">\n<font size=\"4\"> Example of a residual network with 34 parameter layers (ResNet34) vs VGG-19 with 19 layers as reference model.  </font>  \n</div> \n\n<img align=\"center\" src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/arch.jpg\" data-canonical-src=\"https://raw.githubusercontent.com/kabartay/kaggle-siim-isic-melanoma-classification/master/materials/arch.jpg\" width=\"670\" height=\"1520\" />","metadata":{"execution":{"iopub.status.busy":"2021-06-05T23:40:18.015199Z","iopub.execute_input":"2021-06-05T23:40:18.015577Z","iopub.status.idle":"2021-06-05T23:40:18.020793Z","shell.execute_reply.started":"2021-06-05T23:40:18.015547Z","shell.execute_reply":"2021-06-05T23:40:18.018907Z"}}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 2. Libraries </h1>","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 2.1 Load Required Libraries </h1>","metadata":{}},{"cell_type":"code","source":"import os\nimport re\nimport glob\nimport pathlib\nimport time\nimport math\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nimport cv2\n\nimport PIL\nfrom PIL import Image\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.utils import class_weight\n\nfrom collections import Counter\n\nfrom warnings import filterwarnings\nfilterwarnings('ignore')\n\nSEED=123\nnp.random.seed(SEED)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:11:55.773521Z","iopub.execute_input":"2023-09-30T17:11:55.773791Z","iopub.status.idle":"2023-09-30T17:11:56.672218Z","shell.execute_reply.started":"2023-09-30T17:11:55.773767Z","shell.execute_reply":"2023-09-30T17:11:56.671407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 2.2 Load TensorFlow </h1>","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.models import Model,Sequential\nfrom tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow.keras.layers import Dropout, BatchNormalization\nfrom tensorflow.keras.layers import (\n    Input, Dense, Conv2D, Flatten, Activation, \n    MaxPooling2D, AveragePooling2D, ZeroPadding2D, GlobalAveragePooling2D, GlobalMaxPooling2D, add\n)\n\nfrom tensorflow.python.keras.callbacks import EarlyStopping, ModelCheckpoint\nfrom tensorflow.keras.preprocessing import image\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nfrom tensorflow.keras.utils import plot_model\n\nfrom tensorflow.keras.applications.vgg19 import VGG19\nfrom tensorflow.keras.applications.vgg19 import preprocess_input\nfrom tensorflow.keras.applications.resnet50 import ResNet50\nfrom tensorflow.keras.applications.resnet50 import preprocess_input\nfrom tensorflow.keras.applications.inception_v3 import InceptionV3\n#from tensorflow.keras.applications.inception_v3 import preprocess_input","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:11:56.673708Z","iopub.execute_input":"2023-09-30T17:11:56.674271Z","iopub.status.idle":"2023-09-30T17:12:01.384568Z","shell.execute_reply.started":"2023-09-30T17:11:56.674236Z","shell.execute_reply":"2023-09-30T17:12:01.383732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 3. Configs </h1>","metadata":{}},{"cell_type":"code","source":"CFG = dict(\n        batch_size        =  16,     # 8; 16; 32; 64; bigger batch size => moemry allocation issue\n        epochs            =  20,   # 5; 10; 20;\n        verbose           =   1,    # 0; 1\n        workers           =   4,    # 1; 2; 3\n\n        optimizer         = 'adam', # 'SGD', 'RMSprop'\n\n        RANDOM_STATE      =  123,   \n    \n        # Path to save a model\n        path_model        = '../working/',\n\n        # Images sizes\n        img_size          = 224, \n        img_height        = 224, \n        img_width         = 224, \n\n        # Images augs\n        ROTATION          = 180.0,\n        ZOOM              =  10.0,\n        ZOOM_RANGE        =  [0.9,1.1],\n        HZOOM             =  10.0,\n        WZOOM             =  10.0,\n        HSHIFT            =  10.0,\n        WSHIFT            =  10.0,\n        SHEAR             =   5.0,\n        HFLIP             = True,\n        VFLIP             = True,\n\n        # Postprocessing\n        label_smooth_fac  =  0.00,  # 0.01; 0.05; 0.1; 0.2;    \n)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:01.386273Z","iopub.execute_input":"2023-09-30T17:12:01.386770Z","iopub.status.idle":"2023-09-30T17:12:01.393210Z","shell.execute_reply.started":"2023-09-30T17:12:01.386733Z","shell.execute_reply":"2023-09-30T17:12:01.392284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 4. Paths </h1>","metadata":{"execution":{"iopub.status.busy":"2021-06-05T23:43:09.018882Z","iopub.execute_input":"2021-06-05T23:43:09.019279Z","iopub.status.idle":"2021-06-05T23:43:09.761092Z","shell.execute_reply.started":"2021-06-05T23:43:09.019246Z","shell.execute_reply":"2021-06-05T23:43:09.759959Z"}}},{"cell_type":"code","source":"BASEPATH = \"../input/siim-isic-melanoma-classification\"\ndf_train_full = pd.read_csv(os.path.join(BASEPATH, 'train.csv'))\ndf_test  = pd.read_csv(os.path.join(BASEPATH, 'test.csv'))\ndf_sub   = pd.read_csv(os.path.join(BASEPATH, 'sample_submission.csv'))","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:01.394756Z","iopub.execute_input":"2023-09-30T17:12:01.395260Z","iopub.status.idle":"2023-09-30T17:12:01.532609Z","shell.execute_reply.started":"2023-09-30T17:12:01.395225Z","shell.execute_reply":"2023-09-30T17:12:01.531773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_path = '../input/siim-isic-melanoma-classification/jpeg/train'\n#test_path  = '../input/siim-isic-melanoma-classification/jpeg/test'\n\n# Dataset ready for Keras load from directories (structured with respect to classes)\ntrain_path = '../input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaborationn/Train'\ntest_path  = '../input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration/Test'","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:01.533828Z","iopub.execute_input":"2023-09-30T17:12:01.534293Z","iopub.status.idle":"2023-09-30T17:12:01.538749Z","shell.execute_reply.started":"2023-09-30T17:12:01.534258Z","shell.execute_reply":"2023-09-30T17:12:01.537923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dir = pathlib.Path(train_path)\ntest_dir  = pathlib.Path(test_path)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:01.540046Z","iopub.execute_input":"2023-09-30T17:12:01.540640Z","iopub.status.idle":"2023-09-30T17:12:01.548619Z","shell.execute_reply.started":"2023-09-30T17:12:01.540601Z","shell.execute_reply":"2023-09-30T17:12:01.547752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can load images by\n- `.flow_from_directory()` using information from subdirectories which has names from labels (we need to prepare data for that, 9 classes give 9 subdirs). Check [here](https://keras.io/api/preprocessing/image/#flowfromdataframe-method).\n- `.flow_from_dataframe()` using information about labels from dataframe. Check [here](https://keras.io/api/preprocessing/image/). ","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 5. Dataset </h1>\n<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 5.1 Description </h1>","metadata":{}},{"cell_type":"markdown","source":"This set consists of **2357** images of **malignant** and **benign** oncological diseases, which were formed from [The International Skin Imaging Collaboration (ISIC)](https://www.isic-archive.com/).    \n   - All images were sorted according to the classification taken with ISIC, and all subsets were divided into the same number of images, with the exception of melanomas and moles, whose images are slightly dominant.\n\nThe data set contains the following diseases:  \n- actinic keratosis\n- basal cell carcinoma\n- dermatofibroma\n- melanoma\n- nevus\n- pigmented benign keratosis\n- seborrheic keratosis\n- squamous cell carcinoma\n- vascular lesion","metadata":{}},{"cell_type":"code","source":"classes=[\n    'pigmented benign keratosis',\n    'melanoma',\n    'vascular lesion',\n    'actinic keratosis',\n    'squamous cell carcinoma',\n    'basal cell carcinoma',\n    'seborrheic keratosis',\n    'dermatofibroma',\n    'nevus'\n]","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:01.549892Z","iopub.execute_input":"2023-09-30T17:12:01.550558Z","iopub.status.idle":"2023-09-30T17:12:01.558990Z","shell.execute_reply.started":"2023-09-30T17:12:01.550524Z","shell.execute_reply":"2023-09-30T17:12:01.558185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 5.2 EDA </h1>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\n# Display the first few rows of the dataset\nprint(df_train_full.head())\n\n# Check the dimensions of the dataset\nprint(\"Dimensions of the dataset:\", df_train_full.shape)\n\n# Calculate basic statistics for numerical columns\nprint(\"Summary statistics for numerical columns:\")\nprint(df_train_full.describe())\n\n# Check for missing values\nmissing_values = df_train_full.isnull().sum()\nprint(\"Missing values in the dataset:\")\nprint(missing_values)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:01.561688Z","iopub.execute_input":"2023-09-30T17:12:01.561991Z","iopub.status.idle":"2023-09-30T17:12:01.615277Z","shell.execute_reply.started":"2023-09-30T17:12:01.561954Z","shell.execute_reply":"2023-09-30T17:12:01.614407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n\n# Set the style for Seaborn plots\nsns.set(style=\"whitegrid\")\n\n# Visualize the distribution of 'sex' column\nplt.figure(figsize=(8, 6))\nsns.countplot(data=df_train_full, x='sex', palette='Set2')\nplt.title('Distribution of Sex')\nplt.xlabel('Sex')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:01.617035Z","iopub.execute_input":"2023-09-30T17:12:01.617532Z","iopub.status.idle":"2023-09-30T17:12:01.961461Z","shell.execute_reply.started":"2023-09-30T17:12:01.617496Z","shell.execute_reply":"2023-09-30T17:12:01.960602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize the distribution of 'anatom_site_general_challenge' column\nplt.figure(figsize=(12, 6))\nsns.countplot(data=df_train_full, x='anatom_site_general_challenge', palette='Set3')\nplt.title('Distribution of Anatom Site')\nplt.xlabel('Anatom Site')\nplt.ylabel('Count')\nplt.xticks(rotation=45)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:01.962847Z","iopub.execute_input":"2023-09-30T17:12:01.963360Z","iopub.status.idle":"2023-09-30T17:12:02.139255Z","shell.execute_reply.started":"2023-09-30T17:12:01.963320Z","shell.execute_reply":"2023-09-30T17:12:02.138480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Explore the relationship between 'sex' and 'target' (benign or malignant)\nplt.figure(figsize=(8, 6))\nsns.countplot(data=df_train_full, x='sex', hue='benign_malignant', palette='pastel')\nplt.title('Relationship between Sex and Benign/Malignant')\nplt.xlabel('Sex')\nplt.ylabel('Count')\nplt.legend(title='Benign/Malignant')\nplt.show()\n\n# Explore the relationship between 'age_approx' and 'target'\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=df_train_full, x='target', y='age_approx', palette='Set2')\nplt.title('Relationship between Age and Melanoma (Target)')\nplt.xlabel('Target (0: Benign, 1: Malignant)')\nplt.ylabel('Age')\nplt.show()\n\n# Explore the relationship between 'anatom_site_general_challenge' and 'target'\nplt.figure(figsize=(12, 6))\nsns.countplot(data=df_train_full, x='anatom_site_general_challenge', hue='benign_malignant', palette='Set3')\nplt.title('Relationship between Anatom Site and Benign/Malignant')\nplt.xlabel('Anatom Site')\nplt.ylabel('Count')\nplt.xticks(rotation=45)\nplt.legend(title='Benign/Malignant')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:02.140471Z","iopub.execute_input":"2023-09-30T17:12:02.140954Z","iopub.status.idle":"2023-09-30T17:12:02.637028Z","shell.execute_reply.started":"2023-09-30T17:12:02.140916Z","shell.execute_reply":"2023-09-30T17:12:02.636243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count number of images in each set.\nimg_count_train = len(list(train_dir.glob('*/*.jpg')))\nimg_count_test  = len(list(test_dir.glob('*/*.jpg')))\nprint('{} train images'.format(img_count_train))\nprint('{} test  images'.format(img_count_test))","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:02.638273Z","iopub.execute_input":"2023-09-30T17:12:02.638769Z","iopub.status.idle":"2023-09-30T17:12:02.678909Z","shell.execute_reply.started":"2023-09-30T17:12:02.638732Z","shell.execute_reply":"2023-09-30T17:12:02.678054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Define the data directory\ndata_directory = '/kaggle/input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration'\n\n# Define the ImageDataGenerator \ndatagen = ImageDataGenerator(\n    rescale=1.0/255.0,      # Rescale pixel values to the range [0, 1]\n    rotation_range=20,      # Randomly rotate images by up to 20 degrees\n    width_shift_range=0.2,  # Randomly shift images horizontally by up to 20% of the width\n    height_shift_range=0.2, # Randomly shift images vertically by up to 20% of the height\n    horizontal_flip=True,   # Randomly flip images horizontally\n    validation_split=0.2    # Split the data into training (80%) and validation (20%)\n)\n\n# Create the training data generator\ntrain_generator = datagen.flow_from_directory(\n    data_directory,\n    target_size=(224, 224), # Resize images to 224x224 pixels (adjust as needed)\n    batch_size=32,          # Batch size for training\n    class_mode='binary',    # Set class_mode to 'binary' or 'categorical' based on your problem\n    subset='training'       # Specify that this is the training set\n)\n\n# Create the validation data generator\nvalidation_generator = datagen.flow_from_directory(\n    data_directory,\n    target_size=(224, 224), # Resize images to 224x224 pixels (should match training)\n    batch_size=32,          # Batch size for validation\n    class_mode='binary',    \n    subset='validation'     \n)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:02.680060Z","iopub.execute_input":"2023-09-30T17:12:02.680417Z","iopub.status.idle":"2023-09-30T17:12:03.195789Z","shell.execute_reply.started":"2023-09-30T17:12:02.680367Z","shell.execute_reply":"2023-09-30T17:12:03.194735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 6. Keras image data processing </h1>","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 7. Model </h1>","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.applications import ResNet50\nfrom tensorflow.keras.layers import Dense, GlobalAveragePooling2D\nfrom tensorflow.keras.models import Model\n\n# Define the data directory\ndata_directory = '/kaggle/input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration'\n\n# Create an ImageDataGenerator with data augmentation settings\ndatagen = ImageDataGenerator(\n    rescale=1.0/255.0,          # Rescale pixel values to [0, 1]\n    rotation_range=20,          # Randomly rotate images by up to 20 degrees\n    width_shift_range=0.2,      # Randomly shift images horizontally by up to 20% of the width\n    height_shift_range=0.2,     # Randomly shift images vertically by up to 20% of the height\n    horizontal_flip=True,       # Randomly flip images horizontally\n    zoom_range=0.2,             # Randomly zoom into or out of the image by 20%\n    brightness_range=[0.8, 1.2] # Randomly adjust brightness by a factor in this range\n)\n\n# Create a data generator from the directory with augmentation\ndata_generator = datagen.flow_from_directory(\n    data_directory,\n    target_size=(224, 224),     # Target size for resizing images\n    batch_size=32,              # Batch size\n    class_mode='binary',        # Set to 'binary' for binary classification, 'categorical' for multiclass\n    shuffle=True                 # Shuffle the data\n)\n\n# Load a pre-trained ResNet50 model without the top (classification) layer\nbase_model = ResNet50(weights='imagenet', include_top=False)\n\n# Add a custom top layer for binary classification\nx = base_model.output\nx = GlobalAveragePooling2D()(x)\nx = Dense(1024, activation='relu')(x)\npredictions = Dense(1, activation='sigmoid')(x)\n\n# Create the final model\nmodel = Model(inputs=base_model.input, outputs=predictions)\n\n# Compile the model\nmodel.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n\n# Train the model using the data generator\nmodel.fit(data_generator, epochs=10, steps_per_epoch=len(data_generator))","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:12:03.197157Z","iopub.execute_input":"2023-09-30T17:12:03.197686Z","iopub.status.idle":"2023-09-30T17:24:13.890822Z","shell.execute_reply.started":"2023-09-30T17:12:03.197647Z","shell.execute_reply":"2023-09-30T17:24:13.889323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the accuracy\naccuracy = model.evaluate(validation_generator)[1]","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:24:13.892827Z","iopub.execute_input":"2023-09-30T17:24:13.893547Z","iopub.status.idle":"2023-09-30T17:24:24.046763Z","shell.execute_reply.started":"2023-09-30T17:24:13.893496Z","shell.execute_reply":"2023-09-30T17:24:24.045958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Accuracy: {accuracy * 100:.2f}%')","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:24:24.048137Z","iopub.execute_input":"2023-09-30T17:24:24.048680Z","iopub.status.idle":"2023-09-30T17:24:24.054357Z","shell.execute_reply.started":"2023-09-30T17:24:24.048642Z","shell.execute_reply":"2023-09-30T17:24:24.053463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the file path where you want to save the model\nmodel_file_path = '/kaggle/working/.h5'\n\n# Save the trained model to the specified file path\nmodel.save(model_file_path)\n\n# Print a message to confirm that the model has been saved\nprint(f\"Model saved to {model_file_path}\")","metadata":{"execution":{"iopub.status.busy":"2023-09-30T18:05:50.009132Z","iopub.execute_input":"2023-09-30T18:05:50.009653Z","iopub.status.idle":"2023-09-30T18:06:11.307893Z","shell.execute_reply.started":"2023-09-30T18:05:50.009610Z","shell.execute_reply":"2023-09-30T18:06:11.306984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 8. Class weights </h1>","metadata":{}},{"cell_type":"code","source":"# Class weights\nclass_weights = class_weight.compute_class_weight('balanced',\n                                                  np.unique(train_generator.classes), \n                                                  train_generator.classes) \n\nunique_class_weights = np.unique(train_generator.classes)\nclass_weights_dict   = { unique_class_weights[i]: w for i,w in enumerate(class_weights) }\n\nprint('\\nCLASS WEIGHTS: {}\\n'.format(class_weights))\nprint(np.unique(train_generator.classes))\nprint(train_generator.classes)\nprint(unique_class_weights)\nprint(Counter(train_generator.classes).keys())   # equals to list(set(x))\nprint(Counter(train_generator.classes).values()) # counts the elements' frequency","metadata":{"execution":{"iopub.status.busy":"2023-09-30T17:24:24.055691Z","iopub.execute_input":"2023-09-30T17:24:24.056172Z","iopub.status.idle":"2023-09-30T17:24:24.071301Z","shell.execute_reply.started":"2023-09-30T17:24:24.056139Z","shell.execute_reply":"2023-09-30T17:24:24.070323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 8. Model </h1>","metadata":{"execution":{"iopub.status.busy":"2021-06-06T15:25:49.684302Z","iopub.execute_input":"2021-06-06T15:25:49.684653Z","iopub.status.idle":"2021-06-06T15:25:49.689686Z","shell.execute_reply.started":"2021-06-06T15:25:49.684615Z","shell.execute_reply":"2021-06-06T15:25:49.688632Z"}}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 8.1 Build model with ResNet50 </h1>","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 8.2 Visualize model with ResNet50 </h1>","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 9 Fit model </h1>","metadata":{}},{"cell_type":"code","source":"import os\nimport re\nimport glob\nimport pathlib\nimport time\nimport math\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nimport cv2\n\nimport PIL\nfrom PIL import Image\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.utils import class_weight\n\nfrom collections import Counter\n\nfrom warnings import filterwarnings\nfilterwarnings('ignore')\n\nSEED = 123\nnp.random.seed(SEED)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T18:00:09.769441Z","iopub.execute_input":"2023-09-30T18:00:09.769782Z","iopub.status.idle":"2023-09-30T18:00:09.779414Z","shell.execute_reply.started":"2023-09-30T18:00:09.769754Z","shell.execute_reply":"2023-09-30T18:00:09.778540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.models import Model, Sequential\nfrom tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow.keras.layers import Dropout, BatchNormalization\nfrom tensorflow.keras.layers import (\n    Input, Dense, Conv2D, Flatten, Activation, \n    MaxPooling2D, AveragePooling2D, ZeroPadding2D, GlobalAveragePooling2D, GlobalMaxPooling2D, add\n)\n\nfrom tensorflow.python.keras.callbacks import EarlyStopping, ModelCheckpoint\nfrom tensorflow.keras.preprocessing import image\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nfrom tensorflow.keras.utils import plot_model\n\nfrom tensorflow.keras.applications.vgg19 import VGG19\nfrom tensorflow.keras.applications.vgg19 import preprocess_input\nfrom tensorflow.keras.applications.resnet50 import ResNet50\nfrom tensorflow.keras.applications.resnet50 import preprocess_input\nfrom tensorflow.keras.applications.inception_v3 import InceptionV3","metadata":{"execution":{"iopub.status.busy":"2023-09-30T18:00:09.878013Z","iopub.execute_input":"2023-09-30T18:00:09.878346Z","iopub.status.idle":"2023-09-30T18:00:09.884517Z","shell.execute_reply.started":"2023-09-30T18:00:09.878317Z","shell.execute_reply":"2023-09-30T18:00:09.883542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 10. Visualize performance </h1>","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:LightSeaGreen; font-family:newtimeroman; font-size:200%; text-align:left;\"> 12. Submit predictions </h1>","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pandas as pd\nfrom tensorflow.keras.models import load_model\n\n# Load the trained model\nmodel = load_model('/kaggle/working/.h5')","metadata":{"execution":{"iopub.status.busy":"2023-09-30T18:07:28.057128Z","iopub.execute_input":"2023-09-30T18:07:28.057482Z","iopub.status.idle":"2023-09-30T18:07:35.808739Z","shell.execute_reply.started":"2023-09-30T18:07:28.057452Z","shell.execute_reply":"2023-09-30T18:07:35.807830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pandas as pd\n\n# Define the directory containing test images\ntest_image_dir = '/kaggle/input/skin-cancer9-classesisic/Skin cancer ISIC The International Skin Imaging Collaboration/Test'\n\n# List to store predictions and corresponding image filenames\npredictions = []\nimage_filenames = []\n\n# Iterate through test images\nfor image_filename in os.listdir(test_image_dir):\n    if image_filename.endswith('.jpg'):\n        image_path = os.path.join(test_image_dir, image_filename)\n        # Load and preprocess the image\n        img = cv2.imread(image_path)\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)  # Convert to RGB\n        img = cv2.resize(img, (224, 224))  # Resize to match model input size\n        img = img / 255.0  # Normalize pixel values to [0, 1]\n        img = np.expand_dims(img, axis=0)  # Add batch dimension\n        \n        # Make a prediction\n        prediction = model.predict(img)\n        predictions.append(prediction[0][0])  \n        image_filenames.append(image_filename)  \n\nif len(predictions) == len(image_filenames):\n    # Create a DataFrame for predictions\n    df_predictions = pd.DataFrame({\n        'image_name': image_filenames,\n        'target': predictions\n    })\n\n    # Save the predictions to a CSV file\n    df_predictions.to_csv('submission.csv', index=False)\nelse:\n    print(\"Error: Predictions and image filenames have different lengths.\")","metadata":{"execution":{"iopub.status.busy":"2023-09-30T18:12:24.770867Z","iopub.execute_input":"2023-09-30T18:12:24.771181Z","iopub.status.idle":"2023-09-30T18:12:24.782778Z","shell.execute_reply.started":"2023-09-30T18:12:24.771145Z","shell.execute_reply":"2023-09-30T18:12:24.781906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nfrom IPython.display import display, Image\n\n# Define the directory path where your images are located\nimage_dir = '/kaggle/input/skin-cancer9-classesisic'\n\n# List all image files in the directory\nimage_files = [f for f in os.listdir(image_dir) if f.endswith('.jpg')]\n\n# Display each image\nfor image_file in image_files:\n    image_path = os.path.join(image_dir, image_file)\n    display(Image(filename=image_path))","metadata":{"execution":{"iopub.status.busy":"2023-09-30T18:15:49.158088Z","iopub.execute_input":"2023-09-30T18:15:49.158452Z","iopub.status.idle":"2023-09-30T18:15:49.163948Z","shell.execute_reply.started":"2023-09-30T18:15:49.158422Z","shell.execute_reply":"2023-09-30T18:15:49.162919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Visualize some sample images from the training dataset\nsample_images, sample_labels = next(train_generator)  # Get a batch of images and labels\n\n# Define a function to display images with their labels\ndef plot_images(images, labels, num_images=5):\n    plt.figure(figsize=(12, 6))\n    for i in range(num_images):\n        plt.subplot(1, num_images, i + 1)\n        plt.imshow(images[i])\n        plt.title(f'Label: {labels[i]}')\n        plt.axis('off')\n\n# Display the sample images\nplot_images(sample_images, sample_labels, num_images=5)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-30T18:19:13.465027Z","iopub.execute_input":"2023-09-30T18:19:13.465347Z","iopub.status.idle":"2023-09-30T18:19:15.071220Z","shell.execute_reply.started":"2023-09-30T18:19:13.465317Z","shell.execute_reply":"2023-09-30T18:19:15.070487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install tensorflow\n!pip install tensorflow-hub\n!pip install tf-slim\n!pip install opencv-python","metadata":{"execution":{"iopub.status.busy":"2023-09-30T19:06:48.680806Z","iopub.execute_input":"2023-09-30T19:06:48.681613Z","iopub.status.idle":"2023-09-30T19:07:11.784364Z","shell.execute_reply.started":"2023-09-30T19:06:48.681559Z","shell.execute_reply":"2023-09-30T19:07:11.783286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install tensorflow==2.5\n!pip install opencv-python-headless\n\n!git clone https://github.com/tensorflow/models.git\n\nimport os\nos.chdir('/kaggle/working/models/research')\n!protoc object_detection/protos/*.proto --python_out=.\n\nwith open('object_detection/packages/tf2/setup.py', 'a') as f:\n    f.write(\"\\\"tf-models-official\\\",\\n\")\n\n!python -m pip install .\n!python object_detection/builders/model_builder_tf2_test.py\n!python /kaggle/working/models/research/object_detection/model_main_tf2.py \\\n   --pipeline_config_path=PATH_TO_CONFIG_FILE \\\n   --model_dir=MODEL_DIR \\\n   --num_train_steps=NUM_TRAIN_STEPS \\\n   --sample_1_of_n_eval_examples=1 \\\n   --alsologtostderr\n!python /kaggle/working/models/research/object_detection/exporter_main_v2.py \\\n   --input_type=image_tensor \\\n   --pipeline_config_path=PATH_TO_CONFIG_FILE \\\n   --trained_checkpoint_dir=MODEL_DIR \\\n   --output_directory=EXPORT_DIR","metadata":{"execution":{"iopub.status.busy":"2023-09-30T19:11:58.231581Z","iopub.execute_input":"2023-09-30T19:11:58.231909Z","iopub.status.idle":"2023-09-30T19:13:15.570557Z","shell.execute_reply.started":"2023-09-30T19:11:58.231878Z","shell.execute_reply":"2023-09-30T19:13:15.569494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.layers import Input, Dense, GlobalAveragePooling2D\nfrom tensorflow.keras.models import Model","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:28:22.250290Z","iopub.execute_input":"2023-09-30T20:28:22.250811Z","iopub.status.idle":"2023-09-30T20:28:22.256352Z","shell.execute_reply.started":"2023-09-30T20:28:22.250756Z","shell.execute_reply":"2023-09-30T20:28:22.255481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_shape = (224, 224, 3)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:28:40.935357Z","iopub.execute_input":"2023-09-30T20:28:40.935719Z","iopub.status.idle":"2023-09-30T20:28:40.939753Z","shell.execute_reply.started":"2023-09-30T20:28:40.935687Z","shell.execute_reply":"2023-09-30T20:28:40.938570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_layer = Input(shape=input_shape)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:29:46.095452Z","iopub.execute_input":"2023-09-30T20:29:46.095784Z","iopub.status.idle":"2023-09-30T20:29:46.101199Z","shell.execute_reply.started":"2023-09-30T20:29:46.095755Z","shell.execute_reply":"2023-09-30T20:29:46.100353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_model = tf.keras.applications.ResNet50(\n    weights='imagenet', \n    include_top=False, \n    input_tensor=input_layer\n)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:30:09.890047Z","iopub.execute_input":"2023-09-30T20:30:09.890367Z","iopub.status.idle":"2023-09-30T20:30:11.299984Z","shell.execute_reply.started":"2023-09-30T20:30:09.890335Z","shell.execute_reply":"2023-09-30T20:30:11.299075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = GlobalAveragePooling2D()(base_model.output)\nx = Dense(1024, activation='relu')(x)\noutput_layer = Dense(1, activation='sigmoid')(x)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:30:29.390543Z","iopub.execute_input":"2023-09-30T20:30:29.390868Z","iopub.status.idle":"2023-09-30T20:30:29.416523Z","shell.execute_reply.started":"2023-09-30T20:30:29.390839Z","shell.execute_reply":"2023-09-30T20:30:29.415716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = Model(inputs=input_layer, outputs=output_layer)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:30:41.724318Z","iopub.execute_input":"2023-09-30T20:30:41.724685Z","iopub.status.idle":"2023-09-30T20:30:41.744572Z","shell.execute_reply.started":"2023-09-30T20:30:41.724652Z","shell.execute_reply":"2023-09-30T20:30:41.743795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:31:27.148631Z","iopub.execute_input":"2023-09-30T20:31:27.148964Z","iopub.status.idle":"2023-09-30T20:31:27.164154Z","shell.execute_reply.started":"2023-09-30T20:31:27.148935Z","shell.execute_reply":"2023-09-30T20:31:27.163316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(train_generator, epochs=10, steps_per_epoch=len(train_generator))","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:31:51.595525Z","iopub.execute_input":"2023-09-30T20:31:51.595846Z","iopub.status.idle":"2023-09-30T20:41:18.314616Z","shell.execute_reply.started":"2023-09-30T20:31:51.595817Z","shell.execute_reply":"2023-09-30T20:41:18.313675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy = model.evaluate(validation_generator)[1]\nprint(f'Accuracy: {accuracy * 100:.2f}%')","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:41:39.612718Z","iopub.execute_input":"2023-09-30T20:41:39.613030Z","iopub.status.idle":"2023-09-30T20:41:50.074031Z","shell.execute_reply.started":"2023-09-30T20:41:39.613000Z","shell.execute_reply":"2023-09-30T20:41:50.073065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.save('/kaggle/working/.h5')","metadata":{"execution":{"iopub.status.busy":"2023-09-30T20:42:23.389870Z","iopub.execute_input":"2023-09-30T20:42:23.390189Z","iopub.status.idle":"2023-09-30T20:42:47.809691Z","shell.execute_reply.started":"2023-09-30T20:42:23.390158Z","shell.execute_reply":"2023-09-30T20:42:47.808770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Select a batch of images and labels from the validation generator\nimages, labels = validation_generator.next()\n\n# Choose an index to select a specific image from the batch (e.g., index 0)\nindex = 0\n\n# Get the selected image and label\nimage = images[index]\nlabel = labels[index]\n\n# Convert the label to a string (0 or 1)\nlabel_str = \"Melanoma\" if label == 1 else \"Non-Melanoma\"\n\n# Display the image\nplt.imshow(image)\nplt.title(f\"Label: {label_str}\")\nplt.axis('off')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:06:39.811924Z","iopub.execute_input":"2023-09-30T21:06:39.812257Z","iopub.status.idle":"2023-09-30T21:06:40.527300Z","shell.execute_reply.started":"2023-09-30T21:06:39.812224Z","shell.execute_reply":"2023-09-30T21:06:40.526460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\ndef plot_images(images, labels, predictions=None):\n    plt.figure(figsize=(15, 15))\n    for i in range(min(9, len(images))):\n        plt.subplot(3, 3, i + 1)\n        plt.imshow(images[i])\n        plt.title(f\"Label: {labels[i]}\")\n        if predictions is not None:\n            plt.xlabel(f\"Predicted: {predictions[i][0]:.2f}\")\n        plt.axis(\"off\")\n    plt.show()\nsample_images, sample_labels = validation_generator.next()\npredictions = model.predict(sample_images)\nplot_images(sample_images, sample_labels, predictions)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:08:57.280191Z","iopub.execute_input":"2023-09-30T21:08:57.280553Z","iopub.status.idle":"2023-09-30T21:08:59.121233Z","shell.execute_reply.started":"2023-09-30T21:08:57.280520Z","shell.execute_reply":"2023-09-30T21:08:59.120496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\ndataset_path = '/kaggle/input/siim-isic-melanoma-classification'\ntrain_csv_path = f'{dataset_path}/train.csv'\ndf_train = pd.read_csv(train_csv_path)\nclass_names = df_train['diagnosis'].unique()\nprint(\"Class Names:\", class_names)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:16:56.984985Z","iopub.execute_input":"2023-09-30T21:16:56.985362Z","iopub.status.idle":"2023-09-30T21:16:57.033193Z","shell.execute_reply.started":"2023-09-30T21:16:56.985330Z","shell.execute_reply":"2023-09-30T21:16:57.032189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport os\nfrom PIL import Image\nimport matplotlib.pyplot as plt\ndataset_path = '/kaggle/input/siim-isic-melanoma-classification'\ntrain_csv_path = f'{dataset_path}/train.csv'\ndf_train = pd.read_csv(train_csv_path)\nclass_names = df_train['diagnosis'].unique()\nclass_dict = {class_name: i for i, class_name in enumerate(class_names)}\nprint(\"Class Dictionary:\", class_dict)\n\n# Function to display images\ndef show_images(images, labels, class_dict):\n    plt.figure(figsize=(12, 12))\n    for i in range(len(images)):\n        plt.subplot(4, 4, i + 1)\n        plt.imshow(images[i])\n        plt.title(f\"Class: {labels[i]} ({class_dict[labels[i]]})\")\n        plt.axis('off')\n    plt.show()\nsample_images = []\nsample_labels = []\n\nfor idx, row in df_train.head(4).iterrows():\n    image_id = row['image_name'] + '.jpg'\n    image_path = os.path.join(dataset_path, 'jpeg/train', image_id)\n    image = Image.open(image_path)\n    sample_images.append(image)\n    sample_labels.append(row['diagnosis'])\n\nshow_images(sample_images, sample_labels, class_dict)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:18:36.667143Z","iopub.execute_input":"2023-09-30T21:18:36.667483Z","iopub.status.idle":"2023-09-30T21:18:40.956562Z","shell.execute_reply.started":"2023-09-30T21:18:36.667449Z","shell.execute_reply":"2023-09-30T21:18:40.955670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport os\nfrom PIL import Image\nimport matplotlib.pyplot as plt\ndataset_path = '/kaggle/input/siim-isic-melanoma-classification'\ntrain_csv_path = f'{dataset_path}/train.csv'\ndf_train = pd.read_csv(train_csv_path)\nclass_names = df_train['diagnosis'].unique()\nclass_dict = {class_name: i for i, class_name in enumerate(class_names)}\nprint(\"Class Dictionary:\", class_dict)\ndef show_images(images, labels, class_dict):\n    plt.figure(figsize=(12, 12))\n    for i in range(len(images)):\n        plt.subplot(4, 4, i + 1)\n        plt.imshow(images[i])\n        plt.title(f\"Class: {labels[i]} ({class_dict.get(labels[i], 'unknown')})\")\n        plt.axis('off')\n    plt.show()\nsample_images = []\nsample_labels = []\n\nfor idx, row in df_train.head(4).iterrows():\n    image_id = row['image_name'] + '.jpg'\n    image_path = os.path.join(dataset_path, 'jpeg/train', image_id)\n    image = Image.open(image_path)\n    sample_images.append(image)\n    sample_labels.append(row['diagnosis'])\n\nshow_images(sample_images, sample_labels, class_dict)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:20:49.573510Z","iopub.execute_input":"2023-09-30T21:20:49.573856Z","iopub.status.idle":"2023-09-30T21:20:53.608321Z","shell.execute_reply.started":"2023-09-30T21:20:49.573825Z","shell.execute_reply":"2023-09-30T21:20:53.607536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport os\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom tensorflow.keras.layers import Input, Dense, GlobalAveragePooling2D\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:25:20.138547Z","iopub.execute_input":"2023-09-30T21:25:20.138893Z","iopub.status.idle":"2023-09-30T21:25:20.143407Z","shell.execute_reply.started":"2023-09-30T21:25:20.138862Z","shell.execute_reply":"2023-09-30T21:25:20.142471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_path = '/kaggle/input/siim-isic-melanoma-classification'\ntrain_csv_path = f'{dataset_path}/train.csv'\ndf_train = pd.read_csv(train_csv_path)\nclass_names = df_train['diagnosis'].unique()\nclass_dict = {class_name: i for i, class_name in enumerate(class_names)}\nprint(\"Class Dictionary:\", class_dict)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:25:34.366347Z","iopub.execute_input":"2023-09-30T21:25:34.366720Z","iopub.status.idle":"2023-09-30T21:25:34.415648Z","shell.execute_reply.started":"2023-09-30T21:25:34.366686Z","shell.execute_reply":"2023-09-30T21:25:34.414625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_images(images, labels, class_dict):\n    plt.figure(figsize=(12, 12))\n    for i in range(len(images)):\n        plt.subplot(4, 4, i + 1)\n        plt.imshow(images[i])\n        plt.title(f\"Class: {labels[i]} ({class_dict.get(labels[i], 'unknown')})\")\n        plt.axis('off')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:26:03.139782Z","iopub.execute_input":"2023-09-30T21:26:03.140094Z","iopub.status.idle":"2023-09-30T21:26:03.145101Z","shell.execute_reply.started":"2023-09-30T21:26:03.140064Z","shell.execute_reply":"2023-09-30T21:26:03.144210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_images = []\nsample_labels = []\n\nfor idx, row in df_train.head(4).iterrows():\n    image_id = row['image_name'] + '.jpg'\n    image_path = os.path.join(dataset_path, 'jpeg/train', image_id)\n    image = Image.open(image_path)\n    sample_images.append(image)\n    sample_labels.append(row['diagnosis'])\n\nshow_images(sample_images, sample_labels, class_dict)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:26:26.163327Z","iopub.execute_input":"2023-09-30T21:26:26.163736Z","iopub.status.idle":"2023-09-30T21:26:30.308736Z","shell.execute_reply.started":"2023-09-30T21:26:26.163699Z","shell.execute_reply":"2023-09-30T21:26:30.307896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_images = []\nsample_labels = []\n\nfor idx, row in df_train.head(4).iterrows():\n    image_id = row['image_name'] + '.jpg'\n    image_path = os.path.join(dataset_path, 'jpeg/train', image_id)\n    image = Image.open(image_path)\n    sample_images.append(image)\n    sample_labels.append(row['diagnosis'])\n\nshow_images(sample_images, sample_labels, class_dict)","metadata":{"execution":{"iopub.status.busy":"2023-09-30T21:26:52.616072Z","iopub.execute_input":"2023-09-30T21:26:52.616440Z","iopub.status.idle":"2023-09-30T21:26:56.821613Z","shell.execute_reply.started":"2023-09-30T21:26:52.616398Z","shell.execute_reply":"2023-09-30T21:26:56.820811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.layers import Input, Conv2D, MaxPooling2D, Flatten, Dense\nfrom tensorflow.keras.models import Model\nimport numpy as np\nimport cv2\nimport matplotlib.pyplot as plt\ndef generate_synthetic_dataset(num_samples=1000, image_size=(128, 128)):\n    images = np.zeros((num_samples, *image_size, 3), dtype=np.uint8)\n    labels = np.zeros((num_samples, 4), dtype=np.float32)  # [x, y, width, height]\n\n    for i in range(num_samples):\n        x = np.random.randint(0, image_size[0] - 32)\n        y = np.random.randint(0, image_size[1] - 32)\n        width = np.random.randint(20, 50)\n        height = np.random.randint(20, 50)\n        image = np.zeros((*image_size, 3), dtype=np.uint8)\n        image[y:y+height, x:x+width] = [255, 255, 255]\n        images[i] = image\n        labels[i] = [x, y, width, height]\n\n    return images, labels\nimages, labels = generate_synthetic_dataset(num_samples=1000)\nimages = images / 255.0\nsplit_ratio = 0.8\nsplit_index = int(len(images) * split_ratio)\n\ntrain_images, test_images = images[:split_index], images[split_index:]\ntrain_labels, test_labels = labels[:split_index], labels[split_index:]\ninput_shape = (*images.shape[1:],)\ninput_layer = Input(shape=input_shape)\nx = Conv2D(32, (3, 3), activation='relu')(input_layer)\nx = MaxPooling2D((2, 2))(x)\nx = Conv2D(64, (3, 3), activation='relu')(x)\nx = MaxPooling2D((2, 2))(x)\nx = Flatten()(x)\nx = Dense(128, activation='relu')(x)\noutput_layer = Dense(4, activation='linear')(x)  \n\nmodel = Model(inputs=input_layer, outputs=output_layer)\nmodel.compile(optimizer='adam', loss='mse')  \nmodel.fit(train_images, train_labels, epochs=10, batch_size=32, validation_split=0.2)\nloss = model.evaluate(test_images, test_labels)\nprint(f'Test Loss: {loss}')\nsample_image = test_images[0]\nsample_image = np.expand_dims(sample_image, axis=0)\npredicted_bbox = model.predict(sample_image)[0]\nx, y, width, height = predicted_bbox\nsample_image = sample_image[0] * 255.0  \nsample_image = sample_image.astype(np.uint8)\ncv2.rectangle(sample_image, (int(x), int(y)), (int(x+width), int(y+height)), (255, 0, 0), 2)\n\nplt.imshow(sample_image)\nplt.title('Predicted Bounding Box')\nplt.axis('off')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-01T00:38:52.845907Z","iopub.execute_input":"2023-10-01T00:38:52.846252Z","iopub.status.idle":"2023-10-01T00:39:14.024122Z","shell.execute_reply.started":"2023-10-01T00:38:52.846223Z","shell.execute_reply":"2023-10-01T00:39:14.023261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\n\n# Load the saved model from a file\nloaded_model = tf.keras.models.load_model('/kaggle/working/.h5')","metadata":{"execution":{"iopub.status.busy":"2023-09-30T23:50:07.183227Z","iopub.execute_input":"2023-09-30T23:50:07.183579Z","iopub.status.idle":"2023-09-30T23:50:16.969518Z","shell.execute_reply.started":"2023-09-30T23:50:07.183546Z","shell.execute_reply":"2023-09-30T23:50:16.968604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images, labels = generate_synthetic_dataset(num_samples=1000)\nimages = images / 255.0\nsplit_ratio = 0.8\nsplit_index = int(len(images) * split_ratio)\n\ntrain_images, test_images = images[:split_index], images[split_index:]\ntrain_labels, test_labels = labels[:split_index], labels[split_index:]\n\n# Define the model\ninput_shape = (*images.shape[1:],)\ninput_layer = Input(shape=input_shape)\nx = Conv2D(32, (3, 3), activation='relu')(input_layer)\nx = MaxPooling2D((2, 2))(x)\nx = Conv2D(64, (3, 3), activation='relu')(x)\nx = MaxPooling2D((2, 2))(x)\nx = Flatten()(x)\nx = Dense(128, activation='relu')(x)\noutput_layer = Dense(4, activation='linear')(x)\n\nmodel = Model(inputs=input_layer, outputs=output_layer)\nmodel.compile(optimizer='adam', loss='mse')  \n\n# Train the model\nhistory = model.fit(train_images, train_labels, epochs=10, batch_size=32, validation_split=0.2)\n\n# Evaluate the model on the test set and calculate accuracy\nloss = model.evaluate(test_images, test_labels)\naccuracy = 100.0 - loss  # You can consider this as a simple accuracy metric\n\nprint(f'Test Loss: {loss}')\nprint(f'Test Accuracy: {accuracy:.2f}%')","metadata":{"execution":{"iopub.status.busy":"2023-10-01T00:42:30.442273Z","iopub.execute_input":"2023-10-01T00:42:30.442613Z","iopub.status.idle":"2023-10-01T00:42:38.938867Z","shell.execute_reply.started":"2023-10-01T00:42:30.442578Z","shell.execute_reply":"2023-10-01T00:42:38.937923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the trained model to a file\nmodel.save('/kaggle/working/.h5')","metadata":{"execution":{"iopub.status.busy":"2023-10-01T00:43:58.902405Z","iopub.execute_input":"2023-10-01T00:43:58.902764Z","iopub.status.idle":"2023-10-01T00:43:59.863237Z","shell.execute_reply.started":"2023-10-01T00:43:58.902710Z","shell.execute_reply":"2023-10-01T00:43:59.862272Z"},"trusted":true},"execution_count":null,"outputs":[]}]}