{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":20270,"databundleVersionId":1222630,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Melanoma Detection Using a CNN\nCancer is a common illness caused by cells in the body growing out of control. It can show up in different ways, such as Breast Cancer, Lung Cancer, and Pancreatic Cancer.\n\nSkin Cancer, particularly Melanoma, is often treatable, but its prognosis can turn severe without early detection. Therefore, it is crucial to distinguish between malignant and benign tumors.\n\nThe proposed model employs Convolutional Neural Networks for image classification, demonstrating the potential to achieve favorable outcomes in identifying cancer types.","metadata":{}},{"cell_type":"markdown","source":"****1 Data Preparation****\n\nData preparation in melanoma detection involves organizing and pre-processing the data used to train and test the model for identifying melanoma. This process is crucial to ensure the model's effectiveness. Here are the key steps involved in data preparation for melanoma detection.\n1. Import all important python libraries\n2. Reading the dataset\n     * Analyzing the data(shape,head,tail,info etc)\n     * Check for Duplicates\n     * Missing value calculation\n3. Data Reduction\n4. Feature Engineering\n5. Creating Features\n6. Data Cleaning/Wrangling\n7. EDA (trends,patterns, outliers,insigihts,missing values)\n8. Statistics summary(describe)\n9. EDA Univariant Analysis\n10. Data Trasnformation\n11. EDA Bivariant Analysis\n12. EDA Multivariant Analysis (Heat Map)\n13. Impute missing values (Mean,Mode & Median)","metadata":{}},{"cell_type":"markdown","source":"Import all important python libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd                                       # to read csv file \nimport numpy as np                                        #numerical computing library\nimport seaborn as sns                                     #data visualization library\nimport matplotlib.pyplot as plt                           #data visualization library\nimport os                                                 #related to file and directory manipulation\nfrom PIL import Image\n#from sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:01:13.07851Z","iopub.execute_input":"2024-04-17T16:01:13.079447Z","iopub.status.idle":"2024-04-17T16:01:15.931193Z","shell.execute_reply.started":"2024-04-17T16:01:13.079406Z","shell.execute_reply":"2024-04-17T16:01:15.929406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By using the OS module acess the images","metadata":{}},{"cell_type":"code","source":"file_path= '/kaggle/input/siim-isic-melanoma-classification/jpeg/train'\nname_class = os.listdir(file_path)\nname_class[1:11]","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:01:20.826329Z","iopub.execute_input":"2024-04-17T16:01:20.827735Z","iopub.status.idle":"2024-04-17T16:01:21.864277Z","shell.execute_reply.started":"2024-04-17T16:01:20.827695Z","shell.execute_reply":"2024-04-17T16:01:21.862374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the number of rows and columns for subplots\nnum_rows = 5\nnum_cols = 3\n\n# Create a subplot grid\nfig, axes = plt.subplots(num_rows, num_cols, figsize=(15, 15))\n\n# Iterate over the file names and display images\nfor i, file_name in enumerate(name_class[:num_rows*num_cols]):\n    # Load image using PIL\n    img_path = os.path.join(file_path, file_name)\n    img = Image.open(img_path)\n    \n    # Plot the image\n    axes[i // num_cols, i % num_cols].imshow(img)\n    axes[i // num_cols, i % num_cols].set_title(file_name)\n    axes[i // num_cols, i % num_cols].axis('off')\n\n# Adjust layout and display the plot\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:01:25.61147Z","iopub.execute_input":"2024-04-17T16:01:25.612363Z","iopub.status.idle":"2024-04-17T16:01:55.192821Z","shell.execute_reply.started":"2024-04-17T16:01:25.612315Z","shell.execute_reply":"2024-04-17T16:01:55.19108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"file_path_test= '/kaggle/input/siim-isic-melanoma-classification/jpeg/test'\nname_class = os.listdir(file_path_test)\nname_class[1:11]","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:06:30.043577Z","iopub.execute_input":"2024-04-17T16:06:30.044052Z","iopub.status.idle":"2024-04-17T16:06:30.88553Z","shell.execute_reply.started":"2024-04-17T16:06:30.044008Z","shell.execute_reply":"2024-04-17T16:06:30.884489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the number of rows and columns for subplots\nnum_rows = 5\nnum_cols = 3\n\n# Create a subplot grid\nfig, axes = plt.subplots(num_rows, num_cols, figsize=(15, 15))\n\n# Iterate over the file names and display images\nfor i, file_name in enumerate(name_class[:num_rows*num_cols]):\n    # Load image using PIL\n    img_path = os.path.join(file_path_test, file_name)\n    img = Image.open(img_path)\n    \n    # Plot the image\n    axes[i // num_cols, i % num_cols].imshow(img)\n    axes[i // num_cols, i % num_cols].set_title(file_name)\n    axes[i // num_cols, i % num_cols].axis('off')\n\n# Adjust layout and display the plot\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:06:35.419153Z","iopub.execute_input":"2024-04-17T16:06:35.419543Z","iopub.status.idle":"2024-04-17T16:07:01.942533Z","shell.execute_reply.started":"2024-04-17T16:06:35.419514Z","shell.execute_reply":"2024-04-17T16:07:01.940689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pathlib import Path\n\n# Specify the paths to the directories containing the train and test images\ntrain_dir_path = \"/kaggle/input/siim-isic-melanoma-classification/jpeg/train\"\ntest_dir_path = \"/kaggle/input/siim-isic-melanoma-classification/jpeg/test\"\n\n# Convert the string paths to Path objects\ntrain_path = Path(train_dir_path)\ntest_path = Path(test_dir_path)\n\n# Count the number of images in the train and test directories using glob\nimage_count_train = len(list(train_path.glob('*.jpg')))\nprint(\"Number of images in train directory:\", image_count_train)\n\nimage_count_test = len(list(test_path.glob('*.jpg')))\nprint(\"Number of images in test directory:\", image_count_test)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:19:44.969716Z","iopub.execute_input":"2024-04-17T16:19:44.970166Z","iopub.status.idle":"2024-04-17T16:19:45.289253Z","shell.execute_reply.started":"2024-04-17T16:19:44.970117Z","shell.execute_reply":"2024-04-17T16:19:45.288404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Preprocessing of train.csv**","metadata":{}},{"cell_type":"markdown","source":"**Reading the dataset**\n\nBy using the command pd.read_csv fetch the test.csv dataset and process it\nReturn the first five items of dataset after processing using the .head()","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/train.csv')\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:20:56.831531Z","iopub.execute_input":"2024-04-17T16:20:56.831982Z","iopub.status.idle":"2024-04-17T16:20:56.910812Z","shell.execute_reply.started":"2024-04-17T16:20:56.831949Z","shell.execute_reply":"2024-04-17T16:20:56.909622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analyzing the data**\n\nchecking the shape, tail, describe and info of the dataset","metadata":{}},{"cell_type":"code","source":"train_data.tail()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:21:01.314421Z","iopub.execute_input":"2024-04-17T16:21:01.314809Z","iopub.status.idle":"2024-04-17T16:21:01.330332Z","shell.execute_reply.started":"2024-04-17T16:21:01.31478Z","shell.execute_reply":"2024-04-17T16:21:01.3292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:21:05.891114Z","iopub.execute_input":"2024-04-17T16:21:05.892028Z","iopub.status.idle":"2024-04-17T16:21:05.929255Z","shell.execute_reply.started":"2024-04-17T16:21:05.891991Z","shell.execute_reply":"2024-04-17T16:21:05.92813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.describe()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:21:09.822019Z","iopub.execute_input":"2024-04-17T16:21:09.822442Z","iopub.status.idle":"2024-04-17T16:21:09.844318Z","shell.execute_reply.started":"2024-04-17T16:21:09.82241Z","shell.execute_reply":"2024-04-17T16:21:09.843245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.size","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:21:23.211668Z","iopub.execute_input":"2024-04-17T16:21:23.212059Z","iopub.status.idle":"2024-04-17T16:21:23.219717Z","shell.execute_reply.started":"2024-04-17T16:21:23.212025Z","shell.execute_reply":"2024-04-17T16:21:23.218322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:22:46.019207Z","iopub.execute_input":"2024-04-17T16:22:46.019672Z","iopub.status.idle":"2024-04-17T16:22:46.027481Z","shell.execute_reply.started":"2024-04-17T16:22:46.019641Z","shell.execute_reply":"2024-04-17T16:22:46.026198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Check for Duplicates**","metadata":{}},{"cell_type":"code","source":"# Check for duplicates in the 'image_name' column\nimage_name_duplicates = train_data['image_name'].duplicated()\nprint(image_name_duplicates)\n# Print the rows containing duplicate 'image_name'\nprint(\"Duplicate Rows based on 'image_name':\")\nprint(train_data[image_name_duplicates])\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:22:49.347773Z","iopub.execute_input":"2024-04-17T16:22:49.348159Z","iopub.status.idle":"2024-04-17T16:22:49.365287Z","shell.execute_reply.started":"2024-04-17T16:22:49.34813Z","shell.execute_reply":"2024-04-17T16:22:49.363956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Missing value calculation**  or **Data Cleaning**","metadata":{}},{"cell_type":"code","source":"# Check for missing values in the entire DataFrame\nmissing_values = train_data.isnull().sum()\n# Print the columns with missing values\nprint(\"Columns with Missing Values:\")\nprint(missing_values[missing_values > 0])","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:22:54.207555Z","iopub.execute_input":"2024-04-17T16:22:54.207974Z","iopub.status.idle":"2024-04-17T16:22:54.239951Z","shell.execute_reply.started":"2024-04-17T16:22:54.207942Z","shell.execute_reply":"2024-04-17T16:22:54.238647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Unique values for the columns with the missing values**","metadata":{}},{"cell_type":"code","source":"print(train_data['sex'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:22:58.140635Z","iopub.execute_input":"2024-04-17T16:22:58.141721Z","iopub.status.idle":"2024-04-17T16:22:58.149316Z","shell.execute_reply.started":"2024-04-17T16:22:58.141683Z","shell.execute_reply":"2024-04-17T16:22:58.148145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_data['age_approx'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:01.003282Z","iopub.execute_input":"2024-04-17T16:23:01.003732Z","iopub.status.idle":"2024-04-17T16:23:01.012351Z","shell.execute_reply.started":"2024-04-17T16:23:01.003699Z","shell.execute_reply":"2024-04-17T16:23:01.011137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_data['anatom_site_general_challenge'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:05.208208Z","iopub.execute_input":"2024-04-17T16:23:05.208654Z","iopub.status.idle":"2024-04-17T16:23:05.216445Z","shell.execute_reply.started":"2024-04-17T16:23:05.208622Z","shell.execute_reply":"2024-04-17T16:23:05.215236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"****Checking for the outliers before handling the missing values****","metadata":{}},{"cell_type":"markdown","source":"before start to handling the missing values it is always advisable to check the Outliers. Because the outliers can affect the accuracy of the imputation of the missing values.","metadata":{}},{"cell_type":"markdown","source":"**some common techniques to detect outliers in your dataset:**\n\n**Visual Inspection:**\n\n**Boxplots**: Boxplots are graphical representations that display the distribution of data along a number line, highlighting the median, quartiles, and potential outliers.\n\n**Scatter plots**: Scatter plots can help identify outliers by visualizing the relationship between two variables. Outliers may appear as points that are far away from the main cluster of data points.\n\n**Statistical Methods**:\n\n**Z-Score**: Calculate the Z-score for each data point, which measures how many standard deviations an observation is away from the mean. Points with Z-scores beyond a certain threshold (commonly 2.5 or 3) can be considered outliers.\n\n**IQR (Interquartile Range)**: Calculate the interquartile range (IQR) for each numerical variable, which is the difference between the 75th and 25th percentiles. Points outside a defined range (e.g., 1.5 times the IQR above the third quartile or below the first quartile) can be flagged as outliers.","metadata":{}},{"cell_type":"code","source":"# Select a specific variable/column for analysis\nvariable = 'age_approx'  \n\n# Extract the data for the selected variable\ndata = train_data[variable]\n\n# Calculate quartiles and IQR\nQ1 = data.quantile(0.25)\nQ3 = data.quantile(0.75)\nIQR = Q3 - Q1\n\n# Find outliers using IQR\nlower_bound = Q1 - 1.5 * IQR\nupper_bound = Q3 + 1.5 * IQR\noutliers = data[(data < lower_bound) | (data > upper_bound)]\n\n\n\n# Plot IQR chart\nsns.boxplot(x=data)\nplt.title(\"IQR Chart of \" + variable)\nplt.xlabel(variable)\nplt.show()\n\n# Plot histogram\nsns.histplot(data, kde=True)\nplt.title(\"Distribution of \" + variable)\nplt.xlabel(variable)\nplt.ylabel(\"Count\")\nplt.show()\n\n# Print results\nprint(\"Outliers:\")\nprint(outliers)\nprint()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:09.906203Z","iopub.execute_input":"2024-04-17T16:23:09.906639Z","iopub.status.idle":"2024-04-17T16:23:10.816712Z","shell.execute_reply.started":"2024-04-17T16:23:09.906609Z","shell.execute_reply":"2024-04-17T16:23:10.815553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"there is no outliers for the AGE and we are good to go for the imputation of missing values.","metadata":{}},{"cell_type":"markdown","source":"**impute missing values with the mode (most frequent value) for the categorical variables**\n\nIdentifying outliers for categorical variables such as \"anatom_site_general_challenge\" is not a conventional task because outliers are typically associated with numerical data points. However, we can still perform most frequent value to identify unusual or potentially erroneous values in categorical columns. ","metadata":{}},{"cell_type":"code","source":"# Calculate the frequency counts of each category in the column\nfrequency_counts = train_data['anatom_site_general_challenge'].value_counts()\nprint(frequency_counts)\n# Find the most frequent category (mode)\nmost_frequent_category = frequency_counts.index[0]\nprint(\"the mode is\", most_frequent_category)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:17.099958Z","iopub.execute_input":"2024-04-17T16:23:17.101045Z","iopub.status.idle":"2024-04-17T16:23:17.113689Z","shell.execute_reply.started":"2024-04-17T16:23:17.10101Z","shell.execute_reply":"2024-04-17T16:23:17.112549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frequency_counts_gender = train_data['sex'].value_counts()\nprint(frequency_counts_gender)\n# Find the most frequent category (mode)\nmost_frequent_category_sex = frequency_counts_gender.index[0]\nprint(\"the mode is\", most_frequent_category_sex)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:21.88044Z","iopub.execute_input":"2024-04-17T16:23:21.880837Z","iopub.status.idle":"2024-04-17T16:23:21.892989Z","shell.execute_reply.started":"2024-04-17T16:23:21.880809Z","shell.execute_reply":"2024-04-17T16:23:21.892144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Handling or replacing the missing values of each column**","metadata":{}},{"cell_type":"markdown","source":"replace the non-numerical column of \"sex\" and \"anatom_site_general_challenge\" with **most frequent value or mode** and the numeric column \"age_approx\" with the **median**","metadata":{}},{"cell_type":"code","source":"# Replace missing values with the most frequent category\ntrain_data['anatom_site_general_challenge'].fillna(most_frequent_category, inplace=True)\ntrain_data['sex'].fillna(most_frequent_category_sex, inplace=True)\n\n# Replace missing values with the median\nmedian_age = 50\ntrain_data['age_approx'].fillna(median_age, inplace=True)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:32.60491Z","iopub.execute_input":"2024-04-17T16:23:32.60573Z","iopub.status.idle":"2024-04-17T16:23:32.619896Z","shell.execute_reply.started":"2024-04-17T16:23:32.60569Z","shell.execute_reply":"2024-04-17T16:23:32.61911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"check the successful handling of missing values","metadata":{}},{"cell_type":"code","source":"missing_values_after = train_data['anatom_site_general_challenge'].isnull().sum()\nprint(\"Number of missing values after replacement for the anatomy site:\", missing_values_after)\n\nmissing_values_after1 = train_data['sex'].isnull().sum()\nprint(\"Number of missing values after replacement for the sex:\", missing_values_after1)\n\nmissing_values_after2 = train_data['age_approx'].isnull().sum()\nprint(\"Number of missing values after replacement for the age_approx:\", missing_values_after2)\n\n\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:37.941752Z","iopub.execute_input":"2024-04-17T16:23:37.942415Z","iopub.status.idle":"2024-04-17T16:23:37.957529Z","shell.execute_reply.started":"2024-04-17T16:23:37.942376Z","shell.execute_reply":"2024-04-17T16:23:37.95636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for missing values in the entire DataFrame after handling of missing values\nmissing_values_after = train_data.isnull().sum()\n# Print the columns with missing values\nprint(\"Columns with Missing Values:\")\nprint(missing_values_after[missing_values_after > 0])","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:44.113871Z","iopub.execute_input":"2024-04-17T16:23:44.114984Z","iopub.status.idle":"2024-04-17T16:23:44.145277Z","shell.execute_reply.started":"2024-04-17T16:23:44.114946Z","shell.execute_reply":"2024-04-17T16:23:44.143901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**EDA (Exploratory Data Analysis)**","metadata":{}},{"cell_type":"markdown","source":"**Univariate Analysis**","metadata":{}},{"cell_type":"code","source":"# For example, plot histograms of numerical columns\ntrain_data.hist(figsize=(10, 8))\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-16T05:36:22.428978Z","iopub.execute_input":"2024-04-16T05:36:22.429819Z","iopub.status.idle":"2024-04-16T05:36:22.960124Z","shell.execute_reply.started":"2024-04-16T05:36:22.42978Z","shell.execute_reply":"2024-04-16T05:36:22.958808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the distribution of the target variable\ntarget_counts = train_data['target'].value_counts()\n\n# Plotting the pie chart\nplt.figure(figsize=(2, 2))\nplt.pie(target_counts, labels=target_counts.index, autopct='%1.1f%%', startangle=140)\nplt.title('Distribution of Target Variable')\nplt.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle\nplt.show()\n#or \n# Calculate the distribution of the benign_malignant variable\ntarget_counts1 = train_data['benign_malignant'].value_counts()\n\n# Plotting the pie chart\nplt.figure(figsize=(2, 2))\nplt.pie(target_counts1, labels=target_counts1.index, autopct='%1.1f%%', startangle=140)\nplt.title('Distribution of benign_malignant Variable')\nplt.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle\nplt.show()\n\n# Calculate the distribution of the age_approx variable\ntarget_counts2 = train_data['age_approx'].value_counts()\n\n# Plotting the pie chart\nplt.figure(figsize=(14, 5))\nplt.pie(target_counts2, labels=target_counts2.index, autopct='%1.1f%%', startangle=140)\nplt.title('Distribution of age_approx Variable')\nplt.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-04-16T05:36:22.96181Z","iopub.execute_input":"2024-04-16T05:36:22.962161Z","iopub.status.idle":"2024-04-16T05:36:23.520041Z","shell.execute_reply.started":"2024-04-16T05:36:22.962131Z","shell.execute_reply":"2024-04-16T05:36:23.517899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the frequency of each category in the 'anatom_site_general_challenge' column\nanatom_site_counts = train_data['anatom_site_general_challenge'].value_counts()\n\n# Plotting the bar plot\nplt.figure(figsize=(6, 2))\nsns.barplot(x=anatom_site_counts.index, y=anatom_site_counts.values, palette=\"viridis\")\nplt.xticks(rotation=45, ha='right')\nplt.xlabel('Anatomical Site')\nplt.ylabel('Frequency')\nplt.title('Distribution of Anatomical Site')\nplt.show()\n\n# Count the frequency of each category in the 'diagnosis' column\nanatom_site_counts1 = train_data['diagnosis'].value_counts()\n\n# Plotting the bar plot\nplt.figure(figsize=(6, 2))\nsns.barplot(x=anatom_site_counts1.index, y=anatom_site_counts1.values, palette=\"viridis\")\nplt.xticks(rotation=45, ha='right')\nplt.xlabel('diagnosis')\nplt.ylabel('Frequency')\nplt.title('Distribution of diagnosis')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-16T05:36:23.52272Z","iopub.execute_input":"2024-04-16T05:36:23.523851Z","iopub.status.idle":"2024-04-16T05:36:24.131618Z","shell.execute_reply.started":"2024-04-16T05:36:23.523785Z","shell.execute_reply":"2024-04-16T05:36:24.130096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Bivariate Analysis**","metadata":{}},{"cell_type":"code","source":"sns.scatterplot(x='age_approx', y='target', data=train_data)\nplt.xlabel('Age')\nplt.ylabel('Target')\nplt.title('Scatter plot of Age vs. Target')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-16T05:36:24.13328Z","iopub.execute_input":"2024-04-16T05:36:24.133702Z","iopub.status.idle":"2024-04-16T05:36:24.520934Z","shell.execute_reply.started":"2024-04-16T05:36:24.133668Z","shell.execute_reply":"2024-04-16T05:36:24.519318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Multivariate Analysis**","metadata":{}},{"cell_type":"code","source":"# Selecting the columns for the pair plot\ncolumns_for_pairplot = ['age_approx', 'anatom_site_general_challenge', 'sex']\n\n# Creating the pair plot\n#sns.pairplot(train_data.sample(1000), hue='age_approx', vars=columns_for_pairplot)\n#plt.title('Pair Plot of Sex, Anatomical Site, and Age')\n#plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-16T05:40:47.669112Z","iopub.execute_input":"2024-04-16T05:40:47.669643Z","iopub.status.idle":"2024-04-16T05:40:47.67606Z","shell.execute_reply.started":"2024-04-16T05:40:47.669595Z","shell.execute_reply":"2024-04-16T05:40:47.674403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Check whether the target column is balanced or not**","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\ncounts= train_data.target.value_counts()\nsns.barplot(x=counts.index, y= counts)\nplt.xlabel('target')\nplt.xticks(rotation=90);","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:23:58.0395Z","iopub.execute_input":"2024-04-17T16:23:58.039939Z","iopub.status.idle":"2024-04-17T16:23:58.273289Z","shell.execute_reply.started":"2024-04-17T16:23:58.039906Z","shell.execute_reply":"2024-04-17T16:23:58.27239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nWhen using a Convolutional Neural Network (CNN) for detection tasks, such as melanoma detection, addressing class imbalance is crucial for the model to effectively learn and generalize across both classes. Here's how we can approach class imbalance with CNNs:\n\n**Resampling Techniques:**\n\n**Oversampling the minority class:**\nThis involves generating synthetic samples for the minority class to increase its representation in the dataset. Techniques like Synthetic Minority Over-sampling Technique (SMOTE) or its variants can be used to create synthetic samples.\n\n**Undersampling the majority class:**\nThis involves randomly removing samples from the majority class to balance the dataset. However, undersampling may lead to loss of information and might not be suitable for complex datasets.\n\nFor CNNs, oversampling the minority class is often preferred over undersampling because it helps in retaining more information and does not reduce the size of the dataset, which is important for CNNs to learn complex patterns","metadata":{}},{"cell_type":"code","source":"pip install imbalanced-learn","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:05.652268Z","iopub.execute_input":"2024-04-17T16:24:05.652721Z","iopub.status.idle":"2024-04-17T16:24:21.726926Z","shell.execute_reply.started":"2024-04-17T16:24:05.65269Z","shell.execute_reply":"2024-04-17T16:24:21.725446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from imblearn.over_sampling import SMOTE  # Synthetic Minority Over-sampling Technique\n\n# Extract target variable\ny = train_data['target']\n\n# Apply SMOTE to oversample the minority class in the target variable\nsmote = SMOTE(random_state=42)\n_, y_smote = smote.fit_resample(train_data[['target']], y)\n\n# y_smote now contains the oversampled target variable (y) with balanced classes\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:24.98182Z","iopub.execute_input":"2024-04-17T16:24:24.98224Z","iopub.status.idle":"2024-04-17T16:24:25.604079Z","shell.execute_reply.started":"2024-04-17T16:24:24.982204Z","shell.execute_reply":"2024-04-17T16:24:25.603084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Assuming y_smote is the target variable after applying SMOTE\ncounts_after = y_smote.value_counts()\n\n# Plot the distribution of the target variable after applying SMOTE\nsns.barplot(x=counts_after.index, y=counts_after)\nplt.xlabel('target')\nplt.ylabel('Count')\nplt.xticks(rotation=90)\nplt.title('Distribution of Target Variable after SMOTE')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:30.49484Z","iopub.execute_input":"2024-04-17T16:24:30.495461Z","iopub.status.idle":"2024-04-17T16:24:30.758939Z","shell.execute_reply.started":"2024-04-17T16:24:30.495428Z","shell.execute_reply":"2024-04-17T16:24:30.75787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*********************************************************************************************************************************************************\n\n*********************************************************************************************************************************************************","metadata":{}},{"cell_type":"markdown","source":"**Preprocessing of test.csv**","metadata":{}},{"cell_type":"code","source":"test_data = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/test.csv')\ntest_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:36.833095Z","iopub.execute_input":"2024-04-17T16:24:36.833512Z","iopub.status.idle":"2024-04-17T16:24:36.880822Z","shell.execute_reply.started":"2024-04-17T16:24:36.833481Z","shell.execute_reply":"2024-04-17T16:24:36.879642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analyzing the data**\n\nchecking the shape, tail, describe and info of the dataset","metadata":{}},{"cell_type":"code","source":"test_data.tail()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:42.280602Z","iopub.execute_input":"2024-04-17T16:24:42.281075Z","iopub.status.idle":"2024-04-17T16:24:42.295951Z","shell.execute_reply.started":"2024-04-17T16:24:42.281039Z","shell.execute_reply":"2024-04-17T16:24:42.294816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.info()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:47.038908Z","iopub.execute_input":"2024-04-17T16:24:47.039739Z","iopub.status.idle":"2024-04-17T16:24:47.057233Z","shell.execute_reply.started":"2024-04-17T16:24:47.039701Z","shell.execute_reply":"2024-04-17T16:24:47.056031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.describe()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:51.518641Z","iopub.execute_input":"2024-04-17T16:24:51.519438Z","iopub.status.idle":"2024-04-17T16:24:51.535404Z","shell.execute_reply.started":"2024-04-17T16:24:51.519401Z","shell.execute_reply":"2024-04-17T16:24:51.534276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.size","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:55.470741Z","iopub.execute_input":"2024-04-17T16:24:55.471805Z","iopub.status.idle":"2024-04-17T16:24:55.47848Z","shell.execute_reply.started":"2024-04-17T16:24:55.471767Z","shell.execute_reply":"2024-04-17T16:24:55.477127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:24:58.135303Z","iopub.execute_input":"2024-04-17T16:24:58.135771Z","iopub.status.idle":"2024-04-17T16:24:58.143304Z","shell.execute_reply.started":"2024-04-17T16:24:58.135739Z","shell.execute_reply":"2024-04-17T16:24:58.14216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Check for duplicates**","metadata":{}},{"cell_type":"code","source":"# Check for duplicates in the 'image_name' column\nimage_name_duplicates = test_data['image_name'].duplicated()\nprint(image_name_duplicates)\n# Print the rows containing duplicate 'image_name'\nprint(\"Duplicate Rows based on 'image_name':\")\nprint(test_data[image_name_duplicates])","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:01.466802Z","iopub.execute_input":"2024-04-17T16:25:01.467212Z","iopub.status.idle":"2024-04-17T16:25:01.479139Z","shell.execute_reply.started":"2024-04-17T16:25:01.467183Z","shell.execute_reply":"2024-04-17T16:25:01.477894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"****Missing value calculation or Data Cleaning****","metadata":{}},{"cell_type":"code","source":"# Check for missing values in the entire DataFrame\nmissing_values1 = test_data.isnull().sum()\n# Print the columns with missing values\nprint(\"Columns with Missing Values:\")\nprint(missing_values1[missing_values1 > 0])","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:05.83474Z","iopub.execute_input":"2024-04-17T16:25:05.835552Z","iopub.status.idle":"2024-04-17T16:25:05.848845Z","shell.execute_reply.started":"2024-04-17T16:25:05.835519Z","shell.execute_reply":"2024-04-17T16:25:05.847437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**impute missing values with the mode (most frequent value) for the categorical variables**\nCalculate the frequency counts of the colum","metadata":{}},{"cell_type":"code","source":"frequency_counts1 = test_data['anatom_site_general_challenge'].value_counts() \nprint(frequency_counts1)\nmost_frequent_category2 = frequency_counts1.index[0] \nprint(\"the mode is\", most_frequent_category2)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:10.041185Z","iopub.execute_input":"2024-04-17T16:25:10.041594Z","iopub.status.idle":"2024-04-17T16:25:10.051224Z","shell.execute_reply.started":"2024-04-17T16:25:10.041564Z","shell.execute_reply":"2024-04-17T16:25:10.050064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Replace missing values with the most frequent category**","metadata":{}},{"cell_type":"code","source":"test_data['anatom_site_general_challenge'].fillna(most_frequent_category, inplace=True)\n\n# Check for missing values in the entire DataFrame\nmissing_values_test_after = test_data.isnull().sum()\n# Print the columns with missing values\nprint(\"Columns with Missing Values:\")\nprint(missing_values_test_after[missing_values_test_after > 0])","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:17.164124Z","iopub.execute_input":"2024-04-17T16:25:17.164552Z","iopub.status.idle":"2024-04-17T16:25:17.180565Z","shell.execute_reply.started":"2024-04-17T16:25:17.164522Z","shell.execute_reply":"2024-04-17T16:25:17.179412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**check specifically**","metadata":{}},{"cell_type":"code","source":"missing_values_after_test = test_data['anatom_site_general_challenge'].isnull().sum()\nprint(\"Number of missing values after replacement for the anatomy site:\", missing_values_after_test)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:21.322318Z","iopub.execute_input":"2024-04-17T16:25:21.323348Z","iopub.status.idle":"2024-04-17T16:25:21.331298Z","shell.execute_reply.started":"2024-04-17T16:25:21.323296Z","shell.execute_reply":"2024-04-17T16:25:21.330138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*********************************************************************************************************************************************************\n*********************************************************************************************************************************************************","metadata":{}},{"cell_type":"markdown","source":"**Apply One Code Encoding for both Train asnd Test data**","metadata":{}},{"cell_type":"markdown","source":"**One Hot Encoding**\n\nTransforming all categorical features to numerical.\n\nstep 1: sex, anatomy, diagnosis need to be encoded.\n\nstep 2: benign_malignant column will be dropped, as the information is already in the target column.","metadata":{}},{"cell_type":"code","source":"# SKLearn\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.preprocessing import OneHotEncoder\n\n# === TRAIN ===\nto_encode = ['sex', 'anatom_site_general_challenge', 'diagnosis']\nencoded_all = []\n\nlabel_encoder = LabelEncoder()\n\nfor column in to_encode:\n    encoded = label_encoder.fit_transform(train_data[column])\n    encoded_all.append(encoded)\n    \ntrain_data['sex'] = encoded_all[0]  # 1 for male and 0 for female\ntrain_data['anatom_site_general_challenge'] = encoded_all[1]\ntrain_data['diagnosis'] = encoded_all[2]\n\nif 'benign_malignant' in train_data.columns : train_data.drop(['benign_malignant'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:26.204542Z","iopub.execute_input":"2024-04-17T16:25:26.20574Z","iopub.status.idle":"2024-04-17T16:25:26.248485Z","shell.execute_reply.started":"2024-04-17T16:25:26.2057Z","shell.execute_reply":"2024-04-17T16:25:26.24753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:30.301238Z","iopub.execute_input":"2024-04-17T16:25:30.301644Z","iopub.status.idle":"2024-04-17T16:25:30.317657Z","shell.execute_reply.started":"2024-04-17T16:25:30.301616Z","shell.execute_reply":"2024-04-17T16:25:30.316284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_data['anatom_site_general_challenge'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:34.901967Z","iopub.execute_input":"2024-04-17T16:25:34.903231Z","iopub.status.idle":"2024-04-17T16:25:34.908698Z","shell.execute_reply.started":"2024-04-17T16:25:34.903185Z","shell.execute_reply":"2024-04-17T16:25:34.907593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_data['diagnosis'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:39.895009Z","iopub.execute_input":"2024-04-17T16:25:39.895478Z","iopub.status.idle":"2024-04-17T16:25:39.901953Z","shell.execute_reply.started":"2024-04-17T16:25:39.895445Z","shell.execute_reply":"2024-04-17T16:25:39.900737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_data['sex'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:42.877964Z","iopub.execute_input":"2024-04-17T16:25:42.879262Z","iopub.status.idle":"2024-04-17T16:25:42.885045Z","shell.execute_reply.started":"2024-04-17T16:25:42.879222Z","shell.execute_reply":"2024-04-17T16:25:42.884264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# === TEST ===\nto_encode = ['sex', 'anatom_site_general_challenge']\nencoded_all = []\n\nlabel_encoder = LabelEncoder()\n\nfor column in to_encode:\n    encoded = label_encoder.fit_transform(test_data[column])\n    encoded_all.append(encoded)\n    \ntest_data['sex'] = encoded_all[0]\ntest_data['anatom_site_general_challenge'] = encoded_all[1]","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:55.153586Z","iopub.execute_input":"2024-04-17T16:25:55.153997Z","iopub.status.idle":"2024-04-17T16:25:55.163345Z","shell.execute_reply.started":"2024-04-17T16:25:55.153967Z","shell.execute_reply":"2024-04-17T16:25:55.16208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:25:59.974748Z","iopub.execute_input":"2024-04-17T16:25:59.975227Z","iopub.status.idle":"2024-04-17T16:25:59.990509Z","shell.execute_reply.started":"2024-04-17T16:25:59.975171Z","shell.execute_reply":"2024-04-17T16:25:59.988945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(test_data['anatom_site_general_challenge'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:26:03.972714Z","iopub.execute_input":"2024-04-17T16:26:03.973143Z","iopub.status.idle":"2024-04-17T16:26:03.981126Z","shell.execute_reply.started":"2024-04-17T16:26:03.973112Z","shell.execute_reply":"2024-04-17T16:26:03.979597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(test_data['sex'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:26:06.813445Z","iopub.execute_input":"2024-04-17T16:26:06.813882Z","iopub.status.idle":"2024-04-17T16:26:06.821388Z","shell.execute_reply.started":"2024-04-17T16:26:06.813849Z","shell.execute_reply":"2024-04-17T16:26:06.820144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the files\ntrain_data.to_csv('train_clean.csv', index=False)\ntest_data.to_csv('test_clean.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:26:12.475643Z","iopub.execute_input":"2024-04-17T16:26:12.476077Z","iopub.status.idle":"2024-04-17T16:26:12.675931Z","shell.execute_reply.started":"2024-04-17T16:26:12.476042Z","shell.execute_reply":"2024-04-17T16:26:12.674671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*********************************************************************************************************************************************************\n\n*********************************************************************************************************************************************************","metadata":{}},{"cell_type":"markdown","source":"**IMAGES**","metadata":{}},{"cell_type":"markdown","source":"There are 2 types of images containing the same information:\n\n.dcm files: DICOM files. It's saved in the \"Digital Imaging and Communications in Medicine\" format. It contains an image from a medical scan, such as an ultrasound or MRI + information about the patient.\n\n.jpeg files: the DICOM files converted into .jpeg format\n\n.tfrec files: The TFRecord file format is a simple record-oriented binary format for ML training data.","metadata":{}},{"cell_type":"markdown","source":"**check the image balance**","metadata":{}},{"cell_type":"markdown","source":"Check if images in .dcm and .jpeg format have the same number of observations as in train_df and test_df.","metadata":{}},{"cell_type":"code","source":"print('Train .dcm number of images:', len(list(os.listdir('../input/siim-isic-melanoma-classification/train'))), '\\n' +\n      'Test .dcm number of images:', len(list(os.listdir('../input/siim-isic-melanoma-classification/test'))), '\\n' +\n      'Train .jpeg number of images:', len(list(os.listdir('../input/siim-isic-melanoma-classification/jpeg/train'))), '\\n' +\n      'Test .jpeg number of images:', len(list(os.listdir('../input/siim-isic-melanoma-classification/jpeg/test'))), '\\n' +\n      '-----------------------', '\\n' +\n      'There is the same number of images as in train/ test .csv datasets')","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:26:21.078514Z","iopub.execute_input":"2024-04-17T16:26:21.078926Z","iopub.status.idle":"2024-04-17T16:26:22.57259Z","shell.execute_reply.started":"2024-04-17T16:26:21.078897Z","shell.execute_reply":"2024-04-17T16:26:22.571392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**image shape**","metadata":{}},{"cell_type":"code","source":"#shapes_train = []\n\n#for k, path in enumerate(train_data['path_jpeg']):\n #   image = Image.open(path)\n  #  shapes_train.append(image.size)\n    \n   # if k >= 100: break\n        \n#shapes_train = pd.DataFrame(data = shapes_train, columns = ['H', 'W'], dtype='object')\n#shapes_train['Size'] = '[' + shapes_train['H'].astype(str) + ', ' + shapes_train['W'].astype(str) + ']'","metadata":{"execution":{"iopub.status.busy":"2024-04-16T06:38:00.035763Z","iopub.execute_input":"2024-04-16T06:38:00.036753Z","iopub.status.idle":"2024-04-16T06:38:00.043236Z","shell.execute_reply.started":"2024-04-16T06:38:00.036697Z","shell.execute_reply":"2024-04-16T06:38:00.041696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set Color Palettes for the notebook\ncolors_nude = ['#e0798c','#65365a','#da8886','#cfc4c4','#dfd7ca']\nsns.palplot(sns.color_palette(colors_nude))\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:26:33.040748Z","iopub.execute_input":"2024-04-17T16:26:33.041182Z","iopub.status.idle":"2024-04-17T16:26:33.151154Z","shell.execute_reply.started":"2024-04-17T16:26:33.041151Z","shell.execute_reply":"2024-04-17T16:26:33.149254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\n#plt.figure(figsize = (16, 6))\n#\n#a = sns.countplot(shapes_train['Size'], palette=colors_nude)\n\n#for p in a.patches:\n #   a.annotate(format(p.get_height(), ','), \n  #         (p.get_x() + p.get_width() / 2., \n   #         p.get_height()), ha = 'center', va = 'center', \n    #       xytext = (0, 4), textcoords = 'offset points')\n    \n#plt.title('100 Images Shapes', fontsize=16)\n#sns.despine(left=True, bottom=True);","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:26:38.105873Z","iopub.execute_input":"2024-04-17T16:26:38.106287Z","iopub.status.idle":"2024-04-17T16:26:38.112915Z","shell.execute_reply.started":"2024-04-17T16:26:38.106257Z","shell.execute_reply":"2024-04-17T16:26:38.111562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%matplotlib inline\nimport matplotlib.image as mpimg\nfrom tabulate import tabulate\nimport missingno as msno \nfrom IPython.display import display_html\nfrom PIL import Image\nimport gc\nimport cv2\nimport pydicom # for DICOM images\nfrom skimage.transform import resize\n\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T16:26:44.709625Z","iopub.execute_input":"2024-04-17T16:26:44.710054Z","iopub.status.idle":"2024-04-17T16:26:45.320961Z","shell.execute_reply.started":"2024-04-17T16:26:44.710023Z","shell.execute_reply":"2024-04-17T16:26:45.319664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**DICOM Images:\nMalignant vs Benign Images**","metadata":{}},{"cell_type":"markdown","source":"Add Image Path to the csv file in order to access the data ","metadata":{}},{"cell_type":"code","source":"# === DICOM ===\n# Create the paths\npath_train = '/kaggle/input/siim-isic-melanoma-classification' + '/train/' + train_data['image_name'] + '.dcm'\npath_test = '/kaggle/input/siim-isic-melanoma-classification' + '/test/' + test_data['image_name'] + '.dcm'\n\n# Append to the original dataframes\ntrain_data['path_dicom'] = path_train\ntest_data['path_dicom'] = path_test\n\n# === JPEG ===\n# Create the paths\npath_train = '/kaggle/input/siim-isic-melanoma-classification' + '/jpeg/train/' + train_data['image_name'] + '.jpg'\npath_test = '/kaggle/input/siim-isic-melanoma-classification' + '/jpeg/test/' + test_data['image_name'] + '.jpg'\n\n# Append to the original dataframes\ntrain_data['path_jpeg'] = path_train\ntest_data['path_jpeg'] = path_test","metadata":{"execution":{"iopub.status.busy":"2024-04-17T17:58:23.080946Z","iopub.execute_input":"2024-04-17T17:58:23.081476Z","iopub.status.idle":"2024-04-17T17:58:23.134074Z","shell.execute_reply.started":"2024-04-17T17:58:23.08144Z","shell.execute_reply":"2024-04-17T17:58:23.13282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the files\ntrain_data.to_csv('train_clean.csv', index=False)\ntest_data.to_csv('test_clean.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-17T17:58:28.53175Z","iopub.execute_input":"2024-04-17T17:58:28.533148Z","iopub.status.idle":"2024-04-17T17:58:29.121205Z","shell.execute_reply.started":"2024-04-17T17:58:28.5331Z","shell.execute_reply":"2024-04-17T17:58:29.11986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"split the training for the train and validation","metadata":{}},{"cell_type":"code","source":"# Perform train-validation split\nfrom sklearn.model_selection import train_test_split\ntrain_df, val_df = train_test_split(train_data, test_size=0.2, random_state=42)\n\n# Apply data augmentation to train dataset (if needed)\n\n# Now, you can use train_df and val_df for training and validation, respectively\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T17:59:13.957548Z","iopub.execute_input":"2024-04-17T17:59:13.958825Z","iopub.status.idle":"2024-04-17T17:59:13.981922Z","shell.execute_reply.started":"2024-04-17T17:59:13.958776Z","shell.execute_reply":"2024-04-17T17:59:13.980608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_images(data, n = 5, rows=1, cols=5, title='Default'):\n    plt.figure(figsize=(16,4))\n\n    for k, path in enumerate(data['path_dicom'][:n]):\n        image = pydicom.read_file(path)\n        image = image.pixel_array\n        \n        # image = resize(image, (200, 200), anti_aliasing=True)\n\n        plt.suptitle(title, fontsize = 16)\n        plt.subplot(rows, cols, k+1)\n        plt.imshow(image)\n        plt.axis('off')","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:02:11.210641Z","iopub.execute_input":"2024-04-17T18:02:11.211247Z","iopub.status.idle":"2024-04-17T18:02:11.221695Z","shell.execute_reply.started":"2024-04-17T18:02:11.211206Z","shell.execute_reply":"2024-04-17T18:02:11.219661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show Benign Samples\nshow_images(train_df[train_df['target'] == 0], n=10, rows=2, cols=5, title='Benign Sample')","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:02:22.28095Z","iopub.execute_input":"2024-04-17T18:02:22.281447Z","iopub.status.idle":"2024-04-17T18:02:40.690279Z","shell.execute_reply.started":"2024-04-17T18:02:22.281413Z","shell.execute_reply":"2024-04-17T18:02:40.688564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show Malignant Samples\nshow_images(train_df[train_df['target'] == 1], n=10, rows=2, cols=5, title='Malignant Sample')","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:03:19.578666Z","iopub.execute_input":"2024-04-17T18:03:19.57916Z","iopub.status.idle":"2024-04-17T18:03:36.210466Z","shell.execute_reply.started":"2024-04-17T18:03:19.579126Z","shell.execute_reply":"2024-04-17T18:03:36.209082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Data Augementation**","metadata":{}},{"cell_type":"code","source":"# Select a small sample of the .jpeg image paths\nimage_list = train_df.sample(16)['path_jpeg']\nimage_list = image_list.reset_index()['path_jpeg']\n\n# Show the sample\nplt.figure(figsize=(16,6))\nplt.suptitle(\"Original images view\", fontsize = 16)\n    \nfor k, path in enumerate(image_list):\n    image = mpimg.imread(path)\n        \n    plt.subplot(4, 4, k+1)\n    plt.imshow(image)\n    plt.axis('off')","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:06:43.835097Z","iopub.execute_input":"2024-04-17T18:06:43.836256Z","iopub.status.idle":"2024-04-17T18:07:17.017876Z","shell.execute_reply.started":"2024-04-17T18:06:43.836216Z","shell.execute_reply":"2024-04-17T18:07:17.016646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pydicom\nimport cv2\nimport matplotlib.pyplot as plt\n\nfig, axes = plt.subplots(nrows=2, ncols=4, figsize=(16, 6))\nplt.suptitle(\"Grey\", fontsize=16)\n\nfor i in range(0, min(2*4, len(train_df['path_dicom']))):\n    data = pydicom.read_file(train_df['path_dicom'].iloc[i])\n    image = data.pixel_array\n    \n    # Transform to B&W\n    # The function converts an input image from one color space to another.\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n    image = cv2.resize(image, (200, 200))\n    \n    x = i // 4\n    y = i % 4\n    axes[x, y].imshow(image, cmap=plt.cm.bone) \n    axes[x, y].axis('off')\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:12:50.844962Z","iopub.execute_input":"2024-04-17T18:12:50.845469Z","iopub.status.idle":"2024-04-17T18:12:52.951748Z","shell.execute_reply.started":"2024-04-17T18:12:50.845435Z","shell.execute_reply":"2024-04-17T18:12:52.950421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Ben Graham: greyscale + Gaussian Blur**","metadata":{}},{"cell_type":"markdown","source":"With Gaussian Blur","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(nrows=2, ncols=4, figsize=(16,6))\nplt.suptitle(\"With Gaussian Blur\", fontsize=16)\n\nfor i, row in enumerate(train_df.iterrows()):\n    if i >= 2*4:\n        break\n    _, data = row\n    image = pydicom.read_file(data['path_dicom']).pixel_array\n    \n    # Transform to B&W\n    # The function converts an input image from one color space to another.\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n    image = cv2.resize(image, (200,200))\n    image = cv2.addWeighted(image, 4, cv2.GaussianBlur(image, (0,0) ,256/10), -4, 128)\n    \n    x = i // 4\n    y = i % 4\n    axes[x, y].imshow(image, cmap=plt.cm.bone) \n    axes[x, y].axis('off')\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:18:14.140984Z","iopub.execute_input":"2024-04-17T18:18:14.14149Z","iopub.status.idle":"2024-04-17T18:18:16.796702Z","shell.execute_reply.started":"2024-04-17T18:18:14.141457Z","shell.execute_reply":"2024-04-17T18:18:16.795579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Length of train_df:\", len(train_df))\nprint(\"Length of train_df:\", len(val_df))","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:23:53.028416Z","iopub.execute_input":"2024-04-17T18:23:53.028841Z","iopub.status.idle":"2024-04-17T18:23:53.035353Z","shell.execute_reply.started":"2024-04-17T18:23:53.028812Z","shell.execute_reply":"2024-04-17T18:23:53.034031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Without Gaussian Blur","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(nrows=2, ncols=6, figsize=(16,6))\nplt.suptitle(\"Without Gaussian Blur\", fontsize=16)\n\nfor i in range(0, min(2*6, len(train_df))):\n    data = pydicom.read_file(train_df['path_dicom'].iloc[i])  # Use .iloc[i] instead of [i]\n    image = data.pixel_array\n    \n    # Transform to B&W\n    # The function converts an input image from one color space to another.\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n    image = cv2.resize(image, (200,200))\n    \n    x = i // 6\n    y = i % 6\n    axes[x, y].imshow(image, cmap=plt.cm.bone) \n    axes[x, y].axis('off')\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:22:55.864836Z","iopub.execute_input":"2024-04-17T18:22:55.865757Z","iopub.status.idle":"2024-04-17T18:22:59.816239Z","shell.execute_reply.started":"2024-04-17T18:22:55.86572Z","shell.execute_reply":"2024-04-17T18:22:59.814888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Hue, Saturation, Brightness","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(nrows=2, ncols=6, figsize=(16,6))\nplt.suptitle(\"Hue, Saturation, Brightness\", fontsize=16)\n\nfor i in range(0, min(2*6, len(train_df))):\n    data = pydicom.read_file(train_df['path_dicom'].iloc[i])  # Use .iloc[i] instead of [i]\n    image = data.pixel_array\n    \n    # Transform to B&W\n    # The function converts an input image from one color space to another.\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2HLS)\n    image = cv2.resize(image, (200,200))\n    \n    x = i // 6\n    y = i % 6\n    axes[x, y].imshow(image, cmap=plt.cm.bone) \n    axes[x, y].axis('off')\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:26:15.414554Z","iopub.execute_input":"2024-04-17T18:26:15.415112Z","iopub.status.idle":"2024-04-17T18:26:19.110764Z","shell.execute_reply.started":"2024-04-17T18:26:15.415071Z","shell.execute_reply":"2024-04-17T18:26:19.10955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"LUV Color Space","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(nrows=2, ncols=6, figsize=(16,6))\nplt.suptitle(\"LUV Color Space\", fontsize=16)\n\nfor i in range(0, min(2*6, len(train_df))):\n    data = pydicom.read_file(train_df['path_dicom'].iloc[i])  # Use .iloc[i] instead of [i]\n    image = data.pixel_array\n    \n    # Transform to B&W\n    # The function converts an input image from one color space to another.\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2LUV)\n    image = cv2.resize(image, (200,200))\n    \n    x = i // 6\n    y = i % 6\n    axes[x, y].imshow(image, cmap=plt.cm.bone) \n    axes[x, y].axis('off')\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:27:35.655478Z","iopub.execute_input":"2024-04-17T18:27:35.655984Z","iopub.status.idle":"2024-04-17T18:27:39.717573Z","shell.execute_reply.started":"2024-04-17T18:27:35.655949Z","shell.execute_reply":"2024-04-17T18:27:39.716583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Torchvision.transforms","metadata":{}},{"cell_type":"code","source":"# Necessary Imports\nimport torch\nfrom torch.utils.data import DataLoader, Dataset\nimport torchvision.transforms as transforms\nimport torchvision","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:27:52.955111Z","iopub.execute_input":"2024-04-17T18:27:52.955613Z","iopub.status.idle":"2024-04-17T18:27:52.961263Z","shell.execute_reply.started":"2024-04-17T18:27:52.955572Z","shell.execute_reply":"2024-04-17T18:27:52.960312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create PyTorch Dataset Object\nclass DatasetExample(Dataset):\n    def __init__(self, image_list, transforms=None):\n        self.image_list = image_list\n        self.transforms = transforms\n    \n    # To get item's length\n    def __len__(self):\n        return (len(self.image_list))\n    \n    # For indexing\n    def __getitem__(self, i):\n        # Read in image\n        image = plt.imread(self.image_list[i])\n        image = Image.fromarray(image).convert('RGB')        \n        image = np.asarray(image).astype(np.uint8)\n        if self.transforms is not None:\n            image = self.transforms(image)\n            \n        return torch.tensor(image, dtype=torch.float)\n    \n\n    # Predefined Show Images Function\ndef show_transform(image, title=\"Default\"):\n    plt.figure(figsize=(16,6))\n    plt.suptitle(title, fontsize = 16)\n    \n    # Unnormalize\n    image = image / 2 + 0.5  \n    npimg = image.numpy()\n    npimg = np.clip(npimg, 0., 1.)\n    plt.imshow(np.transpose(npimg, (1, 2, 0)))\n    plt.show()  \n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:34:55.817088Z","iopub.execute_input":"2024-04-17T18:34:55.818047Z","iopub.status.idle":"2024-04-17T18:34:55.829311Z","shell.execute_reply.started":"2024-04-17T18:34:55.81801Z","shell.execute_reply":"2024-04-17T18:34:55.82806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Crop ✂","metadata":{}},{"cell_type":"code","source":"# Transform\ntransform = transforms.Compose([\n     transforms.ToPILImage(),\n     transforms.Resize((244, 244)),\n     transforms.CenterCrop((180, 180)),\n     transforms.ToTensor(),\n     transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5)),\n     ])\n\n# Create the dataset\npytorch_dataset = DatasetExample(image_list=image_list, transforms=transform)\npytorch_dataloader = DataLoader(dataset=pytorch_dataset, batch_size=32, shuffle=True)\n\n# Select the data\nimages = next(iter(pytorch_dataloader))\n \n# show images\nshow_transform(torchvision.utils.make_grid(images, nrow=6), title=\"Crop\")","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:35:02.443719Z","iopub.execute_input":"2024-04-17T18:35:02.444147Z","iopub.status.idle":"2024-04-17T18:35:09.486779Z","shell.execute_reply.started":"2024-04-17T18:35:02.444118Z","shell.execute_reply":"2024-04-17T18:35:09.485564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. ColorJitter 🌫","metadata":{}},{"cell_type":"code","source":"# Transform\ntransform = transforms.Compose([\n     transforms.ToPILImage(),\n     transforms.Resize((244, 244)),\n     transforms.ColorJitter(brightness=0.7, contrast=0.7, saturation=0.7, hue=0.5),\n     transforms.ToTensor(),\n     transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5)),\n     ])\n\n# Create the dataset\npytorch_dataset = DatasetExample(image_list=image_list, transforms=transform)\npytorch_dataloader = DataLoader(dataset=pytorch_dataset, batch_size=12, shuffle=True)\n\n# Select the data\nimages = next(iter(pytorch_dataloader))\n \n# show images\nshow_transform(torchvision.utils.make_grid(images, nrow=6), title=\"Color Jitter\")","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:35:20.908488Z","iopub.execute_input":"2024-04-17T18:35:20.908927Z","iopub.status.idle":"2024-04-17T18:35:26.624678Z","shell.execute_reply.started":"2024-04-17T18:35:20.908895Z","shell.execute_reply":"2024-04-17T18:35:26.623405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. RandomGreyscale","metadata":{}},{"cell_type":"code","source":"# Transform\ntransform = transforms.Compose([\n     transforms.ToPILImage(),\n     transforms.Resize((244, 244)),\n     transforms.RandomGrayscale(p=0.7),\n     transforms.ToTensor(),\n     transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5)),\n     ])\n\n# Create the dataset\npytorch_dataset = DatasetExample(image_list=image_list, transforms=transform)\npytorch_dataloader = DataLoader(dataset=pytorch_dataset, batch_size=12, shuffle=True)\n\n# Select the data\nimages = next(iter(pytorch_dataloader))\n \n# show images\nshow_transform(torchvision.utils.make_grid(images, nrow=6), title=\"Random Greyscale\")","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:35:36.497513Z","iopub.execute_input":"2024-04-17T18:35:36.497948Z","iopub.status.idle":"2024-04-17T18:35:42.398437Z","shell.execute_reply.started":"2024-04-17T18:35:36.497918Z","shell.execute_reply":"2024-04-17T18:35:42.39715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"RandomVerticalFlip","metadata":{}},{"cell_type":"code","source":"# Transform\ntransform = transforms.Compose([\n     transforms.ToPILImage(),\n     transforms.Resize((244, 244)),\n     transforms.RandomVerticalFlip(p=0.7),\n     transforms.ToTensor(),\n     transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5)),\n     ])\n\n# Create the dataset\npytorch_dataset = DatasetExample(image_list=image_list, transforms=transform)\npytorch_dataloader = DataLoader(dataset=pytorch_dataset, batch_size=12, shuffle=True)\n\n# Select the data\nimages = next(iter(pytorch_dataloader))\n \n# show images\nshow_transform(torchvision.utils.make_grid(images, nrow=6), title=\"Random Vertical Flip\")","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:35:49.307764Z","iopub.execute_input":"2024-04-17T18:35:49.308175Z","iopub.status.idle":"2024-04-17T18:35:54.156575Z","shell.execute_reply.started":"2024-04-17T18:35:49.308145Z","shell.execute_reply":"2024-04-17T18:35:54.1554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Hair Removal ✂**","metadata":{}},{"cell_type":"code","source":"def hair_remove(image):\n    # convert image to grayScale\n    grayScale = cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n\n    # kernel for morphologyEx\n    kernel = cv2.getStructuringElement(1,(17,17))\n\n    # apply MORPH_BLACKHAT to grayScale image\n    blackhat = cv2.morphologyEx(grayScale, cv2.MORPH_BLACKHAT, kernel)\n\n    # apply thresholding to blackhat\n    _,threshold = cv2.threshold(blackhat,10,255,cv2.THRESH_BINARY)\n\n    # inpaint with original image and threshold image\n    final_image = cv2.inpaint(image,threshold,1,cv2.INPAINT_TELEA)\n\n    return final_image\n\n# Select a small sample of the .jpeg image paths\n# We select some hairy photos on purpose\nhairy_photos = train_data[train_data[\"sex\"] == 1].reset_index().iloc[[12, 14, 17, 22, 33, 34]]\nimage_list = hairy_photos['path_jpeg']\nimage_list = image_list.reset_index()['path_jpeg']\n\n\n# Show the Augmented Images\nplt.figure(figsize=(16,3))\nplt.suptitle(\"Original Hairy Images\", fontsize = 16)\n    \nfor k, path in enumerate(image_list):\n    image = mpimg.imread(path)\n    image = cv2.resize(image,(244, 244))\n        \n    plt.subplot(1, 6, k+1)\n    plt.imshow(image)\n    plt.axis('off')","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:36:02.125534Z","iopub.execute_input":"2024-04-17T18:36:02.125941Z","iopub.status.idle":"2024-04-17T18:36:03.867753Z","shell.execute_reply.started":"2024-04-17T18:36:02.125911Z","shell.execute_reply":"2024-04-17T18:36:03.866613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show the sample\nplt.figure(figsize=(16,3))\nplt.suptitle(\"Non Hairy Images\", fontsize = 16)\n    \nfor k, path in enumerate(image_list):\n    image = mpimg.imread(path)\n    image = cv2.resize(image,(244, 244))\n    image = hair_remove(image)\n        \n    plt.subplot(1, 6, k+1)\n    plt.imshow(image)\n    plt.axis('off')","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:36:12.027401Z","iopub.execute_input":"2024-04-17T18:36:12.027828Z","iopub.status.idle":"2024-04-17T18:36:14.633756Z","shell.execute_reply.started":"2024-04-17T18:36:12.027798Z","shell.execute_reply":"2024-04-17T18:36:14.632436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cropped_width = images.shape[2]\ncropped_height = images.shape[3]\n\nprint(\"Width of cropped images:\", cropped_width)\nprint(\"Height of cropped images:\", cropped_height)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:36:20.775224Z","iopub.execute_input":"2024-04-17T18:36:20.775684Z","iopub.status.idle":"2024-04-17T18:36:20.782693Z","shell.execute_reply.started":"2024-04-17T18:36:20.775637Z","shell.execute_reply":"2024-04-17T18:36:20.781274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*******************************************************************************************************\n*******************************************************************************************************","metadata":{}},{"cell_type":"code","source":"import pathlib\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten, Conv2D, MaxPooling2D,Input\nfrom tensorflow.keras.layers import Dense","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:38:00.005259Z","iopub.execute_input":"2024-04-17T18:38:00.005733Z","iopub.status.idle":"2024-04-17T18:38:00.01269Z","shell.execute_reply.started":"2024-04-17T18:38:00.005699Z","shell.execute_reply":"2024-04-17T18:38:00.011327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Rescaling\nfrom tensorflow.keras import Sequential\n\ninput_shape = (244, 244, 3)  # Define the input shape for 180x180 images\n\nmodel = Sequential()\n\n# Adding Rescaling layer for input preprocessing\nmodel.add(Rescaling(1./255, input_shape=input_shape))\n\n# CNN Layers\nmodel.add(Conv2D(32,(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\nmodel.add(Conv2D(32,(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\nmodel.add(Conv2D(16,(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\nmodel.add(Conv2D(16,(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\nmodel.add(Conv2D(16,(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\nmodel.add(Conv2D(16,(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\n\n# Flatten layer\nmodel.add(Flatten())\n\n# Dense Layers\nmodel.add(Dense(50, activation='relu'))\nmodel.add(Dense(100, activation='relu'))\nmodel.add(Dense(200, activation='relu'))\n\n# Output layer\nmodel.add(Dense(2, activation='softmax'))\n\n# Compile the model, add loss function and optimizer\n# model.compile(...)\n\n# Optionally, print the model summary\n# model.summary()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:42:52.796733Z","iopub.execute_input":"2024-04-17T18:42:52.797193Z","iopub.status.idle":"2024-04-17T18:42:53.101924Z","shell.execute_reply.started":"2024-04-17T18:42:52.797162Z","shell.execute_reply":"2024-04-17T18:42:53.100833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:44:04.738665Z","iopub.execute_input":"2024-04-17T18:44:04.73915Z","iopub.status.idle":"2024-04-17T18:44:04.784786Z","shell.execute_reply.started":"2024-04-17T18:44:04.739117Z","shell.execute_reply":"2024-04-17T18:44:04.783679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(loss=\"sparse_categorical_crossentropy\", \n              optimizer=\"adam\",\n              metrics=[\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:56:05.413168Z","iopub.execute_input":"2024-04-17T18:56:05.413637Z","iopub.status.idle":"2024-04-17T18:56:05.423298Z","shell.execute_reply.started":"2024-04-17T18:56:05.413604Z","shell.execute_reply":"2024-04-17T18:56:05.422026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(\n    train_df,\n    epochs=10,\n    validation_data=val_df\n)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:56:10.669529Z","iopub.execute_input":"2024-04-17T18:56:10.669937Z","iopub.status.idle":"2024-04-17T18:56:10.791905Z","shell.execute_reply.started":"2024-04-17T18:56:10.669907Z","shell.execute_reply":"2024-04-17T18:56:10.7905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the number of steps per epoch\nsamples = 33126\nbatch_size = 32\nsteps_per_epoch = samples // batch_size\n\nprint(\"Steps per epoch:\", steps_per_epoch)\n\n# Assuming train_df and val_df are your training and validation data generators respectively\nhistory = model.fit(\n    train_df,\n    steps_per_epoch=steps_per_epoch,\n    epochs=10,\n    validation_data=val_df\n)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-17T18:53:20.6447Z","iopub.execute_input":"2024-04-17T18:53:20.645111Z","iopub.status.idle":"2024-04-17T18:53:20.853792Z","shell.execute_reply.started":"2024-04-17T18:53:20.645081Z","shell.execute_reply":"2024-04-17T18:53:20.852148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*******************************************************************************************************\n\n*******************************************************************************************************","metadata":{}},{"cell_type":"markdown","source":"*********************************************************************************************************************************************************\n\n*********************************************************************************************************************************************************","metadata":{}},{"cell_type":"markdown","source":"**CROPPED IMAGE RESIZING**","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Display the resized image using Matplotlib\nplt.imshow(resized_image)\nplt.axis('off')  # Turn off axis labels\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-16T06:38:37.141926Z","iopub.execute_input":"2024-04-16T06:38:37.142466Z","iopub.status.idle":"2024-04-16T06:38:37.372831Z","shell.execute_reply.started":"2024-04-16T06:38:37.142423Z","shell.execute_reply":"2024-04-16T06:38:37.371103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*******************************************************************************************************\n\n*******************************************************************************************************","metadata":{}}]}