{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":20270,"databundleVersionId":1222630,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"TEAM13:\n                                                                 SKIN CANCER DETECTION(MELANOMA)\n# Problem Statement:\nSkin cancer is the most prevalent type of cancer. Melanoma, specifically, is responsible for 75% of skin cancer deaths, despite being the least common skin cancer, but if caught early, most melanomas can be cured with minor surgery. In this competition, the task is to use images of skin lesions within the same patient and determine which are likely to represent a melanoma.The goal of this project is to design, develop, and evaluate a robust machine learning algorithm capable of accurately detecting melanoma from dermatological images. The algorithm should demonstrate high sensitivity and specificity, ensuring reliable detection while minimizing false positives and false negatives. Ultimately, this project aims to contribute to improving early detection rates, facilitating timely interventions, and ultimately saving lives in the fight against melanoma.","metadata":{}},{"cell_type":"markdown","source":"# Objective\nThe objective of this competition is to identify melanoma in images of skin lesions. In particular, we need to use images within the same patient and determine which are likely to represent a melanoma. In other words, we need to create a model which should predict the probability whether the lesion in the image is malignantor benign.Value 0 denotes benign, and 1 indicates malignant.Acquire and curate a diverse dataset of dermatological images encompassing a wide range of skin types, lesion sizes, and melanoma stages to ensure the robustness and generalizability of the developed model.Implement rigorous validation methodologies, including cross-validation and independent testing, to evaluate the performance of the developed model in terms of sensitivity, specificity, accuracy, and area under the receiver operating characteristic (ROC) curve.","metadata":{}},{"cell_type":"markdown","source":"# Importing the necessary libraries","metadata":{}},{"cell_type":"markdown","source":"#Importing necessary libraries in Python allows you to access pre-built functions and tools for various tasks. Common libraries like NumPy for numerical operations, Pandas for data manipulation, Matplotlib for data visualization, and Scikit-learn for machine learning are often imported. This can be achieved using 'import' statements followed by the library name.","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file \nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reading Dataset\nReading dataset involves accessing and extracting information stored in a structured format, this process often includes loading the dataset into a programming environment or tool, such as Python with libraries like Pandas.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ntest_df = pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analyzing the data","metadata":{}},{"cell_type":"markdown","source":"Analyzing data involves systematically examining, interpreting, and deriving insights from raw information to uncover patterns, trends, correlations, and meaningful relationships within the data. This process typically follows a structured approach to ensure accuracy and reliability in the findings.","metadata":{}},{"cell_type":"markdown","source":"\n\"Display the shape\" typically refers to showing or presenting a visual representation of a geometric figure or form.","metadata":{}},{"cell_type":"code","source":"train_df.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\"Display the head\" typically refers to displaying the first few lines of a file or command output in a terminal or command prompt","metadata":{}},{"cell_type":"code","source":"train_df.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\"Displaying the tail\" typically refers to showing the last few lines of a text file, usually in a command-line environment. ","metadata":{}},{"cell_type":"code","source":"train_df.tail(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Displaying information about a dataset typically involves providing details such as the dataset's structure, dimensions, types of variables, summary statistics, and possibly some sample data points. ","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are some missing values in some of the columns.","metadata":{}},{"cell_type":"code","source":"duplicates = train_df.duplicated().sum()\nprint(\"\\nNumber of duplicate rows:\", duplicates)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preprocessing\n\n Data preprocessing is a crucial step in the data analysis and machine learning pipeline. It   involves cleaning, transforming, and organizing raw data into a format suitable for analysis or model training.","metadata":{}},{"cell_type":"markdown","source":"# Finding missing value\n\nDetecting missing values in a dataset intended for skin cancer detection, specifically for melanoma, involves crucial steps to ensure the integrity of the data used for analysis.","metadata":{}},{"cell_type":"code","source":"missing_values = train_df.isnull().sum()\nprint(missing_values)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Impute missing values for numeric columns with mean\nnumeric_columns = train_df.select_dtypes(include=['float64', 'int64']).columns\ntrain_df[numeric_columns] = train_df[numeric_columns].fillna(train_df[numeric_columns].mean())\n\n# Impute missing values for categorical columns with mode\ncategorical_columns = train_df.select_dtypes(include=['object']).columns\ntrain_df[categorical_columns] = train_df[categorical_columns].fillna(train_df[categorical_columns].mode().iloc[0])\n\n# Verify if there are any missing values left\nmissing_values_after_imputation = train_df.isnull().sum()\nprint(missing_values_after_imputation)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA Univariate\nExploratory Data Analysis (EDA) involves examining and summarizing individual variables in a dataset to gain insights into their distributions, patterns, and characteristics. When conducting EDA for skin cancer detection, particularly for melanoma, univariate analysis focuses on understanding each variable in isolation.","metadata":{}},{"cell_type":"markdown","source":"#Plot for Distribution of Gender\n\n","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(6, 6))\nsns.countplot(x='sex', data=train_df)\nplt.title('Distribution of Gender')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Plot for Distribution of Age","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.histplot(train_df['age_approx'], bins=30, kde=True)\nplt.title('Distribution of Age')\nplt.xlabel('Age')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA Multivariate\n\nExploratory Data Analysis (EDA) for multivariate skin cancer detection, particularly for melanoma, involves analyzing various factors and features to identify patterns, correlations, and potential indicators of the disease. ","metadata":{}},{"cell_type":"markdown","source":"**Plot for Beningn vs Malignant**","metadata":{}},{"cell_type":"code","source":"sns.catplot(x='target',y='benign_malignant', hue='sex',data=z,kind='bar')\nplt.ylabel('Count')\nplt.xlabel('benign:0 vs malignant:1')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Distribution of Ages w.r.t Target**","metadata":{}},{"cell_type":"code","source":"#plot of age that were diagnosed as benign\nsns.kdeplot(train_df.loc[train_df['target'] == 0, 'age_approx'], label = 'Benign',shade=True)\n\n#plot of age that were diagnosed as malignant\nsns.kdeplot(train_df.loc[train_df['target'] == 1, 'age_approx'], label = 'Malignant',shade=True)\n\n# Labeling of plot\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Image Data Analysis\n\nImage data analysis for skin cancer detection, particularly melanoma, involves utilizing various computational techniques to extract meaningful information from medical images such as dermoscopy images or histopathological slides. ","metadata":{}},{"cell_type":"code","source":"# Defining data path\nIMAGE_PATH = \"../input/siim-isic-melanoma-classification/\"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images = train_df['image_name'].values\n\n# Extract 9 random images from it\nrandom_images = [np.random.choice(images+'.jpg') for i in range(9)]\n\n# Location of the image dir\nimg_dir = IMAGE_PATH+'/jpeg/train'\n\nprint('Display Random Images')\n\nplt.figure(figsize=(10,8))\n\nfor i in range(9):\n    plt.subplot(3, 3, i + 1)\n    img = plt.imread(os.path.join(img_dir, random_images[i]))\n    plt.imshow(img, cmap='gray')\n    plt.axis('off')\n    \nplt.tight_layout()   ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"benign = train_df[train_df['benign_malignant']=='benign']\nmalignant = train_df[train_df['benign_malignant']=='malignant']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing Images with benign","metadata":{}},{"cell_type":"code","source":"images = benign['image_name'].values\n\nrandom_images = [np.random.choice(images+'.jpg') for i in range(9)]\n\n# Location of the image dir\nimg_dir = IMAGE_PATH+'/jpeg/train'\n\nprint('Benign Images')\n\n# Adjust the size of your images\nplt.figure(figsize=(10,8))\n\nfor i in range(9):\n    plt.subplot(3, 3, i + 1)\n    img = plt.imread(os.path.join(img_dir, random_images[i]))\n    plt.imshow(img, cmap='gray')\n    plt.axis('off')\n    \nplt.tight_layout()   ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing Images with Malignant","metadata":{}},{"cell_type":"code","source":"images = malignant['image_name'].values\n\n# Extract 9 random images from it\nrandom_images = [np.random.choice(images+'.jpg') for i in range(9)]\n\n# Location of the image dir\nimg_dir = IMAGE_PATH+'/jpeg/train'\n\nprint('Malignant Images')\n\n# Adjust the size of your images\nplt.figure(figsize=(10,8))\n\n# Iterate and plot random images\nfor i in range(9):\n    plt.subplot(3, 3, i + 1)\n    img = plt.imread(os.path.join(img_dir, random_images[i]))\n    plt.imshow(img, cmap='gray')\n    plt.axis('off')\n    \n# Adjust subplot parameters to give specified padding\nplt.tight_layout() ","metadata":{"execution":{"iopub.status.busy":"2024-04-02T13:59:56.834004Z","iopub.execute_input":"2024-04-02T13:59:56.834412Z","iopub.status.idle":"2024-04-02T13:59:56.871637Z","shell.execute_reply.started":"2024-04-02T13:59:56.834368Z","shell.execute_reply":"2024-04-02T13:59:56.87016Z"},"trusted":true},"execution_count":null,"outputs":[]}]}