{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <mark>UBC SenNet + HOA - Hacking the Human Vasculature in 3D - EDA</mark>\n<span style=\"font-size:22px;color:purple\"> Thank you for having a look at my notebook - advice and feedback always welcomed!</span>\n\n\n<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌 Dataset Link: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/data\">https://www.kaggle.com/competitions/blood-vessel-segmentation/data</a>\n</div>\n\n\n## **Exploratory Data Analysis (EDA)** \nEDA is a crucial step in understanding and preparing your data for any data analysis or machine learning project, including UBC Ovarian Cancer Subtype Classification and Outlier Detection (UBC-OCEAN). Here's a step-by-step guide on how to perform an EDA for this dataset:\n\n**Overview**\n\nThe goal of the UBC Ovarian Cancer subtypE clAssification and outlier detectioN (UBC-OCEAN) competition is to classify ovarian cancer subtypes. You will build a model trained on the world's most extensive ovarian cancer dataset of histopathology images obtained from more than 20 medical centers.\n\n\n**Data Collection:**\n\nBegin by obtaining the UBC-OCEAN dataset, which should include information on ovarian cancer subtypes and possibly outlier detection data. Ensure you have a clear understanding of the dataset's structure and the meaning of each variable.\nData Loading:\n\nImport the dataset into your preferred data analysis environment, such as Python with libraries like pandas, numpy, and matplotlib/seaborn for visualization.\n\n**Data Loading:**\n\nImport the dataset into your preferred data analysis environment, such as Python with libraries like pandas, numpy, and matplotlib/seaborn for visualization.\n\n\n\n\n\n### **Initial Exploration:**\n\n**1 - Start by examining the basic characteristics of the data:**\n\n    Check the first few rows using df.head().\n    Check the data types and missing values using df.info().\n    Calculate basic statistics using df.describe().\n    Data Cleaning:\n\n**2 - Handle missing values, outliers, and duplicates:**\n\n    Use techniques like imputation for missing values.\n    Identify and deal with outliers appropriately.\n    Remove duplicate rows if necessary.\n    \n**3 - Data Visualization:**\n\n    Create visualizations to gain insights into the data:\n    Histograms and box plots for numerical features.\n    Bar plots for categorical features.\n    Correlation matrix and scatter plots to understand relationships between variables.\n    \n**4 - Feature Analysis:**\n\n    Explore relationships between features and the target variable(s) for classification and outlier detection.\n    Visualize how different features vary across different subtypes or classes.\n    Use box plots, violin plots, or swarm plots to compare feature distributions.\n\n**5 - Outlier Detection:**\n\n    If your dataset contains information related to outlier detection, perform a dedicated EDA for this aspect:\n    Visualize outliers using scatter plots or box plots.\n    Apply statistical methods or machine learning techniques to identify outliers.\n\n**6 - Dimensionality Reduction (optional):**\n\n    If the dataset has many features, consider dimensionality reduction techniques like Principal Component Analysis (PCA) to reduce the number of variables while preserving important information.\n\n**8 - Summary and Insights:**\n\n    Summarize your findings from the EDA, including any patterns, trends, or anomalies observed.\n    Document any data preprocessing steps applied.\n\n**7 - Next Steps:**\n\n    Based on your EDA findings, plan your next steps, which may include feature engineering, model selection, and further data preprocessing.\n\n\nRemember that EDA is an iterative process, and you may need to revisit these steps as you delve deeper into the dataset and develop your machine learning or data analysis models.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom skimage import io\nimport os\nimport seaborn as sns\nimport cv2\nimport random\nimport os\nimport glob\nimport imageio\nimport gc\nimport math\nimport copy\nimport time\nimport numpy as np\nfrom matplotlib.image import imread \nimport matplotlib.pyplot as plt\nfrom matplotlib import animation, rc\nrc('animation', html='jshtml')\n\n# Pytorch Imports\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nimport torch.nn.functional as F\nfrom torch.optim import lr_scheduler\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.cuda import amp\nimport torchvision\n\n# Utils\nimport joblib\nfrom tqdm import tqdm\nfrom collections import defaultdict\n\n# Sklearn Imports\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import StratifiedKFold\n\n# For Image Models\nimport timm\n\n# Albumentations for augmentations\nimport albumentations as A\nfrom albumentations.pytorch import ToTensorV2\n\n# For colored terminal text\nfrom colorama import Fore, Back, Style\nb_ = Fore.BLUE\nsr_ = Style.RESET_ALL\n\n#import warnings\n#warnings.filterwarnings(\"ignore\")\n\n# For descriptive error messages\nos.environ['CUDA_LAUNCH_BLOCKING'] = \"1\"","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:35:57.115086Z","iopub.execute_input":"2023-11-10T16:35:57.115477Z","iopub.status.idle":"2023-11-10T16:35:57.127017Z","shell.execute_reply.started":"2023-11-10T16:35:57.115447Z","shell.execute_reply":"2023-11-10T16:35:57.125669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"directory_path_train = '/kaggle/input/blood-vessel-segmentation/train'\ndef list_files_in_directory(directory_path = directory_path_train):\n    direc = []\n    for root, dirs, files in os.walk(directory_path):\n        for dire in dirs:\n            if dire in [\"labels\",\"images\"]:\n                continue \n            file_path = os.path.join(root, dire)\n            direc.append(file_path)\n            print(file_path)\n    return direc\n            \n\ntrain_folders = list_files_in_directory(directory_path_train)","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:35:57.429933Z","iopub.execute_input":"2023-11-10T16:35:57.430370Z","iopub.status.idle":"2023-11-10T16:35:59.718927Z","shell.execute_reply.started":"2023-11-10T16:35:57.430337Z","shell.execute_reply":"2023-11-10T16:35:59.717715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"def count_total_img(folders = train_folders):\n    sub_f = [\"images\",\"labels\"] \n    path = []\n    total_files = []\n    random_dir = random.choice(folders)\n    for dire in folders:\n        for subf in sub_f:\n            if (dire == \"/kaggle/input/blood-vessel-segmentation/train/kidney_3_dense\") & (subf == \"images\"):\n                continue \n            \n            _dir = dire + \"/\" + subf\n            total_sample = len(os.listdir(_dir))\n            print(f\"{_dir}: {total_sample}\")\n            path.append(_dir)\n            total_files.append(total_sample)\n    obj = {\n        \"path\": path,\n        \"total_files\":total_files\n    }\n    return obj\n            \ntrain_file_dir = count_total_img()","metadata":{}},{"cell_type":"code","source":"def count_total_img(folders = train_folders):\n    sub_f = [\"images\",\"labels\"] \n    path = []\n    total_files = []\n    random_dir = random.choice(folders)\n    for dire in folders:\n        for subf in sub_f:\n            if (dire == \"/kaggle/input/blood-vessel-segmentation/train/kidney_3_dense\") & (subf == \"images\"):\n                continue \n            \n            _dir = dire + \"/\" + subf\n            total_sample = len(os.listdir(_dir))\n            print(f\"{_dir}: {total_sample}\")\n            path.append(_dir)\n            total_files.append(total_sample)\n    obj = {\n        \"path\": path,\n        \"total_files\":total_files\n    }\n    return obj\n            \ntrain_file_dir = count_total_img()","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:35:59.721132Z","iopub.execute_input":"2023-11-10T16:35:59.721611Z","iopub.status.idle":"2023-11-10T16:35:59.741759Z","shell.execute_reply.started":"2023-11-10T16:35:59.721571Z","shell.execute_reply":"2023-11-10T16:35:59.740632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_random_img(folders = train_file_dir):\n    \"\"\"\n    will display random slice image from data and its corresponding mask\n    \"\"\"\n    _paths = list(zip(folders['path'], folders['total_files']))\n    _path_tup = random.choice(_paths)\n  \n    split_text = _path_tup[0].split(\"/\")\n\n    img_no = random.choice(range(_path_tup[1]))\n    img_no = f\"{img_no:04}\"\n    random_img_no = str(img_no)\n    _IMG_PATH =  _path_tup[0] + '/' + random_img_no +\".tif\" \n    IMG_PATH = _IMG_PATH.replace(\"labels\",\"images\")\n    LABEL_PATH = IMG_PATH.replace(\"images\",\"labels\")\n    \n    \n    if \"kidney_3_dense\" in split_text:\n        IMG_PATH = IMG_PATH\n        LABEL_PATH = IMG_PATH.replace(\"kidney_3_dense\",\"kidney_3_sparse\").replace(\"labels\",\"images\")\n        \n    \n    try:\n        print(IMG_PATH)\n        print(LABEL_PATH)\n        _slice = imread(IMG_PATH)\n        _mask = imread(LABEL_PATH)\n        \n        plt.figure(figsize=(10, 5))\n        plt.subplot(1, 2, 1)\n        plt.imshow(_slice)\n        plt.title(f'3D image slice: {img_no}')\n\n        plt.subplot(1, 2, 2)\n        plt.imshow(_mask)\n        plt.title(f'Mask: {img_no}')\n\n        plt.show()\n    except Exception as e:\n        \n        print(f\"An error occurred:{e}\")\n        ","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:35:59.743360Z","iopub.execute_input":"2023-11-10T16:35:59.744038Z","iopub.status.idle":"2023-11-10T16:35:59.754584Z","shell.execute_reply.started":"2023-11-10T16:35:59.743964Z","shell.execute_reply":"2023-11-10T16:35:59.753480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Visualization 📈 \n\n### Randomly Visualize Images","metadata":{}},{"cell_type":"code","source":"def create_animation(ims):\n    fig=plt.figure(figsize=(15,6))\n    plt.axis('off')\n    im=plt.imshow(ims[0])\n    #im=plt.imshow(cv2.cvtColor(ims[0],cv2.COLOR_BGR2RGB))\n    plt.close()\n    def animate_func(i):\n        im.set_array(ims[i])\n        #im.set_array(cv2.cvtColor(ims[i],cv2.COLOR_BGR2RGB))\n        return [im]\n    return animation.FuncAnimation(fig, animate_func, frames=len(ims), interval=1000/20)   ","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:35:59.756687Z","iopub.execute_input":"2023-11-10T16:35:59.757485Z","iopub.status.idle":"2023-11-10T16:35:59.767549Z","shell.execute_reply.started":"2023-11-10T16:35:59.757452Z","shell.execute_reply":"2023-11-10T16:35:59.766257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **blood-vessel-segmentation/train/kidney_1_dense**","metadata":{}},{"cell_type":"code","source":"ipaths=[]\nlpaths=[]\nfor dirname, _, filenames in os.walk('/kaggle/input/blood-vessel-segmentation/train/kidney_1_dense'):\n    for filename in filenames:\n        if dirname.split('/')[-1]=='images':\n            ipaths+=[(os.path.join(dirname, filename))]\n        elif dirname.split('/')[-1]=='labels':\n            lpaths+=[(os.path.join(dirname, filename))]\nipaths.sort()\nlpaths.sort()\n\nimages=[]\nfor i in range(len(ipaths)):\n    img1 = cv2.imread(ipaths[i])\n    #img1 = cv2.imread(lpaths[i], cv2.IMREAD_UNCHANGED)\n    img1=np.rot90(img1)\n    img1=cv2.resize(img1,dsize=None,fx=0.2,fy=0.2)\n    images+=[img1]\n    \ncreate_animation(images)","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:36:00.024709Z","iopub.execute_input":"2023-11-10T16:36:00.025779Z","iopub.status.idle":"2023-11-10T16:40:01.257650Z","shell.execute_reply.started":"2023-11-10T16:36:00.025738Z","shell.execute_reply":"2023-11-10T16:40:01.254702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **blood-vessel-segmentation/train/kidney_1_voi**","metadata":{}},{"cell_type":"code","source":"ipaths=[]\nlpaths=[]\nfor dirname, _, filenames in os.walk('/kaggle/input/blood-vessel-segmentation/train/kidney_1_voi'):\n    for filename in filenames:\n        if dirname.split('/')[-1]=='images':\n            ipaths+=[(os.path.join(dirname, filename))]\n        elif dirname.split('/')[-1]=='labels':\n            lpaths+=[(os.path.join(dirname, filename))]\nipaths.sort()\nlpaths.sort()\n\nimages=[]\nfor i in range(len(ipaths)):\n    img1 = cv2.imread(ipaths[i])\n    #img1 = cv2.imread(lpaths[i], cv2.IMREAD_UNCHANGED)\n    img1=np.rot90(img1)\n    img1=cv2.resize(img1,dsize=None,fx=0.2,fy=0.2)\n    images+=[img1]\n    \ncreate_animation(images)","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:40:31.816406Z","iopub.execute_input":"2023-11-10T16:40:31.816823Z","iopub.status.idle":"2023-11-10T16:40:43.641323Z","shell.execute_reply.started":"2023-11-10T16:40:31.816790Z","shell.execute_reply":"2023-11-10T16:40:43.639774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(5):\n    display_random_img(train_file_dir)","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:40:01.259836Z","iopub.execute_input":"2023-11-10T16:40:01.260216Z","iopub.status.idle":"2023-11-10T16:40:04.926163Z","shell.execute_reply.started":"2023-11-10T16:40:01.260185Z","shell.execute_reply":"2023-11-10T16:40:04.925150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Model Building**\n","metadata":{}},{"cell_type":"code","source":"CONFIG = {\n    \"seed\": 42,\n    \"img_size\": 512,\n    \"model_name\": \"tf_efficientnet_b0_ns\",\n    \"num_classes\": 5,\n    \"valid_batch_size\": 64,\n    \"device\": torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\"),\n}","metadata":{"execution":{"iopub.status.busy":"2023-11-10T16:35:36.453637Z","iopub.status.idle":"2023-11-10T16:35:36.454235Z","shell.execute_reply.started":"2023-11-10T16:35:36.453918Z","shell.execute_reply":"2023-11-10T16:35:36.453952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_seed(seed=42):\n    '''Sets the seed of the entire notebook so results are the same every time we run.\n    This is for REPRODUCIBILITY.'''\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    # When running on the CuDNN backend, two further options must be set\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n    # Set a fixed value for the hash seed\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    \nset_seed(CONFIG['seed'])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Training Configuration**  ⚙️\n\nA training configuration is a vital component of any machine learning or deep learning project, serving as the blueprint that outlines the parameters, settings, and conditions under which a model is trained. This essential document encapsulates the entire training process, providing clarity and reproducibility to the development and deployment of AI systems. Here, we delve into the key elements and considerations that constitute a comprehensive training configuration:\n\n**Model Architecture:** The configuration specifies the architecture of the neural network or machine learning model being trained. It outlines the layers, nodes, and connections that define the model's structure. This includes details like the type of layers (e.g., convolutional, recurrent), activation functions, and any custom layers or modifications.\n\n**Data Preparation:** It outlines the methods and procedures for data preprocessing, augmentation, and normalization. This may include data scaling, one-hot encoding, or image augmentation techniques. Proper data preparation is critical for model convergence and performance.\n\n**Hyperparameters:** The training configuration specifies hyperparameters, which are settings that control the learning process. This includes parameters like learning rate, batch size, epochs, and optimization algorithms (e.g., Adam, SGD). Tinkering with these hyperparameters can significantly impact model training outcomes.\n\n**Loss Function:** The choice of the loss function is pivotal to training. This component of the configuration details the specific loss metric that the model optimizes during training, aligning it with the objectives of the project (e.g., mean squared error for regression, cross-entropy for classification).\n\n**Metrics for Evaluation:** The configuration lists the evaluation metrics used to assess the model's performance during and after training. Common metrics include accuracy, F1-score, mean absolute error (MAE), and mean squared error (MSE).\n\n**Regularization Techniques:** If applicable, regularization techniques such as dropout, L1, or L2 regularization are specified in the configuration to prevent overfitting.\n\n**Checkpointing:** Configuration may include settings for model checkpointing, which periodically saves the model's weights and progress during training. This is essential for resuming training or selecting the best model for deployment.\n\n**Early Stopping:** Parameters for early stopping, based on validation metrics, are often included. This helps prevent overtraining by halting training when the model's performance on validation data plateaus or deteriorates.\n\n**Hardware and Environment:** It mentions the hardware resources utilized during training, including CPU, GPU, or TPUs, as well as the software environment, such as the version of deep learning frameworks (e.g., TensorFlow, PyTorch) and the operating system.\n\n**Batch Processing:** Configuration can also include information on distributed training, specifying whether training is done in a single batch or in mini-batches, and whether it spans multiple GPUs or nodes.\n\n**Data Splits:** The division of data into training, validation, and test sets is outlined in the configuration to ensure proper model assessment and generalization.\n\n**Documentation:** A well-documented training configuration is crucial for reproducibility. It should contain comments and explanations for each parameter and decision made, enabling easy sharing and replication of the training process.\n\nIn summary, a training configuration is the comprehensive roadmap that guides the development of machine learning and deep learning models. It plays a pivotal role in achieving reproducibility, scalability, and the successful deployment of AI systems, ensuring that the model can be trained consistently and effectively for various applications.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Create Model**","metadata":{}},{"cell_type":"markdown","source":"<!DOCTYPE html>\n<html>\n<head>\n</head>\n<body>\n    <h1>EfficientNet</h1>\n    <p>\n        EfficientNet is a family of convolutional neural network (CNN) architectures designed for efficient and effective deep learning in computer vision tasks. It was introduced in 2019 by researchers at Google AI. EfficientNet models are known for their exceptional performance in image classification tasks while being computationally efficient, making them suitable for a wide range of applications.\n    </p>\n    <p>\n        Key features of EfficientNet include:\n    </p>\n    <ul>\n        <li>Compound Scaling: EfficientNet employs a novel compound scaling method that balances the network's depth, width, and resolution to achieve optimal performance without a significant increase in computational cost.</li>\n        <li>Efficient Building Blocks: The architecture incorporates efficient building blocks like depthwise separable convolutions and squeeze-and-excitation blocks to reduce the number of parameters and computational overhead.</li>\n        <li>Variants: EfficientNet comes in various variants (e.g., EfficientNet-B0, B1, B2, ..., B7) that offer different trade-offs between model size and accuracy, allowing users to choose the one that suits their specific requirements.</li>\n        <li>State-of-the-Art Performance: EfficientNet models have achieved top performance in benchmark datasets such as ImageNet, outperforming many previous CNN architectures with smaller model sizes.</li>\n    </ul>\n    <p>\n        EfficientNet has become a popular choice in the field of computer vision due to its ability to achieve impressive results with fewer parameters, making it practical for deployment on resource-constrained devices and applications.\n    </p>\n    <p>\n        If you are working on image classification or related tasks, considering EfficientNet as part of your deep learning architecture can lead to efficient and accurate results.\n    </p>\n</body>\n</html>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dears,**\n\n**I hope this message finds you well. I am excited to share that I am participating in the UBC SenNet + HOA - Hacking the Human Vasculature in 3D competition, and I need your support!**\n\n**As part of this competition, I have conducted an in-depth Exploratory Data Analysis (EDA) to gain crucial insights into the dataset, which is a fundamental step in developing effective solutions for cancer subtype classification and outlier detection. Now, I am reaching out to request your vote and support for my EDA submission.**\n\n**Your vote can make a significant difference in this competition and help me advance to the next stages. Here's how you can support me:**\n\n\n\n**1. Cast Your Vote:**\n\nVisit the competition platform and find my EDA submission.\nClick on the \"Vote\" or \"Support\" button to cast your vote.\n\n**2. Share with Your Network:**\n\nSpread the word among your friends, family, and colleagues who may be interested in supporting my work.\n\n**3. Provide Feedback:**\n\nIf you have any feedback or suggestions on my EDA, please feel free to share them with me. Your input is valuable and can help me improve.\nI am committed to making a positive impact in the field of cancer research, and your support will bring me one step closer to achieving that goal.\n\nThank you for taking the time to read this message, and I genuinely appreciate your support in this competition. Together, we can contribute to the fight against ovarian cancer and advance the field of data-driven healthcare.\n\nIf you have any questions or need more information about my EDA, please don't hesitate to reach out to me. Your support means the world to me!\n\nWarm regards,\nJeferson S. Pazze","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}