{"metadata":{"kernelspec":{"display_name":"Python 3 (ipykernel)","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.8.13"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Week 3: CNN Cancer Detection Kaggle Mini-Project","metadata":{}},{"cell_type":"markdown","source":"### For this week's mini-project, you will participate in this Kaggle competition: \n\n#### Histopathologic Cancer Detection\n\nThis Kaggle competition is a binary image classification problem where you will identify metastatic cancer in small image patches taken from larger digital pathology scans.\n\nThe project has 125 total points. The instructions summarize the criteria you will use to guide your submission and review others' submissions. Note: to receive total points for this section, the learner doesn't need to have a top-performing score on the challenge. As a mini-project to complete as a weekly assignment, we don't expect you to iterate over your project until you have a model capable of winning the challenge. The iterative process takes time, so please start early to get better-quality results and reports. The learner needs to show a score that reasonably reflects that they completed the rubric parts of this project. The grades are more based on the quality and depth of the analysis, not just on a better Kaggle score.\n\n##### Deliverable 1 — \nA Jupyter notebook with a description of the problem/data, exploratory data analysis (EDA) procedure, analysis (model building and training), result, and discussion/conclusion. \n\nSuppose your work becomes so large that it doesn’t fit into one notebook (or you think it will be less readable by having one large notebook). In that case, you can make several notebooks or scripts in a GitHub repository (as deliverable 3) and submit a report-style notebook or pdf instead. \n\nIf your project doesn’t fit into Jupyter notebook format (E.g., you built an app that uses ML), write your approach as a report and submit it in a pdf form. \n\n##### Deliverable 2 — \nA public project GitHub repository with your work (please also include the GitHub repo URL in your notebook/report).\n\n##### Deliverable 3 — \nA screenshot of your position on the Kaggle competition leaderboard for your top-performing model.\n\n","metadata":{}},{"cell_type":"markdown","source":"### import libraries","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.figure_factory as ff\nimport plotly.graph_objects as go\nfrom IPython.display import display\nfrom PIL import Image\nfrom sklearn.metrics import classification_report, confusion_matrix\nfrom sklearn.metrics import f1_score, roc_curve, auc\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.utils.class_weight import compute_class_weight\nimport tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Dropout, BatchNormalization\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau\nfrom tensorflow.keras.applications import EfficientNetB0\nfrom tensorflow.keras.regularizers import l2\nfrom kerastuner.tuners import BayesianOptimization\nfrom pandarallel import pandarallel\n\nfrom tensorflow.keras.layers import LeakyReLU\nfrom tensorflow.keras.regularizers import l2\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pandarallel.initialize()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### path definitions","metadata":{}},{"cell_type":"code","source":"def initialize_paths():\n    \n    \"\"\"\n    this function gets all paths to use in this study.\n    \"\"\"\n\n    base_path = r\"/Users/flaviab/Downloads/histopathologic-cancer-detection/train\"\n    csv_path = r\"/Users/flaviab/Downloads/histopathologic-cancer-detection/train_labels.csv\"\n    base_path_test = r\"/Users/flaviab/Downloads/histopathologic-cancer-detection/test\"\n    path = r\"/Users/flaviab/Downloads/histopathologic-cancer-detection/train/0a0ce1220f56a48bf14615f80bc4c684244c909d.tif\"\n    output_directory = r\"/Users/flaviab/Downloads/histopathologic-cancer-detection/outputs\"\n    \n    return base_path, csv_path, base_path_test, path, output_directory\n\n\nbase_path, csv_path, base_path_test, path, output_directory = initialize_paths()\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### get data","metadata":{}},{"cell_type":"code","source":"class DataPreparation:\n    \n    def __init__(self, train_base_path: str, csv_path: str, test_base_path: str):\n        self.train_base_path = train_base_path\n        self.csv_path = csv_path\n        self.test_base_path = test_base_path\n        \n    def _make_path(self, id_str: str, base_path: str) -> str:\n        \n        \"\"\"\n        generate the full file path based on the given id.\n        \"\"\"\n        return os.path.join(base_path, f\"{id_str}.tif\")\n    \n    def prepare_train_labels(self) -> pd.DataFrame:\n        \n        \"\"\"\n        prepare train labels dataframe with proper filename paths and labels.\n        \"\"\"\n        df = pd.read_csv(self.csv_path)\n        df[\"filename\"] = df[\"id\"].apply(lambda x: self._make_path(x, self.train_base_path))\n        df[\"label\"] = df[\"label\"].astype(str)\n        return df\n    \n    def prepare_test_data(self) -> pd.DataFrame:\n        \n        \"\"\"\n        prepare test data dataframe with proper filename paths.\n        \"\"\"\n        df = pd.DataFrame()\n        df['filename'] = [os.path.join(self.test_base_path, filename) for filename in os.listdir(self.test_base_path)]\n        df[\"id\"] = [filename[:-4] for filename in os.listdir(self.test_base_path)]\n        return df\n\n\ndata_preparer = DataPreparation(base_path, csv_path, base_path_test)\ntrain_data = data_preparer.prepare_train_labels()\ntest_data = data_preparer.prepare_test_data()\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Instructions: Step 1\n    - Brief description of the problem and data (5 pts) \n    - Briefly describe the challenge problem and NLP. Describe the size, dimension, structure, etc., of the data. ","metadata":{}},{"cell_type":"markdown","source":"##### Histopathologic Cancer Detection Challenge\n\nThe problem: This challenge is about finding specific cancer cells in detailed scans of lymph nodes. These cancer cells can move from their original spot to other places in the body, often starting with the lymph nodes. Catching them early can really help a patient's chances.\n\nNatural Language Processing (NLP):\nNLP is a part of artificial intelligence that helps computers understand and respond to human language. \n\nThe image we just printed below shows cells from a tissue, sized 96x96 pixels, and uses the RGB scale, that captures color information in images. This color data can be crucial as different classes might have distinct color patterns. \n","metadata":{}},{"cell_type":"code","source":"def load_and_display_image(path: str):\n    \n    \"\"\"\n    load an image from the provided path and display it in the Jupyter notebook.\n    \"\"\"\n\n    with Image.open(path) as img:\n        img_array = np.array(img)\n        print(f\"Image Shape = {img_array.shape}\")\n        display(img)\n\nload_and_display_image(path)\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Instructions: Step 2\n\n#### Exploratory Data Analysis (EDA) — Inspect, Visualize and Clean the Data (15 pts)\n\n- Show a few visualizations like histograms. \n- Describe any data cleaning procedures. \n- Based on your EDA, what is your plan of analysis? ","metadata":{}},{"cell_type":"markdown","source":"EDA Plan:\n\n0. Load Data\n1. EDA: check values distribuition and general info \n2. EDA: check missing values\n3. EDA: Check duplicated lines\n4. EDA: Check unbalanced data (cancer/not cancer)\n5. EDA: Plot data\n6. EDA: Plot part of the images to briefly analyze them and get possible insights\n","metadata":{}},{"cell_type":"markdown","source":"##### checking missing values ","metadata":{}},{"cell_type":"code","source":"train_data.info()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.dtypes","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['label'] = train_data['label'].astype(str)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.describe()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### checking for duplicated rows","metadata":{}},{"cell_type":"code","source":"train_data.loc[train_data.duplicated()]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### checking for unbalanced data","metadata":{}},{"cell_type":"code","source":"train_data.label.value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### plot data","metadata":{}},{"cell_type":"code","source":"def plot_label_distribution(train_data: pd.DataFrame) -> None:\n    \n    \"\"\"\n    Plots the distribution of categories in the given training dataset.\n    \"\"\"\n    \n    plot_data = train_data.groupby('label').agg(no_samples = ('id','count')).reset_index()\n    plot_data['%_labels'] = round((plot_data.no_samples/plot_data.no_samples.sum())*100,2)\n\n    fig = px.bar(plot_data, x=\"label\", y=\"%_labels\", \n                 title='Distribution of categories in the training dataset', \n                 text='%_labels')\n\n    fig.show()\n\n\nplot_label_distribution(train_data)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"even though the data is slightly unbalanced, there's no reason to implement big techniquies. ","metadata":{}},{"cell_type":"markdown","source":"##### check images","metadata":{}},{"cell_type":"code","source":"def load_subset_images(df, label, num_images=50):\n    \n    \"\"\"\n    get images of both labels (cancer/not cancer) and display it to human analysis.\n    \"\"\"\n    \n    filenames = df[df['label'] == label]['filename'].head(num_images).tolist()\n    return [Image.open(filename) for filename in filenames]\n\nlabel_0_images = load_subset_images(train_data, '0')\nlabel_1_images = load_subset_images(train_data, '1')\n\ndef display_side_by_side(images_0, images_1, label_0=\"Not Cancer\", label_1=\"Cancer\"):\n    \n    \"\"\"\n    fix display to better comparison.\n    \"\"\"\n    \n    total_cols = 5 \n    total_rows = max(len(images_0), len(images_1))\n    \n    fig, axs = plt.subplots(total_rows, total_cols, figsize=(15, 5*total_rows))\n    \n    for i in range(total_rows):\n        for j in range(total_cols):\n            axs[i, j].axis('off')\n            if j < 3 and i < len(images_0):  \n                axs[i, j].imshow(images_0[i])\n                axs[i, j].set_title(label_0)\n            elif j >= 3 and i < len(images_1): \n                axs[i, j].imshow(images_1[i % 2])\n                axs[i, j].set_title(label_1)\n\n    plt.tight_layout()\n    plt.show()\n\ndisplay_side_by_side(label_0_images, label_1_images)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Given the results in the images shown below, the cancer and not cancer images have complex features and somehow similar that would make hard to detect each one by just the human eye. The images alone aren't enough to help me get better insight to add the model.","metadata":{}},{"cell_type":"markdown","source":"##### create generator to load data in batches","metadata":{}},{"cell_type":"code","source":"train_data['label'] = train_data['label'].astype(str)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def create_data_generators(dataframe, test_dataframe, target_size=(96, 96), batch_size=50, validation_split=0.25):\n\n    \"\"\"\n    create train, validation, and test generators from dataframes.\n\n    \"\"\"\n\n    datagen = ImageDataGenerator(\n        rescale=1./255,\n        rotation_range=40,\n        width_shift_range=0.2,\n        height_shift_range=0.2,\n        shear_range=0.2,\n        zoom_range=0.2,\n        horizontal_flip=True,\n        fill_mode='nearest',\n        brightness_range=[0.7,1.3],\n        validation_split=validation_split\n    )\n    \n    data_train = datagen.flow_from_dataframe(\n        dataframe=dataframe,\n        x_col=\"filename\",\n        y_col=\"label\",\n        target_size=target_size,\n        color_mode=\"rgb\",\n        batch_size=batch_size,\n        class_mode=\"binary\",\n        subset=\"training\",\n        validate_filenames=False\n    )\n    \n    data_val = datagen.flow_from_dataframe(\n        dataframe=dataframe,\n        x_col=\"filename\",\n        y_col=\"label\",\n        target_size=target_size,\n        color_mode=\"rgb\",\n        batch_size=batch_size,\n        class_mode=\"binary\",\n        subset=\"validation\",\n        validate_filenames=False\n    )\n\n    data_test = ImageDataGenerator(rescale=1./255).flow_from_dataframe(\n        dataframe=test_dataframe,\n        x_col=\"filename\",\n        target_size=target_size,\n        color_mode=\"rgb\",\n        batch_size=batch_size,\n        class_mode=None, \n        shuffle=False,\n        validate_filenames=False\n    )\n    \n    return data_train, data_val, data_test\n\n\ndata_train, data_val, data_test = create_data_generators(train_data, test_data)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Hyperparameter search - test with no training improvement strategy","metadata":{}},{"cell_type":"markdown","source":"### Instructions: Step 3\n\n    - Describe the architecture of your model and why you think that particular architecture would be suitable for this problem. \n    - Compare multiple architectures and adjust hyperparameters.","metadata":{}},{"cell_type":"markdown","source":"##### Architecture's description:\n\n    1. Baseline:\n    - Starts with two conv2D layers having 32 filters and then 64 filters;\n    - Each of these layers is followed up by a pooling step using maxpooling2D;\n    - The data is then flattened and pushed through a dense layer packed with 128 neurons;\n    - At the end, there's just one neuron giving the final output, and it uses a sigmoid activation;\n    \n    I've set up a baseline before training a more complex model to get a starting point. It was helpful to give a general idea of how much progress the solution made as it moved forward. This baseline was made to provide a simpler solution to the problem with more basic concepts of neural networls.\n\n    2. Deep Model:\n    - This one takes inspiration from the baseline but goes deeper with added conv2d and maxpooling2D layers;\n    - For smoother learning, it's got batch normalization, and to keep things in check and avoid overfitting, there's a dropout layer in place;\n\n    3. Transfer Learning Model:\n    - This model it's based on the resnet50, a pretrained model, but without its usual top layer;\n    - After processing through resnet50, the data is flattened and goes through a dense layer with 128 neurons;\n    - Finally, a single neuron, activated by a sigmoid function, wraps things up;\n    \n    Resnet was picked because is a deep neural network with \"shortcuts\" features that help it learn better and avoid some common pitfalls, hence it's fit to help spotting the tricky patterns in cancer detection images. Since resmet is trained with data from previous big datasets, it gets a good start. Plus, its design keeps it from easy overfit, which is desirable in this case. All in all, resnet's depth and prior knowledge make it a top choice for diving deep into complex cancer images.","metadata":{}},{"cell_type":"code","source":"leaky_relu = LeakyReLU(alpha=0.01)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ndef build_model(hp):\n    \"\"\"\n    This snippet defines a function to build a neural network model. \n    Depending on a choice parameter, it constructs a baseline, deep neural, or a transfer learning resnet based model.\n    The code uses Adam optimizer, and returns the configured model.\n    \"\"\"\n    \n    model_choice = hp.Choice('model_type', ['baseline', 'deep', 'transfer'])\n\n    model = Sequential() \n    \n    if model_choice == 'baseline':\n        model.add(Conv2D(32, (3,3), activation=leaky_relu, kernel_regularizer=l2(0.01), input_shape=(96, 96, 3)))\n        model.add(MaxPooling2D(2,2))\n        model.add(Conv2D(64, (3,3), activation=leaky_relu, kernel_regularizer=l2(0.01)))\n        model.add(MaxPooling2D(2,2))\n        model.add(Flatten())\n        model.add(Dense(128, activation='relu'))\n        model.add(Dense(1, activation='sigmoid'))\n        \n    \n    elif model_choice == 'deep':\n        model.add(Conv2D(32, (3,3), activation=leaky_relu, kernel_regularizer=l2(0.01), input_shape=(96, 96, 3)))\n        model.add(MaxPooling2D(2,2))\n        model.add(BatchNormalization())\n        model.add(Conv2D(64, (3,3), activation=leaky_relu))\n        model.add(MaxPooling2D(2,2))\n        model.add(BatchNormalization())\n        model.add(Conv2D(128, (3,3), activation=leaky_relu))\n        model.add(MaxPooling2D(2,2))\n        model.add(BatchNormalization())\n        model.add(Conv2D(256, (3,3), activation=leaky_relu))\n        model.add(MaxPooling2D(2,2))\n        model.add(BatchNormalization())\n        model.add(Flatten())\n        model.add(Dense(128, activation='relu'))\n        model.add(Dropout(0.5))\n        model.add(Dense(1, activation='sigmoid'))\n    \n    else:  # 'transfer'\n        base_model = EfficientNetB0(weights='imagenet', include_top=False, input_shape=(96, 96, 3))\n        model.add(base_model)\n        model.add(Flatten())\n        model.add(Dense(128, activation='relu'))\n        model.add(Dense(1, activation='sigmoid'))\n\n    model.compile(optimizer=tf.keras.optimizers.Adam(hp.Float('learning_rate', 1e-4, 1e-2, sampling='log')),\n                  loss='binary_crossentropy',\n                  metrics=['accuracy'])\n\n    return model\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def compute_class_weights(data_train):\n    \n    \"\"\"\n    Compute class weights based on the classes of the training data.\n    \"\"\"\n    \n    weights = class_weight.compute_class_weight('balanced', np.unique(data_train.classes), data_train.classes)\n    class_weights = {0: weights[0], 1: weights[1]}\n    \n    return class_weights\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tuner = BayesianOptimization(\n    build_model,\n    objective= 'val_accuracy',\n    max_trials= 5,\n    num_initial_points= 2,\n    directory= output_directory,\n    project_name= 'histopathologic_cancer_detection_bayesian'\n)\n\ntuner.search(data_train, epochs=10, validation_data=data_val)\nbest_model = tuner.get_best_models(num_models=1)[0]\nloss, accuracy = best_model.evaluate(data_test, test_labels)\n\nprint(f\"Best model accuracy on test set: {accuracy:.4f}\")\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Hyperparameter search - test implementing improvement strategy","metadata":{}},{"cell_type":"markdown","source":"### Instructions: Step 4 - Results and analysis (35 pts)\n\n    - Perform hyperparameter tuning, try different architectures for comparison, apply techniques to improve training or performance, and discuss what helped.\n\n    - Includes results with tables and figures. There is a breakdown of why something worked well or not, troubleshooting, and a summary of the hyperparameter optimization procedure.","metadata":{}},{"cell_type":"markdown","source":"    As seen above, the code without performance implementations took days to run. This approach turned out to be inefficient and undesirable. For this reason, some changes were made to the code.\n    \n    The new approach sets up a few training enhancements: \n\n    1. Early stopping halts training if validation accuracy stalls for 3 epochs;\n    2. Learning rate reduction halves the learning rate if validation loss doesn't improve for 2 epochs;\n    3. Reduced number of trials and epochs;\n    \n    \n    In both cases, the bayesian optimization technique for hyperparameter tuning was picked for its intelligent way of selecting the next set of hyperparameters to test based on prior results, reducing unnecessary trials and thus saving time and CPU/GPU memory, which was very needed.","metadata":{}},{"cell_type":"code","source":"%%time\n\nearly_stopping = EarlyStopping(monitor='val_accuracy', patience=10, restore_best_weights=True)\nlr_reduction = ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=2, verbose=1, min_lr=1e-6)\n\ntuner2 = BayesianOptimization(\n    build_model,\n    objective= 'val_accuracy',\n    max_trials= 5,\n    num_initial_points= 2,\n    directory= output_directory,\n    project_name= 'histopathologic_cancer_detection_bayesian'\n)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ntuner2.search(data_train, epochs=10, validation_data=data_val, callbacks=[early_stopping, lr_reduction])\nbest_model = tuner2.get_best_models(num_models=1)[0]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\npredictions = best_model.predict(data_test)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### model evaluation","metadata":{}},{"cell_type":"code","source":"%%time\n\ndef optimize_and_evaluate_model(build_model_function, \n                                data_train, \n                                data_val, \n                                data_test,\n                                output_directory,\n                                max_trials=4, \n                                epochs=6,\n                                num_initial_points=2, \n                                project_name='histopathologic_cancer_detection_bayesian'):\n    \"\"\"\n    Optimize a Keras model using Bayesian optimization and evaluate its performance.\n    \"\"\"\n\n    tuner2 = BayesianOptimization(\n        build_model_function,\n        objective='val_accuracy',\n        max_trials=max_trials,\n        num_initial_points=num_initial_points,\n        directory=output_directory,\n        project_name=project_name\n    )\n\n    tuner2.search(data_train, epochs=epochs, validation_data=data_val)\n    best_model = tuner2.get_best_models(num_models=1)[0]\n    \n    val_loss, val_accuracy = best_model.evaluate(data_val)\n    print(f\"Validation Accuracy: {val_accuracy*100:.2f}%\")\n    print(f\"Validation Loss: {val_loss:.4f}\")\n\n    return best_model, val_loss, val_accuracy, tuner2\n\nbest_model, val_loss, val_accuracy, tuner2 = optimize_and_evaluate_model(build_model, data_train, data_val, data_test, output_directory)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ndef evaluate_and_plot_results(model, validation_data):\n    \"\"\"\n    Evaluate model on validation data and visualize the results.\n    \n    Parameters:\n    - model: Keras model.\n    - validation_data: Validation data generator.\n    \n    Returns:\n    None. (Displays plots and prints results.)\n    \"\"\"\n    \n    # Predict on validation data\n    val_predictions = model.predict(validation_data)\n    binary_predictions = np.round(val_predictions).astype(int)\n    true_labels = validation_data.labels\n\n    # Create confusion matrix\n    matrix = confusion_matrix(true_labels, binary_predictions)\n    \n    # Plotly heatmap for confusion matrix\n    z = matrix.tolist()\n    x = ['Predicted 0', 'Predicted 1']\n    y = ['Actual 0', 'Actual 1']\n    fig = ff.create_annotated_heatmap(z, x=x, y=y, colorscale='Blues', showscale=False)\n    fig.update_layout(title='Confusion Matrix')\n    \n    fig.show()\n\n    # Classification report\n    report = classification_report(true_labels, binary_predictions, output_dict=True)\n    print(classification_report(true_labels, binary_predictions))\n    \n    # Plotly bar chart for classification report\n    report_data = [{'label': label, **metrics} for label, metrics in report.items() if label.isdigit()]\n    df = pd.DataFrame.from_dict(report_data)\n    \n    fig = go.Figure()\n    for metric in ['precision', 'recall', 'f1-score']:\n        fig.add_trace(go.Bar(x=df['label'], y=df[metric], name=metric))\n    fig.update_layout(title='Classification Report', \n                      xaxis_title='Class', \n                      yaxis_title='Score', \n                      yaxis=dict(range=[0,1]))\n    fig.show()\n\nevaluate_and_plot_results(best_model, data_val)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels = np.concatenate([data_train[i][1] for i in range(len(data_train))])\nunique, counts = np.unique(train_labels, return_counts=True)\ntotal_samples = len(train_labels)\nn_classes = len(unique)\nclass_weights = {class_label: total_samples / (n_classes * count) for class_label, count in zip(unique, counts)}\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"true_labels = data_val.labels","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\nval_predictions = best_model.predict(data_val)\nthresholds = np.arange(0, 1.05, 0.05)\nf1_scores = [f1_score(true_labels, np.where(val_predictions > t, 1, 0)) for t in thresholds]\noptimal_threshold = thresholds[np.argmax(f1_scores)]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### evaluate train x validation","metadata":{}},{"cell_type":"code","source":"%%time \n\ndef plot_tuner_results(tuner, data_val, val_predictions):\n    \n    \"\"\"\n    Plot training and validation accuracies and losses, and the ROC curve based on a given tuner's trials.\n    \"\"\"\n    \n    all_train_accuracies, all_val_accuracies, all_train_losses, all_val_losses = [], [], [], []\n\n    for trial in tuner.oracle.trials.values():\n        all_train_accuracies.append(trial.metrics.get_history('accuracy')[0].value)\n        all_val_accuracies.append(trial.metrics.get_history('val_accuracy')[0].value)\n        all_train_losses.append(trial.metrics.get_history('loss')[0].value)\n        all_val_losses.append(trial.metrics.get_history('val_loss')[0].value)\n\n    print(all_train_accuracies)\n    \n    epochs = list(range(len(all_train_accuracies)))\n\n    # Plotting Accuracies\n    \n    epochs = list(range(len(all_train_accuracies)))\n    plt.figure(figsize=(10, 5))\n    plt.plot(all_train_accuracies, label=\"Training Accuracies\", marker='o')\n    plt.plot(all_val_accuracies, label=\"Validation Accuracies\", marker='o')\n    plt.xlabel('Epochs')\n    plt.ylabel('Accuracy')\n    plt.title('Training vs Validation Accuracies')\n    plt.legend()\n    plt.grid(True, which='both', linestyle='--', linewidth=0.5)\n    plt.tight_layout()\n    plt.show()\n\n    # Plotting Losses\n    \n    plt.figure(figsize=(10, 5))\n    plt.plot(all_train_losses, label=\"Training Losses\", marker='o')\n    plt.plot(all_val_losses, label=\"Validation Losses\", marker='o')\n    plt.xlabel('Epochs')\n    plt.ylabel('Loss')\n    plt.title('Training vs Validation Losses')\n    plt.legend()\n    plt.grid(True, which='both', linestyle='--', linewidth=0.5)\n    plt.tight_layout()\n    plt.show()\n\n\nplot_tuner_results(tuner2, data_val, val_predictions)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"conclusion: the model is overfitting after 3 epochs.","metadata":{}},{"cell_type":"markdown","source":"### retraining","metadata":{}},{"cell_type":"code","source":"%%time\n\nbest_model.fit(data_train, epochs=2, validation_data=data_val, class_weight=class_weights, callbacks=[early_stopping, lr_reduction])\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"evaluate_and_plot_results(best_model, data_val)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### results","metadata":{}},{"cell_type":"markdown","source":"Given the problem is a serious issue, recall is the first option metric to evaluate the model, since I rather get the true positives labels right.\nThat being said, the recall results for Label 1 (cancer cells) was alarmingly low. The model is missing a substantial amount (74%) of the cancerous cells, which could have serious health implications in a real-world scenario.\nThe model needs improvement.","metadata":{}},{"cell_type":"markdown","source":"## model improvement test","metadata":{}},{"cell_type":"markdown","source":"- activation function changed from relu to leaky_relu;\n- increased number of trials;\n- added brightness and contrast adjustments to improve the image's interpretation; \n- changed learning rate reduction to monitor accuracy, instead of loss;\n- changed binary_predictions method;","metadata":{}},{"cell_type":"code","source":"%%time\n\nearly_stopping = EarlyStopping(monitor='val_accuracy', patience=5, restore_best_weights=True)\nlr_reduction = ReduceLROnPlateau(monitor='val_accuracy', factor=0.5, patience=2, verbose=1, min_lr=1e-6)\nbinary_predictions = np.where(val_predictions > optimal_threshold, 1, 0)\n\n\ntuner3 = BayesianOptimization(\n    build_model,\n    objective= 'val_accuracy',\n    max_trials= 8,\n    num_initial_points= 2,\n    directory= output_directory,\n    project_name= 'histopathologic_cancer_detection_bayesian'\n)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ntuner3.search(data_train, epochs=10, validation_data=data_val, callbacks=[early_stopping, lr_reduction])\nbest_model = tuner3.get_best_models(num_models=1)[0]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\npredictions = best_model.predict(data_test)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ndef optimize_and_evaluate_model(build_model_function, \n                                data_train, \n                                data_val, \n                                data_test,\n                                output_directory,\n                                max_trials=4, \n                                epochs=6,\n                                num_initial_points=2, \n                                project_name='histopathologic_cancer_detection_bayesian'):\n    \"\"\"\n    Optimize a Keras model using Bayesian optimization and evaluate its performance.\n    \"\"\"\n\n    tuner3 = BayesianOptimization(\n        build_model_function,\n        objective='val_accuracy',\n        max_trials=max_trials,\n        num_initial_points=num_initial_points,\n        directory=output_directory,\n        project_name=project_name\n    )\n\n    tuner3.search(data_train, epochs=epochs, validation_data=data_val)\n    best_model = tuner3.get_best_models(num_models=1)[0]\n    \n    val_loss, val_accuracy = best_model.evaluate(data_val)\n    print(f\"Validation Accuracy: {val_accuracy*100:.2f}%\")\n    print(f\"Validation Loss: {val_loss:.4f}\")\n\n    return best_model, val_loss, val_accuracy, tuner3\n\nbest_model, val_loss, val_accuracy, tuner3 = optimize_and_evaluate_model(build_model, data_train, data_val, data_test, output_directory)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### evaluate again","metadata":{}},{"cell_type":"code","source":"%%time \n\nval_predictions = best_model.predict(data_val)\nthresholds = np.arange(0, 1.05, 0.05)\nf1_scores = [f1_score(true_labels, np.where(val_predictions > t, 1, 0)) for t in thresholds]\noptimal_threshold = thresholds[np.argmax(f1_scores)]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\nplot_tuner_results(tuner3, data_val, val_predictions)\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### model train with 3 epochs and ","metadata":{}},{"cell_type":"code","source":"%%time\n\nbest_model.fit(data_train, epochs=3, validation_data=data_val, class_weight=class_weights, callbacks=[early_stopping, lr_reduction])\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nevaluate_and_plot_results(best_model, data_val)\n","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Results","metadata":{}},{"cell_type":"markdown","source":"The final test showed a slightly better result than the others in improving the recall metric, which is the one chosen to evaluate the model. \nDespite this, the model still has many areas that need improvement, requiring more computational resources to continue being refined.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}