{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"},{"sourceId":7007089,"sourceType":"datasetVersion","datasetId":4028277},{"sourceId":7007234,"sourceType":"datasetVersion","datasetId":4028376},{"sourceId":7093730,"sourceType":"datasetVersion","datasetId":4088171},{"sourceId":7179150,"sourceType":"datasetVersion","datasetId":4149221},{"sourceId":7179152,"sourceType":"datasetVersion","datasetId":4149222},{"sourceId":7187871,"sourceType":"datasetVersion","datasetId":4155623},{"sourceId":7187872,"sourceType":"datasetVersion","datasetId":4155624}],"dockerImageVersionId":30587,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ***Histopathic Cancer Detection Submission Notebook***\n## DSCI 598\n Jeffery Boczkaja |\n Sara Bronson |\n Shantel Johnson  ","metadata":{}},{"cell_type":"markdown","source":"# Section 1: Importing Packages\nHere we will import the packages necessary for creating our submission file.","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom tensorflow import keras\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nimport matplotlib.image as mpimg\nimport pickle\nfrom keras.models import load_model","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-17T04:35:39.308946Z","iopub.execute_input":"2023-12-17T04:35:39.309569Z","iopub.status.idle":"2023-12-17T04:35:51.236154Z","shell.execute_reply.started":"2023-12-17T04:35:39.309537Z","shell.execute_reply":"2023-12-17T04:35:51.235300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 2: Data Loading and Initial Setup\nThis section involves importing the test dataset and the sample submission CSV file. The test dataset contains the images to be analyzed while the sample submission file provides a template for our prediction outputs.","metadata":{}},{"cell_type":"code","source":"test = '/kaggle/input/histopathologic-cancer-detection/test/'\ndf = pd.read_csv(f'../input/histopathologic-cancer-detection/sample_submission.csv', dtype=str)\n\ndf.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 3: Data Preprocessing and Preparation for Modeling\nIn this section we add a new column to our DataFrame for the file paths of test images. We then set up an ImageDataGenerator for image rescaling ensuring that the pixel values are normalized. We prepare a test data loader to batch-process the images, setting parameters like batch size and image dimensions for model input.","metadata":{}},{"cell_type":"code","source":"df['path'] = df.id + '.tif'","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df['path'].head())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_datagen = ImageDataGenerator(rescale=1/255)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BATCH_SIZE = 64\n\ntest_loader = test_datagen.flow_from_dataframe(\n    dataframe = df,\n    directory = test,\n    x_col = 'path',\n    batch_size = BATCH_SIZE,\n    shuffle = False,\n    class_mode = None,\n    target_size = (96,96)\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 4: Model Loading and Prediction\nThis section begins by verifying the presence of the saved model file. We then load the pre-trained Keras model and its training history for analysis. We use this model to make predictions on the test dataset generating probabilities that indicate the likelihood of cancer presence in each image.","metadata":{}},{"cell_type":"code","source":"# Check if the file exists\nos.path.exists('/kaggle/input/xlxmodel/')","metadata":{"execution":{"iopub.status.busy":"2023-12-17T04:36:01.495306Z","iopub.execute_input":"2023-12-17T04:36:01.496347Z","iopub.status.idle":"2023-12-17T04:36:01.503583Z","shell.execute_reply.started":"2023-12-17T04:36:01.496313Z","shell.execute_reply":"2023-12-17T04:36:01.502692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kt_model = load_model('/kaggle/input/xlxmodel/HCD_SB_6_Model (4).h5')\n\n# Load the history\nwith open('/kaggle/input/hcd-take6/HCD_SB_6_Model_history.pkl', 'rb') as file:\n    loaded_history = pickle.load(file)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_probs = kt_model.predict(test_loader)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 5: Creating and Saving Submission File\nHere we reload the sample submission file and update it with the model's predictions. The label column is filled with the predicted probabilities. We save this updated DataFrame as a CSV file which will be used the submission.","metadata":{}},{"cell_type":"code","source":"submission = pd.read_csv(\"/kaggle/input/histopathologic-cancer-detection/sample_submission.csv\")\nsubmission.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.label = test_probs\nsubmission.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', header=True, index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}