{"metadata":{"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":20604,"databundleVersionId":1357052,"sourceType":"competition"}],"dockerImageVersionId":30004,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.7.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Revision version of Pulmonary Fibrosis Analysis with LSTM\n\nThis notebook is copied from SIMO333's notebook:https://www.kaggle.com/code/simo333/pulmonary-fibrosis-img-data-2-0.  \n    \nThe problem of losses of nan was fixed.  \n\nLearned a lot  from this notebook.Thanks to SIMO333.","metadata":{}},{"cell_type":"code","source":"from os import listdir\nimport pandas as pd\nimport numpy as np\nimport pydicom as dicom\nfrom pydicom import dcmread\nimport matplotlib.pyplot as plt\nimport tqdm\nfrom scipy import ndimage\nfrom skimage import measure, morphology, segmentation\nfrom skimage.segmentation import clear_border, watershed   # for recent version of skimage\nimport tensorflow as tf\nimport keras    #this line for tensorflow 2.1x and keras 2.1x","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","scrolled":true,"execution":{"iopub.status.busy":"2024-03-02T08:36:23.498262Z","iopub.execute_input":"2024-03-02T08:36:23.498634Z","iopub.status.idle":"2024-03-02T08:36:29.750094Z","shell.execute_reply.started":"2024-03-02T08:36:23.498590Z","shell.execute_reply":"2024-03-02T08:36:29.749250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"file_path = \"/kaggle/input/osic-pulmonary-fibrosis-progression/\"\nlistdir(file_path)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.752303Z","iopub.execute_input":"2024-03-02T08:36:29.752601Z","iopub.status.idle":"2024-03-02T08:36:29.761319Z","shell.execute_reply.started":"2024-03-02T08:36:29.752571Z","shell.execute_reply":"2024-03-02T08:36:29.760526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(file_path + \"train.csv\")\ntest_df = pd.read_csv(file_path + \"test.csv\")\nsub_df = pd.read_csv(file_path + \"sample_submission.csv\")\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.762869Z","iopub.execute_input":"2024-03-02T08:36:29.763202Z","iopub.status.idle":"2024-03-02T08:36:29.814197Z","shell.execute_reply.started":"2024-03-02T08:36:29.763172Z","shell.execute_reply":"2024-03-02T08:36:29.813259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.815507Z","iopub.execute_input":"2024-03-02T08:36:29.815806Z","iopub.status.idle":"2024-03-02T08:36:29.829452Z","shell.execute_reply.started":"2024-03-02T08:36:29.815777Z","shell.execute_reply":"2024-03-02T08:36:29.828276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.833466Z","iopub.execute_input":"2024-03-02T08:36:29.833945Z","iopub.status.idle":"2024-03-02T08:36:29.845929Z","shell.execute_reply.started":"2024-03-02T08:36:29.833897Z","shell.execute_reply":"2024-03-02T08:36:29.844955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# detect those rows with duplicated values of \"Patient\" & \"Weeks\"\nduplicate_data = train_df[train_df.duplicated(subset=[\"Patient\", \"Weeks\"], keep=False)]\nduplicate_data","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.849291Z","iopub.execute_input":"2024-03-02T08:36:29.849599Z","iopub.status.idle":"2024-03-02T08:36:29.872905Z","shell.execute_reply.started":"2024-03-02T08:36:29.849569Z","shell.execute_reply":"2024-03-02T08:36:29.872101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop duplicated rows\ntrain_df.drop_duplicates(subset=[\"Patient\", \"Weeks\"], keep=\"last\", inplace=True)\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.874192Z","iopub.execute_input":"2024-03-02T08:36:29.874473Z","iopub.status.idle":"2024-03-02T08:36:29.891672Z","shell.execute_reply.started":"2024-03-02T08:36:29.874445Z","shell.execute_reply":"2024-03-02T08:36:29.890793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df[[\"Patient\", \"Weeks\"]] = sub_df[\"Patient_Week\"].str.split(\"_\", expand=True)\nsub_df = sub_df[[\"Patient\", \"Weeks\", \"Patient_Week\"]]\nsub_df = sub_df.merge(test_df.drop(\"Weeks\", axis=1), on=\"Patient\")\n\ntrain_df[\"Source\"] = \"train\"\nsub_df[\"Source\"] = \"test\"\n\n# dataset = train_df._append([sub_df])   # original\ndataset = pd.concat([train_df, sub_df])  # for recent version of pandas\ndataset.reset_index(drop=True, inplace=True)\ndataset.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.893042Z","iopub.execute_input":"2024-03-02T08:36:29.893315Z","iopub.status.idle":"2024-03-02T08:36:29.935557Z","shell.execute_reply.started":"2024-03-02T08:36:29.893287Z","shell.execute_reply":"2024-03-02T08:36:29.934548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset[\"FVC_ave\"] = (dataset[\"FVC\"] ) / dataset[\"Percent\"] * 100\ndataset.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.936879Z","iopub.execute_input":"2024-03-02T08:36:29.937193Z","iopub.status.idle":"2024-03-02T08:36:29.972973Z","shell.execute_reply.started":"2024-03-02T08:36:29.937163Z","shell.execute_reply":"2024-03-02T08:36:29.972009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def baseline_week(df):\n    df = df.copy()\n    df[\"Weeks\"] = df[\"Weeks\"].astype(int)\n    df.loc[df[\"Source\"] == \"test\", \"min_weeks\"] = np.nan\n    df[\"min_weeks\"] = df.groupby(\"Patient\")[\"Weeks\"].transform(\"min\")\n    df[\"baseline_week\"] = df[\"Weeks\"] - df[\"min_weeks\"]\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.974312Z","iopub.execute_input":"2024-03-02T08:36:29.974602Z","iopub.status.idle":"2024-03-02T08:36:29.981746Z","shell.execute_reply.started":"2024-03-02T08:36:29.974573Z","shell.execute_reply":"2024-03-02T08:36:29.981014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = baseline_week(dataset)\ndataset.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:29.983083Z","iopub.execute_input":"2024-03-02T08:36:29.983367Z","iopub.status.idle":"2024-03-02T08:36:30.015077Z","shell.execute_reply.started":"2024-03-02T08:36:29.983341Z","shell.execute_reply":"2024-03-02T08:36:30.014288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_baseline_fvc(df):\n    df = df.copy()\n    base = df.loc[df[\"Weeks\"] == df[\"min_weeks\"]].copy()\n    base = df[[\"Patient\", \"FVC\"]].copy()\n    base.columns = [\"Patient\", \"base_fvc\"]\n    base[\"no\"] = 1\n    base[\"no\"] = base.groupby(\"Patient\")[\"no\"].transform(\"cumsum\")\n    base = base[base.no == 1]\n    base.drop(\"no\", axis=1, inplace=True)\n    df = df.merge(base, on = \"Patient\", how = \"left\")\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.016513Z","iopub.execute_input":"2024-03-02T08:36:30.016889Z","iopub.status.idle":"2024-03-02T08:36:30.024871Z","shell.execute_reply.started":"2024-03-02T08:36:30.016856Z","shell.execute_reply":"2024-03-02T08:36:30.024086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = get_baseline_fvc(dataset)\ndataset.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.026359Z","iopub.execute_input":"2024-03-02T08:36:30.026734Z","iopub.status.idle":"2024-03-02T08:36:30.077825Z","shell.execute_reply.started":"2024-03-02T08:36:30.026696Z","shell.execute_reply":"2024-03-02T08:36:30.076708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset[\"Sex\"] = pd.Categorical(dataset[\"Sex\"])\ndataset[\"Sex\"] = dataset.Sex.cat.codes\ndataset[\"SmokingStatus\"] = pd.Categorical(dataset[\"SmokingStatus\"])\ndataset[\"SmokingStatus\"] = dataset.SmokingStatus.cat.codes\ndataset.tail()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.079402Z","iopub.execute_input":"2024-03-02T08:36:30.080001Z","iopub.status.idle":"2024-03-02T08:36:30.109071Z","shell.execute_reply.started":"2024-03-02T08:36:30.079954Z","shell.execute_reply":"2024-03-02T08:36:30.108110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"true_pat_res = test_df.Patient.unique()\ntrue_pat_res.sort()\n\ntrue_result = train_df.loc[train_df[\"Patient\"].isin(true_pat_res)].copy()\ntrue_result.info()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.110560Z","iopub.execute_input":"2024-03-02T08:36:30.110863Z","iopub.status.idle":"2024-03-02T08:36:30.127778Z","shell.execute_reply.started":"2024-03-02T08:36:30.110833Z","shell.execute_reply":"2024-03-02T08:36:30.126637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset.drop_duplicates(subset=[\"Patient\", \"Weeks\"], keep=\"last\", inplace=True)\ntrain_df = dataset.loc[dataset[\"Source\"] == \"train\"].copy()\n\ntest_df = dataset.loc[dataset[\"Source\"] == \"test\"].copy()\ntrain_df.drop(\"Source\", axis=1, inplace=True)\ntest_df.drop(\"Source\", axis=1, inplace=True)\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.129293Z","iopub.execute_input":"2024-03-02T08:36:30.129718Z","iopub.status.idle":"2024-03-02T08:36:30.173036Z","shell.execute_reply.started":"2024-03-02T08:36:30.129677Z","shell.execute_reply":"2024-03-02T08:36:30.170539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.drop([\"Patient_Week\", \"Percent\", \"min_weeks\"], axis=1, inplace=True)\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.174705Z","iopub.execute_input":"2024-03-02T08:36:30.175144Z","iopub.status.idle":"2024-03-02T08:36:30.194472Z","shell.execute_reply.started":"2024-03-02T08:36:30.175094Z","shell.execute_reply":"2024-03-02T08:36:30.193152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.196478Z","iopub.execute_input":"2024-03-02T08:36:30.196859Z","iopub.status.idle":"2024-03-02T08:36:30.218244Z","shell.execute_reply.started":"2024-03-02T08:36:30.196806Z","shell.execute_reply":"2024-03-02T08:36:30.217465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.drop([\"Percent\", \"Patient_Week\", \"min_weeks\"], axis=1, inplace=True)\ntest_df","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.219252Z","iopub.execute_input":"2024-03-02T08:36:30.219508Z","iopub.status.idle":"2024-03-02T08:36:30.241181Z","shell.execute_reply.started":"2024-03-02T08:36:30.219482Z","shell.execute_reply":"2024-03-02T08:36:30.240362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.info()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.242625Z","iopub.execute_input":"2024-03-02T08:36:30.242948Z","iopub.status.idle":"2024-03-02T08:36:30.260084Z","shell.execute_reply.started":"2024-03-02T08:36:30.242897Z","shell.execute_reply":"2024-03-02T08:36:30.258774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# if file_path == \"../OSIC/osic-pulmonary//\":\n#     train_df[\"dcm_path\"] = file_path + \"train/\" + train_df.Patient + \"/\"\n# else:\n#     train_df[\"dcm_path\"] = file_path + \"train/\" + train_df.StudyInstanceUID + \"/\" + train_df.SeriesInstanceUID\ntrain_df[\"dcm_path\"] = file_path + \"train/\" + train_df.Patient + \"/\"\nprint(train_df[\"dcm_path\"])","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.261623Z","iopub.execute_input":"2024-03-02T08:36:30.262103Z","iopub.status.idle":"2024-03-02T08:36:30.271990Z","shell.execute_reply.started":"2024-03-02T08:36:30.262054Z","shell.execute_reply":"2024-03-02T08:36:30.270609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df[\"dcm_path\"] = file_path + \"test/\" + test_df.Patient + \"/\"","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.273610Z","iopub.execute_input":"2024-03-02T08:36:30.274025Z","iopub.status.idle":"2024-03-02T08:36:30.283730Z","shell.execute_reply.started":"2024-03-02T08:36:30.273983Z","shell.execute_reply":"2024-03-02T08:36:30.282894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_id = []\npatient_path = []\npatients = train_df.Patient.unique()\n\nfor patient in patients:\n    patient_id.append(patient)\n    if file_path == \"/kaggle/input/osic-pulmonary-fibrosis-progression/\":\n        path = train_df[train_df.Patient == patient].dcm_path.values[0]\n    else:\n        path = train_df[train_df.StudyInstanceUID == patient].dcm_path.values[0]\n    ex_dcm = listdir(path)[0]\n    patient_path.append(path)\n    ds = dcmread(path + \"/\" + ex_dcm)\n\npatient_df = pd.DataFrame(data=patient_id, columns=[\"patient\"])\npatient_df.loc[:, \"patient_path\"] = patient_path\npatient_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:30.285541Z","iopub.execute_input":"2024-03-02T08:36:30.285908Z","iopub.status.idle":"2024-03-02T08:36:54.703844Z","shell.execute_reply.started":"2024-03-02T08:36:30.285870Z","shell.execute_reply":"2024-03-02T08:36:54.703089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_scan(dcm_path):\n    files = listdir(dcm_path)\n    # file_no = [np.int_(file.split(\".\")[0]) for file in files] # original\n    file_no = [np.int_(file.split(\".\")[0]) for file in files]   # for recent version of numpy\n    sorted_files = np.sort(file_no)[::-1]\n    slices = [dcmread(dcm_path  + \"/\" + str(file_no) + \".dcm\") for file_no in sorted_files]\n            \n    return slices","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2024-03-02T08:36:54.705003Z","iopub.execute_input":"2024-03-02T08:36:54.705284Z","iopub.status.idle":"2024-03-02T08:36:54.711211Z","shell.execute_reply.started":"2024-03-02T08:36:54.705255Z","shell.execute_reply":"2024-03-02T08:36:54.710442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def convert_to_hu(slices):\n    image = np.stack([s.pixel_array for s in slices])\n    image = image.astype(np.int16)\n    \n    image[image == -2000] = 0\n    \n    intercept = slices[0].RescaleIntercept\n    slope = slices[0].RescaleSlope\n    \n    if slope != 1:\n        image = slope * image.astype(np.float64)\n        image = image.astype(np.int16)\n        \n    image += np.int16(intercept)\n    \n    return np.array(image, dtype=np.int16)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2024-03-02T08:36:54.712839Z","iopub.execute_input":"2024-03-02T08:36:54.713256Z","iopub.status.idle":"2024-03-02T08:36:54.726909Z","shell.execute_reply.started":"2024-03-02T08:36:54.713217Z","shell.execute_reply":"2024-03-02T08:36:54.726271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ex = train_df.dcm_path.values[0]\nprint(ex)\nscans = load_scan(ex)\nhu_scans = convert_to_hu(scans)\n\nplt.figure()\nplt.imshow(hu_scans[12], cmap=plt.cm.gray)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:54.728279Z","iopub.execute_input":"2024-03-02T08:36:54.728530Z","iopub.status.idle":"2024-03-02T08:36:55.430514Z","shell.execute_reply.started":"2024-03-02T08:36:54.728505Z","shell.execute_reply":"2024-03-02T08:36:55.429593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def resample(image, scan, new_spacing=[1,1,1]):\n    spacing = np.array([scan[0].SliceThickness] + list(scan[0].PixelSpacing), dtype=np.float32)\n    resize_factor = spacing / new_spacing\n    new_real_shape = image.shape / resize_factor\n    new_shape = np.round(new_real_shape)\n    real_resize_factor = new_shape / image.shape\n    new_spacing = spacing / real_resize_factor\n    \n    image = ndimage.interpolation.zoom(image, real_resize_factor, mode=\"nearest\")\n    \n    return image, new_spacing","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:55.431812Z","iopub.execute_input":"2024-03-02T08:36:55.432125Z","iopub.status.idle":"2024-03-02T08:36:55.439797Z","shell.execute_reply.started":"2024-03-02T08:36:55.432092Z","shell.execute_reply":"2024-03-02T08:36:55.438637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"im_res, spacing = resample(hu_scans, scans, [1,1,1])\nhu_scans.shape, im_res.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:36:55.441220Z","iopub.execute_input":"2024-03-02T08:36:55.441603Z","iopub.status.idle":"2024-03-02T08:37:02.679120Z","shell.execute_reply.started":"2024-03-02T08:36:55.441564Z","shell.execute_reply":"2024-03-02T08:37:02.678147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(11, 7))\nax[0].imshow(hu_scans[20], cmap=plt.cm.gray)\nax[1].imshow(im_res[1], cmap=plt.cm.gray)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:02.680466Z","iopub.execute_input":"2024-03-02T08:37:02.680819Z","iopub.status.idle":"2024-03-02T08:37:03.078885Z","shell.execute_reply.started":"2024-03-02T08:37:02.680788Z","shell.execute_reply":"2024-03-02T08:37:03.078099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_multival_vals(feature):\n    if type(feature) == dicom.multival.MultiValue:\n        return np.int(feature[0])\n    else:\n        return np.int(feature)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:03.080323Z","iopub.execute_input":"2024-03-02T08:37:03.080591Z","iopub.status.idle":"2024-03-02T08:37:03.085558Z","shell.execute_reply.started":"2024-03-02T08:37:03.080564Z","shell.execute_reply":"2024-03-02T08:37:03.084591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def generate_markers(image):\n    marker_internal = image < -400\n    marker_internal = segmentation.clear_border(marker_internal)\n    marker_internal_labels = measure.label(marker_internal)\n    areas = [r.area for r in measure.regionprops(marker_internal_labels)]\n    areas.sort()\n    if len(areas) > 2:\n        for region in measure.regionprops(marker_internal_labels):\n            if region.area < areas[-2]:\n                for coordinates in region.coords:\n                    marker_internal_labels[coordinates[0], coordinates[1]] = 0\n    marker_internal = marker_internal_labels > 0\n    \n    external_a = ndimage.binary_dilation(marker_internal, iterations=10)\n    external_b = ndimage.binary_dilation(marker_internal, iterations=55)\n    marker_external = external_b ^ external_a\n    \n    marker_watershed = np.zeros((image.shape), dtype=np.int_)\n    marker_watershed += marker_internal * 255\n    marker_watershed += marker_external * 128\n    \n    return marker_internal, marker_external, marker_watershed","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:03.086822Z","iopub.execute_input":"2024-03-02T08:37:03.087114Z","iopub.status.idle":"2024-03-02T08:37:03.098367Z","shell.execute_reply.started":"2024-03-02T08:37:03.087086Z","shell.execute_reply":"2024-03-02T08:37:03.097485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_internal, patient_external, patient_watershed = generate_markers(hu_scans[13])\n\nfig, ax = plt.subplots(1, 3, figsize=(17, 6))\nax[0].set_title(\"Internel marker\")\nax[0].imshow(patient_internal, cmap=\"gray\")\nax[1].set_title(\"External Marker\")\nax[1].imshow(patient_external, cmap=\"gray\")\nax[2].set_title(\"watershed image\")\nax[2].imshow(patient_watershed, cmap=\"gray\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:03.106698Z","iopub.execute_input":"2024-03-02T08:37:03.107015Z","iopub.status.idle":"2024-03-02T08:37:03.652026Z","shell.execute_reply.started":"2024-03-02T08:37:03.106985Z","shell.execute_reply":"2024-03-02T08:37:03.651107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def separate_lungs(image):\n    marker_internal, marker_external, marker_watershed = generate_markers(image)\n    \n    sobel_filtered_dx = ndimage.sobel(image, 0)\n    sobel_filtered_dy = ndimage.sobel(image, 1)\n    sobel_gradient = np.hypot(sobel_filtered_dx, sobel_filtered_dy)\n    sobel_gradient *= 255.0 / np.max(sobel_gradient)\n    \n    wd = watershed(sobel_gradient, marker_watershed)   # for recent version of skimage\n    outline = ndimage.morphological_gradient(wd, size=(3, 3))\n    \n    outline = outline.astype(bool)\n    \n    blackhat_structure = [[0, 0, 1, 1, 1, 0, 0],\n                          [0, 1, 1, 1, 1, 1, 0],\n                          [1, 1, 1, 1, 1, 1, 1],\n                          [1, 1, 1, 1, 1, 1, 1],\n                          [1, 1, 1, 1, 1, 1, 1],\n                          [0, 1, 1, 1, 1, 1, 0],\n                          [0, 0, 1, 1, 1, 0, 0]]\n    \n    blackhat_structure = ndimage.iterate_structure(blackhat_structure, iterations=7)\n    outline += ndimage.black_tophat(outline, structure=blackhat_structure)\n    \n    lungfilter = np.bitwise_or(marker_internal, outline)\n    lungfilter = ndimage.morphology.binary_closing(lungfilter, structure=np.ones((5, 5)), iterations = 3)\n    \n    segmented = np.where(lungfilter == 1, image, -2000*np.ones((image.shape)))\n    \n    return segmented","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:03.653778Z","iopub.execute_input":"2024-03-02T08:37:03.654132Z","iopub.status.idle":"2024-03-02T08:37:03.669407Z","shell.execute_reply.started":"2024-03-02T08:37:03.654098Z","shell.execute_reply":"2024-03-02T08:37:03.668489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_segmented = separate_lungs(hu_scans[12])","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:03.670894Z","iopub.execute_input":"2024-03-02T08:37:03.671270Z","iopub.status.idle":"2024-03-02T08:37:06.045009Z","shell.execute_reply.started":"2024-03-02T08:37:03.671239Z","shell.execute_reply":"2024-03-02T08:37:06.043960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(7, 7))\nplt.title(\"Segmented Lung\")\nplt.imshow(train_segmented, cmap=plt.cm.gray)\n\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.046320Z","iopub.execute_input":"2024-03-02T08:37:06.046630Z","iopub.status.idle":"2024-03-02T08:37:06.250105Z","shell.execute_reply.started":"2024-03-02T08:37:06.046599Z","shell.execute_reply":"2024-03-02T08:37:06.249273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def img_hu_processing(patient_df):\n    \n    for i, patient in enumerate(tqdm.tqdm(patient_df[\"patient\"].values)):\n        try:\n            path = patient_df.loc[patient_df[\"patient\"] == patient].patient_path.values[0]\n            scans = load_scan(path)\n            n = len(scans)\n            if n >= 30:\n                m = int(n/10.0)\n                scans = scans[int(n*0.1):int(n*0.9):int(m)*2]\n                hu_scans = convert_to_hu(scans)\n            else:\n                hu_scans = convert_to_hu(scans)\n                \n            for patient in path:\n                b = hu_scans\n                np.savez(\"imgs.npz\", b)   \n        except Exception as e:\n            continue\n            ","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.251375Z","iopub.execute_input":"2024-03-02T08:37:06.251768Z","iopub.status.idle":"2024-03-02T08:37:06.260667Z","shell.execute_reply.started":"2024-03-02T08:37:06.251724Z","shell.execute_reply":"2024-03-02T08:37:06.259937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.262068Z","iopub.execute_input":"2024-03-02T08:37:06.262372Z","iopub.status.idle":"2024-03-02T08:37:06.284802Z","shell.execute_reply.started":"2024-03-02T08:37:06.262344Z","shell.execute_reply":"2024-03-02T08:37:06.283865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"column_indices = {name: i for i, name in enumerate(train_df.columns)}\n\nn = len(train_df)\n\ntrain_ds = train_df[0:int(n*0.7)]\nval_ds = train_df[int(n*0.7):int(n*0.9)]\ntest_ds = train_df[int(n*0.9):]\n\n# num_features = train_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.286833Z","iopub.execute_input":"2024-03-02T08:37:06.287244Z","iopub.status.idle":"2024-03-02T08:37:06.295019Z","shell.execute_reply.started":"2024-03-02T08:37:06.287202Z","shell.execute_reply":"2024-03-02T08:37:06.294213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_ds","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.296311Z","iopub.execute_input":"2024-03-02T08:37:06.296581Z","iopub.status.idle":"2024-03-02T08:37:06.326063Z","shell.execute_reply.started":"2024-03-02T08:37:06.296554Z","shell.execute_reply":"2024-03-02T08:37:06.325031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is the patch for \"all loss is nan\" in original notebook\ntrain_ds.drop('Patient', axis=1, inplace=True)\nval_ds.drop('Patient', axis=1, inplace=True)\ntest_ds.drop('Patient', axis=1, inplace=True)\n\ntrain_ds.drop('dcm_path', axis=1, inplace=True)\nval_ds.drop('dcm_path', axis=1, inplace=True)\ntest_ds.drop('dcm_path', axis=1, inplace=True)\ntrain_ds","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.327460Z","iopub.execute_input":"2024-03-02T08:37:06.327795Z","iopub.status.idle":"2024-03-02T08:37:06.366268Z","shell.execute_reply.started":"2024-03-02T08:37:06.327763Z","shell.execute_reply":"2024-03-02T08:37:06.365398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_features = train_ds.shape[1]","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.367647Z","iopub.execute_input":"2024-03-02T08:37:06.368033Z","iopub.status.idle":"2024-03-02T08:37:06.372490Z","shell.execute_reply.started":"2024-03-02T08:37:06.367991Z","shell.execute_reply":"2024-03-02T08:37:06.371627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_ds","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.373788Z","iopub.execute_input":"2024-03-02T08:37:06.374169Z","iopub.status.idle":"2024-03-02T08:37:06.397232Z","shell.execute_reply.started":"2024-03-02T08:37:06.374131Z","shell.execute_reply":"2024-03-02T08:37:06.396399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_ds","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.398690Z","iopub.execute_input":"2024-03-02T08:37:06.399017Z","iopub.status.idle":"2024-03-02T08:37:06.422654Z","shell.execute_reply.started":"2024-03-02T08:37:06.398981Z","shell.execute_reply":"2024-03-02T08:37:06.421700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe().transpose()\n\ntrain_mean = train_ds.mean(numeric_only=True)\ntrain_std = train_ds.std(numeric_only=True)\n\ntrain_df = (train_ds - train_mean) / train_std\nval_df = (val_ds - train_mean) / train_std\ntest_df = (test_ds - train_mean) / train_std","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.424056Z","iopub.execute_input":"2024-03-02T08:37:06.424479Z","iopub.status.idle":"2024-03-02T08:37:06.465395Z","shell.execute_reply.started":"2024-03-02T08:37:06.424429Z","shell.execute_reply":"2024-03-02T08:37:06.464589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.466872Z","iopub.execute_input":"2024-03-02T08:37:06.467267Z","iopub.status.idle":"2024-03-02T08:37:06.483471Z","shell.execute_reply.started":"2024-03-02T08:37:06.467228Z","shell.execute_reply":"2024-03-02T08:37:06.482367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.484990Z","iopub.execute_input":"2024-03-02T08:37:06.485391Z","iopub.status.idle":"2024-03-02T08:37:06.505055Z","shell.execute_reply.started":"2024-03-02T08:37:06.485359Z","shell.execute_reply":"2024-03-02T08:37:06.503964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.506631Z","iopub.execute_input":"2024-03-02T08:37:06.507032Z","iopub.status.idle":"2024-03-02T08:37:06.526124Z","shell.execute_reply.started":"2024-03-02T08:37:06.507000Z","shell.execute_reply":"2024-03-02T08:37:06.525039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class WindowGenerator():\n    def __init__(self, input_width, label_width, shift,\n                 train_df=train_df, val_df=val_df, test_df=test_df,\n                 label_columns=None):\n        \n        self.train_df = train_df\n        self.val_df = val_df\n        self.test_df = test_df\n        self.label_columns = label_columns\n        if label_columns is not None:\n            self.label_columns_indices = {name: i for i, name in enumerate(label_columns)}\n            \n        self.column_indices = {name: i for i, name in enumerate(train_df.columns)}\n        \n        self.input_width = input_width\n        self.label_width = label_width\n        self.shift = shift\n        \n        self.total_window_size = input_width + shift\n        \n        self.input_slice = slice(0, input_width)\n        self.input_indices = np.arange(self.total_window_size)[self.input_slice]\n        \n        self.label_start = self.total_window_size - self.label_width\n        self.labels_slice = slice(self.label_start, None)\n        self.label_indices = np.arange(self.total_window_size)[self.labels_slice]\n        \n    def __repr__(self):\n        return \"\\n\".join([\n            f\"TotalWindowSpread: {self.total_window_size}\",\n            f\"Total Indices: {self.input_indices}\",\n            f\"Label Indices: {self.label_indices}\",\n            f\"Label Name: {self.label_columns}\"])","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.527442Z","iopub.execute_input":"2024-03-02T08:37:06.527754Z","iopub.status.idle":"2024-03-02T08:37:06.541439Z","shell.execute_reply.started":"2024-03-02T08:37:06.527723Z","shell.execute_reply":"2024-03-02T08:37:06.540411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"w1 = WindowGenerator(input_width=30, label_width=1, shift=30, label_columns=[\"FVC\"])\nw1","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.542598Z","iopub.execute_input":"2024-03-02T08:37:06.542891Z","iopub.status.idle":"2024-03-02T08:37:06.559287Z","shell.execute_reply.started":"2024-03-02T08:37:06.542861Z","shell.execute_reply":"2024-03-02T08:37:06.558298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"w2 = WindowGenerator(input_width=5, label_width=1, shift=1, label_columns=[\"FVC\"])\nw2","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.560487Z","iopub.execute_input":"2024-03-02T08:37:06.560839Z","iopub.status.idle":"2024-03-02T08:37:06.571665Z","shell.execute_reply.started":"2024-03-02T08:37:06.560809Z","shell.execute_reply":"2024-03-02T08:37:06.570694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def split_window(self, features):\n    inputs = features[:, self.input_slice, :]\n    labels = features[:, self.labels_slice, :]\n    if self.label_columns is not None:\n        labels = tf.stack(\n            [labels[:, :, self.column_indices[name]] for name in self.label_columns],\n            axis=-1)\n        \n    inputs.set_shape([None, self.input_width, None])\n    labels.set_shape([None, self.label_width, None])\n    \n    return inputs, labels\n\nWindowGenerator.split_window = split_window","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.573109Z","iopub.execute_input":"2024-03-02T08:37:06.573488Z","iopub.status.idle":"2024-03-02T08:37:06.582237Z","shell.execute_reply.started":"2024-03-02T08:37:06.573449Z","shell.execute_reply":"2024-03-02T08:37:06.581407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ex_window = tf.stack([np.array(train_df[:w2.total_window_size]),\n                      np.array(train_df[:+w2.total_window_size]),\n                      np.array(train_df[:+w2.total_window_size])])\n\n\nex_inputs, ex_labels = w2.split_window(ex_window)\n\nprint(f\"Window shape: {ex_window.shape}\")\nprint(f\"Inputs shape: {ex_inputs.shape}\")\nprint(f\"Labels shape: {ex_labels.shape}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:06.583794Z","iopub.execute_input":"2024-03-02T08:37:06.584158Z","iopub.status.idle":"2024-03-02T08:37:09.094224Z","shell.execute_reply.started":"2024-03-02T08:37:06.584122Z","shell.execute_reply":"2024-03-02T08:37:09.093439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"w2.example = ex_inputs, ex_labels","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:09.095410Z","iopub.execute_input":"2024-03-02T08:37:09.095689Z","iopub.status.idle":"2024-03-02T08:37:09.099496Z","shell.execute_reply.started":"2024-03-02T08:37:09.095661Z","shell.execute_reply":"2024-03-02T08:37:09.098654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Plot the predictions in the figure.","metadata":{}},{"cell_type":"code","source":"def plot(self, model=None, plot_column=\"FVC\", max_subplots=3):\n    inputs, labels = self.example\n    plt.figure(figsize=(13, 8))\n    plot_column_index = self.column_indices[plot_column]\n    max_n = min(max_subplots, len(inputs))\n    for n in range(max_n):\n        plt.subplot(3, 1, n+1)\n        plt.ylabel(f\"{plot_column} [normed]\")\n        plt.plot(self.input_indices, inputs[n, :, plot_column_index],\n                 label=\"Inputs\", marker=\".\", zorder=-10)\n        \n        if self.label_columns:\n            label_column_index = self.label_columns_indices.get(plot_column, None)\n        else:\n            label_column_index = plot_column_index\n            \n        if label_column_index is None:\n            continue\n            \n        plt.scatter(self.label_indices, labels[n, :, label_column_index],\n                    edgecolors=\"k\", label=\"Labels\", c=\"#2ca02c\", s=64)\n        if model is not None:\n            predictions = model(inputs)\n            plt.scatter(self.label_indices, predictions[n, :, label_column_index],\n                        marker=\"X\", edgecolors=\"k\", label=\"Predictions\",\n                        c=\"#ff7f0e\", s=64)\n        \n        if n == 0:\n            plt.legend()\n        \n    plt.xlabel(\"Weeks\")\n    \nWindowGenerator.plot = plot","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:09.100708Z","iopub.execute_input":"2024-03-02T08:37:09.101013Z","iopub.status.idle":"2024-03-02T08:37:09.115536Z","shell.execute_reply.started":"2024-03-02T08:37:09.100984Z","shell.execute_reply":"2024-03-02T08:37:09.114732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"w2.plot()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:09.117018Z","iopub.execute_input":"2024-03-02T08:37:09.117400Z","iopub.status.idle":"2024-03-02T08:37:09.559595Z","shell.execute_reply.started":"2024-03-02T08:37:09.117356Z","shell.execute_reply":"2024-03-02T08:37:09.558608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def create_dataset(self, data):\n    data = np.array(data, dtype=np.float32)\n    ds = keras.preprocessing.timeseries_dataset_from_array(\n        data=data, targets=None, sequence_length=self.total_window_size,\n        sequence_stride=1, shuffle=False, batch_size=32)\n    \n    ds = ds.map(self.split_window)\n    \n    return ds\n\nWindowGenerator.create_dataset = create_dataset","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:09.560822Z","iopub.execute_input":"2024-03-02T08:37:09.561155Z","iopub.status.idle":"2024-03-02T08:37:09.567808Z","shell.execute_reply.started":"2024-03-02T08:37:09.561124Z","shell.execute_reply":"2024-03-02T08:37:09.566866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Here, the @property decorator is used to transform a class method into an attribute that can be accessed.","metadata":{}},{"cell_type":"code","source":"@property\ndef train(self):\n    return self.create_dataset(self.train_df)\n\n@property\ndef val(self):\n    return self.create_dataset(self.val_df)\n\n@property\ndef test(self):\n    return self.create_dataset(self.test_df)\n\n@property\ndef example(self):\n    result = getattr(self, \"_example\", None)\n    if result is None:\n        result = next(iter(self.train))\n        self._example = result\n    return result\n\nWindowGenerator.train = train\nWindowGenerator.val = val\nWindowGenerator.test = test\nWindowGenerator.example = example","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:09.569316Z","iopub.execute_input":"2024-03-02T08:37:09.569698Z","iopub.status.idle":"2024-03-02T08:37:09.579058Z","shell.execute_reply.started":"2024-03-02T08:37:09.569648Z","shell.execute_reply":"2024-03-02T08:37:09.578127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"w2.train.element_spec","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:09.580398Z","iopub.execute_input":"2024-03-02T08:37:09.580802Z","iopub.status.idle":"2024-03-02T08:37:09.957052Z","shell.execute_reply.started":"2024-03-02T08:37:09.580770Z","shell.execute_reply":"2024-03-02T08:37:09.956198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for example_inputs, example_labels in w2.train.take(1):\n    print(f\"exInput shape: {example_inputs.shape}\")\n    print(f\"exLabel shape: {example_labels.shape}\")\n    ","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:09.958364Z","iopub.execute_input":"2024-03-02T08:37:09.958636Z","iopub.status.idle":"2024-03-02T08:37:10.107831Z","shell.execute_reply.started":"2024-03-02T08:37:09.958609Z","shell.execute_reply":"2024-03-02T08:37:10.106833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"singlestep_wind = WindowGenerator(input_width=1, label_width=1, shift=1, label_columns=[\"FVC\"])\nsinglestep_wind","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:10.109492Z","iopub.execute_input":"2024-03-02T08:37:10.109798Z","iopub.status.idle":"2024-03-02T08:37:10.116814Z","shell.execute_reply.started":"2024-03-02T08:37:10.109766Z","shell.execute_reply":"2024-03-02T08:37:10.115864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for example_inputs, example_labels in singlestep_wind.train.take(1):\n    print(f\"Inputs: {example_inputs.shape}\")\n    print(f\"ouputs: {example_labels.shape}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:10.118489Z","iopub.execute_input":"2024-03-02T08:37:10.118877Z","iopub.status.idle":"2024-03-02T08:37:10.213898Z","shell.execute_reply.started":"2024-03-02T08:37:10.118834Z","shell.execute_reply":"2024-03-02T08:37:10.212878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Baseline(keras.Model):\n    def __init__(self, label_index=None):\n        super().__init__()\n        self.label_index = label_index\n        \n    def call(self, inputs):\n        if self.label_index is None:\n            return inputs\n        \n        result = inputs[:, :, self.label_index]\n        return result[:, :, tf.newaxis]","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:10.215781Z","iopub.execute_input":"2024-03-02T08:37:10.216103Z","iopub.status.idle":"2024-03-02T08:37:10.222744Z","shell.execute_reply.started":"2024-03-02T08:37:10.216072Z","shell.execute_reply":"2024-03-02T08:37:10.221559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"baseline = Baseline(label_index=column_indices[\"FVC\"])\n\nbaseline.compile(loss=keras.losses.MeanSquaredError(),\n                 metrics=[keras.metrics.Accuracy(), \"mse\"])\n\nval_performance = {}\nperformance = {}\nval_performance[\"Baseline\"] = baseline.evaluate(singlestep_wind.val)\nperformance[\"Baseline\"] = baseline.evaluate(singlestep_wind.test, verbose=0)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:37:10.224040Z","iopub.execute_input":"2024-03-02T08:37:10.224322Z","iopub.status.idle":"2024-03-02T08:37:10.755287Z","shell.execute_reply.started":"2024-03-02T08:37:10.224293Z","shell.execute_reply":"2024-03-02T08:37:10.754321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wide_wind = WindowGenerator(input_width=50, label_width=50, shift=1, label_columns=[\"FVC\"])\nwide_wind","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:52:30.362595Z","iopub.execute_input":"2024-03-02T08:52:30.362970Z","iopub.status.idle":"2024-03-02T08:52:30.369937Z","shell.execute_reply.started":"2024-03-02T08:52:30.362940Z","shell.execute_reply":"2024-03-02T08:52:30.369040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Input shape:\", singlestep_wind.example[0].shape)\nprint(\"output shape:\", baseline(singlestep_wind.example[0]).shape)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:52:32.327202Z","iopub.execute_input":"2024-03-02T08:52:32.327588Z","iopub.status.idle":"2024-03-02T08:52:32.334632Z","shell.execute_reply.started":"2024-03-02T08:52:32.327551Z","shell.execute_reply":"2024-03-02T08:52:32.333555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wide_wind.plot(baseline)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:52:34.291257Z","iopub.execute_input":"2024-03-02T08:52:34.291733Z","iopub.status.idle":"2024-03-02T08:52:34.825368Z","shell.execute_reply.started":"2024-03-02T08:52:34.291683Z","shell.execute_reply":"2024-03-02T08:52:34.824328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### LSTM model","metadata":{}},{"cell_type":"code","source":"max_epochs = 100\n\ndef compile_and_fit(model, window, patience=5):\n    early_stop = keras.callbacks.EarlyStopping(monitor=\"val_loss\",\n                                                  patience=patience,\n                                                  mode=\"min\")\n    \n    model.compile(loss=keras.losses.MeanSquaredError(),\n                  optimizer=keras.optimizers.Adam(),\n                  metrics=[\"mse\",])\n    \n    history = model.fit(window.train,\n                        epochs = max_epochs,\n                        validation_data = window.val, \n                        callbacks=[early_stop]\n                       )\n    \n    return history","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:52:53.938583Z","iopub.execute_input":"2024-03-02T08:52:53.938977Z","iopub.status.idle":"2024-03-02T08:52:53.946680Z","shell.execute_reply.started":"2024-03-02T08:52:53.938944Z","shell.execute_reply":"2024-03-02T08:52:53.945547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lstm_1 = keras.models.Sequential([\n    keras.layers.LSTM(units=64, return_sequences=True),\n    keras.layers.Dense(units=1)])","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:04.235735Z","iopub.execute_input":"2024-03-02T08:53:04.236125Z","iopub.status.idle":"2024-03-02T08:53:04.252093Z","shell.execute_reply.started":"2024-03-02T08:53:04.236089Z","shell.execute_reply":"2024-03-02T08:53:04.251144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Lstm Input shape:\", wide_wind.example[0].shape)\nprint(\"Lstm Output shape:\", lstm_1(wide_wind.example[0]).shape)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:06.279384Z","iopub.execute_input":"2024-03-02T08:53:06.279723Z","iopub.status.idle":"2024-03-02T08:53:06.518827Z","shell.execute_reply.started":"2024-03-02T08:53:06.279695Z","shell.execute_reply":"2024-03-02T08:53:06.517999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = compile_and_fit(lstm_1, wide_wind)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:12.494409Z","iopub.execute_input":"2024-03-02T08:53:12.494794Z","iopub.status.idle":"2024-03-02T08:53:19.499101Z","shell.execute_reply.started":"2024-03-02T08:53:12.494755Z","shell.execute_reply":"2024-03-02T08:53:19.498345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_performance[\"LSTM\"] = lstm_1.evaluate(wide_wind.val)\nperformance[\"LSTM\"] = lstm_1.evaluate(wide_wind.test, verbose=0)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:29.673150Z","iopub.execute_input":"2024-03-02T08:53:29.673584Z","iopub.status.idle":"2024-03-02T08:53:29.898521Z","shell.execute_reply.started":"2024-03-02T08:53:29.673550Z","shell.execute_reply":"2024-03-02T08:53:29.897766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wide_wind.plot(lstm_1)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:32.690410Z","iopub.execute_input":"2024-03-02T08:53:32.690750Z","iopub.status.idle":"2024-03-02T08:53:33.184332Z","shell.execute_reply.started":"2024-03-02T08:53:32.690720Z","shell.execute_reply":"2024-03-02T08:53:33.183268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output_steps = 50\nmulti_window = WindowGenerator(input_width=50,\n                               label_width=output_steps,\n                               shift=output_steps)\n\nmulti_window.plot()\nmulti_window","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:40.555334Z","iopub.execute_input":"2024-03-02T08:53:40.555743Z","iopub.status.idle":"2024-03-02T08:53:41.046308Z","shell.execute_reply.started":"2024-03-02T08:53:40.555705Z","shell.execute_reply":"2024-03-02T08:53:41.045435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class MultiStepLastBaseline(keras.Model):\n    def call(self, inputs):\n        return tf.tile(inputs[:, -1:, :], [1, output_steps, 1])\n    \n\nlast_baseline = MultiStepLastBaseline()\nlast_baseline.compile(loss=keras.losses.MeanSquaredError(),\n                      metrics=[\"mse\", keras.metrics.Accuracy()])\n\n\nmulti_val_performance = {}\nmulti_performance = {}\n\nmulti_val_performance[\"Last\"] = last_baseline.evaluate(multi_window.val)\nmulti_performance[\"Last\"] = last_baseline.evaluate(multi_window.test, verbose=0)\nmulti_window.plot(last_baseline)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:46.039024Z","iopub.execute_input":"2024-03-02T08:53:46.039371Z","iopub.status.idle":"2024-03-02T08:53:46.812360Z","shell.execute_reply.started":"2024-03-02T08:53:46.039342Z","shell.execute_reply":"2024-03-02T08:53:46.811028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class RepeatBaseline(keras.Model):\n    def call(self, inputs):\n        return inputs\n\n    \nrepeat_baseline = RepeatBaseline()\nrepeat_baseline.compile(loss=keras.losses.MeanSquaredError(),\n                        metrics=[\"mse\",])\n\nmulti_val_performance[\"Repeat\"] = repeat_baseline.evaluate(multi_window.val)\nmulti_performance[\"Repeat\"] = repeat_baseline.evaluate(multi_window.test, verbose=0)\nmulti_window.plot(repeat_baseline)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:52.579033Z","iopub.execute_input":"2024-03-02T08:53:52.579385Z","iopub.status.idle":"2024-03-02T08:53:53.308001Z","shell.execute_reply.started":"2024-03-02T08:53:52.579355Z","shell.execute_reply":"2024-03-02T08:53:53.307054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"multi_lstm_1 = keras.models.Sequential([\n    keras.layers.LSTM(32, return_sequences=False),\n    keras.layers.Dense(output_steps*num_features),\n    keras.layers.Reshape([output_steps, num_features])])\n\nhistory = compile_and_fit(multi_lstm_1, multi_window)\n\n# IPython.display.clear_output()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:53:59.841047Z","iopub.execute_input":"2024-03-02T08:53:59.841424Z","iopub.status.idle":"2024-03-02T08:54:05.426767Z","shell.execute_reply.started":"2024-03-02T08:53:59.841389Z","shell.execute_reply":"2024-03-02T08:54:05.425968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"multi_val_performance[\"LSTM\"] = multi_lstm_1.evaluate(multi_window.val)\nmulti_performance[\"LSTM\"] = multi_lstm_1.evaluate(multi_window.test, verbose=0)\nmulti_window.plot(multi_lstm_1)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:09.579464Z","iopub.execute_input":"2024-03-02T08:54:09.579805Z","iopub.status.idle":"2024-03-02T08:54:10.248259Z","shell.execute_reply.started":"2024-03-02T08:54:09.579776Z","shell.execute_reply":"2024-03-02T08:54:10.247274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class FeedBack(keras.Model):\n    def __init__(self, units, output_steps):\n        super().__init__()\n        self.output_steps = output_steps\n        self.units = units\n        self.lstm_cell = keras.layers.LSTMCell(units)\n        self.lstm_rnn = keras.layers.RNN(self.lstm_cell, return_state=True)\n        self.dense = keras.layers.Dense(num_features)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:15.029321Z","iopub.execute_input":"2024-03-02T08:54:15.029709Z","iopub.status.idle":"2024-03-02T08:54:15.036676Z","shell.execute_reply.started":"2024-03-02T08:54:15.029672Z","shell.execute_reply":"2024-03-02T08:54:15.035720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feedback_lstm = FeedBack(units=32, output_steps=output_steps)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:18.472406Z","iopub.execute_input":"2024-03-02T08:54:18.472751Z","iopub.status.idle":"2024-03-02T08:54:18.486694Z","shell.execute_reply.started":"2024-03-02T08:54:18.472722Z","shell.execute_reply":"2024-03-02T08:54:18.485793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def warmup(self, inputs):\n    x, *state = self.lstm_rnn(inputs)\n    \n    prediction = self.dense(x)\n    \n    return prediction, state\n\nFeedBack.warmup = warmup","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:19.838909Z","iopub.execute_input":"2024-03-02T08:54:19.839299Z","iopub.status.idle":"2024-03-02T08:54:19.844996Z","shell.execute_reply.started":"2024-03-02T08:54:19.839266Z","shell.execute_reply":"2024-03-02T08:54:19.844013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"prediction, state = feedback_lstm.warmup(multi_window.example[0])\nprediction.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:21.933852Z","iopub.execute_input":"2024-03-02T08:54:21.934227Z","iopub.status.idle":"2024-03-02T08:54:21.994389Z","shell.execute_reply.started":"2024-03-02T08:54:21.934193Z","shell.execute_reply":"2024-03-02T08:54:21.993510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def call(self, inputs, training=None):\n    predictions = []\n    \n    prediction, state = self.warmup(inputs)\n    \n    predictions.append(prediction)\n    \n    for n in range(1, self.output_steps):\n        x = prediction\n        x, state = self.lstm_cell(x, states=state, training=training)\n        \n        prediction = self.dense(x)\n        predictions.append(prediction)\n        \n    predictions = tf.stack(predictions)\n    predictions = tf.transpose(predictions, [1, 0, 2])\n    return predictions\n    \nFeedBack.call = call","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:24.932970Z","iopub.execute_input":"2024-03-02T08:54:24.933297Z","iopub.status.idle":"2024-03-02T08:54:24.941496Z","shell.execute_reply.started":"2024-03-02T08:54:24.933269Z","shell.execute_reply":"2024-03-02T08:54:24.940525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Batch, Time, Features:\", feedback_lstm(multi_window.example[0]).shape)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:27.227744Z","iopub.execute_input":"2024-03-02T08:54:27.228139Z","iopub.status.idle":"2024-03-02T08:54:27.312109Z","shell.execute_reply.started":"2024-03-02T08:54:27.228108Z","shell.execute_reply":"2024-03-02T08:54:27.311168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = compile_and_fit(feedback_lstm, multi_window)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:28.836938Z","iopub.execute_input":"2024-03-02T08:54:28.837406Z","iopub.status.idle":"2024-03-02T08:54:53.640847Z","shell.execute_reply.started":"2024-03-02T08:54:28.837361Z","shell.execute_reply":"2024-03-02T08:54:53.640041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"multi_val_performance[\"AR LSTM\"] = feedback_lstm.evaluate(multi_window.val)\nmulti_performance[\"AR LSTM\"] = feedback_lstm.evaluate(multi_window.test, verbose=0)\nmulti_window.plot(feedback_lstm)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:53.644491Z","iopub.execute_input":"2024-03-02T08:54:53.644789Z","iopub.status.idle":"2024-03-02T08:54:54.609580Z","shell.execute_reply.started":"2024-03-02T08:54:53.644760Z","shell.execute_reply":"2024-03-02T08:54:54.608795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for name, value in multi_performance.items():\n    print(f\"{name:8s}: {value[1]:0.4f}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:54:54.610772Z","iopub.execute_input":"2024-03-02T08:54:54.611066Z","iopub.status.idle":"2024-03-02T08:54:54.616548Z","shell.execute_reply.started":"2024-03-02T08:54:54.611037Z","shell.execute_reply":"2024-03-02T08:54:54.615516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:38:01.860060Z","iopub.execute_input":"2024-03-02T08:38:01.860376Z","iopub.status.idle":"2024-03-02T08:38:01.879167Z","shell.execute_reply.started":"2024-03-02T08:38:01.860337Z","shell.execute_reply":"2024-03-02T08:38:01.878298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:38:01.880530Z","iopub.execute_input":"2024-03-02T08:38:01.880874Z","iopub.status.idle":"2024-03-02T08:38:01.899718Z","shell.execute_reply.started":"2024-03-02T08:38:01.880845Z","shell.execute_reply":"2024-03-02T08:38:01.898874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def score(FVC_true, FVC_pred, sigma):\n    sigma_clipped = np.max(sigma, 70)\n    delta = np.abs(FVC_true - FVC_pred)\n    delta = np.min(delta, 1000)\n    sq_2 = np.sqrt(2)\n    metric = (delta / sigma_clipped) * sq_2 + (np.log(sigma_clipped * sq_2))\n    \n    return np.mean(metric)","metadata":{"execution":{"iopub.status.busy":"2024-03-02T08:38:01.901052Z","iopub.execute_input":"2024-03-02T08:38:01.901384Z","iopub.status.idle":"2024-03-02T08:38:01.909586Z","shell.execute_reply.started":"2024-03-02T08:38:01.901337Z","shell.execute_reply":"2024-03-02T08:38:01.908843Z"},"trusted":true},"execution_count":null,"outputs":[]}]}