{"cells":[{"metadata":{"_uuid":"9f53f0c55752e96bfc1373e632314d42ce0fc076"},"cell_type":"markdown","source":"**Summary**\n\nPython generators can be used to handle large amounts of data when only a relatively small amount of RAM is available.\n\nThis kernel uses all 25,684 full size images to train and validate a simple keras cnn. The code is able to run in a kaggle kernel (with GPU) without exceeding the 14GB RAM limit.\n\nTo achieve this I used custom python generators with a batch size of 10.\n\nUsing Generators may be helpful to those who want to take part in this interesting competition but can't because their code keeps crashing due to limited memory.\n\nThe pickled input data used in this notebook was created in Part 1. A simple explanation of how a generator works is included in the appendix section.\n<hr>"},{"metadata":{"trusted":true,"_uuid":"b17898014ec51febdf54b48c1a56b642045fcda5"},"cell_type":"code","source":"from numpy.random import seed\nseed(101)\nfrom tensorflow import set_random_seed\nset_random_seed(101)\n\nimport pandas as pd\nimport numpy as np\nimport math\nimport pydicom\nimport pylab\nimport os\nimport pickle\n\nfrom sklearn.model_selection import train_test_split\nfrom skimage.transform import resize\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n\n# Don't Show Warning Messages\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport gc; gc.enable()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"af99b6f8f87904b0eb05c4cf0477de4ebb4d079e"},"cell_type":"code","source":"os.listdir('../input')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3d1da875428771d4604ba91ca0cc0d0b958f85f8"},"cell_type":"markdown","source":"### What will be X, the model input?"},{"metadata":{"_uuid":"b1a6ad8d0b3eb7f8884ff8451c111d4d30bf3379"},"cell_type":"markdown","source":"The model input will be the images as 2D numpy arrays.\n\nThe size of each image is 1024x1024. Each image will  be read from the folder and converted into a 1024x1024 numpy array. Because there are 25,684 unique images, trying to read them all into a single numpy array will exceed the 14GB of RAM that is available on kaggle kernels (when GPU is on). \n\nTo solve this problem we will only read a batch of 10 images at a time into memory. Then we will feed this batch to the model. Once the model has eaten up a batch, that batch will disapper from memory. To achieve this memory saving batch effect we will create something called a Generator. When feeding the model we will use fit_generator() and predict_generator() instead of the usual fit() and predict() methods. \n\nNote that  keras needs an input that has the following shape: <br>\n(num_samples, image_size, image_size, num_channels). \n\nThe input batch will have the shape (10, 1024, 1024, 1). Also, in order to use fit_generator() the training and validation generators that we build must be able to loop infinitely. The appendix  has an example of how to do this - as well as quick explanation of what python generators are.\n"},{"metadata":{"_uuid":"8608667a3fc48074e9ad0420e74bb60e4a46c3ce"},"cell_type":"markdown","source":"### What will be y, the target?"},{"metadata":{"_uuid":"fb47184e62aa19dfb7ea5fe222c040a394a99ac6"},"cell_type":"markdown","source":"We will treat this as a regression problem.\n\nTo keep this example simple, we will set up the model to predict a max of 2 bounding boxes.\n\nThe model will need to predict an output consisting of 10 columns.\n\n**This is the list of output columns:**\n\npredictions = <br>\n['conf_1', 'x_1', 'y_1', 'width_1', 'height_1', <br>\n'conf_2', 'x_2', 'y_2', 'width_2', 'height_2'] \n\n**Key:**\n\n* 'conf_1' = binary, bounding box confidence score\n* 'x_1' = x coordinate (top left corner of bounding box)\n* 'y_1' = y coordinate (top left corner of bounding box)\n* 'width_1' = width of bounding box\n* 'height_1' = height of bounding box\n"},{"metadata":{"_uuid":"54e6359ff49aa108acbc722744d97d282dd08b4d"},"cell_type":"markdown","source":"### What will be the loss function?"},{"metadata":{"_uuid":"dd4e4e39668d2a7b806d06fb2950c5faf37ba886"},"cell_type":"markdown","source":"We will use the 'mse' loss function because this is a regression task. \n"},{"metadata":{"_uuid":"7808523434d1739e631760132b22731aac5c5dd8"},"cell_type":"markdown","source":"### What steps will we follow?"},{"metadata":{"_uuid":"6093eb4bf6cd6ede778f5585ab81e7caf3794535"},"cell_type":"markdown","source":"1. Pre-process the data.<br>\nThis has been done in a seperate kernel. We will use the output of that kernel.\n2. Create the output dataframe, df_y\n3. Train test split<br>\n4. Build the Generators for the training data, validation data and the prediction data.\n5. Create the model architecture, train the model and make predictions.\n6. Post process the predictions to get them into submission format.\n7. Create the submission csv file."},{"metadata":{"_uuid":"5d5ff95fbe855cec13d0511d5738f06353b16032"},"cell_type":"markdown","source":"### Load the pre processed data"},{"metadata":{"trusted":true,"_uuid":"4c24c97d3f19545fcf6d66b5d9de8013f7cfefd6"},"cell_type":"code","source":"# load the pickled dataframes\n\ndf_train = pickle.load(open('../input/python-generators-to-reduce-ram-usage-part-1/dftrain.pickle','rb'))\ndf_test = pickle.load(open('../input/python-generators-to-reduce-ram-usage-part-1/dftest.pickle','rb'))\n\n\nprint(df_train.shape)\nprint(df_test.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dc83edf929e17a69671efef33834966a28f036f9"},"cell_type":"markdown","source":"### Let's start building df_y"},{"metadata":{"trusted":true,"_uuid":"a1f3122cc9cba570b8ff415b2c036d26d1b8a4b6"},"cell_type":"code","source":"# Source: https://www.kaggle.com/peterchang77/exploratory-data-analysis\n\ndef parse_data(df):\n    \"\"\"\n    Method to read a CSV file (Pandas dataframe) and parse the \n    data into the following nested dictionary:\n\n      parsed = {\n        \n        'patientId-00': {\n            'dicom': path/to/dicom/file,\n            'label': either 0 or 1 for normal or pnuemonia, \n            'boxes': list of box(es)\n        },\n        'patientId-01': {\n            'dicom': path/to/dicom/file,\n            'label': either 0 or 1 for normal or pnuemonia, \n            'boxes': list of box(es)\n        }, ...\n\n      }\n\n    \"\"\"\n    # --- Define lambda to extract coords in list [y, x, height, width]\n    extract_box = lambda row: [row['y'], row['x'], row['height'], row['width']]\n\n    parsed = {}\n    for n, row in df.iterrows():\n        # --- Initialize patient entry into parsed \n        pid = row['patientId']\n        if pid not in parsed:\n            parsed[pid] = {\n                'dicom': '../input/stage_1_train_images/%s.dcm' % pid,\n                'label': row['Target'],\n                'boxes': []}\n\n        # --- Add box if opacity is present\n        if parsed[pid]['label'] == 1:\n            parsed[pid]['boxes'].append(extract_box(row))\n\n    return parsed","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"eb53a5ef30080ae4ec5d3ef09ba1d8afcbe87e29"},"cell_type":"code","source":"# define a function to output a row containing all box info incl. confidence scores\n\ndef create_bounding_rows(df_train):\n    \n    \"\"\"\n    Takes each patientId and creates a row of combined bounding boxes and \n    also includes their confidence scores. All patientId's are \n    included in one matrix.\n    This fuction is based on a max of 4 bounding boxes per patientId.\n    Output: Numpy matrix of shape(len(df_train), 20) \n    \"\"\"\n    \n    # read in the dataframe that will be parsed by the function parse_data(df)\n    df_boxes = \\\n    pd.read_csv('../input/rsna-pneumonia-detection-challenge/stage_1_train_labels.csv')\n\n    \n    # set the length depending on how many bounding boxes we want the model to output\n    length = 20\n    \n    h = np.ones(20)\n    k = np.zeros(20)\n\n    # create an empty numpy matrix matching the size of the output matrix\n    y = np.zeros((len(df_train),length))\n\n    # run the function\n    # this must be here because this must be run each time this script is run or\n    # the resulting matrix will have errors.\n    parsed = parse_data(df_boxes)\n\n\n    for i in range(0,len(df_train)):\n\n        # get the patientId\n        patientId = df_train.loc[i, 'patientId']\n\n        # extract the bounding boxes for a particular patient\n        box = parsed[patientId]['boxes']\n        if len(box) == 0:\n\n            # the first row becomes a dummy row of ones this must be deleted later\n            # k is an array of zeros\n            h = np.vstack((h,k))\n\n        if len(box) != 0:\n\n\n            # insert 1 as the first entry in each bounding box\n            # the 1 represents confidence for that bounding box\n            a=[]\n            for i in range(0,len(box)):\n                box[i].insert(0,1)\n                a = a + box[i]\n\n            # calculate how much padding to add\n            b = length - len(a)\n\n            # pad the list because not all lists have 4 bounding boxes\n            # we want all lists to have the same length\n            for i in range(0,b):\n                a.insert(len(a),0)\n\n            # reshape to horizontal because the above code makes the list vertical\n            a = np.array(a).reshape(1,length)\n            \n            # stack\n            h = np.vstack((h,a))\n\n    # delete the first row because we added this row just to make the code run\n    h = np.delete(h, 0, axis=0)\n    \n    return h\n\n\n# call the function\nbox_rows = create_bounding_rows(df_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c7094495e9ecf618e22fa036abab2c71b9c9214c"},"cell_type":"code","source":"# concat box_rows with df_y\n\n# put box_rows in a dataframe\ndf_y = pd.DataFrame(box_rows)\n\n# rename the columns in df_box_rows\nnew_names = ['conf_1', 'x_1', 'y_1', 'width_1', 'height_1',\n           'conf_2', 'x_2', 'y_2', 'width_2', 'height_2',\n           'conf_3', 'x_3', 'y_3', 'width_3', 'height_3',\n           'conf_4', 'x_4', 'y_4', 'width_4', 'height_4']\n\ndf_y.columns = new_names\n\n# Let's choose only the first two bounding boxes for each sample\ndf_y = df_y[['conf_1', 'x_1', 'y_1', 'width_1', 'height_1',\n           'conf_2', 'x_2', 'y_2', 'width_2', 'height_2']]\n\n# add the patientId column to df_y\ndf_y['patientId'] = df_train['patientId']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a0b698e7198c17e1b3a632bc1b193edd6a390c3b"},"cell_type":"code","source":"df_y.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"837479608d49ec5371c74cc7ade8756f0636e989"},"cell_type":"code","source":"df_y.head(2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f0f79fb5e9c396ca07d4bdf0b80917b00150df4d"},"cell_type":"markdown","source":"### Train_Test_Split"},{"metadata":{"trusted":true,"_uuid":"7a128e9675a27e9e47d46684f487f4808a1bab9c"},"cell_type":"code","source":"# shuffle df_y\nfrom sklearn.utils import shuffle\n\ndf_y = shuffle(df_y)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"41c3e656fdfb5839ab8301130a28ca351ae95309"},"cell_type":"code","source":"df_train_images, df_val_images = train_test_split(df_y, test_size=0.20,\n                                                   random_state=5)\n\nprint(df_train_images.shape)\nprint(type(df_train_images))\nprint(df_val_images.shape)\nprint(type(df_val_images))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e06d75c72d8bff3cd642e0219dc89255923b5e1f"},"cell_type":"code","source":"# Reset the index of df_train_images and df_val_images.\n\n# We do this because we are going to loop through these dataframes in the next step so\n# we need the index to be sequential, starting from 0.\n\ndf_train_images.reset_index(inplace=True)\n\ndf_val_images.reset_index(inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f0e93020796243745d06d7ab5ee102342e8fef6c"},"cell_type":"code","source":"# Create a version without any unnecessary columns.\n\ndf_train = df_train_images.drop(['index', 'patientId'], axis=1)\n\ndf_val = df_val_images.drop(['index','patientId'], axis=1)\n\n# check that we have only 10 columns\nprint(df_train.shape)\nprint(df_val.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"42ba04aa50e62410135398a62428a9200998d261"},"cell_type":"code","source":"df_train_images.head(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b2514e649297d7025e3e032aeaf9a4b89a1e5c0a"},"cell_type":"code","source":"df_val_images.head(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"40066f79d4538df3d34bf68e57b2c41f999fdfa4"},"cell_type":"code","source":"df_train.head(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"71b985b8863aa0433eba703a341be3293ee11c4b"},"cell_type":"code","source":"df_val.head(1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eb8a73912a7a1b968f8bdd78d15027683a2c6c9b"},"cell_type":"markdown","source":"### Create the Data Generators"},{"metadata":{"_uuid":"cff4b184b0e7aabee68adb9d44057a23d6d4b8b9"},"cell_type":"markdown","source":"If you find Generators a bit confusing, I've included a simple explanation in the Appendix."},{"metadata":{"trusted":true,"_uuid":"4521b42f9e53dd62bb1d5c95e5f7d5b0c4d700d1"},"cell_type":"code","source":"# We have 20547 train images and 5137 validation images.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"895ebc1a7e969adf4fd35138b4e3f0e3a0a077b4"},"cell_type":"markdown","source":"### [1] Train Generator"},{"metadata":{"trusted":true,"_uuid":"4cf975dd11f53ccad6d177421561e070866622ab"},"cell_type":"code","source":"def train_generator(df_train_images, df_train, batch_size, num_rows, num_cols):\n    \n    '''\n    Input: Dataframes, df_train_images and df_train\n    \n    Outputs one batch (X_train, y_train) on each iteration of the for loop.\n    \n    X_train:\n    Reads images from a folder, converts the images to a numpy array \n    with shape: (batch_size, num_rows, num_cols, 1)\n    \n    y_train:\n    Takes data from a pandas dataframe. Converts the data into a numpy array\n    with shape (batch_size, num_rows, num_cols, 1)\n    \n    '''\n    \n    \n    while True: \n\n        batch = []\n        k = 0\n\n\n        # note that we are rounding down.\n        num_batches = math.ceil(df_train_images.shape[0]/batch_size)\n\n        # create an empty numpy array matching the number of images\n        image_array = np.zeros((batch_size,num_rows,num_cols))\n\n\n\n        # this loop runs only once each time the next() function is called.\n        for i in range(0,num_batches): # 20547 rows in train_images. we are using only 20000 of them\n\n            if i < num_batches-1:\n\n                # [1] Create X_train\n\n                # carve out 1000 rows of the 'patientId' column\n                batch = list(df_train_images['patientId'][k:(i+1)*batch_size])\n\n                #for patientId in batch:\n                for j in range(0,len(batch)):\n                    patientId = batch[j]\n\n\n                    path = \\\n                '../input/rsna-pneumonia-detection-challenge/stage_1_train_images/%s.dcm' % patientId\n\n                    dcm_data = pydicom.read_file(path)\n\n                    # get the image as a numpy array\n                    image = dcm_data.pixel_array\n\n                    # resize the image\n                    small_image = resize(image,(num_rows,num_cols))\n\n                    # add the image to the empty numpy array\n                    image_array[j,:,:] = small_image\n\n                # reshape the array and normalize\n                X_train = image_array.reshape(batch_size,num_rows,num_cols,1)/255\n\n                # [2] Create y_train\n\n                # note: Here we use df_train instead of df_train_images\n                # because we don't want the output to have the patientId column.\n\n                # carve out 1000 rows\n                y_train = df_train[k:(i+1)*batch_size]\n\n                # convert to a numpy array\n                y_train = y_train.values\n\n            # to cater for the last batch i.e. the fractional part\n            if i == num_batches-1: \n\n                batch_size_fractional = df_train.shape[0] - (batch_size*(num_batches-1)) # -1\n\n                # create an empty numpy array matching the number of images\n                image_array = np.zeros((batch_size_fractional,num_rows,num_cols))\n\n                # select rows from the tail of df_test upwards\n                batch1 = list(df_train_images['patientId'][-batch_size_fractional:]) #1000\n\n                #for patientId in batch:\n                for j in range(0,len(batch1)):\n                    patientId = batch1[j]\n\n                    path = \\\n            '../input/rsna-pneumonia-detection-challenge/stage_1_train_images/%s.dcm' % patientId\n\n                    dcm_data = pydicom.read_file(path)\n\n                    # get the image as a numpy array\n                    image = dcm_data.pixel_array\n\n                    # resize the image\n                    small_image = resize(image,(num_rows,num_cols))\n\n                    # add the image to the empty numpy array\n                    image_array[j,:,:] = small_image\n\n                # reshape the array and normalize\n                X_train = image_array.reshape(batch_size_fractional,num_rows,num_cols,1)/255\n\n                # [2] Create y_train\n\n                # note: Here we use df_val instead of df_val_images\n                # because we don't want the output to have the patientId column.\n\n                # carve out 1000 rows\n                y_train = df_train[-batch_size_fractional:]\n\n                # convert to a numpy array\n                y_train = y_train.values\n\n\n            k = k + batch_size\n\n            # For testing the generator so we can see how many batches it outputs\n            # by calling next(). Uncomment the next line for testing.\n            #print(i)\n\n            # Keras requires a tuple in the form (inputs,targets)\n            yield (X_train.astype(np.float32), y_train)\n            \n    \n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e0e63e3ca389bced814e7d05d69836bc2c77f0e6"},"cell_type":"markdown","source":"### [2] Validation Generator"},{"metadata":{"trusted":true,"_uuid":"1381ba27da9ff7de153412960d55c83eba3b6155"},"cell_type":"code","source":"def val_generator(df_val_images, df_val, batch_size, num_rows, num_cols):\n    \n    '''\n    Input: Dataframes, df_val_images and df_val\n    \n    Outputs one batch (X_val, y_val) on each iteration of the for loop.\n    \n    X_val:\n    Reads images from a folder, converts the images to a numpy array \n    with shape: (batch_size, num_rows, num_cols, 1)\n    \n    y_val:\n    Takes data from a pandas dataframe. Converts the data into a numpy array\n    with shape (batch_size, num_rows, num_cols, 1)\n    \n    '''\n    \n    \n    while True: \n\n        batch = []\n        k = 0\n\n        # note that we are rounding up.\n        num_batches = math.ceil(df_val_images.shape[0]/batch_size)\n\n        # Create an empty numpy array that matches the batch size.\n        image_array = np.zeros((batch_size,num_rows,num_cols))\n\n\n         # this loop runs only once each time the next() function is called.\n        for i in range(0,num_batches): \n            \n            if i < num_batches-1:\n\n                # [1] Create X_train\n\n                # carve out a batch of rows of the 'patientId' column\n                batch = list(df_val_images['patientId'][k:(i+1)*batch_size])\n\n                #for patientId in batch:\n                for j in range(0,len(batch)):\n                    patientId = batch[j]\n\n                    path = \\\n            '../input/rsna-pneumonia-detection-challenge/stage_1_train_images/%s.dcm' % patientId\n\n                    dcm_data = pydicom.read_file(path)\n\n                    # get the image as a numpy array\n                    image = dcm_data.pixel_array\n\n                    # resize the image\n                    small_image = resize(image,(num_rows,num_cols))\n\n                    # add the image to the empty numpy array\n                    image_array[j,:,:] = small_image\n\n                # reshape the array and normalize\n                X_val = image_array.reshape(batch_size,num_rows,num_cols,1)/255\n\n                # [2] Create y_train\n\n                # note: Here we use df_val instead of df_val_images\n                # because we don't want the output to have the patientId column.\n\n                # carve out 1000 rows\n                y_val = df_val[k:(i+1)*batch_size]\n\n                # convert to a numpy array\n                y_val = y_val.values\n\n             # to cater for the last batch i.e. the fractional part\n            if i == num_batches-1: \n\n                batch_size_fractional = df_val.shape[0] - (batch_size*(num_batches-1)) \n\n                # create an empty numpy array matching the number of images\n                image_array = np.zeros((batch_size_fractional,num_rows,num_cols))\n\n                # select rows from the tail of df_test upwards\n                batch1 = list(df_val_images['patientId'][-batch_size_fractional:]) \n\n                #for patientId in batch:\n                for j in range(0,len(batch1)):\n                    patientId = batch1[j]\n\n                    path = \\\n            '../input/rsna-pneumonia-detection-challenge/stage_1_train_images/%s.dcm' % patientId\n\n                    dcm_data = pydicom.read_file(path)\n\n                    # get the image as a numpy array\n                    image = dcm_data.pixel_array\n\n                    # resize the image\n                    small_image = resize(image,(num_rows,num_cols))\n\n                    # add the image to the empty numpy array\n                    image_array[j,:,:] = small_image\n\n                # reshape the array and normalize\n                X_val = image_array.reshape(batch_size_fractional,num_rows,num_cols,1)/255\n\n                # [2] Create y_train\n\n                # note: Here we use df_val instead of df_val_images\n                # because we don't want the output to have the patientId column.\n\n                # carve out a batch of rows\n                y_val = df_val[-batch_size_fractional:]\n\n                # convert to a numpy array\n                y_val = y_val.values\n\n\n            k = k + batch_size\n\n            # For testing the generator so we can see how many batches it outputs\n            # by calling next().\n            #print(i)\n\n            # Keras requires a tuple in the form (inputs,targets)\n            yield (X_val.astype(np.float32), y_val)\n            \n           \n    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"345610915aef2da14c2e7586f374ab52cad4031a"},"cell_type":"markdown","source":"### [3] Test Generator"},{"metadata":{"trusted":true,"_uuid":"595c7d7d114fbfaefe4354a3ffeeebe707df6d7e"},"cell_type":"code","source":"df_test.head(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1c3992488cc899d0b2a7471b6e31d6aef2f64dbe"},"cell_type":"code","source":"df_test.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"262855c41155779b7333bdafe50d713335fbe834"},"cell_type":"code","source":"# There are 1000 rows in df_test i.e. 1000 test images","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ec4c5afc03f7eb545a2d29c4017375d5ee728c02"},"cell_type":"code","source":"def test_generator(df_test, batch_size, num_rows, num_cols):\n    \n    \"\"\"\n    Input: Dataframe df_test.\n    \n    Outputs one batch (X_test) on each iteration of the for loop.\n    \n    X_test:\n    Reads images from a folder, converts the images to a numpy array \n    with shape: (batch_size, num_rows, num_cols, 1)\n    \n    \"\"\"\n\n    batch = []\n    k = 0\n    \n    # note that we are rounding up.\n    num_batches = math.ceil(df_test.shape[0]/batch_size)\n\n    # create an empty numpy array matching the number of images\n    image_array = np.zeros((batch_size,num_rows,num_cols))\n    \n    # this loop runs only once each time the next() function is called.\n    for i in range(0,num_batches):\n        \n        if i < num_batches-1:\n        \n            # [1] Create X_test\n\n            # carve out a batch of rows of the 'patientId' column\n            batch = list(df_test['patientId'][k:(i+1)*batch_size]) #1000\n\n            #for patientId in batch:\n            for j in range(0,len(batch)):\n                patientId = batch[j]\n\n                path = \\\n        '../input/rsna-pneumonia-detection-challenge/stage_1_test_images/%s.dcm' % patientId\n\n                dcm_data = pydicom.read_file(path)\n\n                # get the image as a numpy array\n                image = dcm_data.pixel_array\n\n                # resize the image\n                small_image = resize(image,(num_rows,num_cols))\n\n                # add the image to the empty numpy array\n                image_array[j,:,:] = small_image\n\n            # reshape the array and normalize\n            X_test = image_array.reshape(batch_size,num_rows,num_cols,1)/255\n            \n        # to cater for the last batch i.e. the fractional part\n        if i == num_batches-1: \n            \n            batch_size_fractional = df_test.shape[0] - (batch_size*(num_batches - 1))\n            \n            # create an empty numpy array matching the number of images\n            image_array = np.zeros((batch_size_fractional,num_rows,num_cols))\n            \n            # select rows from the tail of df_test upwards\n            batch = list(df_test['patientId'][-batch_size_fractional:]) #1000\n\n            \n            for j in range(0,len(batch)):\n                patientId = batch[j]\n\n                path = \\\n        '../input/rsna-pneumonia-detection-challenge/stage_1_test_images/%s.dcm' % patientId\n\n                dcm_data = pydicom.read_file(path)\n\n                # get the image as a numpy array\n                image = dcm_data.pixel_array\n\n                # resize the image\n                small_image = resize(image,(num_rows,num_cols))\n\n                # add the image to the empty numpy array\n                image_array[j,:,:] = small_image\n\n            # reshape the array and normalize\n            X_test = image_array.reshape(batch_size_fractional,num_rows,num_cols,1)/255\n            \n        \n        # For testing the generator so we can see how many batches it outputs\n        # by calling next(). Uncomment the next line for testing.\n        #print(i)\n        \n        k = k + batch_size\n        \n        # Keras requires a tuple in the form (inputs,targets)\n        yield (X_test.astype(np.float32))\n    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7724c2c97499ec7269a3a2175e1d53b615c15e36"},"cell_type":"markdown","source":"### Check the Generators"},{"metadata":{"_uuid":"8a2f31318d00b9d2ffd7ed01a48ed66706e2af23"},"cell_type":"markdown","source":"In this section I ran a few checks to see if the generators were performing as expected or if there were errors in the code. To save memory I've commented out this code."},{"metadata":{"trusted":true,"_uuid":"d322ac19612d4371c6dbc83e607ec494f2bc250b"},"cell_type":"code","source":"# train_generator\n\n#train_gen = \\\n#train_generator(df_train_images, df_train, batch_size=50, num_rows=500, num_cols=500)\n\n#val_gen = \\\n#val_generator(df_val_images, df_val, batch_size=10, num_rows=500, num_cols=500)\n\n#test_gen = \\\n #test_generator(df_test, batch_size=1000, num_rows=500, num_cols=500)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"041d126d89b75c7bda12ceef3181cb3ae7db34d8"},"cell_type":"markdown","source":"#### Are the output shapes correct?"},{"metadata":{"trusted":true,"_uuid":"4b81f8fca75b11edb31f0cbbef5d1ff785521999"},"cell_type":"code","source":"# Note: Each time this notebook cell is run, the generators will output only one batch.\n\n# If the generators are working correctly, the following shapes should be output:\n# X_train (10,500,500,1)\n# y_train (10,10)\n# X_val (10,500,500,1)\n# y_val (10,10)\n# X_test(10,500,500,1)\n\n# tuple unpacking\n#X_train, y_train = next(train_gen)\n#X_val, y_val = next(val_gen)\n#X_test = next(test_gen)\n\n#print(X_train.shape)\n#print(X_train.dtype)\n#print(y_train.shape)\n#print(X_val.shape)\n#print(X_val.dtype)\n#print(y_val.shape)\n#print(X_test.shape)\n#print(X_test.dtype)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9ff0dc8cdc83026b6567448b202bb20f32232225"},"cell_type":"markdown","source":"#### Are the generators running the correct number of times?"},{"metadata":{"_uuid":"a3b926c93631a73537d994b911c130821e5e93ca"},"cell_type":"markdown","source":"The train and val generators should loop infinitely through the batches i.e. without a \"StopIteration\" exception being raised. \n\nThe test generator should reach \"StopIteration\" after the last batch.\n\nThe size of the last batch will be smaller than the other batches."},{"metadata":{"trusted":true,"_uuid":"cc28c53fa728b380f208661a657e228c4c1a1b29"},"cell_type":"code","source":"# Check the train_generator()\n#train_gen = \\\n#train_generator(df_train_images, df_train, batch_size=5000,, num_rows=500, num_cols=500)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"65d7923cba514cc6fbfa3fa1c1cd7997e3156db7"},"cell_type":"code","source":"#X_train, y_train = next(train_gen)\n\n#print(X_train.shape)\n#print(y_train.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5fce6ca0dad79d54d906feadead6e6bb3e384985"},"cell_type":"code","source":"# Check the val_generator()\n#val_gen = \\\n#val_generator(df_val_images, df_val, batch_size=2000, num_rows=500, num_cols=500)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0a7b8e98bfa4ee977ccaad7a503804567c7bd594"},"cell_type":"code","source":"#X_val, y_val = next(val_gen)\n\n#print(X_val.shape)\n#print(y_val.shape) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2df5311f93b1a0c54d0e07e3b9e6462c77b15aa9"},"cell_type":"code","source":"# check test_generator\n#test_gen = \\\n#test_generator(df_test, batch_size=300, num_rows=500, num_cols=500)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6ad7ac591bda76e7f140dfdac1d20bf422dba248"},"cell_type":"code","source":"# Uncomment the print() function in test_generator() before running this cell. \n# Each time this cell is run the output should increment by 1. \n# The last number to be output should be 4.\n\n# Remember to re-quote the print() function after this test.\n\n# With 1000 test samples, a batch size of 300, and image size of 500x500\n# these are the shapes we should get with each iteration:\n\n# 0: (300,500,500,1)\n# 1: (300,500,500,1)\n# 2: (300,500,500,1)\n# 3: (100,500,500,1)\n# 4: Error\n\n# With 1000 samples and a batch_size of 300 the generator should only run for 4 loops.\n# Run this cell 5 times. On the fifth time you should get a \"StopIteration\" error.\n\n#X_test = next(test_gen)\n\n#X_test.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b402c5ac68fd181cd206a1c2afa01b8b7b3b68ca"},"cell_type":"markdown","source":"### MODELING "},{"metadata":{"_uuid":"593d727aea29583dde34d559e41485df07b781e1"},"cell_type":"markdown","source":"### Initialize the generators"},{"metadata":{"trusted":true,"_uuid":"0233e1e06e8084e4fae520de6fdc740339b1b8d9"},"cell_type":"code","source":"# get the number of train and val images\n\nprint(df_train.shape)\nprint(df_val.shape)\nprint(df_test.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"24fff79763e45adb79667b12405369ea0f0a7fd1"},"cell_type":"code","source":"########################\n# INPUTS\n\n# Set the batch sizes:\n\ntrain_batch_size = 10\nval_batch_size = 10\ntest_batch_size = 1\n\n# Set the image size:\n\nnum_rows = 1024\nnum_cols = 1024\n\n#########################\n\n# train_generator\ntrain_gen = \\\ntrain_generator(df_train_images, df_train, train_batch_size, num_rows, num_cols)\n\nnum_train_samples = df_train.shape[0]\n\nnum_train_batches = math.ceil(num_train_samples/train_batch_size) # round down\n\n\n# val_generator\nval_gen = \\\nval_generator(df_val_images, df_val, val_batch_size, num_rows, num_cols)\n\nnum_val_samples = df_val.shape[0]\n\nnum_val_batches = math.ceil(num_val_samples/val_batch_size) # round down\n\n# test_generator\ntest_gen = \\\ntest_generator(df_test, test_batch_size, num_rows, num_cols)\n\nnum_test_samples = df_test.shape[0]\n\nnum_test_batches = math.ceil(num_test_samples/test_batch_size) # round up\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"43bdacae1a6642ed87a49ef51b7e92e5c1b5237a"},"cell_type":"markdown","source":"### Set up the Model Architecture"},{"metadata":{"trusted":true,"_uuid":"2688063f45bc11368462abba5e352d1b45076b5e"},"cell_type":"code","source":"from keras.models import Sequential\nfrom keras.layers import Conv2D, MaxPooling2D, BatchNormalization, Dense, Dropout, Flatten\nfrom keras.optimizers import Adam\nfrom keras.callbacks import ModelCheckpoint\n\nmodel = Sequential()\n\nmodel.add(Conv2D(filters=32, kernel_size=(3, 3), activation='relu',\n                        input_shape=(num_rows, num_cols, 1)))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\n\nmodel.add(Conv2D(64, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D((2, 2)))\n\nmodel.add(Conv2D(128, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D((2, 2)))\n\nmodel.add(Conv2D(128, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D((2, 2)))\n\nmodel.add(Flatten())\nmodel.add(Dense(256, activation='relu'))\nmodel.add(Dropout(0.2))\n\nmodel.add(Dense(10, activation='linear'))\n\n\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0492b9f69cf85d973f4b7ea20320cbf4737a510f"},"cell_type":"code","source":"# compile the model\nAdam_opt = Adam(lr=0.0001, beta_1=0.9, beta_2=0.999, epsilon=1e-08, decay=0.0)\nmodel.compile(optimizer=Adam_opt, loss='mse')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"262a5e845d1ea73c8dc39762e95a338d64692ea5"},"cell_type":"code","source":"# Notes: \n# To reduce RAM use it's best to keep 'max_queue_size' small.\n# The test and val generators run infinitely therefore we must set \n# steps_per_epoch=num_train_batches and validation_steps=num_val_batches so \n# that fit_generator() knows when to stop an epoch and to ensure that the model sees\n# the same batch only once.\n\nfilepath = \"model.h5\"\ncheckpoint = ModelCheckpoint(filepath, monitor='val_loss', verbose=1, save_best_only=True, mode='min')\ncallbacks_list = [checkpoint]\n\n\nhistory = model.fit_generator(generator=train_gen, \n                        steps_per_epoch=num_train_batches, \n                        epochs=3, \n                        verbose=1, \n                        callbacks=callbacks_list, \n                        validation_data=val_gen,\n                        validation_steps=num_val_batches, \n                        class_weight=None, \n                        max_queue_size=2, \n                        workers=4,\n                        use_multiprocessing=True, \n                        shuffle=False, \n                        initial_epoch=0)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"35ff73d75e02c42972b6f3d4e8571eb0980c0380"},"cell_type":"markdown","source":"### Plot the Loss Curves"},{"metadata":{"trusted":true,"_uuid":"edb41e10896560a9fb17ace8e01470edf2097124"},"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nloss = history.history['loss']\nval_loss = history.history['val_loss']\nepochs = range(1, len(loss) + 1)\nplt.legend()\nplt.plot(epochs, loss, 'bo', label='Training loss')\nplt.plot(epochs, val_loss, 'b', label='Validation loss')\nplt.title('Training and validation loss')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"_uuid":"9e46b3685615b1c7b12c329319e418b113a92854"},"cell_type":"markdown","source":"### Make a Prediction"},{"metadata":{"trusted":true,"_uuid":"81238e0692da31462a2cfc2b26adf8544a51ca06"},"cell_type":"code","source":"# Initialize the test generator\n# Note: Put the intilization in the same cell as the prediction step because the \n# test generator was not designed to run infinitely. We want the prediction process\n# to always start at the first batch and run only once.\n\n# I keep the test_batch_size=1 just to ensure that nothing strange happens.\n\ntest_gen = \\\ntest_generator(df_test, test_batch_size, num_rows, num_cols)\n\nmodel.load_weights(filepath = 'model.h5')\npredictions = model.predict_generator(test_gen, \n                                      steps=num_test_batches, \n                                      max_queue_size=1, \n                                      workers=1, \n                                      use_multiprocessing=False, \n                                      verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ce2580f7f77b6df8fac2f03700e13e05927ab82d"},"cell_type":"code","source":"predictions.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a81773a233743d282d2d0791bec03b16ecefa429"},"cell_type":"markdown","source":"### See  the Results"},{"metadata":{"trusted":true,"_uuid":"a6d834ca447a4a17095aec4c08b4bef799747ecc"},"cell_type":"code","source":"predictions[1]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"42244f97cd7e4623652e5e0ccf96a32246fa736b"},"cell_type":"markdown","source":"### Process the Predictions"},{"metadata":{"trusted":true,"_uuid":"cbab430c8435bcf4fecd8d13dc631172eb74c418"},"cell_type":"code","source":"# put the predictions into a dataframe\ndf_preds = pd.DataFrame(predictions)\n\n# add column names\nnew_names = ['conf_1', 'x_1', 'y_1', 'width_1', 'height_1',\n       'conf_2', 'x_2', 'y_2', 'width_2', 'height_2']\n\ndf_preds.columns = new_names\n\n# add the patientId column\ndf_preds['patientId'] = df_test['patientId']\n\n# add the PredictionString column\ndf_preds['PredictionString'] = 0","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"70089058aff1c92f7f716f8fb085d56c0c0bd444"},"cell_type":"code","source":"df_preds.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8266cc2a7fa35ebc8314fd4da2882bbd9593ca0c"},"cell_type":"code","source":"# Version 2: Changes were made. See comments below.\n\ndef process_preds(df):\n    \n    limit = 0.5\n    \n    conf_1 = 0\n    conf_2 = 0\n    conf_3 = 0\n    conf_4 = 0\n    \n    string_1 = ''\n    string_2 = ''\n    string_3 = ''\n    string_4 = ''\n    \n    \n    for i in range(0,len(df)):\n        \n        #get the conf scores\n        conf_1 = df.loc[i,'conf_1'] # revised in Version 2\n        conf_2 = df.loc[i,'conf_2'] # revised in Version 2\n        \n        if conf_1 >= limit:\n            string_1 = \\\n            str(conf_1) + ' ' + str(round(df.loc[i,'x_1']))+ ' ' + \\\n            str(round(df.loc[i,'y_1']))+ ' ' + str(round(df.loc[i,'width_1']))+ ' ' + str(round(df.loc[i,'height_1']))\n\n        if conf_2 >= limit:\n            string_2 = \\\n            str(conf_2) + ' ' + str(round(df.loc[i,'x_2']))+ ' ' + \\\n            str(round(df.loc[i,'y_2']))+ ' ' + str(round(df.loc[i,'width_2']))+ ' ' + str(round(df.loc[i,'height_2']))\n\n        df.loc[i,'PredictionString']  = \\\n        string_1 + ' ' + string_2 \n\n    df_submission = df[['patientId', 'PredictionString']]\n    \n    return df_submission\n\n# call the function\ndf_submission = process_preds(df_preds)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"166c4db314a587c4abb4c3077199f5e1afbdfa86"},"cell_type":"code","source":"df_submission.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4f6442c89e2c44ff13d45d121ab9b3fa4db499ed"},"cell_type":"markdown","source":"### Create the submission csv file"},{"metadata":{"trusted":true,"_uuid":"495a534234df67744186d50b95f362f0bc483e90"},"cell_type":"code","source":"\nID = df_preds['patientId']\npreds = df_preds['PredictionString']\n\nsubmission = pd.DataFrame({'patientId':ID, \n                           'PredictionString':preds, \n                          }).set_index('patientId')\n\nsubmission.to_csv('pneu_keras_model.csv', columns=['PredictionString']) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"07b8902c278aeecaeef9b6acd87856b02cb9d908"},"cell_type":"markdown","source":"<hr>"},{"metadata":{"trusted":true,"_uuid":"e7dbe4e43998a327d69bfcaf1f52756c491b646e","collapsed":true},"cell_type":"markdown","source":"### APPENDIX"},{"metadata":{"_uuid":"2ecce22c892da3d8d5e9206a84b2609cd615f94f"},"cell_type":"markdown","source":"### 1. What is a python  generator?"},{"metadata":{"trusted":true,"_uuid":"03c15a4784e503765a8830b1ac644df85908035c"},"cell_type":"code","source":"# This is a simple example of a generator.\ndef my_generator():\n    for i in range(0,3):\n        yield print(i)\n\nmy_gen = my_generator()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bc12cc1f111dee114945d0e02d2bb3f778951640"},"cell_type":"code","source":"# If you run this cell 3 times  you will notice that the output increases by 1 each time.\n# On the 4th iteration there will be a 'StopIteration'.\n\nout_put = next(my_gen)\nout_put","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ba490b9e14368df5b45483a0384961a3d7bb6f0b"},"cell_type":"markdown","source":"Unlike a normal function, the ouput from a generator does not stay in memory permanently. Therefore, it can be used to handle large amounts of image data when only a limited amount of memory is available.\n"},{"metadata":{"_uuid":"4660648e3eb7cf581bad62f7fac07db11b2ed7c3"},"cell_type":"markdown","source":"### How to make a generator run infinitely?"},{"metadata":{"trusted":true,"_uuid":"54fa13209db37db900c4cb5787d4ed4a2daf9eef"},"cell_type":"code","source":"# source: @Liquid_Fire\n#https://stackoverflow.com/questions/3704918/\n    #python-way-to-restart-a-for-loop-similar-to-continue-for-while-loops\n\n# To use fit_generator() in keras our generator needs to loop infinitely.\n# This is how to do that:\n\ndef my_generator():\n\n   \n    while True: \n        \n        for i in range(0,4):\n            \n            yield i\n        \n        \n\ninfinity_gen = my_generator()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"792322779d977367a6446e259d33939c10bea40a"},"cell_type":"code","source":"# if you run this cell you will see that a 'StopIteration' never happens.\nout_put = next(infinity_gen)\nout_put","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"add4531f860574095cf1506021aa1ba0dcf95bdd"},"cell_type":"markdown","source":"In this kernel we used a generator to read batches of X_train images from a folder and also read y_train info from a dataframe. The generator then outputs a batch in the form of a tuple (X_train, y_train)."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"d0d62c2fd289a24c689434d044ddccf3c4c358ad"},"cell_type":"markdown","source":"### 2. Resources\n\nThese are some resources that I found helpful:"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"dbd5ddb35987c8394d839a4866f733debbf1028c"},"cell_type":"markdown","source":"Excellent kernel by @Peter and friends: <br>\nhttps://www.kaggle.com/peterchang77/exploratory-data-analysis\n\nKeras info on fit_generator() that explains what the input format needs to be: <br>\nhttps://keras.io/models/sequential/\n\nBlog post on using generators with keras: <br>\nhttps://adriannunez.github.io/generators-in-keras/\n\nBlog post on using keras for regression: <br>\nhttps://machinelearningmastery.com/how-to-make-classification-and-regression-predictions-for-deep-learning-models-in-keras/\n\nKeras fit_generator() issue: <br>\nhttps://github.com/keras-team/keras/issues/3675"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"4cfed595932e026fe3efac3ff2036f7d9bf2f2b8"},"cell_type":"markdown","source":"<hr>\n\nThank you for reading."},{"metadata":{"trusted":true,"_uuid":"3f21731ddb034f7892331bf7ff86650fcd165273"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}