{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Petals to the Metal Competition Submission #\n\n<blockquote style=\"margin-right:auto; margin-left:auto; background-color: #ebf9ff; padding: 1em; margin:24px;\">\n    <strong>Group 11</strong><br>\nShailja Kartik, Xinyu Li, Paul Mello, and Rodrigo Teixeira\n</blockquote>\n\n# Step 1: Imports #","metadata":{}},{"cell_type":"code","source":"!pip install  efficientnet\n\nimport math, re, os, random\nimport numpy as np\nimport tensorflow as tf\nimport tensorflow_addons as tfa #Image Augmentation\nimport matplotlib.pyplot as plt\nimport efficientnet.tfkeras as efn\n\nfrom matplotlib import pyplot as plt\n\nfrom kaggle_datasets import KaggleDatasets\n\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau # LR_Scheduler\n\nfrom sklearn.metrics import f1_score, precision_score, recall_score, confusion_matrix\n\n\n\n\nprint(\"Tensorflow version \" + tf.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:13.632187Z","iopub.execute_input":"2022-05-04T00:16:13.633444Z","iopub.status.idle":"2022-05-04T00:16:24.097854Z","shell.execute_reply.started":"2022-05-04T00:16:13.633332Z","shell.execute_reply":"2022-05-04T00:16:24.097055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2: Distribution Strategy #\n\nA TPU has eight different *cores* and each of these cores acts as its own accelerator. (A TPU is sort of like having eight GPUs in one machine.) We tell TensorFlow how to make use of all these cores at once through a **distribution strategy**. Run the following cell to create the distribution strategy that we'll later apply to our model.","metadata":{}},{"cell_type":"code","source":"# Detect TPU, return appropriate distribution strategy\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver() \n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy() \n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:24.102419Z","iopub.execute_input":"2022-05-04T00:16:24.102627Z","iopub.status.idle":"2022-05-04T00:16:29.673240Z","shell.execute_reply.started":"2022-05-04T00:16:24.102601Z","shell.execute_reply":"2022-05-04T00:16:29.672191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 3: Loading the Competition Data #\n\n## Get GCS Path ##\n\nWhen used with TPUs, datasets need to be stored in a [Google Cloud Storage bucket](https://cloud.google.com/storage/). You can use data from any public GCS bucket by giving its path just like you would data from `'/kaggle/input'`. The following will retrieve the GCS path for this competition's dataset.","metadata":{}},{"cell_type":"code","source":"GCS_DS_PATH = KaggleDatasets().get_gcs_path('tpu-getting-started')\nprint(GCS_DS_PATH) # what do gcs paths look like?","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:29.674637Z","iopub.execute_input":"2022-05-04T00:16:29.674862Z","iopub.status.idle":"2022-05-04T00:16:30.105901Z","shell.execute_reply.started":"2022-05-04T00:16:29.674836Z","shell.execute_reply":"2022-05-04T00:16:30.104978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Declaring Global Variables ##","metadata":{}},{"cell_type":"code","source":"#Adjust These for File Format\nVAL = 512\nIMAGE_SIZE = [VAL, VAL]\n\nAUTO = tf.data.experimental.AUTOTUNE\n\nCLASSES = ['pink primrose',    'hard-leaved pocket orchid', 'canterbury bells', 'sweet pea',     'wild geranium',     'tiger lily',           'moon orchid',              'bird of paradise', 'monkshood',        'globe thistle',         # 00 - 09\n           'snapdragon',       \"colt's foot\",               'king protea',      'spear thistle', 'yellow iris',       'globe-flower',         'purple coneflower',        'peruvian lily',    'balloon flower',   'giant white arum lily', # 10 - 19\n           'fire lily',        'pincushion flower',         'fritillary',       'red ginger',    'grape hyacinth',    'corn poppy',           'prince of wales feathers', 'stemless gentian', 'artichoke',        'sweet william',         # 20 - 29\n           'carnation',        'garden phlox',              'love in the mist', 'cosmos',        'alpine sea holly',  'ruby-lipped cattleya', 'cape flower',              'great masterwort', 'siam tulip',       'lenten rose',           # 30 - 39\n           'barberton daisy',  'daffodil',                  'sword lily',       'poinsettia',    'bolero deep blue',  'wallflower',           'marigold',                 'buttercup',        'daisy',            'common dandelion',      # 40 - 49\n           'petunia',          'wild pansy',                'primula',          'sunflower',     'lilac hibiscus',    'bishop of llandaff',   'gaura',                    'geranium',         'orange dahlia',    'pink-yellow dahlia',    # 50 - 59\n           'cautleya spicata', 'japanese anemone',          'black-eyed susan', 'silverbush',    'californian poppy', 'osteospermum',         'spring crocus',            'iris',             'windflower',       'tree poppy',            # 60 - 69\n           'gazania',          'azalea',                    'water lily',       'rose',          'thorn apple',       'morning glory',        'passion flower',           'lotus',            'toad lily',        'anthurium',             # 70 - 79\n           'frangipani',       'clematis',                  'hibiscus',         'columbine',     'desert-rose',       'tree mallow',          'magnolia',                 'cyclamen ',        'watercress',       'canna lily',            # 80 - 89\n           'hippeastrum ',     'bee balm',                  'pink quill',       'foxglove',      'bougainvillea',     'camellia',             'mallow',                   'mexican petunia',  'bromelia',         'blanket flower',        # 90 - 99\n           'trumpet creeper',  'blackberry lily',           'common tulip',     'wild rose']                                                                                                                                               # 100 - 102","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-05-04T00:16:30.108855Z","iopub.execute_input":"2022-05-04T00:16:30.109146Z","iopub.status.idle":"2022-05-04T00:16:30.120237Z","shell.execute_reply.started":"2022-05-04T00:16:30.109116Z","shell.execute_reply":"2022-05-04T00:16:30.119271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Unpacking Data ###","metadata":{}},{"cell_type":"code","source":"GCS_PATH_SELECT = { # available image sizes\n    192: GCS_DS_PATH + '/tfrecords-jpeg-192x192',\n    224: GCS_DS_PATH + '/tfrecords-jpeg-224x224',\n    331: GCS_DS_PATH + '/tfrecords-jpeg-331x331',\n    512: GCS_DS_PATH + '/tfrecords-jpeg-512x512'\n}\nGCS_PATH = GCS_PATH_SELECT[IMAGE_SIZE[0]]\n\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train/*.tfrec')\nVALIDATION_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/val/*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/test/*.tfrec') ","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:30.121783Z","iopub.execute_input":"2022-05-04T00:16:30.122469Z","iopub.status.idle":"2022-05-04T00:16:30.407901Z","shell.execute_reply.started":"2022-05-04T00:16:30.122432Z","shell.execute_reply":"2022-05-04T00:16:30.407006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def decode_image(image_data):\n    image = tf.image.decode_jpeg(image_data, channels=3)\n    image = tf.cast(image, tf.float32) / 255.0  # convert image to floats in [0, 1] range\n    image = tf.reshape(image, [*IMAGE_SIZE, 3]) # explicit size needed for TPU\n    return image\n\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    label = tf.cast(example['class'], tf.int32)\n    return image, label # returns a dataset of (image, label) pairs\n\ndef read_unlabeled_tfrecord(example):\n    UNLABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"id\": tf.io.FixedLenFeature([], tf.string),  # shape [] means single element\n        # class is missing, this competitions's challenge is to predict flower classes for the test dataset\n    }\n    example = tf.io.parse_single_example(example, UNLABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    idnum = example['id']\n    return image, idnum # returns a dataset of image(s)\n\ndef load_dataset(filenames, labeled=True, ordered=False):\n    # Read from TFRecords. For optimal performance, reading from multiple files at once and\n    # disregarding data order. Order does not matter since we will be shuffling the data anyway.\n\n    ignore_order = tf.data.Options()\n    if not ordered:\n        ignore_order.experimental_deterministic = False # disable order, increase speed\n\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO) # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order) # uses data as soon as it streams in, rather than in its original order\n    dataset = dataset.map(read_labeled_tfrecord if labeled else read_unlabeled_tfrecord, num_parallel_calls=AUTO)\n    # returns a dataset of (image, label) pairs if labeled=True or (image, id) pairs if labeled=False\n    return dataset","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:30.409289Z","iopub.execute_input":"2022-05-04T00:16:30.409601Z","iopub.status.idle":"2022-05-04T00:16:30.422174Z","shell.execute_reply.started":"2022-05-04T00:16:30.409561Z","shell.execute_reply":"2022-05-04T00:16:30.421225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Augmentation ##","metadata":{}},{"cell_type":"code","source":"def data_augment(image, label):\n    # Thanks to the dataset.prefetch(AUTO)\n    # statement in the next function (below), this happens essentially\n    # for free on TPU. Data pipeline code is executed on the \"CPU\"\n    #image = tf.image.per_image_standardization(image)\n    image = tf.image.random_jpeg_quality(image, 85, 100)\n    image = tf.image.random_contrast(image, .95, 1) \n    #image = tf.image.random_crop(value = image, size = (round(VAL * .8), round(VAL * .8)))\n    image = tf.image.random_flip_up_down(image)\n    image = tf.image.random_flip_left_right(image) # Flipping image randomly\n    image = tf.image.resize(image, IMAGE_SIZE)\n    return image, label  \n\n# Consider the below lines for data augmentation\n# image = tf.image.rgb_to_grayscale(image)\n# image = tf.image.random_crop(value=image, size = [crop_height, crop_width, 3])\n# image = tf.image.random_brightness(image, upperBound)\n# image = tf.image.random_contrast(image, lowerBound, upperBound)\n# image = tf.image.per_image_standardization(image)\n# image = tf.image.random_saturation(image, lowerBound, upperBound)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:30.423426Z","iopub.execute_input":"2022-05-04T00:16:30.423652Z","iopub.status.idle":"2022-05-04T00:16:30.441008Z","shell.execute_reply.started":"2022-05-04T00:16:30.423627Z","shell.execute_reply":"2022-05-04T00:16:30.440285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Split ###","metadata":{}},{"cell_type":"code","source":" def get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n\ndef get_validation_dataset(ordered=False):\n    dataset = load_dataset(VALIDATION_FILENAMES, labeled=True, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.cache()\n    dataset = dataset.prefetch(AUTO)\n    return dataset\n\ndef get_test_dataset(ordered=False):\n    dataset = load_dataset(TEST_FILENAMES, labeled=False, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO)\n    return dataset\n\ndef count_data_items(filenames):\n    # the number of data items is written in the name of the .tfrec\n    # files, i.e. flowers00-230.tfrec = 230 data items\n    n = [int(re.compile(r\"-([0-9]*)\\.\").search(filename).group(1)) for filename in filenames]\n    return np.sum(n)\n\nNUM_TRAINING_IMAGES = count_data_items(TRAINING_FILENAMES)\nNUM_VALIDATION_IMAGES = count_data_items(VALIDATION_FILENAMES)\nNUM_TEST_IMAGES = count_data_items(TEST_FILENAMES)\nprint('Dataset: {} training images, {} validation images, {} unlabeled test images'.format(NUM_TRAINING_IMAGES, NUM_VALIDATION_IMAGES, NUM_TEST_IMAGES))\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-05-04T00:16:30.442517Z","iopub.execute_input":"2022-05-04T00:16:30.443271Z","iopub.status.idle":"2022-05-04T00:16:30.462752Z","shell.execute_reply.started":"2022-05-04T00:16:30.443227Z","shell.execute_reply":"2022-05-04T00:16:30.461819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Size ###","metadata":{}},{"cell_type":"markdown","source":"Creating training set, validation set, and testing set","metadata":{}},{"cell_type":"code","source":"# Define the batch size. This will be 16 with TPU off and 128 (=16*8) with TPU on\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\n\nds_train = get_training_dataset()\nds_valid = get_validation_dataset()\nds_test = get_test_dataset()\n\nprint(\"Training:\", ds_train)\nprint (\"Validation:\", ds_valid)\nprint(\"Test:\", ds_test)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:30.464474Z","iopub.execute_input":"2022-05-04T00:16:30.464785Z","iopub.status.idle":"2022-05-04T00:16:31.093462Z","shell.execute_reply.started":"2022-05-04T00:16:30.464745Z","shell.execute_reply":"2022-05-04T00:16:31.092542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Shape ###","metadata":{}},{"cell_type":"code","source":"np.set_printoptions(threshold=15, linewidth=80)\n\nprint(\"Training data shapes:\")\nfor image, label in ds_train.take(3):\n    print(image.numpy().shape, label.numpy().shape)\nprint(\"Training data label examples:\", label.numpy())\n\nprint(\"Test data shapes:\")\nfor image, idnum in ds_test.take(3):\n    print(image.numpy().shape, idnum.numpy().shape)\nprint(\"Test data IDs:\", idnum.numpy().astype('U')) # U=unicode string","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:31.094578Z","iopub.execute_input":"2022-05-04T00:16:31.094792Z","iopub.status.idle":"2022-05-04T00:16:40.615126Z","shell.execute_reply.started":"2022-05-04T00:16:31.094768Z","shell.execute_reply":"2022-05-04T00:16:40.614177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 4: Explore Data #","metadata":{}},{"cell_type":"markdown","source":"### Displaying Data Methods ###","metadata":{}},{"cell_type":"code","source":"def batch_to_numpy_images_and_labels(data):\n    images, labels = data\n    numpy_images = images.numpy()\n    numpy_labels = labels.numpy()\n    if numpy_labels.dtype == object: # binary string in this case,\n                                     # these are image ID strings\n        numpy_labels = [None for _ in enumerate(numpy_images)]\n    # If no labels, only image IDs, return None for labels (this is\n    # the case for test data)\n    return numpy_images, numpy_labels\n\ndef title_from_label_and_target(label, correct_label):\n    if correct_label is None:\n        return CLASSES[label], True\n    correct = (label == correct_label)\n    return \"{} [{}{}{}]\".format(CLASSES[label], 'OK' if correct else 'NO', u\"\\u2192\" if not correct else '',\n                                CLASSES[correct_label] if not correct else ''), correct\n\ndef display_one_flower(image, title, subplot, red=False, titlesize=16):\n    plt.subplot(*subplot)\n    plt.axis('off')\n    plt.imshow(image)\n    if len(title) > 0:\n        plt.title(title, fontsize=int(titlesize) if not red else int(titlesize/1.2), color='red' if red else 'black', fontdict={'verticalalignment':'center'}, pad=int(titlesize/1.5))\n    return (subplot[0], subplot[1], subplot[2]+1)\n    \ndef display_batch_of_images(databatch, predictions=None):\n    \"\"\"This will work with:\n    display_batch_of_images(images)\n    display_batch_of_images(images, predictions)\n    display_batch_of_images((images, labels))\n    display_batch_of_images((images, labels), predictions)\n    \"\"\"\n    # data\n    images, labels = batch_to_numpy_images_and_labels(databatch)\n    if labels is None:\n        labels = [None for _ in enumerate(images)]\n        \n    # auto-squaring: this will drop data that does not fit into square\n    # or square-ish rectangle\n    rows = int(math.sqrt(len(images)))\n    cols = len(images)//rows\n        \n    # size and spacing\n    FIGSIZE = 13.0\n    SPACING = 0.1\n    subplot=(rows,cols,1)\n    if rows < cols:\n        plt.figure(figsize=(FIGSIZE,FIGSIZE/cols*rows))\n    else:\n        plt.figure(figsize=(FIGSIZE/rows*cols,FIGSIZE))\n    \n    # display\n    for i, (image, label) in enumerate(zip(images[:rows*cols], labels[:rows*cols])):\n        title = '' if label is None else CLASSES[label]\n        correct = True\n        if predictions is not None:\n            title, correct = title_from_label_and_target(predictions[i], label)\n        dynamic_titlesize = FIGSIZE*SPACING/max(rows,cols)*40+3 # magic formula tested to work from 1x1 to 10x10 images\n        subplot = display_one_flower(image, title, subplot, not correct, titlesize=dynamic_titlesize)\n    \n    #layout\n    plt.tight_layout()\n    if label is None and predictions is None:\n        plt.subplots_adjust(wspace=0, hspace=0)\n    else:\n        plt.subplots_adjust(wspace=SPACING, hspace=SPACING)\n    plt.show()\n\n# Write our own implementations for loss, accuracy\ndef display_training_curves(training, validation, title, subplot):\n    if subplot%10==1: # set up the subplots on the first call\n        plt.subplots(figsize=(10,10), facecolor='#F0F0F0')\n        plt.tight_layout()\n    ax = plt.subplot(subplot)\n    ax.set_facecolor('#F8F8F8')\n    ax.plot(training)\n    ax.plot(validation)\n    ax.set_title('model '+ title)\n    ax.set_ylabel(title)\n    #ax.set_ylim(0.28,1.05)\n    ax.set_xlabel('epoch')\n    ax.legend(['train', 'valid.'])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-05-04T00:16:40.616785Z","iopub.execute_input":"2022-05-04T00:16:40.617316Z","iopub.status.idle":"2022-05-04T00:16:40.639258Z","shell.execute_reply.started":"2022-05-04T00:16:40.617268Z","shell.execute_reply":"2022-05-04T00:16:40.638448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds_iter = iter(ds_train.unbatch().batch(16))","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:40.640443Z","iopub.execute_input":"2022-05-04T00:16:40.641120Z","iopub.status.idle":"2022-05-04T00:16:40.661335Z","shell.execute_reply.started":"2022-05-04T00:16:40.641084Z","shell.execute_reply":"2022-05-04T00:16:40.660319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"one_batch = next(ds_iter)\ndisplay_batch_of_images(one_batch)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:40.665689Z","iopub.execute_input":"2022-05-04T00:16:40.665959Z","iopub.status.idle":"2022-05-04T00:16:46.515360Z","shell.execute_reply.started":"2022-05-04T00:16:40.665930Z","shell.execute_reply":"2022-05-04T00:16:46.514519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 5: Define Model #","metadata":{}},{"cell_type":"markdown","source":"### Model Design ###","metadata":{}},{"cell_type":"code","source":"# Epochs\nEPOCHS = 25\n\n# Steps\nSTEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:46.516568Z","iopub.execute_input":"2022-05-04T00:16:46.516902Z","iopub.status.idle":"2022-05-04T00:16:46.520582Z","shell.execute_reply.started":"2022-05-04T00:16:46.516875Z","shell.execute_reply":"2022-05-04T00:16:46.520017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with strategy.scope():\n    transferLearning = efn.EfficientNetB7(weights='imagenet', include_top=False)\n    transferLearning.trainable = False\n    \n    model = tf.keras.Sequential([transferLearning, \n            tf.keras.layers.GlobalAveragePooling2D(),\n            tf.keras.layers.Dense(len(CLASSES), activation='softmax')])","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:16:46.521590Z","iopub.execute_input":"2022-05-04T00:16:46.521898Z","iopub.status.idle":"2022-05-04T00:17:28.909868Z","shell.execute_reply.started":"2022-05-04T00:16:46.521871Z","shell.execute_reply":"2022-05-04T00:17:28.908882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Update Own Model  \n# Transfer Learning, https://keras.io/api/applications/\n# Self Learning Deep CNN https://www.tensorflow.org/api_docs/python/tf/keras/layers\n\n\n\n\n\n\n\n\n\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:17:28.911052Z","iopub.execute_input":"2022-05-04T00:17:28.911290Z","iopub.status.idle":"2022-05-04T00:17:28.916251Z","shell.execute_reply.started":"2022-05-04T00:17:28.911264Z","shell.execute_reply":"2022-05-04T00:17:28.915423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(optimizer = 'adam', loss = 'sparse_categorical_crossentropy', metrics = ['sparse_categorical_accuracy'])\n\n#model.compile(optimizer = 'Ftrl', loss = 'sparse_categorical_crossentropy', metrics = ['sparse_categorical_accuracy'])\n\n#model.compile(optimizer = 'adagrad', loss = 'sparse_categorical_crossentropy', metrics = ['sparse_categorical_accuracy'])\n\n\n# model.summary() Transfer Learning Makes Output Massive","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:17:28.917561Z","iopub.execute_input":"2022-05-04T00:17:28.917886Z","iopub.status.idle":"2022-05-04T00:17:28.990225Z","shell.execute_reply.started":"2022-05-04T00:17:28.917844Z","shell.execute_reply":"2022-05-04T00:17:28.989427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 6: Training #","metadata":{}},{"cell_type":"markdown","source":"### Learning Rate Scheduler's ###","metadata":{}},{"cell_type":"code","source":"# lr factors\ndiscountFactor = 0.25\nbeginningRate = 0.2\n\ndef schedule(epoch):\n    def lr(epoch, beginningRate, discountFactor):\n        return beginningRate * math.exp((-discountFactor) * epoch)\n    return lr(epoch, beginningRate, discountFactor)\n\nlr_callback = tf.keras.callbacks.LearningRateScheduler(schedule, verbose=True)\n\nrng = [i for i in range(EPOCHS)]\ny = [schedule(x) for x in rng]\nplt.plot(rng, y)\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:17:28.991433Z","iopub.execute_input":"2022-05-04T00:17:28.991657Z","iopub.status.idle":"2022-05-04T00:17:28.999479Z","shell.execute_reply.started":"2022-05-04T00:17:28.991632Z","shell.execute_reply":"2022-05-04T00:17:28.998553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lrReducer = ReduceLROnPlateau(monitor = 'val_loss', factor = 0.3, patience = 3, min_lr = 0.0001)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-05-04T00:17:29.001277Z","iopub.execute_input":"2022-05-04T00:17:29.001708Z","iopub.status.idle":"2022-05-04T00:17:29.012344Z","shell.execute_reply.started":"2022-05-04T00:17:29.001662Z","shell.execute_reply":"2022-05-04T00:17:29.011347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# callbacks = [lr_callback]\n# callbacks = [lrReducer]\n\nhistory = model.fit(ds_train, validation_data = ds_valid, epochs = EPOCHS, steps_per_epoch = STEPS_PER_EPOCH, callbacks = [lrReducer])","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:17:29.014291Z","iopub.execute_input":"2022-05-04T00:17:29.014570Z","iopub.status.idle":"2022-05-04T00:33:19.710299Z","shell.execute_reply.started":"2022-05-04T00:17:29.014538Z","shell.execute_reply":"2022-05-04T00:33:19.709385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plotting Loss and Accuracy ###","metadata":{}},{"cell_type":"code","source":"plt.plot(history.history['loss'], label = \"Training Loss\")\nplt.plot(history.history['val_loss'], label = \"Validation Loss\")\nplt.xlabel(\"Epochs\")\nplt.ylabel(\"Loss\")\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:33:19.712258Z","iopub.execute_input":"2022-05-04T00:33:19.712715Z","iopub.status.idle":"2022-05-04T00:33:19.940852Z","shell.execute_reply.started":"2022-05-04T00:33:19.712668Z","shell.execute_reply":"2022-05-04T00:33:19.939853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['sparse_categorical_accuracy'], label = \"Training Accuracy\")\nplt.plot(history.history['val_sparse_categorical_accuracy'], label = \"Validation Accuracy\")\nplt.xlabel(\"Epochs\")\nplt.ylabel(\"Accuracy\")\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:33:19.942375Z","iopub.execute_input":"2022-05-04T00:33:19.942699Z","iopub.status.idle":"2022-05-04T00:33:20.172465Z","shell.execute_reply.started":"2022-05-04T00:33:19.942658Z","shell.execute_reply":"2022-05-04T00:33:20.171460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Architecture ###","metadata":{}},{"cell_type":"code","source":"tf.keras.utils.plot_model(model, show_shapes = True, show_layer_names = True)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:33:20.173729Z","iopub.execute_input":"2022-05-04T00:33:20.173953Z","iopub.status.idle":"2022-05-04T00:33:21.509452Z","shell.execute_reply.started":"2022-05-04T00:33:20.173927Z","shell.execute_reply":"2022-05-04T00:33:21.508403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 7: Evaluate Predictions #","metadata":{}},{"cell_type":"markdown","source":"### Plotting Confusion Matrix, Recall, F1 Score, and Precision ###","metadata":{}},{"cell_type":"code","source":"def display_confusion_matrix(cmat, score, precision, recall):\n    plt.figure(figsize=(15,15))\n    ax = plt.gca()\n    ax.matshow(cmat, cmap='Reds')\n    ax.set_xticks(range(len(CLASSES)))\n    ax.set_xticklabels(CLASSES, fontdict={'fontsize': 7})\n    plt.setp(ax.get_xticklabels(), rotation=45, ha=\"left\", rotation_mode=\"anchor\")\n    ax.set_yticks(range(len(CLASSES)))\n    ax.set_yticklabels(CLASSES, fontdict={'fontsize': 7})\n    plt.setp(ax.get_yticklabels(), rotation=45, ha=\"right\", rotation_mode=\"anchor\")\n    titlestring = \"\"\n    if score is not None:\n        titlestring += 'f1 = {:.3f} '.format(score)\n    if precision is not None:\n        titlestring += '\\nprecision = {:.3f} '.format(precision)\n    if recall is not None:\n        titlestring += '\\nrecall = {:.3f} '.format(recall)\n    if len(titlestring) > 0:\n        ax.text(101, 1, titlestring, fontdict={'fontsize': 18, 'horizontalalignment':'right', 'verticalalignment':'top', 'color':'#804040'})\n    plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-05-04T00:33:21.511373Z","iopub.execute_input":"2022-05-04T00:33:21.511632Z","iopub.status.idle":"2022-05-04T00:33:21.524149Z","shell.execute_reply.started":"2022-05-04T00:33:21.511601Z","shell.execute_reply":"2022-05-04T00:33:21.523331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Confusion Matrix ##","metadata":{}},{"cell_type":"markdown","source":"Displaying the relative accuracy our model has in identifying each class of flower.","metadata":{}},{"cell_type":"code","source":"cmdataset = get_validation_dataset(ordered=True)\nimages_ds = cmdataset.map(lambda image, label: image)\nlabels_ds = cmdataset.map(lambda image, label: label).unbatch()\n\ncm_correct_labels = next(iter(labels_ds.batch(NUM_VALIDATION_IMAGES))).numpy()\ncm_probabilities = model.predict(images_ds)\ncm_predictions = np.argmax(cm_probabilities, axis=-1)\n\nlabels = range(len(CLASSES))\ncmat = confusion_matrix(\n    cm_correct_labels,\n    cm_predictions,\n    labels=labels,\n)\ncmat = (cmat.T / cmat.sum(axis=1)).T # normalize","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:33:21.525162Z","iopub.execute_input":"2022-05-04T00:33:21.525637Z","iopub.status.idle":"2022-05-04T00:33:47.297475Z","shell.execute_reply.started":"2022-05-04T00:33:21.525598Z","shell.execute_reply":"2022-05-04T00:33:47.296556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"score = f1_score(cm_correct_labels, cm_predictions,labels=labels, average='macro')\nprecision = precision_score(cm_correct_labels, cm_predictions, labels=labels, average='macro',)\nrecall = recall_score(cm_correct_labels, cm_predictions, labels=labels, average='macro')\ndisplay_confusion_matrix(cmat, score, precision, recall)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:33:47.298724Z","iopub.execute_input":"2022-05-04T00:33:47.298975Z","iopub.status.idle":"2022-05-04T00:33:52.380114Z","shell.execute_reply.started":"2022-05-04T00:33:47.298946Z","shell.execute_reply":"2022-05-04T00:33:52.379448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visual Validation ##","metadata":{}},{"cell_type":"markdown","source":"Displaying accuracy of our model using validation set as proof","metadata":{}},{"cell_type":"code","source":"dataset = get_validation_dataset()\ndataset = dataset.unbatch().batch(16)\nbatch = iter(dataset)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:33:52.381399Z","iopub.execute_input":"2022-05-04T00:33:52.381814Z","iopub.status.idle":"2022-05-04T00:33:52.423314Z","shell.execute_reply.started":"2022-05-04T00:33:52.381770Z","shell.execute_reply":"2022-05-04T00:33:52.422618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images, labels = next(batch)\nprobabilities = model.predict(images)\npredictions = np.argmax(probabilities, axis=-1)\ndisplay_batch_of_images((images, labels), predictions)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:33:52.424778Z","iopub.execute_input":"2022-05-04T00:33:52.425301Z","iopub.status.idle":"2022-05-04T00:34:05.905191Z","shell.execute_reply.started":"2022-05-04T00:33:52.425258Z","shell.execute_reply":"2022-05-04T00:34:05.904133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 8: Prediction #","metadata":{}},{"cell_type":"code","source":"test_ds = get_test_dataset(ordered=True)\n\nprint('Computing predictions...')\ntest_images_ds = test_ds.map(lambda image, idnum: image)\nprobabilities = model.predict(test_images_ds)\npredictions = np.argmax(probabilities, axis=-1)\nprint(predictions)","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:34:05.906590Z","iopub.execute_input":"2022-05-04T00:34:05.906876Z","iopub.status.idle":"2022-05-04T00:34:29.299110Z","shell.execute_reply.started":"2022-05-04T00:34:05.906842Z","shell.execute_reply":"2022-05-04T00:34:29.298057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Generate Submission File with Predictions ###","metadata":{}},{"cell_type":"code","source":"print('Generating submission.csv file...')\n\n# Get image ids from test set and convert to unicode\ntest_ids_ds = test_ds.map(lambda image, idnum: idnum).unbatch()\ntest_ids = next(iter(test_ids_ds.batch(NUM_TEST_IMAGES))).numpy().astype('U')\n\n# Write the submission file\nnp.savetxt(\n    'submission.csv',\n    np.rec.fromarrays([test_ids, predictions]),\n    fmt=['%s', '%d'],\n    delimiter=',',\n    header='id,label',\n    comments='',\n)\n\n# Look at the first few predictions\n!head submission.csv","metadata":{"execution":{"iopub.status.busy":"2022-05-04T00:34:29.300338Z","iopub.execute_input":"2022-05-04T00:34:29.300573Z","iopub.status.idle":"2022-05-04T00:34:33.065539Z","shell.execute_reply.started":"2022-05-04T00:34:29.300546Z","shell.execute_reply":"2022-05-04T00:34:33.064670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 9: Make a submission #\n\nIf you haven't already, create your own editable copy of this notebook by clicking on the **Copy and Edit** button in the top right corner. Then, submit to the competition by following these steps:\n\n1. Begin by clicking on the blue **Save Version** button in the top right corner of the window.  This will generate a pop-up window.  \n2. Ensure that the **Save and Run All** option is selected, and then click on the blue **Save** button.\n3. This generates a window in the bottom left corner of the notebook.  After it has finished running, click on the number to the right of the **Save Version** button.  This pulls up a list of versions on the right of the screen.  Click on the ellipsis **(...)** to the right of the most recent version, and select **Open in Viewer**.  This brings you into view mode of the same page. You will need to scroll down to get back to these instructions.\n4. Click on the **Output** tab on the right of the screen.  Then, click on the file you would like to submit, and click on the blue **Submit** button to submit your results to the leaderboard.\n\nYou have now successfully submitted to the competition!\n\nIf you want to keep working to improve your performance, select the blue **Edit** button in the top right of the screen. Then you can change your code and repeat the process. There's a lot of room to improve, and you will climb up the leaderboard as you work.\n","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [Learn Discussion forum](https://www.kaggle.com/learn-forum/161321) to chat with other Learners.*","metadata":{}}]}