{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Computer Vision: Petals to the Metal - Flower Classification on TPU**","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## **Step 1: Import the Required Libraries**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport math, re, os\nimport numpy as np\nimport tensorflow as tf\nimport matplotlib.pyplot as plt\nprint(\"Tensorflow version \" + tf.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:00:27.943696Z","iopub.execute_input":"2022-06-16T15:00:27.944047Z","iopub.status.idle":"2022-06-16T15:00:34.526273Z","shell.execute_reply.started":"2022-06-16T15:00:27.943961Z","shell.execute_reply":"2022-06-16T15:00:34.525331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Step 2: Distribution Strategy**\n\nA TPU has eight different *cores* and each of these cores acts as its own accelerator. (A TPU is sort of like having eight GPUs in one machine.) We tell TensorFlow how to make use of all these cores at once through a **distribution strategy**. Run the following cell to create the distribution strategy that will be applied later to our model.","metadata":{}},{"cell_type":"code","source":"# Detect TPU, return appropriate distribution strategy\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver() \n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy() \n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:00:34.528133Z","iopub.execute_input":"2022-06-16T15:00:34.528412Z","iopub.status.idle":"2022-06-16T15:00:40.461790Z","shell.execute_reply.started":"2022-06-16T15:00:34.528359Z","shell.execute_reply":"2022-06-16T15:00:40.460908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Distribution strategy is used when the neural network model is created. Then, TensorFlow will distribute the training among the eight TPU cores by creating eight different *replicas* of the model, one for each core.\n\n## **Step 3: Loading & Preprocessing the Competition Data**\n\n### **Get GCS Path**\n\nWhen used with TPUs, datasets need to be stored in a [Google Cloud Storage bucket](https://cloud.google.com/storage/). The data from any public GCS bucket can be used by giving its path just like data from `'/kaggle/input'` is retrieved. The following will retrieve the GCS path for this competition's dataset.","metadata":{}},{"cell_type":"code","source":"from kaggle_datasets import KaggleDatasets\n\nGCS_DS_PATH = KaggleDatasets().get_gcs_path('tpu-getting-started')\nprint(GCS_DS_PATH) # what do gcs paths look like?","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:00:40.463129Z","iopub.execute_input":"2022-06-16T15:00:40.463456Z","iopub.status.idle":"2022-06-16T15:00:40.913553Z","shell.execute_reply.started":"2022-06-16T15:00:40.463393Z","shell.execute_reply":"2022-06-16T15:00:40.912608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Load & Preprocess Data**\nWhen used with TPUs, datasets are often serialized into [TFRecords](https://www.kaggle.com/ryanholbrook/tfrecords-basics). This is a format convenient for distributing data to each of the TPUs cores. ","metadata":{}},{"cell_type":"code","source":"IMAGE_SIZE = [512, 512]\nGCS_PATH = GCS_DS_PATH + '/tfrecords-jpeg-512x512'\nAUTO = tf.data.experimental.AUTOTUNE\n\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train/*.tfrec')\nVALIDATION_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/val/*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/test/*.tfrec') \n\nCLASSES = ['pink primrose',    'hard-leaved pocket orchid', 'canterbury bells', 'sweet pea',     'wild geranium',     'tiger lily',           'moon orchid',              'bird of paradise', 'monkshood',        'globe thistle',         # 00 - 09\n           'snapdragon',       \"colt's foot\",               'king protea',      'spear thistle', 'yellow iris',       'globe-flower',         'purple coneflower',        'peruvian lily',    'balloon flower',   'giant white arum lily', # 10 - 19\n           'fire lily',        'pincushion flower',         'fritillary',       'red ginger',    'grape hyacinth',    'corn poppy',           'prince of wales feathers', 'stemless gentian', 'artichoke',        'sweet william',         # 20 - 29\n           'carnation',        'garden phlox',              'love in the mist', 'cosmos',        'alpine sea holly',  'ruby-lipped cattleya', 'cape flower',              'great masterwort', 'siam tulip',       'lenten rose',           # 30 - 39\n           'barberton daisy',  'daffodil',                  'sword lily',       'poinsettia',    'bolero deep blue',  'wallflower',           'marigold',                 'buttercup',        'daisy',            'common dandelion',      # 40 - 49\n           'petunia',          'wild pansy',                'primula',          'sunflower',     'lilac hibiscus',    'bishop of llandaff',   'gaura',                    'geranium',         'orange dahlia',    'pink-yellow dahlia',    # 50 - 59\n           'cautleya spicata', 'japanese anemone',          'black-eyed susan', 'silverbush',    'californian poppy', 'osteospermum',         'spring crocus',            'iris',             'windflower',       'tree poppy',            # 60 - 69\n           'gazania',          'azalea',                    'water lily',       'rose',          'thorn apple',       'morning glory',        'passion flower',           'lotus',            'toad lily',        'anthurium',             # 70 - 79\n           'frangipani',       'clematis',                  'hibiscus',         'columbine',     'desert-rose',       'tree mallow',          'magnolia',                 'cyclamen ',        'watercress',       'canna lily',            # 80 - 89\n           'hippeastrum ',     'bee balm',                  'pink quill',       'foxglove',      'bougainvillea',     'camellia',             'mallow',                   'mexican petunia',  'bromelia',         'blanket flower',        # 90 - 99\n           'trumpet creeper',  'blackberry lily',           'common tulip',     'wild rose']                                                                                                                                               # 100 - 102\n\n\ndef decode_image(image_data):\n    image = tf.image.decode_jpeg(image_data, channels=3)\n    image = tf.cast(image, tf.float32) / 255.0  # convert image to floats in [0, 1] range\n    image = tf.reshape(image, [*IMAGE_SIZE, 3]) # explicit size needed for TPU\n    return image\n\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    label = tf.cast(example['class'], tf.int32)\n    return image, label # returns a dataset of (image, label) pairs\n\ndef read_unlabeled_tfrecord(example):\n    UNLABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"id\": tf.io.FixedLenFeature([], tf.string),  # shape [] means single element\n        # class is missing, this competitions's challenge is to predict flower classes for the test dataset\n    }\n    example = tf.io.parse_single_example(example, UNLABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    idnum = example['id']\n    return image, idnum # returns a dataset of image(s)\n\ndef load_dataset(filenames, labeled=True, ordered=False):\n    # Read from TFRecords. For optimal performance, reading from multiple files at once and\n    # disregarding data order. Order does not matter since we will be shuffling the data anyway.\n\n    ignore_order = tf.data.Options()\n    if not ordered:\n        ignore_order.experimental_deterministic = False # disable order, increase speed\n\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO) # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order) # uses data as soon as it streams in, rather than in its original order\n    dataset = dataset.map(read_labeled_tfrecord if labeled else read_unlabeled_tfrecord, num_parallel_calls=AUTO)\n    # returns a dataset of (image, label) pairs if labeled=True or (image, id) pairs if labeled=False\n    return dataset","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:00:40.915769Z","iopub.execute_input":"2022-06-16T15:00:40.915995Z","iopub.status.idle":"2022-06-16T15:00:41.174338Z","shell.execute_reply.started":"2022-06-16T15:00:40.915969Z","shell.execute_reply":"2022-06-16T15:00:41.173532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Create Data Pipelines** \nIn this final step use the `tf.data` API to define an efficient data pipeline for each of the training, validation, and test splits.","metadata":{}},{"cell_type":"code","source":"import random\ndef data_augment(image, label):\n    # Thanks to the dataset.prefetch(AUTO)\n    # statement in the next function (below), this happens essentially\n    # for free on TPU. Data pipeline code is executed on the \"CPU\"\n    # part of the TPU while the TPU itself is computing gradients.\n    \n    image = tf.image.random_flip_left_right(image)\n    image = tf.image.random_flip_up_down(image)\n    #flag =random.randint(1,3)\n    #coef_1 = random.randint(70,90)*0.01\n    #coef_2 = random.randint(70,90)*0.01\n    #if flag ==1:\n        #image = tf.image.random_flip_left_right(image)\n    #elif flag ==2:\n        #image = tf.image.random_flip_up_down(image)\n    #else: \n        #image = tf.image.random_crop(image,[int(IMAGE_SIZE(0)* coef_1), int(IMAGE_SIZE(0)* coef_2),3], seed = 807)\n\n    #image = tf.image.random_saturation(image, 0, 2)\n    #image = tf.image.random_contrast(image, lower=0.2, upper=1.8)\n    #image = tf.image.random_saturation(image, lower=0.5, upper=1.5)\n    #image = tf.image.random_brightness(image, 0.2)\n    #image = tf.image.random_brightness(image, max_delta=63. / 255.)\n    #image = tf.image.per_image_standardization(image)#whiten\n    \n\n    return image, label   \n\ndef get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n\ndef get_validation_dataset(ordered=False):\n    dataset = load_dataset(VALIDATION_FILENAMES, labeled=True, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.cache()\n    dataset = dataset.prefetch(AUTO)\n    return dataset\n\ndef get_test_dataset(ordered=False):\n    dataset = load_dataset(TEST_FILENAMES, labeled=False, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO)\n    return dataset\n\ndef count_data_items(filenames):\n    # the number of data items is written in the name of the .tfrec\n    # files, i.e. flowers00-230.tfrec = 230 data items\n    n = [int(re.compile(r\"-([0-9]*)\\.\").search(filename).group(1)) for filename in filenames]\n    return np.sum(n)\n\nNUM_TRAINING_IMAGES = count_data_items(TRAINING_FILENAMES)\nNUM_VALIDATION_IMAGES = count_data_items(VALIDATION_FILENAMES)\nNUM_TEST_IMAGES = count_data_items(TEST_FILENAMES)\nprint('Dataset: {} training images, {} validation images, {} unlabeled test images'.format(NUM_TRAINING_IMAGES, NUM_VALIDATION_IMAGES, NUM_TEST_IMAGES))\n\n","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:03:57.277752Z","iopub.execute_input":"2022-06-16T15:03:57.278058Z","iopub.status.idle":"2022-06-16T15:03:57.294200Z","shell.execute_reply.started":"2022-06-16T15:03:57.278028Z","shell.execute_reply":"2022-06-16T15:03:57.292938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This next cell will create the datasets that can be used with Keras during training and inference. Notice how the size of the batches to the number of TPU cores is scaled.","metadata":{}},{"cell_type":"code","source":"# Define the batch size. This will be 16 with TPU off and 128 (=16*8) with TPU on\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\n\nds_train = get_training_dataset()\nds_valid = get_validation_dataset()\nds_test = get_test_dataset()\n\nprint(\"Training:\", ds_train)\nprint (\"Validation:\", ds_valid)\nprint(\"Test:\", ds_test)","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:03:57.824216Z","iopub.execute_input":"2022-06-16T15:03:57.824567Z","iopub.status.idle":"2022-06-16T15:03:58.058437Z","shell.execute_reply.started":"2022-06-16T15:03:57.824536Z","shell.execute_reply":"2022-06-16T15:03:58.057501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These datasets are `tf.data.Dataset` objects. Think about a dataset in TensorFlow as a *stream* of data records. The training and validation sets are streams of `(image, label)` pairs.","metadata":{}},{"cell_type":"code","source":"np.set_printoptions(threshold=15, linewidth=80)\n\nprint(\"Training data shapes:\")\nfor image, label in ds_train.take(3):\n    print(image.numpy().shape, label.numpy().shape)\nprint(\"Training data label examples:\", label.numpy())","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:03:58.060042Z","iopub.execute_input":"2022-06-16T15:03:58.060297Z","iopub.status.idle":"2022-06-16T15:04:03.835884Z","shell.execute_reply.started":"2022-06-16T15:03:58.060267Z","shell.execute_reply":"2022-06-16T15:04:03.834654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test set is a stream of `(image, idnum)` pairs; `idnum` here is the unique identifier given to the image that we'll use later when we make our submission as a `csv` file.","metadata":{}},{"cell_type":"code","source":"print(\"Test data shapes:\")\nfor image, idnum in ds_test.take(3):\n    print(image.numpy().shape, idnum.numpy().shape)\nprint(\"Test data IDs:\", idnum.numpy().astype('U')) # U=unicode string","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:04:03.837961Z","iopub.execute_input":"2022-06-16T15:04:03.838339Z","iopub.status.idle":"2022-06-16T15:04:07.798982Z","shell.execute_reply.started":"2022-06-16T15:04:03.838305Z","shell.execute_reply":"2022-06-16T15:04:07.798054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Step 4: Define Model**\n\nNow create a neural network for classifying images by using what's known as **transfer learning**. With transfer learning, the model reuse part of a pretrained model to get a head-start on a new dataset.\n\nFor this project, use a model called **VGG16 or DenseNet201 or InceptionResNetV2** pretrained on [ImageNet](http://image-net.org/)). [other models](https://www.tensorflow.org/api_docs/python/tf/keras/applications) included with Keras can also be experimented with. ([Xception](https://www.tensorflow.org/api_docs/python/tf/keras/applications/Xception) wouldn't be a bad choice.)\n\nThe distribution strategy created earlier contains a [context manager](https://docs.python.org/3/reference/compound_stmts.html#with), `strategy.scope`. This context manager tells TensorFlow how to divide the work of training among the eight TPU cores. When using TensorFlow with a TPU, it's important to define the model in a `strategy.scope()` context.","metadata":{}},{"cell_type":"code","source":"import tensorflow.keras.applications as apps\nhelp(apps)","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:04:07.800365Z","iopub.execute_input":"2022-06-16T15:04:07.800704Z","iopub.status.idle":"2022-06-16T15:04:07.808998Z","shell.execute_reply.started":"2022-06-16T15:04:07.800657Z","shell.execute_reply":"2022-06-16T15:04:07.808080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EPOCHS = 60\n\nwith strategy.scope(): \n    #pretrained_model = tf.keras.applications.VGG16(\n        #weights='imagenet',\n        #include_top=False ,\n        #input_shape=[*IMAGE_SIZE, 3])\n    pretrained_model = tf.keras.applications.densenet.DenseNet201(\n        include_top=False,weights='imagenet',input_tensor=None, \n        input_shape=[*IMAGE_SIZE, 3], pooling='avg')\n    #pretrained_model = tf.keras.applications.xception.Xception(\n        #include_top=False,weights='imagenet',input_tensor=None, \n        #input_shape=[*IMAGE_SIZE, 3],pooling='avg')\n    #pretrained_model = tf.keras.applications.inception_resnet_v2.InceptionResNetV2(\n        #include_top=False,weights='imagenet',input_tensor=None, \n        #input_shape=[*IMAGE_SIZE, 3])\n    pretrained_model.trainable = False\n    \n   \n    model = tf.keras.Sequential([\n        pretrained_model,\n        # ... attach a new head to act as a classifier.\n        # To a base pretrained on VGG16 to extract features from images, uncheck the following code.\n        #tf.keras.layers.GlobalAveragePooling2D(),\n        #tf.keras.layers.BatchNormalization(),\n        tf.keras.layers.Dense(len(CLASSES), activation='softmax')\n    ])","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:04:25.807624Z","iopub.execute_input":"2022-06-16T15:04:25.807938Z","iopub.status.idle":"2022-06-16T15:05:04.775785Z","shell.execute_reply.started":"2022-06-16T15:04:25.807908Z","shell.execute_reply":"2022-06-16T15:05:04.774662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `'sparse_categorical'` versions of the loss and metrics are appropriate for a classification task with more than two labels, like this one.","metadata":{}},{"cell_type":"code","source":"model.compile(\n    optimizer='adam',\n    loss = 'sparse_categorical_crossentropy',\n    metrics=['sparse_categorical_accuracy'],\n)\n\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:05:04.777620Z","iopub.execute_input":"2022-06-16T15:05:04.777889Z","iopub.status.idle":"2022-06-16T15:05:04.889132Z","shell.execute_reply.started":"2022-06-16T15:05:04.777857Z","shell.execute_reply":"2022-06-16T15:05:04.888050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Step 5: Training**\n### **Fit Model** \n\nAnd now we're ready to train the model. After defining a few parameters, we're good to go!\n#### **Learning Rate Scheduler**\nWe'll train this network with a special learning rate schedule.\n\n#### **Early Stopping** \nStop training when a monitored quantity has stopped improving.\n* Arguments:\n    * monitor: Quantity to be monitored.\n    * min_delta: Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement.\n    * patience: Number of epochs with no improvement after which training will be stopped.\n    * verbose: verbosity mode.\n    * mode: One of `{\"auto\", \"min\", \"max\"}`. \n    In `min` mode,training will stop when the quantity monitored has stopped decreasing; \n    in `max` mode it will stop when the quantity monitored has stopped increasing; \n    in `auto` mode, the direction is automatically inferred from the name of the monitored quantity","metadata":{}},{"cell_type":"markdown","source":"#### Learning Rate Schedule for Fine Tuning #\ndef exponential_lr(epoch,\n                   start_lr = 0.00001, min_lr = 0.00001, max_lr = 0.00005,\n                   rampup_epochs = 5, sustain_epochs = 0,\n                   exp_decay = 0.8):\n\n    def lr(epoch, start_lr, min_lr, max_lr, rampup_epochs, sustain_epochs, exp_decay):\n        # linear increase from start to rampup_epochs\n        if epoch < rampup_epochs:\n            lr = ((max_lr - start_lr) /\n                  rampup_epochs * epoch + start_lr)\n        # constant max_lr during sustain_epochs\n        elif epoch < rampup_epochs + sustain_epochs:\n            lr = max_lr\n        # exponential decay towards min_lr\n        else:\n            lr = ((max_lr - min_lr) *\n                  exp_decay**(epoch - rampup_epochs - sustain_epochs) +\n                  min_lr)\n        return lr\n    return lr(epoch,\n              start_lr,\n              min_lr,\n              max_lr,\n              rampup_epochs,\n              sustain_epochs,\n              exp_decay)\n\nlr_callback = tf.keras.callbacks.LearningRateScheduler(exponential_lr, verbose=True)\n\nrng = [i for i in range(EPOCHS)]\ny = [exponential_lr(x) for x in rng]\nplt.plot(rng, y)\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))","metadata":{}},{"cell_type":"code","source":"# Define training epochs\nEPOCHS = 60\nSTEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\n\nfrom tensorflow.keras.callbacks import EarlyStopping\nearly_stop = EarlyStopping(monitor='val_loss', mode='min', verbose=1, patience=25)\nhistory = model.fit(\n    ds_train,\n    validation_data=ds_valid,\n    epochs=EPOCHS,\n    steps_per_epoch=STEPS_PER_EPOCH,\n    callbacks=[early_stop]\n)","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:05:04.891099Z","iopub.execute_input":"2022-06-16T15:05:04.891451Z","iopub.status.idle":"2022-06-16T15:35:11.542865Z","shell.execute_reply.started":"2022-06-16T15:05:04.891391Z","shell.execute_reply":"2022-06-16T15:35:11.541846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Step 6: Model Evaluation**\nNext cell shows how the loss and metrics progressed during training. Thankfully, it converges!","metadata":{}},{"cell_type":"code","source":"history_frame = pd.DataFrame(history.history)\nhistory_frame.loc[:, ['loss', 'val_loss']].plot()\nhistory_frame.loc[:, ['sparse_categorical_accuracy', 'val_sparse_categorical_accuracy']].plot();","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:35:11.545123Z","iopub.execute_input":"2022-06-16T15:35:11.545373Z","iopub.status.idle":"2022-06-16T15:35:12.130016Z","shell.execute_reply.started":"2022-06-16T15:35:11.545344Z","shell.execute_reply":"2022-06-16T15:35:12.128884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['sparse_categorical_accuracy'])\nplt.plot(history.history['val_sparse_categorical_accuracy'])\nplt.title('model accuracy')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'test'], loc='upper left')\nplt.show()\n\n \nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss'])\nplt.title('model loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'test'], loc='upper left')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:35:12.131644Z","iopub.execute_input":"2022-06-16T15:35:12.132169Z","iopub.status.idle":"2022-06-16T15:35:12.588598Z","shell.execute_reply.started":"2022-06-16T15:35:12.132118Z","shell.execute_reply":"2022-06-16T15:35:12.587702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Step 7: Make Test Predictions** \nOnce everything is set fine, the model is ready to make predictions on the test set.","metadata":{}},{"cell_type":"code","source":"test_ds = get_test_dataset(ordered=True)\n\nprint('Computing predictions...')\ntest_images_ds = test_ds.map(lambda image, idnum: image)\nprobabilities = model.predict(test_images_ds)\npredictions = np.argmax(probabilities, axis=-1)\nprint(predictions)","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:35:12.589622Z","iopub.execute_input":"2022-06-16T15:35:12.589846Z","iopub.status.idle":"2022-06-16T15:35:54.802862Z","shell.execute_reply.started":"2022-06-16T15:35:12.589819Z","shell.execute_reply":"2022-06-16T15:35:54.802207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Generating submission.csv file...')\n\n# Get image ids from test set and convert to unicode\ntest_ids_ds = test_ds.map(lambda image, idnum: idnum).unbatch()\ntest_ids = next(iter(test_ids_ds.batch(NUM_TEST_IMAGES))).numpy().astype('U')\n\n# Write the submission file\nnp.savetxt(\n    'submission.csv',\n    np.rec.fromarrays([test_ids, predictions]),\n    fmt=['%s', '%d'],\n    delimiter=',',\n    header='id,label',\n    comments='',\n)","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:35:54.804320Z","iopub.execute_input":"2022-06-16T15:35:54.805354Z","iopub.status.idle":"2022-06-16T15:35:57.425761Z","shell.execute_reply.started":"2022-06-16T15:35:54.805301Z","shell.execute_reply":"2022-06-16T15:35:57.424976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Look at the first few predictions\n!head submission.csv","metadata":{"execution":{"iopub.status.busy":"2022-06-16T15:35:57.427305Z","iopub.execute_input":"2022-06-16T15:35:57.427852Z","iopub.status.idle":"2022-06-16T15:35:58.223786Z","shell.execute_reply.started":"2022-06-16T15:35:57.427817Z","shell.execute_reply":"2022-06-16T15:35:58.222777Z"},"trusted":true},"execution_count":null,"outputs":[]}]}