{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\nimport tensorflow as tf\nimport keras\nimport keras.layers as L\nfrom kaggle_datasets import KaggleDatasets","metadata":{"_uuid":"1e61b8ab-aa51-40b1-9f13-937257a1a39d","_cell_guid":"c0832dae-406b-471e-b12a-23dafeb944ba","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-07-25T16:42:41.159880Z","iopub.execute_input":"2021-07-25T16:42:41.160337Z","iopub.status.idle":"2021-07-25T16:42:41.166729Z","shell.execute_reply.started":"2021-07-25T16:42:41.160299Z","shell.execute_reply":"2021-07-25T16:42:41.165468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2: Distribution Strategy #\n\nA TPU has eight different *cores* and each of these cores acts as its own accelerator. (A TPU is sort of like having eight GPUs in one machine.) We tell TensorFlow how to make use of all these cores at once through a **distribution strategy**. Run the following cell to create the distribution strategy that we'll later apply to our model.","metadata":{"_uuid":"7498f940-2ac5-4446-8aa8-6a82c4c309e4","_cell_guid":"deb9bb2c-b40e-47ff-bb83-df0e0d64df63","trusted":true}},{"cell_type":"code","source":"try:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    print(\"Device:\", tpu.master())\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nexcept:\n    strategy = tf.distribute.get_strategy()\nprint(\"Number of replicas:\", strategy.num_replicas_in_sync)","metadata":{"_uuid":"2cac76d6-783d-4b3e-8970-17aa160efc37","_cell_guid":"971f9080-1586-4e38-90ef-c7817e298edd","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-07-25T16:42:41.168546Z","iopub.execute_input":"2021-07-25T16:42:41.168975Z","iopub.status.idle":"2021-07-25T16:42:46.919448Z","shell.execute_reply.started":"2021-07-25T16:42:41.168906Z","shell.execute_reply":"2021-07-25T16:42:46.918244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We'll use the distribution strategy when we create our neural network model. Then, TensorFlow will distribute the training among the eight TPU cores by creating eight different *replicas* of the model, one for each core.\n\n# Step 3: Loading the Competition Data #\n\n## Get GCS Path ##\n\nWhen used with TPUs, datasets need to be stored in a [Google Cloud Storage bucket](https://cloud.google.com/storage/). You can use data from any public GCS bucket by giving its path just like you would data from `'/kaggle/input'`. The following will retrieve the GCS path for this competition's dataset.","metadata":{"_uuid":"cca99f8f-7b07-49b7-b3f8-5b0e5fe387e8","_cell_guid":"579d283a-3e05-4870-bcbc-ec4ecdebf121","trusted":true}},{"cell_type":"code","source":"GCS_DS_Path = KaggleDatasets().get_gcs_path('tpu-getting-started')\nprint(GCS_DS_Path)","metadata":{"_uuid":"5f15a643-273c-405a-9347-bdd8b6d43026","_cell_guid":"91fd580c-e573-444f-a337-3bc24c710c18","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-07-25T16:42:46.922397Z","iopub.execute_input":"2021-07-25T16:42:46.922741Z","iopub.status.idle":"2021-07-25T16:42:47.350750Z","shell.execute_reply.started":"2021-07-25T16:42:46.922709Z","shell.execute_reply":"2021-07-25T16:42:47.349520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can use data from any public dataset here on Kaggle in just the same way. If you'd like to use data from one of your private datasets, see [here](https://www.kaggle.com/docs/tpu#tpu3pt5).\n\n## Load Data ##\n\nWhen used with TPUs, datasets are often serialized into [TFRecords](https://www.kaggle.com/ryanholbrook/tfrecords-basics). This is a format convenient for distributing data to each of the TPUs cores. We've hidden the cell that reads the TFRecords for our dataset since the process is a bit long. You could come back to it later for some guidance on using your own datasets with TPUs.","metadata":{"_uuid":"320f09b7-465c-4f3c-b8c4-d7b3f8ba9c13","_cell_guid":"fde76afb-d8da-4c7f-8230-e71e5a70d056","trusted":true}},{"cell_type":"code","source":"\nIMAGE_SIZE = [224,224]\nGCS_PATH = GCS_DS_Path + '/tfrecords-jpeg-224x224'\nAUTO = tf.data.experimental.AUTOTUNE\n\ntraining_file = tf.io.gfile.glob(GCS_PATH+'/train/*.tfrec') \ntest_file = tf.io.gfile.glob(GCS_PATH+'/test/*.tfrec')\nvalid_file = tf.io.gfile.glob(GCS_PATH+'/val/*.tfrec')\n","metadata":{"_uuid":"45cf7456-2b10-4773-93db-0ac632642cfc","_cell_guid":"c287a033-eeeb-4327-a785-d7f853c7d334","collapsed":false,"_kg_hide-input":true,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-07-25T16:42:47.352813Z","iopub.execute_input":"2021-07-25T16:42:47.353518Z","iopub.status.idle":"2021-07-25T16:42:47.562425Z","shell.execute_reply.started":"2021-07-25T16:42:47.353463Z","shell.execute_reply":"2021-07-25T16:42:47.561274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndef decode_image(image_data):\n    image = tf.image.decode_jpeg(image_data, channels=3)\n    image = tf.cast(image, tf.float32) / 255.0  # convert image to floats in [0, 1] range\n    image = tf.reshape(image, [*IMAGE_SIZE, 3]) # explicit size needed for TPU\n    return image\n\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"class\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    label = tf.cast(example['class'], tf.int32)\n    return image, label # returns a dataset of (image, label) pairs\n\ndef read_unlabeled_tfrecord(example):\n    UNLABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"id\": tf.io.FixedLenFeature([], tf.string),  # shape [] means single element\n        # class is missing, this competitions's challenge is to predict flower classes for the test dataset\n    }\n    example = tf.io.parse_single_example(example, UNLABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    idnum = example['id']\n    return image, idnum # returns a dataset of image(s)\n\ndef load_dataset(filenames, labeled=True, ordered=False):\n    # Read from TFRecords. For optimal performance, reading from multiple files at once and\n    # disregarding data order. Order does not matter since we will be shuffling the data anyway.\n\n    ignore_order = tf.data.Options()\n    if not ordered:\n        ignore_order.experimental_deterministic = False # disable order, increase speed\n\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO) # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order) # uses data as soon as it streams in, rather than in its original order\n    dataset = dataset.map(read_labeled_tfrecord if labeled else read_unlabeled_tfrecord, num_parallel_calls=AUTO)\n    # returns a dataset of (image, label) pairs if labeled=True or (image, id) pairs if labeled=False\n    return dataset","metadata":{"execution":{"iopub.status.busy":"2021-07-25T16:42:47.563855Z","iopub.execute_input":"2021-07-25T16:42:47.564239Z","iopub.status.idle":"2021-07-25T16:42:47.576914Z","shell.execute_reply.started":"2021-07-25T16:42:47.564205Z","shell.execute_reply":"2021-07-25T16:42:47.575520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create Data Pipelines ##\n\nIn this final step we'll use the `tf.data` API to define an efficient data pipeline for each of the training, validation, and test splits.","metadata":{"_uuid":"5ddc69ad-ff85-47ab-8a02-227ac4d69aef","_cell_guid":"031992bd-7194-497b-a273-123e70931821","trusted":true}},{"cell_type":"code","source":"\ndef data_augment(image, label):\n    # Thanks to the dataset.prefetch(AUTO)\n    # statement in the next function (below), this happens essentially\n    # for free on TPU. Data pipeline code is executed on the \"CPU\"\n    # part of the TPU while the TPU itself is computing gradients.\n    image = tf.image.random_flip_left_right(image)\n    #image = tf.image.random_saturation(image, 0, 2)\n    return image, label   \n\ndef get_training_dataset():\n    dataset = load_dataset(training_file, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() # the training dataset must repeat for several epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prefetch next batch while training (autotune prefetch buffer size)\n    return dataset\n\ndef get_validation_dataset(ordered=False):\n    dataset = load_dataset(valid_file, labeled=True, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.cache()\n    dataset = dataset.prefetch(AUTO)\n    return dataset\n\ndef get_test_dataset(ordered=False):\n    dataset = load_dataset(test_file, labeled=False, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO)\n    return dataset\n\ndef count_data_items(filenames):\n    # the number of data items is written in the name of the .tfrec\n    # files, i.e. flowers00-230.tfrec = 230 data items\n    n = [int(re.compile(r\"-([0-9]*)\\.\").search(filename).group(1)) for filename in filenames]\n    return np.sum(n)\n\n","metadata":{"_uuid":"18a18cee-b83a-4358-9912-2834f8e64e47","_cell_guid":"a35c5e0c-4784-4bf9-b005-7694d45835eb","collapsed":false,"_kg_hide-input":true,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-07-25T16:42:47.578512Z","iopub.execute_input":"2021-07-25T16:42:47.578865Z","iopub.status.idle":"2021-07-25T16:42:47.597192Z","shell.execute_reply.started":"2021-07-25T16:42:47.578830Z","shell.execute_reply":"2021-07-25T16:42:47.595965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This next cell will create the datasets that we'll use with Keras during training and inference. Notice how we scale the size of the batches to the number of TPU cores.","metadata":{"_uuid":"989c8d09-27e2-42f0-845a-db0306bc2750","_cell_guid":"8a5eedb1-ba64-4c32-aa3d-666e474dfb9c","trusted":true}},{"cell_type":"code","source":"# Define the batch size. This will be 16 with TPU off and 128 (=16*8) with TPU on\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\n\nds_train = get_training_dataset()\nds_valid = get_validation_dataset()\nds_test = get_test_dataset()\n\nprint(\"Training:\", ds_train)\nprint (\"Validation:\", ds_valid)\nprint(\"Test:\", ds_test)\n","metadata":{"_uuid":"4e578616-ed6c-4bed-a271-034ccb8b792b","_cell_guid":"d5d6ef31-dbdc-4815-8d75-f8702167cd52","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-07-25T16:42:47.600179Z","iopub.execute_input":"2021-07-25T16:42:47.600745Z","iopub.status.idle":"2021-07-25T16:42:47.700793Z","shell.execute_reply.started":"2021-07-25T16:42:47.600710Z","shell.execute_reply":"2021-07-25T16:42:47.699861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These datasets are `tf.data.Dataset` objects. You can think about a dataset in TensorFlow as a *stream* of data records. The training and validation sets are streams of `(image, label)` pairs.","metadata":{"_uuid":"32154850-3efa-4987-94cc-2d90bff90ab1","_cell_guid":"44e6116c-a94b-4e83-a626-4680a079332e","trusted":true}},{"cell_type":"code","source":"def convblock(filter_size,is_block2=False):\n    model.add(L.Conv2D(filter_size,kernel_size=(3,3),padding='same',activation='relu'))\n    model.add(L.Conv2D(filter_size,kernel_size=(3,3),padding='same',activation='relu'))\n    if is_block2:\n        model.add(L.Conv2D(filter_size,kernel_size=(3,3),padding='same',activation='relu'))\n    model.add(L.MaxPool2D(pool_size=(2,2),strides=(2,2),padding='same'))","metadata":{"execution":{"iopub.status.busy":"2021-07-25T16:42:47.701780Z","iopub.execute_input":"2021-07-25T16:42:47.702076Z","iopub.status.idle":"2021-07-25T16:42:47.709438Z","shell.execute_reply.started":"2021-07-25T16:42:47.702048Z","shell.execute_reply":"2021-07-25T16:42:47.708099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with strategy.scope():\n    model = keras.Sequential()\n    model.add(L.InputLayer(input_shape=(224,224,3)))\n    convblock(64)\n    \n    convblock(128)\n    \n    convblock(256,is_block2=True)\n\n    convblock(512,is_block2=True)\n    \n    convblock(512,is_block2=True)\n    model.add(L.Flatten())\n    model.add(L.Dense(4096,activation='relu'))\n    model.add(L.Dropout(0.5))\n    model.add(L.Dense(4096,activation='relu'))\n    model.add(L.Dropout(0.5))\n    model.add(L.Dense(104, activation='softmax')) # since our dataset have 104 classes","metadata":{"_uuid":"1094a9ce-8776-417e-9892-3cb48d76030c","_cell_guid":"3fb29691-4b71-4b6d-a530-a9d02cb26ff7","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-07-25T16:42:47.712087Z","iopub.execute_input":"2021-07-25T16:42:47.712421Z","iopub.status.idle":"2021-07-25T16:42:48.993613Z","shell.execute_reply.started":"2021-07-25T16:42:47.712391Z","shell.execute_reply":"2021-07-25T16:42:48.992151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2021-07-25T16:42:48.995689Z","iopub.execute_input":"2021-07-25T16:42:48.996069Z","iopub.status.idle":"2021-07-25T16:42:49.016108Z","shell.execute_reply.started":"2021-07-25T16:42:48.996031Z","shell.execute_reply":"2021-07-25T16:42:49.014371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test set is a stream of `(image, idnum)` pairs; `idnum` here is the unique identifier given to the image that we'll use later when we make our submission as a `csv` file.","metadata":{"_uuid":"026749db-76e4-4f23-b975-b1a0cb8b9405","_cell_guid":"cc7b1c9b-4733-4754-a3f0-67f6826b95d6","trusted":true}},{"cell_type":"code","source":"model.compile(\n    optimizer=keras.optimizers.Adam(learning_rate=0.00001),\n    loss = 'sparse_categorical_crossentropy',\n    metrics=['sparse_categorical_accuracy'],\n)","metadata":{"_uuid":"2c0bca10-faeb-4146-ae95-a743e926bc4f","_cell_guid":"8544d180-d7b0-427b-9c66-ec877bca5759","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-07-25T16:42:49.018038Z","iopub.execute_input":"2021-07-25T16:42:49.018474Z","iopub.status.idle":"2021-07-25T16:42:49.548211Z","shell.execute_reply.started":"2021-07-25T16:42:49.018425Z","shell.execute_reply":"2021-07-25T16:42:49.547165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"NUM_TRAINING_IMAGES = 12753\nNUM_TEST_IMAGES = 7382\nSTEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE","metadata":{"execution":{"iopub.status.busy":"2021-07-25T16:42:49.549673Z","iopub.execute_input":"2021-07-25T16:42:49.550026Z","iopub.status.idle":"2021-07-25T16:42:49.555849Z","shell.execute_reply.started":"2021-07-25T16:42:49.549990Z","shell.execute_reply":"2021-07-25T16:42:49.554664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(\n    ds_train,\n    validation_data=ds_valid,\n    epochs=50,steps_per_epoch=STEPS_PER_EPOCH\n)","metadata":{"execution":{"iopub.status.busy":"2021-07-25T16:42:49.557405Z","iopub.execute_input":"2021-07-25T16:42:49.557701Z","iopub.status.idle":"2021-07-25T16:53:22.242091Z","shell.execute_reply.started":"2021-07-25T16:42:49.557673Z","shell.execute_reply":"2021-07-25T16:53:22.241084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_ds = get_test_dataset(ordered=True)\n\nprint('Computing predictions...')\ntest_images_ds = test_ds.map(lambda image, idnum: image)\nprobabilities = model.predict(test_images_ds)\npredictions = np.argmax(probabilities, axis=-1)\nprint(predictions)","metadata":{"execution":{"iopub.status.busy":"2021-07-25T16:53:22.243250Z","iopub.execute_input":"2021-07-25T16:53:22.243537Z","iopub.status.idle":"2021-07-25T16:53:38.683676Z","shell.execute_reply.started":"2021-07-25T16:53:22.243508Z","shell.execute_reply":"2021-07-25T16:53:38.682425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Generating submission.csv file...')\n\n# Get image ids from test set and convert to unicode\ntest_ids_ds = test_ds.map(lambda image, idnum: idnum).unbatch()\ntest_ids = next(iter(test_ids_ds.batch(NUM_TEST_IMAGES))).numpy().astype('U')\n\n# Write the submission file\nnp.savetxt(\n    'submission.csv',\n    np.rec.fromarrays([test_ids, predictions]),\n    fmt=['%s', '%d'],\n    delimiter=',',\n    header='id,label',\n    comments='',\n)\n\n# Look at the first few predictions\n!head submission.csv","metadata":{"execution":{"iopub.status.busy":"2021-07-25T16:53:38.685173Z","iopub.execute_input":"2021-07-25T16:53:38.685522Z","iopub.status.idle":"2021-07-25T16:53:40.212406Z","shell.execute_reply.started":"2021-07-25T16:53:38.685486Z","shell.execute_reply":"2021-07-25T16:53:40.210950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [Learn Discussion forum](https://www.kaggle.com/learn-forum/161321) to chat with other Learners.*","metadata":{"_uuid":"f832d5db-0b0a-4e26-9cd0-b8907970d976","_cell_guid":"24c6838c-f407-433f-bb1b-ecb6ce40d719","trusted":true}}]}