{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":21154,"databundleVersionId":1243559,"sourceType":"competition"},{"sourceId":12481101,"sourceType":"datasetVersion","datasetId":7875270}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## EfficientNet Model","metadata":{}},{"cell_type":"markdown","source":"EfficientNet is a CNN that uses the concept of compound scaling to scale the image width, depth and resolution in a balanced manner to create a well performing yet efficient model given these three components.\n\nTraditional CNN's choose one of the these dimmensions to scale on, while EfficientNet was designed to scale all three of these dimensions simultaneously using what's referred to as a compound coefficient. This was eloquently described in a medium article by saying, \"Instead of randomly scaling up width, depth or resolution, compound scaling uniformly scales each dimension with a certain fixed set of scaling coefficients.\" (Source 2)\n\nEfficientNet scales the width, depth and resolution using the compound coefficient to balance these three which allows the model to perform better and more efficiently. \n\nEfficientNet comes in different variants that vary in size, performance and number of parameters. EfficientNetB0 is considered to be the baseline model for EfficientNet and can scale all the way up to B7 which would be the largest and most complex model. EfficentNet Models:  B0, B1, B2, B3, B4, B5, B6, B7","metadata":{}},{"cell_type":"markdown","source":"##### Import Libraries","metadata":{}},{"cell_type":"code","source":"## Import the libraries\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense\nimport warnings\nfrom matplotlib import pyplot as plt\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:06.483774Z","iopub.execute_input":"2025-07-28T23:58:06.484419Z","iopub.status.idle":"2025-07-28T23:58:10.211050Z","shell.execute_reply.started":"2025-07-28T23:58:06.484369Z","shell.execute_reply":"2025-07-28T23:58:10.210412Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## EfficientNet B0 - Data Balanced and Hyperparameter Tuned","metadata":{}},{"cell_type":"markdown","source":"To improve from our baseline model we augmented and created new data for all classes that had under 25 samples by applying a moderate random rotation, brightness and zoom along with a random flip across the horizontal (horizontal and vertical flips appeared to be problematic for the flower data since some flower images have an inherent vertical position if taken from a profile view.)\n\nIn order to help out model learn more robust features while preventing overfitting, we utilized a single dense layer with 256 neurons with two dropout layers of .3 and .2.\n\nAdditionally, since EfficientNet uses transfer learning from Imagenet, we froze only the bottom layers of the model to preserve some of the general pre-trained features from Imagenet.\n\nThese updates helped to improve the overall F1 score and this combination of hyperparameter tuning and specific data augmentation produced the highest results for the EfficientNet Model.\n\nOther changes:\n- Increased the batch size to 64 to speed up training and process in parallel while processing on Kaggle's GPU's","metadata":{}},{"cell_type":"markdown","source":"#### Data Preparation and Necessary Pre-Processing","metadata":{}},{"cell_type":"code","source":"## Function to decode and normalize the images and standardize to RGB\nimage_size = [224,224]\n\ndef decode_image(image_data):\n    \"\"\"Function to decode the image from the .tfrec\"\"\"\n    ## Converts the raw JPEG file bytes into a 3D tensor, channels=3  indicates RGB, create the shape (height, width, 3), the output is a uint tensor with values 0-255\n    image = tf.image.decode_jpeg(image_data, channels=3)\n    \n    ## Resize all images to the image size specified above [512,512], using bilinear method to smooth the images for efficient processing\n    image = tf.image.resize(image, image_size, method = \"bilinear\")\n    \n    ## Converts the uint to float32, then normalizes the inputs by dividing by the number of pixel values 255\n    ## Removed /255 because EfficientNet expects input [0-255] not [0-1]\n    image = tf.cast(image, tf.float32) \n    \n    ## Takes the image size defined above this function and reshapes it to be the image size [height, width, 3]\n    image = tf.reshape(image, [*image_size,3])\n    return image","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:18.520526Z","iopub.execute_input":"2025-07-28T23:58:18.521498Z","iopub.status.idle":"2025-07-28T23:58:18.526516Z","shell.execute_reply.started":"2025-07-28T23:58:18.521471Z","shell.execute_reply":"2025-07-28T23:58:18.525569Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Function to load the datasets\ndef load_dataset(filenames, labeled=True):\n    \"\"\"Load TFRecord dataset from filenames\"\"\"\n    ## Creates a dataset that reads the files, AUTOTUNE processes them simultneously and TF optimizes the number of readers\n    ## Creates a dataset of the raw binary .tfrec examples\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=tf.data.AUTOTUNE)\n\n    ## dataset.map applies the read_labeled_tfrec function to each example, AUTOTUNE processes them simultneously and TF optimizes the number of readers\n    ## Transforms the raw data to (image_tensor, label_int) pairs\n    if labeled:\n        dataset = dataset.map(read_labeled_tfrec, num_parallel_calls=tf.data.AUTOTUNE)\n    else:\n        dataset = dataset.map(read_unlabeled_tfrec, num_parallel_calls=tf.data.AUTOTUNE)       \n    return dataset","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:20.719873Z","iopub.execute_input":"2025-07-28T23:58:20.720800Z","iopub.status.idle":"2025-07-28T23:58:20.725404Z","shell.execute_reply.started":"2025-07-28T23:58:20.720766Z","shell.execute_reply":"2025-07-28T23:58:20.724737Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Function to return an image, label pair for the training and validation sets\ndef read_labeled_tfrec(input_example):\n    \"\"\"Read and parse the labeled .tfrec\"\"\"\n    ## Tells Tensorflow how to interpret the binary .tfrec data, \"image\" tells TF to expect binary jppeg bytes, \"class\" tells TF to expect integer labels (the flower labels)\n    labeled_tfrec_format = { \n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"class\": tf.io.FixedLenFeature([], tf.int64)\n    }\n    ## Parses the input_example using the format specified above\n    ## Takes raw binary data (input_example) and uses the labeled_tfrec_format to return a dictionary {image bytes, flower label}\n    input_example = tf.io.parse_single_example(input_example, labeled_tfrec_format)\n    \n    ## Process the image - Takes the JPEG bytes from the example image, and normalizes them to a [512,512,3] tensor using the decode_image function\n    image = decode_image(input_example[\"image\"])\n    \n    ## Process the label - input_example['class'] is the flower class ID\n    label = tf.cast(input_example['class'], tf.int32)\n    return image, label","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:22.424107Z","iopub.execute_input":"2025-07-28T23:58:22.424391Z","iopub.status.idle":"2025-07-28T23:58:22.429516Z","shell.execute_reply.started":"2025-07-28T23:58:22.424367Z","shell.execute_reply":"2025-07-28T23:58:22.428532Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Function to return an image without labels for the test set\ndef read_unlabeled_tfrec(input_example):\n    \"\"\"Read and parse the unlabeled .tfrec\"\"\"\n    unlabeled_tfrec_format = { \n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"id\": tf.io.FixedLenFeature([], tf.string)\n    }    \n    input_example = tf.io.parse_single_example(input_example, unlabeled_tfrec_format)\n    \n    ## Process the image - Takes the JPEG bytes from the example image, and normalizes them to a [512,512,3] tensor using the decode_image function\n    image = decode_image(input_example[\"image\"])\n    \n    ## Process the label - input_example['id'] is the image ID\n    image_id = input_example['id']\n    return image, image_id","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:24.899895Z","iopub.execute_input":"2025-07-28T23:58:24.900424Z","iopub.status.idle":"2025-07-28T23:58:24.904938Z","shell.execute_reply.started":"2025-07-28T23:58:24.900400Z","shell.execute_reply":"2025-07-28T23:58:24.904028Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Get filenames for 224x224 images only\nfolder = 'tfrecords-jpeg-224x224'\ntrain_files = tf.io.gfile.glob(f\"/kaggle/input/tpu-getting-started/{folder}/train/*.tfrec\")\nval_files = tf.io.gfile.glob(f\"/kaggle/input/tpu-getting-started/{folder}/val/*.tfrec\")\ntest_files = tf.io.gfile.glob(f\"/kaggle/input/tpu-getting-started/{folder}/test/*.tfrec\")\n\n## Create train, validation and test data set\ntrain_dataset = load_dataset(train_files, labeled=True)\nvalidation_dataset = load_dataset(val_files, labeled=True)\ntest_dataset = load_dataset(test_files, labeled=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:26.727742Z","iopub.execute_input":"2025-07-28T23:58:26.728027Z","iopub.status.idle":"2025-07-28T23:58:27.315840Z","shell.execute_reply.started":"2025-07-28T23:58:26.728007Z","shell.execute_reply":"2025-07-28T23:58:27.314748Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Balance the dataset by creating augmented data only for the classes that have less than the specified number of samples. \ndef quick_balance(dataset, min_samples=25):\n    ## Count the classes\n    class_counts = {}\n    for _, label in dataset:\n        lbl = int(label.numpy())\n        class_counts[lbl] = class_counts.get(lbl, 0) + 1\n    \n    ## Find minority classes\n    minority_classes = [cls for cls, count in class_counts.items() if count < min_samples]\n    print(f\"Boosting {len(minority_classes)} minority classes\")\n    \n    augmentation = tf.keras.Sequential([\n        tf.keras.layers.RandomFlip(\"horizontal\"),\n        tf.keras.layers.RandomRotation(0.1),\n        tf.keras.layers.RandomBrightness(0.1),\n        tf.keras.layers.RandomZoom(0.1)\n    ])\n\n    minority_data = dataset.filter(lambda img, lbl: tf.reduce_any(tf.equal(lbl, minority_classes)))\n    boosted_minorities = minority_data.map(lambda img, lbl: (augmentation(img, training=True), lbl)).repeat(3)\n    \n    ## Concatenates the training data only with the minority classes augmented data\n    return dataset.concatenate(boosted_minorities)\n\n## Apply balancing to training data only\ntrain_dataset = quick_balance(train_dataset, min_samples=25)\n## Keep validation dataset unchanged","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:29.256047Z","iopub.execute_input":"2025-07-28T23:58:29.256573Z","iopub.status.idle":"2025-07-28T23:58:34.791533Z","shell.execute_reply.started":"2025-07-28T23:58:29.256549Z","shell.execute_reply":"2025-07-28T23:58:34.790719Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Shuffling the training and validation sets\n\n## Sets random seed for reproducability\ntf.random.set_seed(42)\n\n## Set shuffle buffer to set how many are shuffled at once\nshuffle_buffer = 500\n\n## Shuffle the training set\ntrain_dataset = train_dataset.shuffle(shuffle_buffer, seed=42, reshuffle_each_iteration=False)\n\n## Shuffle the validation set\nvalidation_dataset = validation_dataset.shuffle(shuffle_buffer, seed=42, reshuffle_each_iteration=False)\n\n## No shuffling of test data as it's not needed","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:37.728647Z","iopub.execute_input":"2025-07-28T23:58:37.729376Z","iopub.status.idle":"2025-07-28T23:58:38.746148Z","shell.execute_reply.started":"2025-07-28T23:58:37.729348Z","shell.execute_reply":"2025-07-28T23:58:38.745540Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Set hyperparameters\n## Sets the batch to 32 as the number of images to be processed at a time durin one forward/backward pass through the model, 32 is a default\n## Larger batches take more memory and are more stable but if they batch is too large the model may generalize poorly\nBATCH_SIZE = 32\n## Tensorflow optimization that automatically determines the most optimal number of parallel processes\nAUTO = tf.data.AUTOTUNE\n## Sets the number of classes to the number of flower categories\nNUM_CLASSES = 104\n\n## Prepare the datasets for training by grouping a batch based on the size set above, and pre-fetches the next batch while the current batch is being processed\ntrain_dataset = train_dataset.batch(BATCH_SIZE).prefetch(AUTO)\nvalidation_dataset = validation_dataset.batch(BATCH_SIZE).prefetch(AUTO)\ntest_dataset = test_dataset.batch(BATCH_SIZE).prefetch(AUTO)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-28T23:58:43.398512Z","iopub.execute_input":"2025-07-28T23:58:43.399103Z","iopub.status.idle":"2025-07-28T23:58:43.412378Z","shell.execute_reply.started":"2025-07-28T23:58:43.399079Z","shell.execute_reply":"2025-07-28T23:58:43.411709Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def create_b0_model():\n    ## Load the pretrained EfficientNetB0 model\n    ## Set include_top to false to remove the 1000 pre-set image net classification layers because we want to use the 104 flower classification labels\n    ## Set weights to use pretrained ImageNet weights through ImageNet transfer learning\n    ## Set the model to expect inputs with the shape (224, 224, 3)\n    ## Sets the pooling to average to help reduce the number of parameters\n    b0_model = tf.keras.applications.EfficientNetB0(\n        include_top=False,\n        weights='imagenet', \n        input_shape=(224, 224, 3),\n        pooling='avg'\n    )\n    \n    ## Unfreeze the top layers for hyperparameter tuning\n    b0_model.trainable = True\n\n    ## Freeze only the bottom layers of the model\n    tuned = int(len(b0_model.layers) *.85)\n    for layer in b0_model.layers[:tuned]:\n        layer.trainable = False\n    \n    ## Create the model\n    ## Set .Dropout() to help regualrize the model and prevent overfitting, essentially say to leave out 20% of the neurons in training\n    ## Pass NUM_CLASSES to specify the number of output to be 104 flower classes, activation softmax indicates all outputs should sum to 1, output represents probability it is that flower\n    model = tf.keras.Sequential([\n        b0_model,\n        ## Add dropout layer with .3 dropout\n        tf.keras.layers.Dropout(0.4),\n        ## Add dense layer\n        tf.keras.layers.Dense(256, activation='relu'),\n        ## Add another dropout layer with .2 dropout\n        tf.keras.layers.Dropout(0.3),\n        ## This is the output layer\n        tf.keras.layers.Dense(NUM_CLASSES, activation='softmax')\n    ])\n    \n    return model\n\n## Create and compile the model\n## Use the adam optimizer for efficiency and ability to adapt learning rates, sets the learning rate of the optimizer\n## Loss is sparse (meaning not one-hot encoded but numerical 0-103, categorical_crossentropy is used for multiclass clasification\n## Metrics track the accuracy\nmodel = create_b0_model()\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n    loss='sparse_categorical_crossentropy', \n    metrics=['accuracy']\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T00:23:05.752382Z","iopub.execute_input":"2025-07-29T00:23:05.753326Z","iopub.status.idle":"2025-07-29T00:23:06.702961Z","shell.execute_reply.started":"2025-07-29T00:23:05.753289Z","shell.execute_reply":"2025-07-29T00:23:06.702114Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Set training parameters\n## Sets epochs\nEPOCHS = 25\n\n## EarlyStopping - Add callback to monitor the validation loss and stop training after 3 epochs if the loss stops improving and reverts the weights back to where the model was still improving and had converged, this helps prevent overfitting\n##ReduceLROnPlateau - Add callback to monitor the validation loss and reduce the learning rate once performance starts to palteau, factor says multiply the learning rate by this value, patience say wait two epochs before reducing again, min_lr says don't go below this learning rate\ncallbacks = [\n    tf.keras.callbacks.EarlyStopping(\n        monitor='val_loss',\n        patience=4,\n        restore_best_weights=True\n    ),\n    tf.keras.callbacks.ReduceLROnPlateau(\n        monitor='val_loss',\n        factor=0.2,\n        patience=3,\n        min_lr=1e-5\n    )\n]\n\n## Train the model using the train_dataset, number of epochs, specifying the validation set as validation_dataset, and point to the callback defined above\nhistory = model.fit(\n    train_dataset,\n    epochs=EPOCHS,\n    validation_data=validation_dataset,\n    callbacks=callbacks\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T00:23:12.316930Z","iopub.execute_input":"2025-07-29T00:23:12.317416Z","iopub.status.idle":"2025-07-29T00:34:21.613874Z","shell.execute_reply.started":"2025-07-29T00:23:12.317393Z","shell.execute_reply":"2025-07-29T00:34:21.613215Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.summary()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Plot the training history\ndef plot_training_history(history):\n    fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))\n    \n    ## Plot the accuracy\n    ax1.plot(history.history['accuracy'], label='Training')\n    ax1.plot(history.history['val_accuracy'], label='Validation')\n    ax1.set_title('Model Accuracy')\n    ax1.set_xlabel('Epoch')\n    ax1.set_ylabel('Accuracy')\n    ax1.legend()\n    \n    ## Plot the loss\n    ax2.plot(history.history['loss'], label='Training')\n    ax2.plot(history.history['val_loss'], label='Validation')\n    ax2.set_title('Model Loss')\n    ax2.set_xlabel('Epoch')\n    ax2.set_ylabel('Loss')\n    ax2.legend()\n    \n    plt.tight_layout()\n    plt.show()\n\n## Show the training results\nplot_training_history(history)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## Function to assemble the models class predictions\ndef get_predictions():\n    image_ids = []\n    predictions = []\n    \n    print(\"Getting predictions for test set...\")\n    batch_count = 0\n    \n    # Get predictions for test set (now properly batched)\n    for batch_images, batch_image_names in test_dataset:\n        batch_count += 1\n        print(f\"Processing batch {batch_count}, batch shape: {batch_images.shape}\")\n        \n        # Get predictions for this batch\n        pred = model.predict(batch_images, verbose=0)\n        pred_labels = tf.argmax(pred, axis=1)\n        \n        # Convert tensor to numpy for easier handling\n        batch_image_names = batch_image_names.numpy()\n        pred_labels = pred_labels.numpy()\n        \n        # Extend our lists\n        image_ids.extend([name.decode('utf-8') for name in batch_image_names])\n        predictions.extend(pred_labels)\n    \n    print(f\"Total predictions made: {len(predictions)}\")\n    return image_ids, predictions","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def create_submission():\n    # Get predictions\n    image_ids, predictions = get_predictions()\n    \n    # Create submission DataFrame\n    submission_df = pd.DataFrame({\n        'id': image_ids,\n        'label': predictions\n    })\n    \n    # Save submission file\n    submission_df.to_csv('submission.csv', index=False)\n    print(\"Submission file created!\")\n    \n# Create submission after training\ncreate_submission()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.save('kaggle/working/EfficientNetB0.h5')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Sources\n1. https://blog.roboflow.com/what-is-efficientnet/\n2. https://arjun-sarkar786.medium.com/understanding-efficientnet-the-most-powerful-cnn-architecture-eaeb40386fad\n3. https://www.kaggle.com/competitions/tpu-getting-started/discussion/311487\n   - This article was so important for learning not to standardize the pixels by dividing by 255 when using efficientnet since this already standardizes the images","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}