{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":21154,"databundleVersionId":1243559,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Download all important import","metadata":{}},{"cell_type":"markdown","source":"# <center>CNN и Transfer learning</center>","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"**Flower Classification using Transfer Learning**\nThere are millions of beautiful flowers bursting around every corner, & we've been constantly awed by the beauty & uniqueness of each flower. Classifying different flowers from one another is indeed a challenging task, as there's a plethora of flowers to classify & flowers can appear similar to each other. However, classifying different flower species will be advantageous in the fields such as the pharmaceutical industry, botany, agricultural, & trade activities, which is why we thought of performing this task.\n\nThe main aim of this project is to solve a Supervised Image Classification problem of classifying the flower types - rose, daisy, dandelion, sunflower, & tulip. In the end, we'll have a trained model, which can predict the class of the flower using the Convolutional Neural Networks (CNN).\n\nThe dataset consists of 5 classes of flower species - rose, daisy, dandelion, sunflower, & tulip, each having more than 800 images.\n\nThis project follows a basic workflow:\n\nExamine and understand data\n* Build an input pipeline\n* Build the model\n* Train the model\n* Test the model\nInitially, the following libraries were imported to use in the data conversion process:\n\nHere we are going to import necessary libraries:\nMatplotlib for data visualization\nNumpy to perform array and matrix operations.\nOs provides functions for interacting with the operating system.\nPIL Python Imaging Library which provides the python interpreter with image editing capabilities.\nTensorFlow provides a collection of workflows to develop and train models.\nkeras is used to make the implementation of neural networks easy.\nSequential the core idea of Sequential API is simply arranging the Keras layers in a sequential order.","metadata":{}},{"cell_type":"markdown","source":"## Импортируем библиотеки.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nimport os\nimport PIL\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras.models import Sequential  # Corrected import\nfrom tensorflow.keras.layers import Dense, Flatten, Dropout  # Corrected import\nfrom tensorflow.keras.optimizers import Adam  # Corrected import\n","metadata":{"execution":{"iopub.status.busy":"2024-04-19T01:42:31.136054Z","iopub.execute_input":"2024-04-19T01:42:31.136992Z","iopub.status.idle":"2024-04-19T01:42:48.491581Z","shell.execute_reply.started":"2024-04-19T01:42:31.136953Z","shell.execute_reply":"2024-04-19T01:42:48.489925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Скачаем данные для классификации.","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\n\n# dataset_url = link of the dataset\ndataset_url = 'https://storage.googleapis.com/download.tensorflow.org/example_images/flower_photos.tgz'\ndata_dir = tf.keras.utils.get_file('flower_photos', origin=dataset_url, untar=True)\ndata_dir = Path(data_dir)\n# print the data_dir\nprint(data_dir)     ","metadata":{"execution":{"iopub.status.busy":"2024-04-19T02:29:04.928204Z","iopub.execute_input":"2024-04-19T02:29:04.928656Z","iopub.status.idle":"2024-04-19T02:29:04.937814Z","shell.execute_reply.started":"2024-04-19T02:29:04.928623Z","shell.execute_reply":"2024-04-19T02:29:04.936415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Посмотрим на папки для классов цветов. Для корректной обработки нужно, чтобы каждый цветок был в своей папке с названием . ","metadata":{}},{"cell_type":"code","source":"import os\nd = set()\nfor dirname, _, filenames in os.walk(data_dir):\n    d.add(dirname)\nfor x in d:\n    print(x)","metadata":{"execution":{"iopub.status.busy":"2024-04-19T02:36:33.155395Z","iopub.execute_input":"2024-04-19T02:36:33.155841Z","iopub.status.idle":"2024-04-19T02:36:33.171072Z","shell.execute_reply.started":"2024-04-19T02:36:33.155810Z","shell.execute_reply":"2024-04-19T02:36:33.169395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Получили классы: Розы, Одуванчик, Тюльпаны, Подсолнухи, Маргаритка.**","metadata":{}},{"cell_type":"code","source":"# одуванчики\ndandelion = list(data_dir.glob('dandelion/*'))\nprint(dandelion[0])\nim = PIL.Image.open(str(dandelion[0]))","metadata":{"execution":{"iopub.status.busy":"2024-04-19T02:50:36.424993Z","iopub.execute_input":"2024-04-19T02:50:36.425450Z","iopub.status.idle":"2024-04-19T02:50:36.440435Z","shell.execute_reply.started":"2024-04-19T02:50:36.425403Z","shell.execute_reply":"2024-04-19T02:50:36.438982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(im)","metadata":{"execution":{"iopub.status.busy":"2024-04-19T02:50:52.823444Z","iopub.execute_input":"2024-04-19T02:50:52.823863Z","iopub.status.idle":"2024-04-19T02:50:53.379652Z","shell.execute_reply.started":"2024-04-19T02:50:52.823834Z","shell.execute_reply":"2024-04-19T02:50:53.378388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"im.size","metadata":{"execution":{"iopub.status.busy":"2024-04-19T02:51:15.976299Z","iopub.execute_input":"2024-04-19T02:51:15.977345Z","iopub.status.idle":"2024-04-19T02:51:15.984648Z","shell.execute_reply.started":"2024-04-19T02:51:15.977307Z","shell.execute_reply":"2024-04-19T02:51:15.983462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Создадим тренировочную и валидационную выборку по всем картинкам из папок выше. Картинки изменим в размере, сделаем 180 x 180.**\n\n\nМодуль tensorflow.keras.utils.image_dataset_from_directory() имеет несколько аргументов, подробнее про каждый можно прочитать в документации к tensorflow.\n* В первый аргумент 'directory=train' передается путь к той самой папке 'train'. \n\n* Дальше идет 'validation_split', который отвечает за разделение файлов на обучающую и тестовую выборку. \n\n* Если не указывать его, то все 100% будут использоваться для обучения, будем делить датасет в пропорциях 80/20, так обучение будет проходить корректнее. \n\n* Соответственно в 'validation_split=0.2' указываем процент, который будет отведен на валидацию, а следующим аргументом subset=\"training\" указываем, что это будет тренировочная выборка. Для проверочный выборки subset=\"validation\". \n\n* Также можно указать параметр 'image_size=(img_height, img_width)', который будет приводить все изображения к одному размеру. \n\n","metadata":{}},{"cell_type":"code","source":"\n# картинки возьмем 180 x 180, тренировочная выборка\nimg_height, img_width =180,180\n\n# take batch size as 32\nbatch_size =32\n\n# Preprocess the training data i.e., train_ds\ntrain_ds = tf.keras.preprocessing.image_dataset_from_directory(\n    data_dir,\n    label_mode ='categorical',\n    validation_split =0.2,\n    subset= 'training',\n    seed =123,\n    image_size=(img_height,img_width),\n    batch_size=batch_size\n    )","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:03:16.524817Z","iopub.execute_input":"2024-04-19T03:03:16.525279Z","iopub.status.idle":"2024-04-19T03:03:17.014660Z","shell.execute_reply.started":"2024-04-19T03:03:16.525248Z","shell.execute_reply":"2024-04-19T03:03:17.013408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preprocess the validation data i.e., val_ds\nval_ds = tf.keras.preprocessing.image_dataset_from_directory(\n    data_dir,\n    label_mode ='categorical',\n    validation_split =0.2,\n    subset= 'validation',\n    seed =123,\n    image_size=(img_height,img_width),\n    batch_size=batch_size\n    )","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:03:18.028341Z","iopub.execute_input":"2024-04-19T03:03:18.028787Z","iopub.status.idle":"2024-04-19T03:03:18.360695Z","shell.execute_reply.started":"2024-04-19T03:03:18.028754Z","shell.execute_reply":"2024-04-19T03:03:18.359704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## метки классов соответствуют названиям папок.","metadata":{}},{"cell_type":"code","source":"class_names = ['daisy', 'dandelion', 'roses', 'sunflowers', 'tulips']\n\n# print the class_names\nprint(class_names)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:12:59.536248Z","iopub.execute_input":"2024-03-27T16:12:59.536697Z","iopub.status.idle":"2024-03-27T16:12:59.544526Z","shell.execute_reply.started":"2024-03-27T16:12:59.536662Z","shell.execute_reply":"2024-03-27T16:12:59.542965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nplt.figure(figsize=(10,10))\nfor images,labels in train_ds.take(1):\n  for i in range(6):\n    ax = plt.subplot(3,3,i+1)\n    plt.imshow(images[i].numpy().astype('uint8'))\n    plt.axis('off')","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:12:20.251482Z","iopub.execute_input":"2024-04-19T03:12:20.252864Z","iopub.status.idle":"2024-04-19T03:12:21.256030Z","shell.execute_reply.started":"2024-04-19T03:12:20.252821Z","shell.execute_reply":"2024-04-19T03:12:21.254177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" ## Создадим модель и обучим. Будем использовать RESNET50 из keras.applications. \n \n Используем сеть RESNET50 с весами обученными на imagenet и Перенос обучения (т.е. берем старые слои без изменения весов, для этого ставим trainable = False для всех старых слоев). \n Убираем слой с классификатором, выставив параметр  include_top = False. ","metadata":{}},{"cell_type":"code","source":"# declare a variable named resnet_model which will be equal to sequential\nresnet_model = Sequential()\n\n# use pretrained_model\npretrained_model = keras.applications.ResNet50(include_top=False, input_shape=(180,180,3), pooling='avg',\n                                                classes=5, weights='imagenet')\n\n\nfor layer in pretrained_model.layers:\n        layer.trainable =False\n\nresnet_model.add(pretrained_model)","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:14:05.514990Z","iopub.execute_input":"2024-04-19T03:14:05.515486Z","iopub.status.idle":"2024-04-19T03:14:09.079906Z","shell.execute_reply.started":"2024-04-19T03:14:05.515450Z","shell.execute_reply":"2024-04-19T03:14:09.078537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Добавим полносвязный слой для классификации, 5 выходов по числу классов.","metadata":{}},{"cell_type":"code","source":"# add flatten\nresnet_model.add(Flatten())\n\n# add a Dense layer of 512 neurons and give activation as 'relu'\nresnet_model.add(Dense(512, activation='relu'))\n\n# add a dense layer of 5 neurons and give activation as 'softmax'\nresnet_model.add(Dense(5, activation='softmax'))","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:25:12.078650Z","iopub.execute_input":"2024-04-19T03:25:12.079074Z","iopub.status.idle":"2024-04-19T03:25:12.089781Z","shell.execute_reply.started":"2024-04-19T03:25:12.079044Z","shell.execute_reply":"2024-04-19T03:25:12.088539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# call the summary attribute of the model\nresnet_model.summary()","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:25:14.585061Z","iopub.execute_input":"2024-04-19T03:25:14.585505Z","iopub.status.idle":"2024-04-19T03:25:14.634091Z","shell.execute_reply.started":"2024-04-19T03:25:14.585470Z","shell.execute_reply":"2024-04-19T03:25:14.632809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compile the model with Adam(lr=0.001) optimizer, use categorical_crossentropy as loss, and metrics will be equal to accuracy\nresnet_model.compile(loss='categorical_crossentropy', optimizer=Adam(learning_rate=0.001), metrics=['accuracy'])    ","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:25:31.338812Z","iopub.execute_input":"2024-04-19T03:25:31.339340Z","iopub.status.idle":"2024-04-19T03:25:31.362966Z","shell.execute_reply.started":"2024-04-19T03:25:31.339305Z","shell.execute_reply":"2024-04-19T03:25:31.359538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Обучим сеть. В history запишем значения accuracy, loss для тренировочной и валидационной выборки. ","metadata":{}},{"cell_type":"code","source":"history = resnet_model.fit(train_ds, validation_data=val_ds, epochs=2)\n# train upto 10 epochs\n# fit the model in history variable","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:25:34.705569Z","iopub.execute_input":"2024-04-19T03:25:34.706018Z","iopub.status.idle":"2024-04-19T04:08:44.383467Z","shell.execute_reply.started":"2024-04-19T03:25:34.705988Z","shell.execute_reply":"2024-04-19T04:08:44.381718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12, 6))\n\n# Plot accuracy\nplt.subplot(1, 2, 1)\nplt.plot(history.history['accuracy'], label='Training Accuracy')\nplt.plot(history.history['val_accuracy'], label='Validation Accuracy')\nplt.title('Model Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n\n# Plot loss\nplt.subplot(1, 2, 2)\nplt.plot(history.history['loss'], label='Training Loss')\nplt.plot(history.history['val_loss'], label='Validation Loss')\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-19T04:28:20.672291Z","iopub.execute_input":"2024-04-19T04:28:20.674003Z","iopub.status.idle":"2024-04-19T04:28:21.442306Z","shell.execute_reply.started":"2024-04-19T04:28:20.673853Z","shell.execute_reply":"2024-04-19T04:28:21.441025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Все то же самое, но со слоем dropout.**","metadata":{}},{"cell_type":"code","source":"# declare a variable named resnet_model which will be equal to sequential\nmodel = Sequential()\n\n# use pretrained_model\npretrained_model = keras.applications.ResNet50(include_top=False, pooling='avg',\n                                                classes=5, weights='imagenet')\n\n\nfor layer in pretrained_model.layers:\n        layer.trainable =False\n\nmodel.add(pretrained_model)\nmodel.add(Dense(256, activation='relu'))\nmodel.add(Dropout(0.6))\nmodel.add(Dense(5, activation='softmax'))\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2024-04-19T04:29:34.634418Z","iopub.execute_input":"2024-04-19T04:29:34.635867Z","iopub.status.idle":"2024-04-19T04:29:36.183640Z","shell.execute_reply.started":"2024-04-19T04:29:34.635815Z","shell.execute_reply":"2024-04-19T04:29:36.182286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(loss='categorical_crossentropy', optimizer=Adam(learning_rate=0.001), metrics=['accuracy'])\n\nresnet2 = model.fit(train_ds, validation_data=val_ds, epochs=1)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-19T04:29:46.882352Z","iopub.execute_input":"2024-04-19T04:29:46.882799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(1, figsize = (12,3))\n\nplt.subplot(121)\nplt.plot(resnet2.history['accuracy'])\nplt.plot(resnet2.history['val_accuracy'])\nplt.title('model accuracy')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validation'])\n\nplt.subplot(122)\nplt.plot(resnet2.history['loss'])\nplt.plot(resnet2.history['val_loss'])\nplt.title('model loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validation'])\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Conclusion**:\nwe embarked on a journey to classify flower species using Transfer Learning with Convolutional Neural Networks (CNNs). Here's a summary of our workflow:\n\n1. **Data Examination:** We started by examining the dataset, which consisted of five classes of flower species: rose, daisy, dandelion, sunflower, and tulip. Each class contained more than 800 images.\n\n2. **Input Pipeline:** We built an input pipeline to preprocess the image data, including resizing, normalization, and augmentation.\n\n3. **Model Building:** Utilizing Transfer Learning, we employed a pre-trained CNN model (such as ResNet or MobileNet) as the base model and added a few additional layers on top for fine-tuning. This allowed us to leverage the pre-trained model's learned features while adapting to our specific task.\n\n4. **Model Training:** We trained our model on the training data while validating its performance on the validation set. We compiled the model with appropriate loss function, optimizer, and metrics and trained it for several epochs.\n\n5. **Model Evaluation:** After training, we evaluated the model's performance on the test set to assess its generalization ability. We visualized the model's accuracy and loss over epochs to understand its training progress.\n\nThroughout this project, we harnessed the power of deep learning to classify flower species accurately. Such models have vast applications in various domains, including agriculture, botany, and pharmaceuticals. By continuing to refine and optimize our models, we can unlock even more insights from image data and contribute to solving real-world challenges.","metadata":{}},{"cell_type":"markdown","source":"# Задания","metadata":{}},{"cell_type":"markdown","source":"1. **Попробовать дообучить несколько  моделей из Keras (Xception, DenseNet, EfficientNet и т.д.)** из списка https://keras.io/api/applications/, сравнить точность c обычной CNN. Число эпох - небольшое,\nчтобы долго не ждать. \n2. **Добавлять слои аугментации, batchnormalization, dropout**\n3. **Попробовать дообучить RESNET50 на наборе Stanford Dogs Dataset** отсюда https://www.kaggle.com/datasets/jessicali9530/stanford-dogs-dataset. \n4. **Сложное задание.** Поучаствуйте в соревновании. \nВ этом соревновании вам предлагается создать модель машинного обучения, которая идентифицирует тип цветов (роза, тюльпан и т.д.) в наборе данных изображений (чуть более 100 типов). \n\n**Задача**: \n\n1. Надо сдать файл с предсказанными метками цветов по [адресу](https://www.kaggle.com/competitions/tpu-getting-started).\n2. Обязательно: использовать KERAS.\n3. Надо использовать предобученные модели из KERAS.APPLICATIONS и перенос обучения. \n\n\n**Пример.** см. прикрепленный файл submission.csv.\n\n\n**Оценка посылки производится по усредненной метрике F1.** Для каждого класса вычисляется метрика F1, потом усредняется. \n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}