{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# # This Python 3 environment comes with many helpful analytics libraries installed\n# # It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# # For example, here's several helpful packages to load\n\n# import numpy as np # linear algebra\n# import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# # Input data files are available in the read-only \"../input/\" directory\n# # For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\n# import os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# # You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# # You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Giới thiệu\nĐây là cuộc thi Plant Pathology 2021. Mục tiêu của cuộc thi này là xác định bệnh của cây dựa trên hình ảnh của lá, có rất nhiều loại bệnh. Giải quyết vấn đề này rất quan trọng vì chẩn đoán bệnh cây sớm có thể tiết kiệm hàng tấn nông sản hàng năm. \nTrong notebook này, chúng ta sẽ khám phá dữ liệu này có những gì, trực quan hóa nó thông qua các hình ảnh, biểu đồ","metadata":{}},{"cell_type":"markdown","source":"**Xác định mục tiêu:**\n\nMục tiêu chính của cuộc thi là phát triển các mô hình dựa trên máy học để phân loại chính xác một hình ảnh lá nhất định từ bộ dữ liệu thử nghiệm cho một loại bệnh cụ thể và xác định một bệnh riêng lẻ từ nhiều triệu chứng bệnh trên một hình ảnh lá đơn","metadata":{}},{"cell_type":"markdown","source":"**Mô tả dữ liệu:**\n\nDữ liệu lưu giữ hình ảnh của cây táo. Lá cây khỏe mạnh và bị nhiễm bệnh.\n\nFiles train.csv - dữ liệu tập huấn luyện.\n\nImage - ID của hình ảnh\n\nLabel - các lớp mục tiêu thể hiện tất cả các bệnh được tìm thấy trong hình ảnh. Những lá không tốt có quá nhiều bệnh để phân loại bằng mắt thường sẽ có lớp phức tạp, và cũng có thể có một tập hợp con của các bệnh được xác định.\n\nsample_submission.csv - Tệp gửi mẫu ở định dạng:\n\n1. image\n2. labels\n\ntrain_images - tập ảnh train.\n\ntest_images - tập ảnh test.","metadata":{}},{"cell_type":"markdown","source":"# Cấu hình","metadata":{}},{"cell_type":"code","source":"import torch\nimport cv2\nimport os\nimport torch.nn as tnn\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport plotly.express as px\nimport seaborn as sns\n\nfrom PIL import Image\nfrom skimage import io, transform\nfrom torchvision.transforms import transforms\nfrom torchvision import utils\nfrom torchvision import datasets\nfrom torch.utils.data import DataLoader, Dataset\nfrom sklearn.preprocessing import MultiLabelBinarizer\nfrom collections import Counter\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Đường dẫn ","metadata":{}},{"cell_type":"code","source":"IMAGE_PATH = \"../input/plant-pathology-2021-fgvc8/train_images/\"\nTEST_IMG_PATH = \"../input/plant-pathology-2021-fgvc8/test_images/\"\nTRAIN_PATH = \"../input/plant-pathology-2021-fgvc8/train.csv\"\nSUB_PATH = \"../input/plant-pathology-2021-fgvc8/sample_submission.csv\"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dữ liệu","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv(TRAIN_PATH)\ntrain_labels","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels['labels'].unique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **Nhận xét:** \nDữ liệu gồm 12 loại nhãn. Trong đó có 6 nhãn chính là  healthy,scab, rust, complex, powdery_mildew và frog_eye_leaf_spot, 6 nhãn còn lại là các nhãn kết hợp từ 6 nhãn chính","metadata":{}},{"cell_type":"markdown","source":"# Thống kê số lượng nhãn","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18,12))\nplt.title(\"Phân phối số lượng ảnh trong các nhãn\",size= 25)\nplt.ylabel(\"Số lượng ảnh\", size=20);\nplt.xlabel(\"Nhãn\", size=20);\nlabels = sns.barplot(train_labels.labels.value_counts().index,train_labels.labels.value_counts())\nfor item in labels.get_xticklabels():\n    item.set_rotation(45)\nplt.savefig('plot.png')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét:\nCác bệnh đơn lẻ chiếm tỉ lệ lớn trong khi các bệnh kết hợp rất hiếm.\n\nKhoảng 51% dữ liệu đầu vào thuộc loại Scab hoặc Healthy. \n\n==> Do đó ta chuyển từ bài toán phân loại một nhãn duy nhất cho ảnh,sang bài toán phân lớp đa nhãn","metadata":{}},{"cell_type":"code","source":"mlb = MultiLabelBinarizer().fit(train_labels.labels.apply(lambda x : x.split()))\nlabels = pd.DataFrame(mlb.transform(train_labels.labels.apply(lambda x : x.split())), columns = mlb.classes_)\n\nlabels = pd.concat([train_labels['image'], labels], axis=1)\nlabels.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Thống kê số lượng ảnh đa nhãn","metadata":{}},{"cell_type":"code","source":"data = ['1','2','3']\nvalue = labels.iloc[:,1:].sum(axis=1).value_counts().values\ncolors = ['mediumturquoise', 'burlywood','sandybrown']\nplt.figure(figsize=(8, 8))\nplt.bar(data, value, color = colors)\nplt.title('Ảnh có nhiều nhãn',fontsize = 14)\nplt.xlabel('Số nhãn',fontsize = 12)\nplt.ylabel('Số lượng ảnh',fontsize = 12)\nplt.savefig('plot2.png')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Nhận xét:**\nTập dữ liệu bị mất cân bằng khá nhiều","metadata":{}},{"cell_type":"markdown","source":"# Kích thước của ảnh","metadata":{}},{"cell_type":"code","source":"img_name = labels.iloc[:,0].tolist()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hs = []\nws = []\nfor i in range(len(img_name)):\n        img = Image.open(IMAGE_PATH+(img_name[i]))\n        h, w = img.size\n        hs.append(h)\n        ws.append(w)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Chiều cao của ảnh","metadata":{}},{"cell_type":"code","source":"labels, values = zip(*Counter(hs).items())\n\nindexes = np.arange(len(labels))\nwidth = 1\n\nplt.bar(indexes, values, width)\nplt.xticks(indexes + width * 0.5, labels)\nplt.savefig('plot4.png')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Chiều rộng của ảnh","metadata":{}},{"cell_type":"code","source":"labels, values = zip(*Counter(ws).items())\n\nindexes = np.arange(len(labels))\nwidth = 1\n\nplt.bar(indexes, values, width)\nplt.xticks(indexes + width * 0.5, labels)\nplt.savefig('plot5.png')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét:\nCác ảnh có kích cỡ không đồng đều, bên cạnh đó dữ liệu ảnh rất lớn\n\n==> Giải pháp: Giảm kích cỡ ảnh, đưa dữ liệu về kích cỡ đồng nhất","metadata":{}},{"cell_type":"markdown","source":"# Dữ liệu theo từng nhãn","metadata":{}},{"cell_type":"code","source":"labels = pd.DataFrame(mlb.transform(train_labels.labels.apply(lambda x : x.split())), columns = mlb.classes_)\n\nlabels = pd.concat([train_labels['image'], labels], axis=1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_batch(path,image_ids, labels):\n    plt.figure(figsize=(16, 12))\n    \n    for ind, (image_id, label) in enumerate(zip(image_ids, labels)):\n        plt.subplot(3, 3, ind + 1)\n        image = cv2.imread(os.path.join(path, image_id))\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n\n        plt.imshow(image)\n        plt.title(f\"Class: {label}\", fontsize=12)\n        plt.axis(\"off\")\n        plt.savefig('plot3.png')\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_s = train_labels.sample(9)\nimage_ids = img_s[\"image\"].values\nlabels_s = img_s[\"labels\"].values\nvisualize_batch(IMAGE_PATH,image_ids,labels_s)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l_complex = labels.loc[labels['complex'] == 1].iloc[:,0].tolist()\nfrog_eye_leaf_spot = labels.loc[labels['frog_eye_leaf_spot'] == 1].iloc[:,0].tolist()\nhealthy = labels.loc[labels['healthy'] == 1].iloc[:,0].tolist()\npowdery_mildew = labels.loc[labels['powdery_mildew'] == 1].iloc[:,0].tolist()\nrust = labels.loc[labels['rust'] == 1].iloc[:,0].tolist()\nscab = labels.loc[labels['scab'] == 1].iloc[:,0].tolist()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Healthy","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(12, 12))\nfor i in range(0,9):\n        img_array = np.array(Image.open(IMAGE_PATH +healthy[i]))\n        fig.add_subplot(3, 3, i+1) \n        plt.imshow(img_array)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Những lá khỏe mạnh là những lá xanh hoàn toàn ","metadata":{}},{"cell_type":"markdown","source":"## Complex","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(12, 12))\nfor i in range(0,9):\n        img_array = np.array(Image.open(IMAGE_PATH +l_complex[i]))\n        fig.add_subplot(3, 3, i+1) \n        plt.imshow(img_array)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Những lá này có màu xanh nhợt nhạt, các đốm vàng nâu","metadata":{}},{"cell_type":"markdown","source":"## Frog Eye Leaf Spot ","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(12, 12))\nfor i in range(0,9):\n        img_array = np.array(Image.open(IMAGE_PATH +frog_eye_leaf_spot[i]))\n        fig.add_subplot(3, 3, i+1) \n        plt.imshow(img_array)","metadata":{"_kg_hide-output":false,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Triệu chứng chẩn đoán rõ nhất của bệnh Frog Eye Leaf Spot là những đốm có góc cạnh với tâm màu xám nhạt và rìa lá màu tím đến nâu đỏ rõ rệt. Không có quầng vàng xung quanh chỗ bệnh. Các đốm lá có thể đơn lẻ hoặc hợp nhất để tạo thành các đốm lớn hơn. ","metadata":{}},{"cell_type":"markdown","source":"## Scab","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(12, 12))\nfor i in range(0,9):\n        img_array = np.array(Image.open(IMAGE_PATH +scab[i]))\n        fig.add_subplot(3, 3, i+1) \n        plt.imshow(img_array)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Trong những hình ảnh trên, chúng ta có thể thấy những chiếc lá nhãn scab có những vết màu nâu lớn và những vết loang lổ khắp mặt lá. ","metadata":{}},{"cell_type":"markdown","source":"## Powdery Mildew","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(12, 12))\nfor i in range(0,9):\n        img_array = np.array(Image.open(IMAGE_PATH +powdery_mildew[i]))\n        fig.add_subplot(3, 3, i+1) \n        plt.imshow(img_array)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Dấu hiệu của các lá bị Powdery Mildew được bao phủ một lớp nấm xám dày đặc như bột phấn hết cả phiến lá. Lớp phấn trắng xuất hiện trên cả thân, cành, hoa làm hoa khô rụng và chết.","metadata":{}},{"cell_type":"markdown","source":"## Rust","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(12, 12))\nfor i in range(0,9):\n        img_array = np.array(Image.open(IMAGE_PATH +rust[i]))\n        fig.add_subplot(3, 3, i+1) \n        plt.imshow(img_array)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thông qua những hình ảnh trên, những chiếc lá bị \"rust\" có một vài đốm màu vàng nâu trên khắp mặt lá.","metadata":{}}]}