{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## <a name=\"Wheat Detection\">Plant Pathology 2021 FGVC8 </a>\n\n#### <a name=\"About_Competition\"> Giới thiệu </a>\n\nTáo là một trong những loại cây ăn quả ôn đới quan trọng nhất trên thế giới. Bệnh cháy lá  là mối đe dọa lớn đối với năng suất và chất lượng chung của vườn táo. Quy trình chẩn đoán bệnh trên vườn táo hiện nay dựa trên việc dò tìm thủ công của con người, tốn nhiều thời gian và chi phí.\n\nMặc dù các mô hình dựa trên thị giác máy tính đã cho thấy nhiều hứa hẹn trong việc xác định bệnh thực vật, nhưng vẫn còn một số hạn chế cần được giải quyết. Sự khác biệt lớn về các triệu chứng hình ảnh của một bệnh đơn lẻ trên các giống táo khác nhau, hoặc các giống mới có nguồn gốc được trồng trọt, là những thách thức lớn đối với việc xác định bệnh dựa trên thị giác máy tính. Những biến thể này phát sinh do sự khác biệt trong môi trường chụp ảnh và tự nhiên, ví dụ, màu sắc lá và hình thái lá, tuổi của các mô bị nhiễm bệnh, nền ảnh không đồng nhất và độ chiếu sáng khác nhau trong quá trình chụp ảnh, v.v.\n\nPlant Pathology 2021-FGVC8 có tập dữ liệu thí điểm gồm 3.651 hình ảnh RGB về bệnh lá trên quả táo.Tập dữ liệu chứa khoảng 23.000 hình ảnh RGB chất lượng cao về các bệnh trên lá táo, bao gồm một tập dữ liệu lớn về bệnh được chuyên gia chú thích. Bộ dữ liệu này phản ánh các tình huống thực tế bằng cách thể hiện các nền không đồng nhất của hình ảnh chiếc lá được chụp ở các giai đoạn trưởng thành khác nhau và vào các thời điểm khác nhau trong ngày trong các cài đặt máy ảnh tiêu cự khác nhau.\n                           \n\n#### <a name=\"Specific Objectives\">Xác định mục tiêu</a>           \n\nMục tiêu chính của cuộc thi là phát triển các mô hình dựa trên máy học để phân loại chính xác một hình ảnh lá nhất định từ bộ dữ liệu thử nghiệm cho một loại bệnh cụ thể và xác định một bệnh riêng lẻ từ nhiều triệu chứng bệnh trên một hình ảnh lá đơn lẻ.\n\n\n#### <a name=\"Yêu cầu\">Yêu cầu</a>           \n\nMục tiêu chính của cuộc thi là phát triển các mô hình dựa trên máy học để phân loại chính xác một hình ảnh lá nhất định từ bộ dữ liệu thử nghiệm cho một loại bệnh cụ thể và xác định một bệnh riêng lẻ từ nhiều triệu chứng bệnh trên một hình ảnh lá đơn lẻ.\n\n\n#### <a name=\"dataset_description\">Mô tả dữ liệu</a>: \n\nDữ liệu lưu giữ hình ảnh của cây táo. Lá cây khỏe mạnh và bị nhiễm bệnh.\n\nFiles train.csv - dữ liệu tập huấn luyện.\n\nImage - ID của hình ảnh\n\nLabel - các lớp mục tiêu thể hiện tất cả các bệnh được tìm thấy trong hình ảnh. Những lá không tốt có quá nhiều bệnh để phân loại bằng mắt thường sẽ có lớp phức tạp, và cũng có thể có một tập hợp con của các bệnh được xác định.\n\n\nsample_submission.csv - Tệp gửi mẫu ở định dạng:\n\n    1. image\n    2. labels\n\ntrain_images - tập huấn luyện.\ntest_images - bộ thử nghiệm. Cuộc thi này có một bộ thử nghiệm ẩn: chỉ có ba hình ảnh được cung cấp ở đây dưới dạng mẫu trong khi 5.000 hình ảnh còn lại sẽ có sẵn trong sổ ghi chép sau khi nó được gửi.\n\nPhân loại Labels:\n*     healthy\n*     complex\n*     frog_eye_leaf_spot\n*     frog_eye_leaf_spot complex\n*     powdery_mildew\n*     powdery_mildew complex\n*     rust\n*     rust complex\n*     rust frog_eye_leaf_spot\n*     scab\n*     scab frog_eye_leaf_spot\n*     scab frog_eye_leaf_spot complex\n","metadata":{}},{"cell_type":"markdown","source":"# Nội dung\n\n* [<font size=4>EDA</font>](#1)\n    * [Chuẩn bị dữ liệu](#1.1)\n    * [Một số ảnh ví dụ từ tập dữ liệu](#1.2)\n    * [Phân phối RBG](#1.3)\n    * [Parallel categories plot](#1.4)\n","metadata":{}},{"cell_type":"markdown","source":"## Importing các thư viện cần thiết","metadata":{}},{"cell_type":"code","source":"import os\nfrom tqdm import tqdm\n\n# Data Processing Libraries \n\nimport pandas as pd \nimport numpy as np \n\n\n# Feature Engineering Libraries\n\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn import preprocessing\n\n# Data Visualisation libraries \n%matplotlib inline\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nimport cv2\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport plotly.figure_factory as ff\nfrom plotly.subplots import make_subplots\nfrom plotly.offline import init_notebook_mode, iplot\ninit_notebook_mode(connected=True)\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n## Image Augmentation \n\n# skimage\nfrom skimage.io import imshow, imread, imsave\nfrom skimage.transform import rotate, AffineTransform, warp,rescale, resize, downscale_local_mean\nfrom skimage import color,data\nfrom skimage.exposure import adjust_gamma\nfrom skimage.util import random_noise\n\n\n# 3D scatter plot\nfrom mpl_toolkits.mplot3d import Axes3D\nfrom matplotlib import cm\nfrom matplotlib import colors\n\n\n#OpenCV-Python\nimport cv2\n\n# imgaug\nimport imageio\nimport imgaug as ia\nimport imgaug.augmenters as iaa\n\n# Albumentations\nimport albumentations as A\n\nSAMPLE_LEN=100","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Chuẩn bị dữ liệu","metadata":{}},{"cell_type":"code","source":"train_image_path = '../input/plant-pathology-2021-fgvc8/train_images'\ntest_image_path = '../input/plant-pathology-2021-fgvc8/test_images'\ntrain_df_path = '../input/plant-pathology-2021-fgvc8/train.csv'\ntest_df_path = '../input/plant-pathology-2021-fgvc8/sample_submission.csv'","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Đọc dữ liệu\ndf_train = pd.read_csv(train_df_path)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#In dữ liệu\ndf_train.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Kích thước dữ liệu\ndf_train.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n\n<!-- #### <a>Đếm số lượng các labels</a>            -->\n### Đếm số lượng các labels","metadata":{}},{"cell_type":"code","source":"#Số lượng của mỗi label\ndf_train['labels'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# sns.histplot(df_train['labels'].value_counts(sort=True))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Lập biểu đồ","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15,12))\nlabels = sns.barplot(df_train.labels.value_counts().index,df_train.labels.value_counts())\nfor item in labels.get_xticklabels():\n    item.set_rotation(45)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"source = df_train['labels'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure(data=[go.Pie(labels=source.index,values=source.values)])\nfig.update_layout(title='Label distribution')\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Kết Luận\n\n- Tập dữ liệu khá không cân bằng theo biểu đồ hình tròn ở trên\n- Chúng tôi sẽ chọn chiến lược lấy mẫu thích hợp để giải quyết vấn đề này (sử dụng albumentation)","metadata":{}},{"cell_type":"markdown","source":"# Một số ảnh ví dụ từ tập dữ liệu","metadata":{}},{"cell_type":"markdown","source":"Chúng tôi sẽ kiểm tra kích thước của 300 hình ảnh đầu tiên\n\nNhư bạn có thể thấy bên dưới, tất cả các hình ảnh có kích thước khác nhau.","metadata":{}},{"cell_type":"code","source":"img_shapes = {}\nfor image_name in tqdm(os.listdir(train_image_path)[:300]):\n    image = cv2.imread(os.path.join(train_image_path, image_name))\n    img_shapes[image.shape] = img_shapes.get(image.shape, 0) + 1\n\nprint(img_shapes)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_batch(path,image_ids, labels):\n    plt.figure(figsize=(16, 12))\n    \n    for ind, (image_id, label) in enumerate(zip(image_ids, labels)):\n        plt.subplot(3, 3, ind + 1)\n        image = cv2.imread(os.path.join(path, image_id))\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n\n        plt.imshow(image)\n        plt.title(f\"Class: {label}\", fontsize=12)\n        plt.axis(\"off\")\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tmp_df = df_train.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\nvisualize_batch(train_image_path,image_ids,labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label_encoder = preprocessing.LabelEncoder()\n  \n# Label encoding.\ndf_train[\"labels_code\"]= label_encoder.fit_transform(df_train[[\"labels\"]])\ndf_train","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = complex\ntmp_df = df_train[df_train[\"labels_code\"] == 0]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = frog_eye_leaf_spot\ntmp_df = df_train[df_train[\"labels_code\"] == 1]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = frog_eye_leaf_spot complex\ntmp_df = df_train[df_train[\"labels_code\"] == 2]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = healthy\ntmp_df = df_train[df_train[\"labels_code\"] == 3]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = powdery_mildew\ntmp_df = df_train[df_train[\"labels_code\"] == 4]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = powdery_mildew complex\ntmp_df = df_train[df_train[\"labels_code\"] == 5]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = rust\ntmp_df = df_train[df_train[\"labels_code\"] == 6]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = rust complex\ntmp_df = df_train[df_train[\"labels_code\"] == 7]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = rust frog_eye_leaf_spot\ntmp_df = df_train[df_train[\"labels_code\"] == 8]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = scab\ntmp_df = df_train[df_train[\"labels_code\"] == 9]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = scab frog_eye_leaf_spot\ntmp_df = df_train[df_train[\"labels_code\"] == 10]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label = scab frog_eye_leaf_spot complex\ntmp_df = df_train[df_train[\"labels_code\"] == 11]\nprint(f\"Total train images for class 0: {tmp_df.shape[0]}\")\n\ntmp_df = tmp_df.sample(9)\nimage_ids = tmp_df[\"image\"].values\nlabels = tmp_df[\"labels\"].values\n\nvisualize_batch(train_image_path, image_ids, labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Chúng đã plot một vài hình ảnh trong training data ở trên (các giá trị RGB có thể được nhìn thấy bằng cách di chuột qua hình ảnh). Các phần màu xanh lá cây của hình ảnh có giá trị màu xanh lam rất thấp, nhưng ngược lại, các phần màu nâu có giá trị màu xanh lam cao. Điều này cho thấy rằng các phần màu xanh lá cây (healthy) của hình ảnh có giá trị màu xanh lam thấp, trong khi các phần unhealthy có nhiều khả năng có giá trị màu xanh lam cao. \n**Điều này có thể cho thấy rằng kênh màu xanh lam có thể là chìa khóa để phát hiện bệnh trên cây trồng**","metadata":{}},{"cell_type":"code","source":"def load_image(image_id):\n    file_path = image_id\n    image = cv2.imread(train_image_path+'/'+ file_path)\n    return cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n\n# Just take 100 sample images with SAMPLE_LEN=100 for RBG Channel Analysis\n\ntrain_images = df_train[\"image\"][:SAMPLE_LEN].apply(load_image)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"red_values = [np.mean(train_images[idx][:, :, 0]) for idx in range(len(train_images))]\ngreen_values = [np.mean(train_images[idx][:, :, 1]) for idx in range(len(train_images))]\nblue_values = [np.mean(train_images[idx][:, :, 2]) for idx in range(len(train_images))]\nvalues = [np.mean(train_images[idx]) for idx in range(len(train_images))]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Phân phối RBG (Tất cả các giá trị kênh)","metadata":{}},{"cell_type":"markdown","source":"Histofram là một biểu diễn đồ họa cho biết tần suất xuất hiện của các giá trị màu khác nhau trong hình ảnh. Trong không gian màu RGB, các giá trị pixel nằm trong khoảng từ 0 đến 255 trong đó 0 là màu đen và 255 là màu trắng. Phân tích biểu đồ có thể giúp chúng ta hiểu được phân bố độ sáng, độ tương phản và cường độ của hình ảnh. Bây giờ chúng ta hãy xem biểu đồ của một mẫu được chọn ngẫu nhiên từ mỗi danh mục.","metadata":{}},{"cell_type":"markdown","source":"# Phân phối Kênh Đỏ ","metadata":{}},{"cell_type":"code","source":"fig = ff.create_distplot([red_values], group_labels=[\"R\"], colors=[\"red\"])\nfig.update_layout(showlegend=False, template=\"simple_white\")\nfig.update_layout(title_text=\"Phân phối Kênh Đỏ\")\nfig.data[0].marker.line.color = 'rgb(0, 0, 0)'\nfig.data[0].marker.line.width = 0.5\nfig","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Quan sát :\nCác giá trị kênh màu đỏ có vẻ gần như phân phối chuẩn, nhưng hơi lệch về bên trái (Độ lệch âm). Điều này cho thấy rằng kênh màu đỏ có xu hướng tập trung nhiều hơn ở các giá trị cao hơn, vào khoảng 100. Có sự thay đổi lớn về giá trị màu đỏ trung bình trên các hình ảnh.","metadata":{}},{"cell_type":"code","source":"fig = ff.create_distplot([green_values], group_labels=[\"G\"], colors=[\"green\"])\nfig.update_layout(showlegend=False, template=\"simple_white\")\nfig.update_layout(title_text=\"Phân phối Kênh Xanh Lá\")\nfig.data[0].marker.line.color = 'rgb(0, 0, 0)'\nfig.data[0].marker.line.width = 0.5\nfig","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Quan sát: \nGiá trị kênh màu xanh lá cây có phân phối đồng đều hơn giá trị kênh màu đỏ nhưng lệch phải, với đỉnh nhỏ hơn. Sự phân bố cũng có độ lệch bên phải (trái ngược với màu đỏ) và chế độ lớn hơn khoảng 160. Điều này cho thấy rằng màu xanh lá cây rõ nét hơn trong những hình ảnh này so với màu đỏ, điều này có ý nghĩa, bởi vì đây là hình ảnh của những chiếc lá!","metadata":{}},{"cell_type":"markdown","source":"# Distribution of Blue Channel Values","metadata":{}},{"cell_type":"code","source":"fig = ff.create_distplot([blue_values], group_labels=[\"B\"], colors=[\"blue\"])\nfig.update_layout(showlegend=False, template=\"simple_white\")\nfig.update_layout(title_text=\"Phân phối Kênh Xanh Lam\")\nfig.data[0].marker.line.color = 'rgb(0, 0, 0)'\nfig.data[0].marker.line.width = 0.5\nfig","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Quan sát: \n\nKênh màu xanh lam có sự phân bố đồng đều nhất trong số ba kênh màu, với độ lệch tối thiểu (lệch một chút sang trái). Kênh màu xanh lam cho thấy sự thay đổi lớn giữa các hình ảnh trong tập dữ liệu.","metadata":{}},{"cell_type":"markdown","source":"# Tất cả các kênh hợp lại","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\n\nfor idx, values in enumerate([red_values, green_values, blue_values]):\n    if idx == 0:\n        color = \"Red\"\n    if idx == 1:\n        color = \"Green\"\n    if idx == 2:\n        color = \"Blue\"\n    fig.add_trace(go.Box(x=[color]*len(values), y=values, name=color, marker=dict(color=color.lower())))\n    \nfig.update_layout(yaxis_title=\"Mean value\", xaxis_title=\"Color channel\",\n                  title=\"Mean value vs. Color channel\", template=\"plotly_white\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = ff.create_distplot([red_values, green_values, blue_values],\n                         group_labels=[\"R\", \"G\", \"B\"],\n                         colors=[\"red\", \"green\", \"blue\"])\nfig.update_layout(title_text=\"Distribution of red channel values\", template=\"simple_white\")\nfig.data[0].marker.line.color = 'rgb(0, 0, 0)'\nfig.data[0].marker.line.width = 0.5\nfig.data[1].marker.line.color = 'rgb(0, 0, 0)'\nfig.data[1].marker.line.width = 0.5\nfig.data[2].marker.line.color = 'rgb(0, 0, 0)'\nfig.data[2].marker.line.width = 0.5\nfig","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image = train_images[10]\nimshow(image)\nprint(image.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3D scatter plot cho ảnh với hệ màu RGB\n","metadata":{}},{"cell_type":"code","source":"r, g, b = cv2.split(image)\nfig = plt.figure()\naxis = fig.add_subplot(1, 1, 1, projection=\"3d\")\n\npixel_colors = image.reshape((np.shape(image)[0]*np.shape(image)[1], 3))\nnorm = colors.Normalize(vmin=-1.,vmax=1.)\nnorm.autoscale(pixel_colors)\npixel_colors = norm(pixel_colors).tolist()\n\naxis.scatter(r.flatten(), g.flatten(), b.flatten(), facecolors=pixel_colors, marker=\".\")\naxis.set_xlabel(\"Red\")\naxis.set_ylabel(\"Green\")\naxis.set_zlabel(\"Blue\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3D scatter plot cho ảnh với hệ màu HSV","metadata":{}},{"cell_type":"code","source":"hsv_image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\nh, s, v = cv2.split(hsv_image)\nfig = plt.figure()\naxis = fig.add_subplot(1, 1, 1, projection=\"3d\")\n\naxis.scatter(h.flatten(), s.flatten(), v.flatten(), facecolors=pixel_colors, marker=\".\")\naxis.set_xlabel(\"Hue\")\naxis.set_ylabel(\"Saturation\")\naxis.set_zlabel(\"Value\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Parallel categories plot","metadata":{}},{"cell_type":"code","source":"df_train['label_list'] = df_train['labels'].str.split(' ')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Distinct List of labels \n\n\n\n*     healthy\n*     complex\n*     rust\n*     frog_eye_leaf_spot\n*     powdery_mildew\n*     scab","metadata":{}},{"cell_type":"code","source":"lbls = ['healthy','complex','rust','frog_eye_leaf_spot','powdery_mildew','scab']\nfor x in lbls:\n    df_train[x]=0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def lbl_lgc(col,lbl_list):\n    if col in lbl_list:\n        res = 1 \n    else:\n        res = 0\n    return res","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lbls = ['healthy','complex','rust','frog_eye_leaf_spot','powdery_mildew','scab']\n\nfor x in lbls:\n    df_train[x] = np.vectorize(lbl_lgc)(x,df_train['label_list'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_lbl_onehot = pd.get_dummies(df_train, columns=[\"labels\"], prefix=[\"LBL\"])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_lbl_onehot.columns","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(35,20))\nfig = px.parallel_categories(df_train[['healthy','complex','rust','frog_eye_leaf_spot','powdery_mildew','scab']], color=\"healthy\", color_continuous_scale=\"sunset\",\\\n                             title=\"Parallel categories plot of targets\")\nfig","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Quan sát: \n\nTrong sơ đồ trên, chúng ta có thể thấy mối quan hệ giữa tất cả 6 loại. Đúng như dự đoán, không thể nào một chiếc lá khỏe mạnh lại có thể bị vảy, gỉ sắt, hay nhiều bệnh. Ngoài ra, mỗi chiếc lá không khỏe mạnh đều có một trong các bệnh vảy, gỉ sắt hoặc nhiều bệnh. Tần suất của mỗi kết hợp có thể được nhìn thấy bằng cách di chuột qua plot.","metadata":{}}]}