{"cells":[{"metadata":{},"cell_type":"markdown","source":"# 1.**Data set overview**"},{"metadata":{},"cell_type":"markdown","source":"The data set is divided into training images in **train_images**, labeled data in **train_tfrecords**, test images in **test_images**, labeled data in **test_tfrecords**, and the corresponding relationship between ****5 category** numbers** and **actual disease categories** found in the **label_num_to_disease_map.json** file"},{"metadata":{},"cell_type":"markdown","source":"# 1.1View the training image, the test image is the same"},{"metadata":{"trusted":true},"cell_type":"code","source":"\nimport os\nimport matplotlib.image as imgplt\nimport matplotlib.pyplot as plt\n\n\ntrain_image_number = len(os.listdir('../input/cassava-leaf-disease-classification/train_images') )\nprint(\"数据集包含了\",train_image_number,\"图像\")\n\n#取出其中一张\nimage_demo = imgplt.imread('../input/cassava-leaf-disease-classification/train_images/1000015157.jpg')\nplt.imshow(image_demo)\n# 图像的尺寸以及通道\nimage_demo.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 1.2 Read the json file information to view the correspondence between the categories in 5"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import json\nwith open ('../input/cassava-leaf-disease-classification/label_num_to_disease_map.json') as file:\n    name = json.loads(file.read())\n    name = {int(k): i for k ,i in name.items()}\nprint(json.dumps(name,indent=4))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"train_tfrecords label information\nThere are a total of 21397 training images. Train_tfrecords is divided into 16 groups. Groups 1-15. Each group contains 1338 images. The 16th group contains 1327 images."},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport seaborn as sn\ndf_train = pd.read_csv('../input/cassava-leaf-disease-classification/train.csv')\n# df_train = df_train[\"label\"].map(name)\ndf_train\nplt.figure(figsize = (8,4))\nsn.countplot(y = 'label',data=df_train)\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is a fairly unbalanced data set. The number of samples for the fourth type of disease exceeds half of the training data set, which will adversely affect the training of the network. We can perform data amplification methods on a small number of samples, such as flipping, enhancing contrast, etc.;\nIn addition, we can use the loss function to excessively punish samples of small categories that are misclassified, and slightly punish large categories."},{"metadata":{},"cell_type":"markdown","source":"# 2.Disease details\n\n2.1 CCB:\nMain characteristics to leverage: angular spots, brown spots with yellow borders, yellow leaves, leaves wilting\n\n\n2.2 CBSD\nMain characteristics to leverage:The infected leaves do not become distorted in shape as occurs with leaves infected by Cassava mosaic disease.\n\n2.3 CGM\nMain characteristics to leverage:yellow patterns, irregular patches of yellow and green, leaf margins distortion, stunted\n\n2.4 CMD\nMain characteristics to leverage: severe shape distortion, mosaic patterns\n\nThere is a strong coupling between the characteristics of the disease, and it is necessary to find more effective classification features","attachments":{}}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}