{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"%matplotlib inline\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plant2021 - Preprocessing\n\nIn this notebook, we downscale the image data for the Plant Pathology 2021 competition. In this way, we obtain image files that are smaller than the original data by a factor of 0.2. The downscaled images are saved in a zip file. In addition, the image files are transformed into segmented images by a k-mean cluster processing. Since it takes a fairly long time to process all the image files, we limited ourselves to 200 images being processed in this notebook.\n\nThe preprocessed images are also available as a kaggle data set.\n\n* Plant2021 - Downscaled Images Dataset\n* [Plant2021 - Segmented Images Dataset](www.kaggle.com/dataset/9cdcc447902d2a313a2c8a3837029baf103fd82287e888b3190ddf1c7a2cfd09)\n \n","metadata":{}},{"cell_type":"markdown","source":"## Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport PIL\nimport shutil\n\nimport skimage.io as io\nimport skimage.feature\nfrom skimage import color\nfrom skimage import segmentation\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom zipfile import ZipFile\nfrom tqdm.notebook import tqdm","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rc('font', size=15)\nplt.rc('axes', titlesize=18)  \nplt.rc('xtick', labelsize=10)  \nplt.rc('ytick', labelsize=10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Config: \n    \"\"\"\n    \"\"\"\n    DATA_PATH = '../input/plant-pathology-2021-fgvc8'\n    ZIP_ARCHIVE = 'downscaled_images'\n    ZIP_ARCHIVE_SEGMENTED = 'segmented_images'\n    SCALE_FACTOR = 0.2\n    REMOVE_FOLDERS = False\n    RANDOM_STATE = 2021\n    MAX_IMAGES_PROCESSED = 200\n    \n    folders = dict({\n        'data': DATA_PATH,\n        'train': os.path.join(DATA_PATH, 'train_images'),\n        'test': os.path.join(DATA_PATH, 'test_images'),\n        'downscaled': os.path.join('./', 'downscaled_images'),\n        'segmented': os.path.join('./', 'segmented_images'),\n    })","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import shutil\n\nif Config.REMOVE_FOLDERS:\n    if os.path.exists(Config.folders['downscaled']):\n        shutil.rmtree(Config.folders['downscaled'])\n\n    if os.path.exists(Config.folders['segmented']):\n        shutil.rmtree(Config.folders['segmented'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not os.path.exists(Config.folders['downscaled']):\n    os.mkdir(Config.folders['downscaled'])\n    \nif not os.path.exists(Config.folders['segmented']):\n    os.mkdir(Config.folders['segmented'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load images labels","metadata":{}},{"cell_type":"code","source":"def read_image_labels(data_path=Config.folders['data']):\n    \"\"\"\n    \"\"\"\n    fname = os.path.join(data_path, 'train.csv')\n    df = pd.read_csv(fname).set_index('image')\n    \n    return df\n\nimg_labels = read_image_labels()\nimg_labels.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_downscaled = len(os.listdir(Config.folders['downscaled']))\nn_segmented= len(os.listdir(Config.folders['segmented']))\n\nprint(f'images           : {img_labels.shape[0]}')\nprint(f'downscaled images: {n_downscaled}')\nprint(f'segmented images : {n_segmented}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_label_info(img_labels):\n    \"\"\"\n    \"\"\"\n    df = img_labels.reset_index().groupby(by='labels').count().reset_index()\n    df.columns = ['disease', 'count']\n    \n    df['%'] = np.round((df['count'] / img_labels.shape[0]), 2) * 100\n    df = df.set_index('disease').sort_values(by='count', ascending=False)\n\n    return df\n\nget_label_info(img_labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_label_counts(img_labels):\n    fig, ax = plt.subplots(figsize=(15, 8))\n    sns.set_style(\"whitegrid\")\n    palette = sns.color_palette(\"Blues_r\", 12)\n\n    sns.countplot(\n        x='labels', \n        palette=palette,\n        data=img_labels,\n        order=img_labels['labels'].value_counts().index,\n    );\n\n    plt.ylabel(\"# of observations\", size=20);\n    plt.xlabel(\"Class names\", size=20)\n\n    plt.xticks(rotation=45)\n    \n    fig.tight_layout()\n    plt.show()\n    \n    \nplot_label_counts(img_labels)    ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create downscaled images","metadata":{}},{"cell_type":"code","source":"def create_downscaled_images(\n    img_labels,\n    folder=Config.folders['downscaled'],\n    zip_archive=Config.ZIP_ARCHIVE, \n) -> None:\n    \"\"\"\n    \"\"\"\n    if not os.path.exists(folder):\n        return\n    \n    already_processed  = pd.Series(os.listdir(folder))\n    labels = img_labels.loc[~img_labels.index.isin(already_processed)]\n    \n    if len(labels.index) == 0:\n        print('No images found to downscale.')\n        return\n\n    labels = labels.head(Config.MAX_IMAGES_PROCESSED)\n    progress = tqdm(enumerate(labels.index), total=labels.shape[0])\n\n    for idx, image_id in progress:\n        fname =  os.path.join(Config.folders['train'], image_id)\n        img = PIL.Image.open(fname)\n\n        scale_factor = Config.SCALE_FACTOR\n        img = img.resize([int(scale_factor * s) for s in img.size])\n\n        fname =  os.path.join(folder, image_id)\n        img.save(fname)\n        \n    \n    # create archive\n    print(f'Make zip file {zip_archive}.zip')\n    shutil.make_archive(\n        zip_archive, \n        'zip', \n        folder\n    )          ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_downscaled_images(img_labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create segmented images","metadata":{}},{"cell_type":"code","source":"def create_segmented_images(\n    img_labels,\n    source_folder=Config.folders['downscaled'],\n    target_folder=Config.folders['segmented'],\n    zip_archive=Config.ZIP_ARCHIVE_SEGMENTED\n) -> None:\n    \"\"\"Segments image using k-means clustering\n    \"\"\"\n    if not os.path.exists(source_folder):\n        return\n    \n    if not os.path.exists(target_folder):\n        return\n\n    already_processed  = pd.Series(os.listdir(target_folder))\n    labels = img_labels.loc[~img_labels.index.isin(already_processed)]\n    \n    if len(labels.index) == 0:\n        print('No images found for segmentation.')\n        return\n    \n    labels = labels.head(Config.MAX_IMAGES_PROCESSED)\n    progress = tqdm(enumerate(labels.index), total=labels.shape[0])\n    \n    for idx, image_id in progress:\n        fname =  os.path.join(source_folder, image_id)\n        img = io.imread(fname)\n        \n        segmentes = segmentation.slic(\n            img, \n            n_segments=1200, \n            compactness=10, \n            sigma=1, \n            start_label=1\n        )\n        \n        seg_img = color.label2rgb(segmentes, img, kind='avg')\n        \n        fname = os.path.join(target_folder, image_id)\n        io.imsave(fname, seg_img)\n        \n    # create archive\n    print(f'Make zip file {zip_archive}.zip')\n    shutil.make_archive(\n        zip_archive, \n        'zip', \n        target_folder\n    )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_segmented_images(img_labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Images","metadata":{}},{"cell_type":"code","source":"def get_already_processed(\n    img_labels: pd.DataFrame\n) -> pd.DataFrame:\n    \"\"\"\n    \"\"\"\n    idx_downscaled = pd.Index(os.listdir(Config.folders['downscaled']))\n    idx_segmented = pd.Index(os.listdir(Config.folders['segmented']))\n\n    already_processed = idx_downscaled.intersection(idx_segmented)\n    labels = img_labels.loc[img_labels.index.isin(already_processed)]\n\n    return labels\n\nimage_labels = get_already_processed(img_labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def filter_by(img_labels:pd.DataFrame, kind:str=None) -> pd.DataFrame:\n    if kind is None:\n        return img_labels\n    \n    return image_labels[image_labels['labels'] == kind]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_image(image_id, kind='downscaled') -> None:\n    \"\"\"Loads an image from file\n    \"\"\"\n    if kind == 'archive':\n        zip_file = f'{Config.ZIP_ARCHIVE}.zip' \n        with ZipFile(zip_file, 'r') as archive:\n             with archive.open(image_id) as file:\n                return np.array(PIL.Image.open(file))\n\n    fname = os.path.join(Config.folders[kind], image_id)\n    return np.array(PIL.Image.open(fname))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_images(image_ids, labels, nrows=1, ncols=4, kind='downscaled') -> None:\n    \"\"\"\n    \"\"\"\n    if labels.shape[0] == 0:\n        return\n    \n    fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(20, 8))\n    for image_id, label, ax in zip(image_ids, labels, axes.flatten()):\n        image = get_image(image_id, kind=kind)\n        io.imshow(image, ax=ax)\n        \n        ax.set_title(f\"Class: {label}\", fontsize=12)\n        ax.get_xaxis().set_visible(False)\n        ax.get_yaxis().set_visible(False)\n        \n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_images(image_labels.index, image_labels.labels, nrows=2, ncols=4)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_images(image_labels.index, image_labels.labels, nrows=2, ncols=4, kind='segmented')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Images by classes ","metadata":{}},{"cell_type":"code","source":"# healthy\nimages = filter_by(image_labels, kind='healthy')\n\nvisualize_images(images.index, images.labels)\nvisualize_images(images.index, images.labels, kind='segmented')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Venturia inaequalis `scab`\nhttps://www.wikiwand.com/en/Apple_scab","metadata":{}},{"cell_type":"code","source":"# scab\nimages = filter_by(image_labels, kind='scab')\n\nvisualize_images(images.index, images.labels)\nvisualize_images(images.index, images.labels, kind='segmented')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pucciniales `rust`\nhttps://www.wikiwand.com/en/Rust_(fungus)","metadata":{}},{"cell_type":"code","source":"# rust\nimages = filter_by(image_labels, kind='rust')\n\nvisualize_images(images.index, images.labels)\nvisualize_images(images.index, images.labels, kind='segmented')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Botryosphaeria obtusa `frog_eye_leaf_spot`\nhttps://www.wikiwand.com/en/Botryosphaeria_obtusa","metadata":{}},{"cell_type":"code","source":"# frog_eye_leaf_spot\nimages = filter_by(image_labels, kind='frog_eye_leaf_spot')\n\nvisualize_images(images.index, images.labels)\nvisualize_images(images.index, images.labels, kind='segmented')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Podosphaera leucotricha `powdery_mildew`\nhttps://www.wikiwand.com/en/Podosphaera_leucotricha","metadata":{}},{"cell_type":"code","source":"# powdery_mildew \nimages = filter_by(image_labels, kind='powdery_mildew')\n\nvisualize_images(images.index, images.labels)\nvisualize_images(images.index, images.labels, kind='segmented')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### `complex`","metadata":{}},{"cell_type":"code","source":"# complex\nimages = filter_by(image_labels, kind='complex')\n\nvisualize_images(images.index, images.labels)\nvisualize_images(images.index, images.labels, kind='segmented')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Read image from archive file","metadata":{}},{"cell_type":"code","source":"image_id = image_labels.iloc[0].name\nimg = get_image(image_id, kind='archive')\n\nio.imshow(img);","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Summary\n\n* The training data consists of a total of 18632 images of apple tree leaves affected by one or more plant diseases (viruses, fungal infections, bacteria, etc.).  \n\n* The images are labeled by the corresponding plant disease.\n\n* There are 12 different classes of plant diseases.\n\n* However, five of these classes represent a grouping of plant diseases. Therefore, there are only six actual classes of plant diseases.\n\n* The leaves without a plant disease is labeled with `healty`.\n\n* The most common plant disease in the dataset is apple scab `scab` with  about 26%.\n\n* About 25% of the data show leaves without plant diseases `healty`.\n\n* 1555 records are assigned to more than one plant disease.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"Thank you for reading. If you find this notebook useful, don't forget to upvote.","metadata":{}}]}