{"cells":[{"metadata":{},"cell_type":"markdown","source":"**What is a herbarium?**\n<br>\nA herbarium is a collection of preserved plants stored, catalogued and arranged systematically for study by both professional taxonomists (scientists who name and identify plants), botanists and amateurs.\n\nThe creation of a herbarium specimen involves the pressing and drying of plants between sheets of paper, a practice that has changed very little since the beginning, 500 years ago. Thanks to this simple technique, most of the characteristics of living plants are visible on the dried plant. The few that are not (e.g. flower colour, scent, height of a tree, vegetation type) are written on the collection label by the collector. Most importantly, the label should tell us where and when the specimen was collected.\n\nA working reference collection\nA herbarium acts like a plant library or vast catalogue with each of our three million specimens providing unique information – where it was found, when it flowered, what it looks like and it’s DNA, which remains intact for many years. DNA is now routinely extracted from herbarium specimens. The most important specimens are called 'types'. The type specimen, chosen by the author of the species name, becomes the physical reference for the new species.\n\nThis unique working reference collection brings species from all over the world together into one place to be discovered, described and compared. The work is disseminated through the writing of Floras (a description of all the plants in a country or region), monographs (a description of plants or fungi within a group, such as a family) and scientific papers. This fundamental research provides an essential baseline for other plant-based research and helps inform conservation practices.\n\n[Click here for further details.[](http://)](https://www.rbge.org.uk/science-and-conservation/herbarium/)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport json\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n\n#To visualise the trend and analyse.\nimport plotly.express as px\nimport plotly.io as pio\npio.templates.default = \"plotly_dark\"\n\nimport plotly.offline as py\nfrom plotly.offline import init_notebook_mode \n\n\npy.init_notebook_mode(connected=True)\n%matplotlib inline\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"Train_data = \"../input/herbarium-2020-fgvc7/nybg2020/train/\"\nTest_data = \"../input/herbarium-2020-fgvc7/nybg2020/test/\"\nMeta_info  = \"metadata.json\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**As the data description can be seen, Dataset is in COCO format and we have to handle it accordingly**<br>\n[For further information on this data fromat click here.](http://cocodataset.org/#format-data)"},{"metadata":{"trusted":true},"cell_type":"code","source":"import codecs\ndef meta_ifo():\n    with codecs.open(Train_data+Meta_info,\"r\",encoding=\"utf-8\",errors=\"ignore\") as f:\n        training_meta_info = json.load(f)\n\n    with codecs.open(Test_data+Meta_info,\"r\",encoding=\"utf-8\",errors=\"ignore\") as f:\n        testing_meta_info = json.load(f)\n        \n    return training_meta_info,testing_meta_info","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_meta_info ,test_meta_info = meta_ifo()\ntrain_meta_info.keys()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **Column Renaming**"},{"metadata":{"trusted":true},"cell_type":"code","source":"annotations = pd.DataFrame(train_meta_info['annotations'])\nannotations.columns = ['category_id', 'id', 'image_id', 'region_id']\n\ncategories = pd.DataFrame(train_meta_info['categories'])\ncategories.columns = ['family', 'genus', 'category_id', 'category_name']\n\nimages = pd.DataFrame(train_meta_info['images'])\nimages.columns = ['image_file_name', 'height', 'image_id', 'license', 'width']\n\nlicenses = pd.DataFrame(train_meta_info['licenses'])\nlicenses.columns = ['licenses_id', 'license_name', 'url']\n\nregions = pd.DataFrame(train_meta_info['regions'])\nregions.columns = ['region_id', 'region_name']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"column_info = {\n                \"categories\":categories.columns,\n                \"annotations\":annotations.columns,\n                \"images\":images.columns,\n                \"licenses\":licenses.columns,\n                \"regions\":regions.columns    \n                }","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataframe = annotations.copy(deep=True)\ndataframe = dataframe.merge(categories,on=\"category_id\",how=\"outer\")\ndataframe = dataframe.merge(images,on=\"image_id\",how=\"outer\")\ndataframe = dataframe.merge(regions,on=\"region_id\",how=\"outer\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataframe.sample(n=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"imageFiles = dataframe.dropna(subset=['image_file_name'])\nimages  = imageFiles['image_file_name'].tolist()\ntrain_images = ['../input/herbarium-2020-fgvc7/nybg2020/train/'+i for i in images]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"imageFiles.tail()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **Training Images Plot**"},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib.image as mpimg\nmax_rows = 5\nmax_cols = 5\npic_index = 0\npic_index += 250\nfig = plt.gcf()\nfig.set_size_inches(max_cols * 5 , max_rows * 5)\n\nfor i, img_path in enumerate(train_images[pic_index - 25:pic_index]):\n    # Set up subplot; subplot indices start at 1\n    sp = plt.subplot(max_rows, max_cols, (i+1))\n    sp.axis('Off')  # Don't show axes (or gridlines)\n    img = mpimg.imread(img_path)\n    plt.imshow(img)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **Exploratory data analysis**"},{"metadata":{"trusted":true},"cell_type":"code","source":"sortedData = dataframe.groupby(by=['category_id'],as_index=False,sort=True)['family'].count().sort_values(['family'], ascending=False)\nsortedData = sortedData.head(n=10000)\nsortedData.columns = [\"Category\",\"Total Specimen\"]\nsortedData.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **Top 100 categories**"},{"metadata":{"trusted":true},"cell_type":"code","source":"df = px.data.gapminder()\n\nfig = px.scatter(sortedData,\n                 x=\"Category\",\n                 y=\"Total Specimen\",\n                 size=\"Total Specimen\",\n                 color=\"Total Specimen\",\n                 hover_name=\"Total Specimen\",\n                 log_x=True,\n                 height=1000,\n                 size_max=60)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"imageFilesCopyDf = imageFiles.copy(deep=True)\nimageFilesCopyDf = imageFiles.groupby([\"height\",\"width\"]).size().reset_index(name='Total')\nimageFilesCopyDf.sort_values(\"Total\",axis=0,ascending=False)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**As it is clearly visible that there are 211 different types of shapes in entire image set.**\n> We need to reshape these images."},{"metadata":{"trusted":true},"cell_type":"code","source":"image_training_dataset = imageFiles[[\"category_id\",\"family\",\"genus\",\"image_file_name\"]]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_training_dataset.sample(n=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split as TTS\ntrain_set , validation_set= TTS(image_training_dataset,test_size=0.2,shuffle=True,random_state=42)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}