{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction\nIn this competition, you are asked to take test images and recognize which landmarks (if any) are depicted in them. The training set is available in the train/ folder, with corresponding landmark labels in train.csv. The test set images are listed in the test/ folder. Each image has a unique id. Since there are a large number of images, each image is placed within three subfolders according to the first three characters of the image id (i.e. image abcdef.jpg is placed in a/b/c/abcdef.jpg).\n","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import numpy as np\nimport math\nimport pandas as pd\nimport glob\nimport os\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom PIL import Image","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"  # Read Dataset","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"PATH = \"../input/landmark-recognition-2020\"\ntrain = pd.read_csv(PATH + \"/train.csv\")\nsample_submission = pd.read_csv(PATH + \"/sample_submission.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Exploratory Data Analysis","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Visualizing some images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"\nimgs = train[train.landmark_id==123]['id'].values\nprint(f\"landmark_id: {123}\")\n_, axs = plt.subplots(2, 5, figsize=(20, 8))\naxs = axs.flatten()\nfor f_name,ax in zip(imgs[:10],axs):\n    img = Image.open(f\"{PATH}/train/{f_name[0]}/{f_name[1]}/{f_name[2]}/{f_name}.jpg\")\n    ax.imshow(img)\n    ax.axis('off')\nplt.show()\n\n\nimgs = train[train.landmark_id==138982]['id'].values\nprint(f\"landmark_id: {138982}\")\n_, axs = plt.subplots(2, 5, figsize=(20, 8))\naxs = axs.flatten()\nfor f_name,ax in zip(imgs[:10],axs):\n    img = Image.open(f\"{PATH}/train/{f_name[0]}/{f_name[1]}/{f_name[2]}/{f_name}.jpg\")\n    ax.imshow(img)\n    ax.axis('off')    \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Data Exploration","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f'Total numbers of images in training set is {train.shape[0]}')\nprint(f'Total numbers of images in test set is {sample_submission.shape[0]}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Top 10 and Bottom 10 Landmarks \n[Source](https://www.kaggle.com/anshuls235/google-landmark-recognition-eda)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"landmarks = train.groupby('landmark_id',as_index=False)['id'].count()\\\n    .sort_values('id',ascending=False).reset_index(drop=True)\nlandmarks.rename(columns={'id':'count'},inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def add_text(ax,fontsize=12):\n    for p in ax.patches:\n        x=p.get_bbox().get_points()[:,0]\n        y=p.get_bbox().get_points()[1,1]\n        ax.annotate('{}'.format(int(y)), (x.mean(), y), ha='center', va='bottom',size=fontsize)\nfig, (ax1,ax2) = plt.subplots(2,1,figsize=(16,8))\nsns.barplot(data=landmarks[:10],x='landmark_id',y='count',ax=ax1,color='#30a2da',\n           order=landmarks[:10]['landmark_id'])\nadd_text(ax1,fontsize=8)\nax1.set_title('Top 50 Landmarks')\nax1.set_ylabel('Number of Images')\nax1.set_xticklabels(ax1.get_xticklabels(), rotation=40, ha=\"right\",size=8)\nsns.barplot(data=landmarks[-10:],x='landmark_id',y='count',ax=ax2,color='#fc4f30')\nax2.set_title('Bottom 50 Landmarks')\nax2.set_ylabel('Number of Images')\nax2.set_xticklabels(ax2.get_xticklabels(), rotation=40, ha=\"right\",size=8)\nplt.tight_layout()\nprint(f\"Number of Landmarks with less than 10 images are {len(landmarks[landmarks['count']<10])}\")\nprint(f\"Number of Landmarks with less than 20 images are {len(landmarks[landmarks['count']<20])}\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Visualizing Data Imbalance ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset = train.landmark_id.value_counts()\ndataset = pd.DataFrame({\"landmark_id\": dataset.index, \"num_of_images\": dataset.values})\nmax_img = dataset.num_of_images.max()\nmin_img = dataset.num_of_images.min()\nprint(f'Total number of classes is: {len(dataset)}')\nprint(f'maximum image for a landmark class is:{max_img}, minimum image for landmark class is:{min_img}')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plot\nplt.scatter(dataset.index, dataset.num_of_images, alpha=0.5)\nplt.title('No. of images Vs. Class')\nplt.xlabel('Classes')\nplt.ylabel('No. of images')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"num_img =50\nper = len(dataset[dataset.num_of_images<num_img])/len(dataset)*100\nprint(f'There are {int(per)}% classes having less than {num_img} images')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is huge class imbalance in dataset. landmark (138982) having highest images which is equal to 6272. Some of clasees having 2 images also.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Inferences\n* Training set contains 1580470 images\n* Test set contains 10345 images\n* Total number of classes is: 81313\n* maximum image for a class is: 6272, \n* minimum image for a class is: 2\n* There are 92% classes having less than 50 images","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"####  If you like it. Feel free to upvote it !!!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **Model Coming Soon !!!**","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}