{"cells":[{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# 1. GLE - EDA(all you need to know)\n- https://www.kaggle.com/rohitsingh9990/glr-eda-all-you-need-to-know"},{"metadata":{},"cell_type":"markdown","source":"## 1. Introduction\nGoogle Landmark Recognition은.. 81K의 클래스가 있다.  \n클래스당 training examples가 많지 않아서 Landmark Recognition을 하는 것은 어려운 일."},{"metadata":{},"cell_type":"markdown","source":"## 2. Preliminaries"},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\n\nimport random\nimport seaborn as sns\nimport cv2\n\n# General packages\nimport pandas as pd\nimport numpy as np\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport PIL\nimport IPython.display as ipd\nimport glob\nimport h5py\nimport plotly.graph_objs as go\nimport plotly.express as px\nfrom PIL import Image\nfrom tempfile import mktemp\n\nfrom bokeh.layouts import column, row\nfrom bokeh.models import ColumnDataSource, LinearAxis, Range1d\nfrom bokeh.models.tools import HoverTool\nfrom bokeh.palettes import BuGn4\nfrom bokeh.plotting import figure, output_notebook, show\nfrom bokeh.transform import cumsum\nfrom math import pi\n\noutput_notebook()\n\nfrom IPython.display import Image, display\nimport warnings\nwarnings.filterwarnings(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir('../input/landmark-recognition-2020/')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 2.1 Loading Data"},{"metadata":{"trusted":true},"cell_type":"code","source":"BASE_PATH = '../input/landmark-recognition-2020'\n\nTRAIN_DIR = f'{BASE_PATH}/train'\nTEST_DIR = f'{BASE_PATH}/test'\n\nprint('Reading data...')\ntrain = pd.read_csv(f'{BASE_PATH}/train.csv')\nsubmission = pd.read_csv(f'{BASE_PATH}/sample_submission.csv')\nprint('Reading data completed')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"다음과 같은 구성\n- ```train.csv``` : id와 targets 포함\n    - ```id``` : 이미지의 id\n    - ```landmark_id``` : target landmark id  \n    \n- training set은 ```train/```폴더에 있고, landmark의 label값들은 ```train.csv```에 있음  \n\n- test set은 ```test/```폴더에 있고, 각 이미지는 unique한 id를 갖고 있음  \n\n> 이미지 수가 겁나 많아서 파일명이 abcdef.jpg이라면 a/b/c/abcdef.jpg의 경로에 저장되어 있음"},{"metadata":{"trusted":true},"cell_type":"code","source":"display(train.head())\nprint(\"Shape of train_data :\", train.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"display(submission.head())\nprint(\"Shape of submission :\", submission.shape)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 3. Let's perform some EDA"},{"metadata":{},"cell_type":"markdown","source":"### 3.1 Target Distribution (Number of images per landmark_id)\n- landmark_id 당 image의 갯수.."},{"metadata":{"trusted":true},"cell_type":"code","source":"# displaying only top 30 landmark\nlandmark = train.landmark_id.value_counts()\nlandmark_df = pd.DataFrame({'landmark_id':landmark.index, 'frequency':landmark.values}).head(30)\n\nlandmark_df['landmark_id'] =   landmark_df.landmark_id.apply(lambda x: f'landmark_id_{x}')\n\nfig = px.bar(landmark_df, x=\"frequency\", y=\"landmark_id\",color='landmark_id', orientation='h',\n             hover_data=[\"landmark_id\", \"frequency\"],\n             height=1000,\n             title='Number of images per landmark_id (Top 30 landmark_ids)')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"어떤 것을 알 수 있을까?\n- landmark_id는 81313개의 유일한 값들이 있다.\n- 2300개 이상의 이미지를 갖고 있는 랜드마크는 단 한개 : landmark_id= 138982\n- 랜드마크당 이미지의 개수는 2장에서 6272장 사이\n- 81313개를 제외하고 나머지 97.5퍼센트의 79298개의 랜드마크는 100장 이하"},{"metadata":{},"cell_type":"markdown","source":"## 4. Let's visualize few images"},{"metadata":{},"cell_type":"markdown","source":"### 4.1 Visualizing random images"},{"metadata":{"trusted":true},"cell_type":"code","source":"import PIL\nfrom PIL import Image, ImageDraw\n\n\ndef display_images(images, title=None): \n    f, ax = plt.subplots(5,5, figsize=(18,22))\n    if title:\n        f.suptitle(title, fontsize = 30)\n\n    for i, image_id in enumerate(images):\n        image_path = os.path.join(TRAIN_DIR, f'{image_id[0]}/{image_id[1]}/{image_id[2]}/{image_id}.jpg')\n        image = Image.open(image_path)\n        \n        ax[i//5, i%5].imshow(image) \n        image.close()       \n        ax[i//5, i%5].axis('off')\n\n        landmark_id = train[train.id==image_id.split('.')[0]].landmark_id.values[0]\n        ax[i//5, i%5].set_title(f\"ID: {image_id.split('.')[0]}\\nLandmark_id: {landmark_id}\", fontsize=\"12\")\n\n    plt.show() ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"samples = train.sample(25).id.values\ndisplay_images(samples)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 4.2 Visualizing landmark with most number of images (landmark_id: 138982)\n- 가장 많은 이미지를 갖고 있는 랜드마크 시각화"},{"metadata":{"trusted":true},"cell_type":"code","source":"samples = train[train.landmark_id == 138982].sample(25).id.values\n\ndisplay_images(samples)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Visualizing landmark with 2nd most number of images (landmark_id : 126637)\n- 이미지가 두번째로 많은 랜드마크 확인"},{"metadata":{"trusted":true},"cell_type":"code","source":"samples = train[train.landmark_id == 126637].sample(25).id.values\n\ndisplay_images(samples)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 4.4 Visualizing landmark with 3rd most number of images(landmark_id:20409)\n- 그 다음 많은 이미지"},{"metadata":{"trusted":true},"cell_type":"code","source":"samples = train[train.landmark_id == 20409].sample(25).id.values\n\ndisplay_images(samples)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 4.5 Visualizing landmark with 4th most number of images (landmark_id: 83144)"},{"metadata":{"trusted":true},"cell_type":"code","source":"samples = train[train.landmark_id == 83144].sample(25).id.values\n\ndisplay_images(samples)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"---"},{"metadata":{},"cell_type":"markdown","source":"# 2. Google Landmark Recognition EDA\n- https://www.kaggle.com/anshuls235/google-landmark-recognition-eda"},{"metadata":{},"cell_type":"markdown","source":"## 라이브러리 import"},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nplt.style.use('fivethirtyeight') ## ???\nimport plotly_express as px\nimport plotly.graph_objects as go\nimport glob\nfrom tqdm.notebook import tqdm\nimport cv2\nimport os\nimport random\nimport seaborn as sns\n\nimport warnings\nwarnings.filterwarnings('ignore')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 데이터를 봅시다"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train = pd.read_csv('/kaggle/input/landmark-recognition-2020/train.csv')\ntest = glob.glob('/kaggle/input/landmark-recognition-2020/test/*/*/*/*.jpg')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Total Train Images: {}'.format(len(df_train))) \nprint('Total Test Images: {}'.format(len(test)))\nprint('Total Unique Landmarks: {}'.format(df_train.landmark_id.nunique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 랜드마크\n### 이미지 개수가 많은 상위 랜드마크와 하위 랜드마크표시"},{"metadata":{"trusted":true},"cell_type":"code","source":"landmarks = df_train.groupby('landmark_id',as_index=False)['id'].count()\\\n    .sort_values('id',ascending=False).reset_index(drop=True)\nlandmarks.rename(columns={'id':'count'},inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def add_text(ax,fontsize=12):\n    for p in ax.patches:\n        x=p.get_bbox().get_points()[:,0]\n        y=p.get_bbox().get_points()[1,1]\n        ax.annotate('{}'.format(int(y)), (x.mean(), y), ha='center', va='bottom',size=fontsize)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, (ax1,ax2) = plt.subplots(2,1,figsize=(16,8))\nsns.barplot(data=landmarks[:50],x='landmark_id',y='count',ax=ax1,color='#30a2da',\n           order=landmarks[:50]['landmark_id'])\nadd_text(ax1,fontsize=8)\nax1.set_title('Top 50 Landmarks')\nax1.set_ylabel('Number of Images')\nax1.set_xticklabels(ax1.get_xticklabels(), rotation=40, ha=\"right\",size=8)\nsns.barplot(data=landmarks[-50:],x='landmark_id',y='count',ax=ax2,color='#fc4f30')\nax2.set_title('Bottom 50 Landmarks')\nax2.set_ylabel('Number of Images')\nax2.set_xticklabels(ax2.get_xticklabels(), rotation=40, ha=\"right\",size=8)\nplt.tight_layout()\nprint(f\"Number of Landmarks with less than 10 images are {len(landmarks[landmarks['count']<10])}\")\nprint(f\"Number of Landmarks with less than 20 images are {len(landmarks[landmarks['count']<20])}\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 랜드마크의 분포 확인"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,4))\nax = sns.distplot(df_train['landmark_id'],bins=500)\nax.set_title('Distribution of Landmarks')\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 랜드마크 이미지 시각화"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_image(id):\n    path = os.path.join('/kaggle/input/landmark-recognition-2020/train',\n                        id[0],id[1],id[2],id+'.jpg')\n    img = cv2.imread(path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    return img","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_data(df,rows,cols):\n    df.reset_index(inplace=True,drop=True)\n    fig = plt.figure(figsize=(24,24))\n    i = 1\n    for r in range(rows):\n        for c in range(cols):\n            id = df.loc[i-1,'id']\n            label = df.loc[i-1,'landmark_id']\n            ax = fig.add_subplot(rows,cols,i)\n            img = get_image(id)\n            ax.set_xticks([])\n            ax.set_yticks([])\n            ax.set_title(label)\n            ax.imshow(img)\n            i+=1\n    return fig","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# 랜덤하게 가져온다\ninds = np.random.choice(df_train.index.tolist(),20)\nfig = show_data(df_train.iloc[inds,:],4,5)\nfig.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"이건 뭐하는 부분이지...  \n- 랜드마크 id를 복사해서 랜덤하게 1000개의 sample을 사용  \n- h, w, c를 0으로 해서 각각의 컬럼 생성  \n- 그리고 랜덤하게 뽑은 1000개의 이미지의 사이즈를 데이터프레임에 넣어준다"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_images = df_train.drop_duplicates(subset=['landmark_id'])\ndf_images = df_images.sample(n=1000,random_state=23)\ndf_images.reset_index(inplace=True,drop=True)\ndf_images['height'] = 0\ndf_images['width'] = 0\ndf_images['channels'] = 0\nfor i in tqdm(range(len(df_images))):\n    img = get_image(df_images.loc[i,'id'])\n    df_images.loc[i,'height'] = img.shape[0]\n    df_images.loc[i,'width'] = img.shape[1]\n    df_images.loc[i,'channels'] = img.shape[2]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"이미지들의 **size**의 분포를 보여주는 것"},{"metadata":{"trusted":true},"cell_type":"code","source":"def img_distribution(df):\n    shape = (np.min(df['width']), np.max(df['width']),\n            np.min(df['height']), np.max(df['height']))\n    fig = px.scatter(df,x='width',y='height')\n    fig.add_shape(\n        x0 = shape[0],\n        x1 = shape[1],\n        y0 = shape[2],\n        y1 = shape[3],\n        fillcolor = 'yellow',\n        opacity=0.3,\n        layer='below'\n    )\n    fig.add_trace(go.Scatter(name='mean',x=[np.mean(df['width'])],y=[np.mean(df['height'])],\n                         marker=dict(color='red',size=10)))\n    #fig.update_traces(marker_line_color='black',marker_line_width=1)\n    fig.update_layout(width=700,height=400,margin=dict(l=0,b=0,r=0,t=40),template='seaborn',\n                 title='Distribution of Image Dimensions', showlegend=False,\n                 xaxis=dict(title='Width', mirror=True, linewidth=2, linecolor='black',showgrid=False),\n                 yaxis=dict(title='Height', mirror=True, linewidth=2, linecolor='black',showgrid=False),\n                 plot_bgcolor='rgb(255,255,255)')\n    return fig","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_distribution(df_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}