{"cells":[{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\nsns.set_palette(\"husl\")\nimport os\nprint(os.listdir(\"../input/\"))\nimport warnings\nwarnings.filterwarnings('ignore')\nimport gc\nfrom pathlib import Path\nfrom PIL import Image\nfrom IPython.display import clear_output\nfrom tqdm import tqdm_notebook as tqdm","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"# Let's recognize artwork attributes from The Metropolitan Museum of Art"},{"metadata":{},"cell_type":"markdown","source":"## Note\nThis is a Kernels-only competition  \nSubmissions to this competition must be made through Kernels. In order for the \"Submit to Competition\" button to be active after a commit, the following conditions must be met:  \n\n* 9 hour runtime limit (including GPU Kernels)  \n* No internet access enabled  \n* Only whitelisted data is allowed  \n* No custom packages  \n* Submission file must be named \"submission.csv\"  \nPlease see the [Kernels-only](https://www.kaggle.com/docs/competitions#kernels-only-FAQ) FAQ for more information on how to submit."},{"metadata":{},"cell_type":"markdown","source":"## Files\nThe filename of each image is its id.  \n* **train.csv** gives the attribute_ids for the train images in **/train**\n* **/test** contains the test images. You must predict the attribute_ids for these images.\n* **sample_submission.csv** contains a submission in the correct format\n* **labels.csv** provides descriptions of the attributes"},{"metadata":{},"cell_type":"markdown","source":"## labels.csv\n* labels.csv provides descriptions of the attributes"},{"metadata":{"trusted":true},"cell_type":"code","source":"labels_df = pd.read_csv(\"../input/labels.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels_df.tail()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(f\"labels.csv have {labels_df.shape[0]} attributes_name.\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's check the number of culture and tag."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"kind_dict = {}\nfor i in range(len(labels_df)):\n    kind, name = labels_df.attribute_name[i].split(\"::\")\n    if(kind in kind_dict.keys()):\n        kind_dict[kind] += 1\n    else:\n        kind_dict[kind] = 1\nfor key, val in kind_dict.items():\n    print(\"The number of {} is {}({:.2%})\".format(key, val, val/len(labels_df)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label_dict = labels_df.attribute_name.to_dict()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## train.csv\n* train.csv gives the attribute_ids for the train images in /train  "},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(\"../input/train.csv\")\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Check the amount of train/test data!"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"test_path = Path(\"../input/test/\")\ntest_num = len(list(test_path.glob(\"*.png\")))\ntrain_num = len(train_df)\nfig, ax = plt.subplots()\nsns.barplot(y=[\"train\", \"test\"], x=[train_num, test_num])\nax.set_title(\"The amount of data\")\nclear_output()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"Test data is small.  \nBecause this competition is 2-stage.  \nWe can see the**The second-stage test set is approximately five times the size of the first.** in [Data Description](https://www.kaggle.com/c/imet-2019-fgvc6/data).  \nSo please care your memory usage in kernel!"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"id_len_dict = {}\nid_num_dict = {}\nfor i in range(train_df.shape[0]):\n    ids = list(map(int, train_df.attribute_ids[i].split()))\n    id_len = len(ids)\n    if(id_len in id_len_dict.keys()):\n        id_len_dict[id_len] += 1\n    else:\n        id_len_dict[id_len] = 1\n    for num in ids:\n        if(num in id_num_dict.keys()):\n            id_num_dict[num] += 1\n        else:\n            id_num_dict[num] = 1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Check the number of attribute_id per image and appearance frequency of attribute!"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(2, 1, figsize=(12, 8))\nsns.barplot(x=list(id_len_dict.keys()), y=list(id_len_dict.values()), ax=ax1)\nax1.set_title(\"The number of attribute_id per image\")\nax2.bar(list(id_num_dict.keys()), list(id_num_dict.values()))\nax2.set_title(\"Appearance frequency of attribute\")\nax2.set_xticks(np.linspace(0, max(id_num_dict.keys()), 10, dtype='int'))\nclear_output()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"id_len_list = sorted(id_len_dict.items(), key=lambda x: -x[1])\nprint(\"The number of attribute_id per image\\n\")\nprint(\"{0:9s}{1:20s}\".format(\"label num\".rjust(9), \"amount\".rjust(20)))\nfor i in id_len_list:\n    print(\"{0:9d}{1:20d}\".format(i[0], i[1]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"id_num_list = sorted(id_num_dict.items(), key=lambda x: -x[1])\nprint(\"Top 10 high appearance frequency attitude\\n\")\nprint(\"{0:4s}{1:15s}{2:30s}\".format(\"rank\".rjust(4), \"num\".rjust(15), \"attitude_name\".rjust(30)))\nfor i in range(10):\n    print(\"{0:3d}.{1:15d}{2:30s}\".format(i+1, id_num_list[i][1], (label_dict[id_num_list[i][0]]).rjust(30)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"id_num_list = sorted(id_num_dict.items(), key=lambda x: x[1])\nprint(\"Top 16 low appearance frequency attitude\\n\")\nprint(\"{0:4s}{1:15s}{2:30s}\".format(\"rank\".rjust(4), \"num\".rjust(15), \"attitude_name\".rjust(50)))\nfor i in range(16):\n    print(\"{0:3d}.{1:15d}{2:50s}\".format(i+1, id_num_list[i][1], (label_dict[id_num_list[i][0]]).rjust(50)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* The number of attribute_id per image  \nMost data have 1-6 labels.  \nBut few data have 7-11 labels.  \nThe data that have 11 labels is only one!!!\n\n\n* Appearance frequency of attribute  \nTop2 high appearance frequency tag is \"men\" and \"women\".  \nTop3 high appearance frequency culture is \"french\", \"italian\" and \"american\".  \nLow appearance frequency attitude is too many.  \nI think we need care this low attitudes.  "},{"metadata":{},"cell_type":"markdown","source":"## train/test images\nLet's show training images."},{"metadata":{"_kg_hide-output":false,"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_path = Path(\"../input/train/\")\nfig, ax = plt.subplots(3, figsize=(10, 20))\nfor i, index in enumerate(np.random.randint(0, len(train_df), 3)):\n    path = (train_path / (train_df.id[index] + \".png\"))\n    img = np.asarray(Image.open(str(path)))\n    ax[i].imshow(img)\n    ids = list(map(int, train_df.attribute_ids[index].split()))\n    for num, attribute_id in enumerate(ids):\n        x_pos = img.shape[1] + 100\n        y_pos = (img.shape[0] - 100) / len(ids) * num + 100\n        ax[i].text(x_pos, y_pos, label_dict[attribute_id], fontsize=20)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":false},"cell_type":"markdown","source":"Let's show test images."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"test_path = Path(\"../input/test/\")\ntest_img_paths = list(test_path.glob(\"*.png\"))\nfig, ax = plt.subplots(3, figsize=(10, 20))\nfor i, path in enumerate(np.random.choice(test_img_paths, 3)):\n    img = np.asarray(Image.open(str(path)))\n    ax[i].imshow(img)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"Check train/test image area and size."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"def check_area_size(folder_path):\n    area_list = []\n    max_width = None\n    min_width = None\n    max_height = None\n    min_height = None\n    img_paths = list(folder_path.glob(\"*.png\"))\n    for path in tqdm(img_paths):\n        img = np.asarray(Image.open(str(path)))\n        shape = img.shape\n        area_list.append(shape[0]*shape[1])\n        if(max_width is None):\n            max_width = (shape[1], path)\n            min_width = (shape[1], path)\n            max_height = (shape[0], path)\n            min_height = (shape[0], path)\n        else:\n            if(max_width[0] < shape[1]):\n                max_width = (shape[1], path)\n            elif(min_width[0] > shape[1]):\n                min_width = (shape[1], path)\n            if(max_height[0] < shape[0]):\n                max_height = (shape[0], path)\n            elif(min_height[0] > shape[0]):\n                min_height = (shape[0], path)\n    return area_list, max_width, min_width, max_height, min_height","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":false,"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_area_list, train_max_width, train_min_width, train_max_height, train_min_height\\\n    = check_area_size(train_path)\nclear_output()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"test max area size is {}\".format(max(train_area_list)))\nprint(\"test min area size is {}\".format(min(train_area_list)))\nprint(\"Max area is {:.2f} times min area\".format(max(train_area_list)/ min(train_area_list)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(f\"max train image width is {train_max_width[0]}\")\nimg = np.asarray(Image.open(str(train_max_width[1])))\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"WTF!? what is this..."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(f\"min train image width is {train_min_width[0]}\")\nimg = np.asarray(Image.open(str(train_min_width[1])))\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(f\"max train image height is {train_max_height[0]}\")\nimg = np.asarray(Image.open(str(train_max_height[1])))\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(f\"min train image height is {train_min_height[0]}\")\nimg = np.asarray(Image.open(str(train_min_height[1])))\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"test_area_list, test_max_width, test_min_width, test_max_height, test_min_height\\\n    = check_area_size(test_path)\nclear_output()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(\"test max area size is {}\".format(max(test_area_list)))\nprint(\"test min area size is {}\".format(min(test_area_list)))\nprint(\"Max area is {:.2f} times min area\".format(max(test_area_list)/ min(test_area_list)))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 6), sharex=True)\nsns.distplot(train_area_list, kde=False, ax=ax1)\nax1.set_title(\"Distribution of the train image area\")\nsns.distplot(test_area_list, kde=False, ax=ax2)\nax2.set_title(\"Distribution of the test image area\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(f\"max test image width is {test_max_width[0]}\")\nimg = np.asarray(Image.open(str(test_max_width[1])))\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(f\"min train image width is {test_min_width[0]}\")\nimg = np.asarray(Image.open(str(test_min_width[1])))\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(f\"max train image height is {test_max_height[0]}\")\nimg = np.asarray(Image.open(str(test_max_height[1])))\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(f\"min train image height is {test_min_height[0]}\")\nimg = np.asarray(Image.open(str(test_min_height[1])))\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Dataset images have big difference in size(10 times over).  \nAnd image that have big difference between width and height(like about 280\\*5314) is exist in dataset.  \nI think we need care that adjust the size."},{"metadata":{"trusted":true},"cell_type":"markdown","source":"# Thank you for watching!\nI hope this will help.  \nPlease tell me if i make mistake.  "},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}