{"cells":[{"metadata":{"_uuid":"533be5c57fc4deae89737a51eca8b3d4126f4c64"},"cell_type":"markdown","source":"# Introduction\n\nI found out that image size is related to the rate (probability) of `new_whale`. Check below example if you are interested.\n\n# Example\n## Import Packages\n\ncv2 is faster, but PIL is easy to distinguish gray scale images and RGB color images.\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nfrom PIL import Image","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2854bc13475247903237db59656655eaed9020b8"},"cell_type":"markdown","source":"## Make function to get image shapes\nmake the list of image shapes. Gray scale images are two dimensional array, Color images are three dimensional array."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"def get_size_list(targets, dir_target):\n\n    result = list()\n\n    for target in tqdm(targets):\n\n        img = np.array(Image.open(os.path.join(dir_target, target)))\n        result.append(str(img.shape))\n\n    return result","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4d43daa422263139e643866ba2b5df1bfcc52678"},"cell_type":"markdown","source":"## Get image shape for each train image\nload `train.csv` and add a column which represents size of images."},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"e3310d564778c602c60a4ea6883125f43f9f7e32"},"cell_type":"code","source":"data = pd.read_csv('../input/train.csv')\ndata['size_info'] = get_size_list(data.Image.tolist(), dir_target='../input/train')\ndata.to_csv('./size_train.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a149607efc8275462115d0dc9c1540d68f49763f"},"cell_type":"markdown","source":"## Group by shape and summerize\nSummerizing number of samples and rate of `new_whale, we can see unnatual bias.\n(700, 1050, 3), (600, 1050, 3) includes 23-25% of new_whales. On the other hand, (600, 1050) includes 89% of new_whale."},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"95ad7d10d043e100957c5048b3eb12653b430f1e"},"cell_type":"code","source":"counts = data.size_info.value_counts()\n\nagg = data.groupby('size_info').Id.agg({'number_sample': len,\n                                        'rate_new_whale': lambda g: np.mean(g == 'new_whale')})\n\nagg = agg.sort_values('number_sample', ascending=False)\nagg.to_csv('result.csv')\nprint(agg.head(20))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"29c254bdef1ca514d38fd5c742f409da53091573"},"cell_type":"markdown","source":"# Conclusion\nIt seems that image size is related to the rate of `new_whale`. Does this feature help us ? Please your comment."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}