{"cells":[{"metadata":{},"cell_type":"markdown","source":"# TAU Vehicle Type Recognition Competition"},{"metadata":{},"cell_type":"markdown","source":"In this competitions we have to classify images of different vehicle types, including cars, bicycles, vans, ambulances, etc. (total 17 categories).\nThe data for the competition consists of training data together with the class labels and test data without the labels. The task is to predict the secret labels for the test data. So, this is a straight forward image classification task.\nThe data has been collected from the [Open Images dataset](https://storage.googleapis.com/openimages/web/index.html); an annotated collection of over 9 million images. We have only subset of openimages, selected to contain only vehicle categories among the total of 600 object classes.\n\nFew important points to note:\n* Any use of external data is not allowed\n* The evaluation metric for this competition is classification accuracy; "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nfrom collections import defaultdict\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom PIL import Image\nfrom tqdm import tqdm_notebook as tqdm","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"!ls -lh ../input/vehicle/","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls -lh ../input/vehicle/train/train","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls -lh ../input/vehicle/test/testset | head -5","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Folder description:\n\n* `train/train` folder contains the training set: a set of images with true labels in the folder names. The the folder contains altogether 28045 files organized in many sub-folders. The sub-folder names are the true classes; i.e., \"Boat\" sub-folder has all boat images, \"Car\" sub-folder has all the car images and so on.\n\n* `test/testset` - folder contains the test set: a set of images without labels. The folder contains altogether 7958 files in a single folder. The file name is the id for the solution's first column; i.e., the predicted class for file \"000000.jpg\" should appear on the first row of your submission.\n\n* `sample_submission.csv` - a sample submission file in the correct format.\n"},{"metadata":{},"cell_type":"markdown","source":"Let's get the training data in a dataframe"},{"metadata":{"trusted":true},"cell_type":"code","source":"'''walks through the train directory, creates a dataframe with class and filepaths for all images present in the train directory'''\n\nroot = '../input/vehicle/train/train/'\ndata = []\nfor category in sorted(os.listdir(root)):\n    for file in sorted(os.listdir(os.path.join(root, category))):\n        data.append((category, os.path.join(root, category,  file)))\n\ndf = pd.DataFrame(data, columns=['class', 'file_path'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len_df = len(df)\nprint(f\"There are {len_df} images\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df['class'].value_counts().plot(kind='bar');\nplt.title('Class counts');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df['class'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like the dataset is highly imbalanced, we have 8695 `Boat` category images and only 51 `Cart` category images, let's plot few images of each category to see what the images look like."},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(25, 16))\nfor num, category in enumerate(sorted(df['class'].unique())):\n    for i, (idx, row) in enumerate(df.loc[df['class'] == category].sample(4).iterrows()):\n        ax = fig.add_subplot(17, 4, num * 4 + i + 1, xticks=[], yticks=[])\n        im = Image.open(row['file_path'])\n        plt.imshow(im)\n        ax.set_title(f'Class: {category}')\nfig.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The images are quite distinguishable. Let's analyse the shape and size of the images"},{"metadata":{"trusted":true},"cell_type":"code","source":"data = defaultdict(lambda: defaultdict(list))\nfor idx, row in tqdm(df.iterrows(), total=len(df)):\n    image = Image.open(row[1])\n    data[row[0]]['width'].append(image.size[0])\n    data[row[0]]['height'].append(image.size[1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_dist(category):\n    '''plot height/width dist curves for given category'''\n    fig, ax = plt.subplots(figsize=(10, 10))\n    sns.distplot(data[category]['height'], color='darkorange', ax=ax).set_title(category, fontsize=16)\n    sns.distplot(data[category]['width'], color='purple', ax=ax).set_title(category, fontsize=16)\n    plt.xlabel('size', fontsize=15)\n    plt.legend(['height', 'width'])\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for category in df['class'].unique():\n    plot_dist(category)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is quite a variation in sizes for almost all categories"},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission = pd.read_csv('../input/vehicle/sample_submission.csv')\nsample_submission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have to predict for 7958 test images"},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## All the best :)"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}