{"cells":[{"metadata":{"_uuid":"6da80d8a411119faa781ce45a52c74859bb27dfb"},"cell_type":"markdown","source":"# Data Exploration for Yelp Image Classification Challenge\n---\nExploring the image data and labels, visualizing random samples of images, and plotting image shape distributions."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom IPython.display import display\nimport matplotlib.pyplot as plt\nimport plotly\nimport plotly.graph_objs as go\nimport cv2\nimport random\nimport os\n%matplotlib inline\n\n# Init plotly for offline plotting\nplotly.offline.init_notebook_mode(connected=True)\n\nprint('Pandas version:', pd.__version__)\nprint('Numpy version:', np.__version__)\nprint('OpenCV version:', cv2.__version__)\nprint('Plotly version:', plotly.__version__)\nprint(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"937dcae1fb44b00180a0a03b656ed8c710ece35e"},"cell_type":"markdown","source":"# Data Loading\n---\nLoad the data. The csv files that map image IDs to business IDs and business IDs to labels"},{"metadata":{"trusted":true,"_uuid":"74760795593117469ffcfad2aff8cc510897e0db","collapsed":true},"cell_type":"code","source":"# Load training data that maps business ID to labels\ntrain = pd.read_csv('../input/train.csv')\ndisplay(train.head())\nprint('Shape of train data:', train.shape)\nprint('Number of unique businesses:', train.shape[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d2dfa63df4794bdfe54d65624700373acd579d8d","collapsed":true},"cell_type":"code","source":"# Load training data that maps photos to business ID\ntrain_photo_to_id = pd.read_csv('../input/train_photo_to_biz_ids.csv')\ndisplay(train_photo_to_id.head())\nprint('Shape of train_photo_to_id:', train_photo_to_id.shape)\nprint('Number of images in training set:', train_photo_to_id.shape[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"885504c154f8f61a4b09e71e9e631d686e837c54","collapsed":true},"cell_type":"code","source":"train_dir = '../input/train_photos'\ntrain_imgs = os.listdir(train_dir)\n\ntest_dir = '../input/test_photos'\ntest_imgs = os.listdir(test_dir)\n\nprint('Number of training images:', len(train_imgs))\nprint('Number of testing images:', len(test_imgs))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"41ee8d97147bd2786a30c0f2fd47fa1a8a762dcd"},"cell_type":"markdown","source":"## Missing Values"},{"metadata":{"trusted":true,"_uuid":"428984c7a4f2a1b94eb34daff5f0bd1be3f3ea05","collapsed":true},"cell_type":"code","source":"# Business id to labels dataframe\nprint('Total number of missing labels:', train['labels'].isnull().sum())\ndisplay(train[train['labels'].isnull()])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9317d5211c3517b386ef3aa8d3629a2a3b7fc1e3"},"cell_type":"markdown","source":"# Visualize some Images\n---\nTake some random images and visualize them with their lables"},{"metadata":{"trusted":true,"_uuid":"2a69fdbafd1073747a1cfefcb8644f98e48bef52","scrolled":false,"collapsed":true},"cell_type":"code","source":"# Randomly sample 8 images\nimgs_samples = random.sample(train_imgs, 8)\n\n# Plot random sample of 8 images\nplt.figure(figsize=(15, 10))\nfor i in range(len(imgs_samples)):\n    # OpenCV2 reads images in BGR format\n    img = cv2.imread(os.path.join(train_dir, imgs_samples[i]))\n    # Switch color channels to RGB to make compatible with matplotlib imshow func\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    # Grab image's business ID and labels\n    business = train_photo_to_id.loc[train_photo_to_id['photo_id'] == int(imgs_samples[i][:-4]), 'business_id']\n    labels = train.loc[train['business_id'] == business.values[0], 'labels']\n    # Annotate each image with image ID, business ID, and labels\n    title = \"Image ID: \" + imgs_samples[i] + ' Business: ' + str(business.values[0]) + '\\nLabels: ' + ''.join(labels.values)\n    # Plot the image\n    plt.subplot(2, 4, i+1)\n    plt.tight_layout(pad=0.4, w_pad=0.5, h_pad=1.0)\n    plt.imshow(img)\n    plt.axis('off')\n    plt.title(title)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9e381c713c42c2f2fce171171ebe8465cf57ea13"},"cell_type":"markdown","source":"## List of Labels\n\n0: good_for_lunch\n\n1: good_for_dinner\n\n2: takes_reservations\n\n3: outdoor_seating\n\n4: restaurant_is_expensive\n\n5: has_alcohol\n\n6: has_table_service\n\n7: ambience_is_classy\n\n8: good_for_kids"},{"metadata":{"_uuid":"598a6f68b94f55c9c9cc6bfc160fef9c0a9cd85c"},"cell_type":"markdown","source":"# Plot Image Size Distribution\n---\nPlot the distribution of image sizes for training and testing datasets. This is useful because if there are images with different pixel shapes then we have to resize them before feeding into a machine learning model."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"9942f4f22c20504ffe78c71be6783a444f496b39"},"cell_type":"code","source":"def load_img_shapes(path_to_img):\n    \"\"\" Return only the shape of an image (width, height, channels) \"\"\"\n    return cv2.imread(path_to_img).shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a70287a68227319319b661a1e5b9468e93dc7096","collapsed":true},"cell_type":"code","source":"# Initialize arrays to hold image sizes\ntrain_shapes = []\ntest_shapes = []\n# Load in training/testing image sizes\nfor i in range(len(train_imgs)):\n    img_path = os.path.join(train_dir, train_imgs[i])\n    train_shapes.append(load_img_shapes(img_path))\nfor i in range(len(test_imgs)):\n    img_path = os.path.join(test_dir, test_imgs[i])\n    test_shapes.append(load_img_shapes(img_path))\n\n# Store training image sizes in dataframe\ndf_train = pd.DataFrame({'Shapes': train_shapes})\ntrain_counts = df_train['Shapes'].value_counts()\n# Store testing image sizes in dataframe\ndf_test = pd.DataFrame({'Shapes': test_shapes})\ntest_counts = df_test['Shapes'].value_counts()\n\nprint(\"Training Image Shapes: First 100\")\nfor i in range(100):\n    print(\"Shape %s counts: %d\" % (train_counts.index[i], train_counts.values[i]))\nprint(\"*\"*50)\nprint(\"Testing Image Shapes: First 100\")\nfor i in range(100):\n    print(\"Shape %s counts: %d\" % (test_counts.index[i], test_counts.values[i]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f30002ef96e7e7ff8ee1269e5b2777e953fe0e9e","collapsed":true},"cell_type":"code","source":"# Create barplot for image sizes distribution (training set)\nx_train = train_counts.index[:100]\nx_train = [str(x) for x in x_train]\ny_train = train_counts.values[:100]\n\n# Only plot first 100 value counts\nx_test = test_counts.index[:100]\nx_test = [str(x) for x in x_test]\ny_test = test_counts.values[:100]\n\n# Create traces\ntraining_trace = go.Bar(x=x_train,\n                        y=y_train,\n                        marker=dict(\n                            color='rgb(158,202,225)',\n                            line=dict(\n                                color='rgb(8,48,107)',\n                                width=1.5),\n                        ),\n                        opacity=0.6,\n                        name='training'\n                       )\ntesting_trace = go.Bar(x=x_test,\n                       y=y_test,\n                       marker=dict(\n                           color='rgb(58,102,245)',\n                           line=dict(\n                               color='rgb(8,48,107)',\n                               width=1.5),\n                       ),\n                       opacity=0.6,\n                       name='testing'\n                      )\n# Create layout\nlayout = go.Layout(font = dict(family = \"Overpass\"),\n                   title = 'Image size distributions (Top is training set, bottom is testing set)',\n#                    xaxis = dict(tickangle=-45, title = \"Image shapes\"),\n#                    xaxis2 = dict(tickangle=-45, title = \"Image shapes\"),\n                  )\n\nfig = plotly.tools.make_subplots(rows=2, cols=1,\n                                 vertical_spacing=0.4\n                                )\nfig.append_trace(training_trace, 1, 1)\nfig.append_trace(testing_trace, 2, 1)\nfig['layout'].update(font = dict(family = \"Overpass\"),\n                     title='Image size distributions (Top is training set, bottom is testing set)',\n                     xaxis = dict(tickangle=-45, title = \"Image shapes\"),\n                     xaxis2 = dict(tickangle=-45, title = \"Image shapes\")\n                    )\nplotly.offline.iplot(fig, validate=False, show_link=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"21ad5c6b9198d59cb2cfb988bf31315733309611"},"cell_type":"markdown","source":"Most common image shape is (357, 500, 3) and all images have 3 color channels (RGB). There are a lot of images with different shapes so there will definitely be a lot of resizing. Probably to have a shape of (357, 500, 3) since that is the most common shape for both training and testing sets."},{"metadata":{"_uuid":"52d6af82847d43fd44a9b3a78049a48b5bed4f09"},"cell_type":"markdown","source":"# Distribution of Classes\n---\nEach image can have multiple classes assigned to it. Here we will look at the number of times each class occurs over the entire training set."},{"metadata":{"trusted":true,"_uuid":"31525d07b2a7798218b31ba44848a396d560d3ed"},"cell_type":"code","source":"# Count all labels in training set\nall_labels = ' '.join(list(train['labels'].fillna('nan').values)).split()\nfrom collections import Counter\nlabel_counts = Counter(all_labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"504cbaf39e02b7e5253d7e410e7387a0ee2e3fb5"},"cell_type":"code","source":"for key in label_counts:\n    print('Label {0} appears {1} times in training dataset'.format(key, label_counts[key]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"74f66f9165577b9474e6940a8ff3e24fd891097b"},"cell_type":"markdown","source":"# Duplicate Values in Dataset\n---\nFind duplicate values in our training dataset if any."},{"metadata":{"trusted":true,"_uuid":"c616a6815e016f7b32d9a16180cc5fc4c4904775"},"cell_type":"code","source":"# train_photo_to_id.groupby(\"photo_id\")\nprint('Number of duplicate photo IDs:', len(train_photo_to_id[train_photo_to_id['photo_id'].duplicated()]))\nprint('Number of duplicate business IDS:', len(train[train['business_id'].duplicated()]))","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}