{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Weakly Supervised Room Classification with YOLOv3 and Snorkel\n\nThis notebook is based on the idea that one of the difficulties in classifying hotels may be that we see very different rooms. We can't usually see both the bedroom and bathroom in the same photo, and some hotels have other spaces which are distinct from either (kitchens, etc). It's not particularly likely that the mapping between a particular bathroom and bedroom features is obvious.","metadata":{}},{"cell_type":"markdown","source":"The idea behind this notebook is to detect objects associated with particular rooms and views, such as toilets and sinks in the bathroom, and beds for the bedroom. We use Snorkel as a framework for handling and combining noisy labels.","metadata":{}},{"cell_type":"code","source":"!pip install -q snorkel pytorchyolo opencv-python torch==1.10.2 torchvision==0.11.3","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:37:47.143131Z","iopub.execute_input":"2022-04-26T10:37:47.144008Z","iopub.status.idle":"2022-04-26T10:39:19.796710Z","shell.execute_reply.started":"2022-04-26T10:37:47.143884Z","shell.execute_reply":"2022-04-26T10:39:19.795635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport cv2\nimport os\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-04-26T10:39:19.801230Z","iopub.execute_input":"2022-04-26T10:39:19.801455Z","iopub.status.idle":"2022-04-26T10:39:20.093528Z","shell.execute_reply.started":"2022-04-26T10:39:19.801428Z","shell.execute_reply":"2022-04-26T10:39:20.092820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import snorkel\nfrom snorkel.labeling import labeling_function, LFAnalysis, PandasLFApplier\nfrom snorkel.preprocess import preprocessor\nfrom snorkel.labeling.model import LabelModel","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:39:20.097833Z","iopub.execute_input":"2022-04-26T10:39:20.099939Z","iopub.status.idle":"2022-04-26T10:39:21.739890Z","shell.execute_reply.started":"2022-04-26T10:39:20.099895Z","shell.execute_reply":"2022-04-26T10:39:21.739090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that the YOLO directory below is in a private dataset. The repository I clone in that is GPLv3 licensed, so I can't republish under the Apache license.\n\nYou can find the script to download the data from https://github.com/eriklindernoren/PyTorch-YOLOv3 . The code is loaded through Pip, because it was painful to get dependencies to work otherwise.","metadata":{}},{"cell_type":"code","source":"YOLO_DIR = '/kaggle/input/pytorch-yolov3/PyTorch-YOLOv3/'\nDATA_DIR = '/kaggle/input/hotel-id-to-combat-human-trafficking-2022-fgvc9/'","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:39:21.741920Z","iopub.execute_input":"2022-04-26T10:39:21.742132Z","iopub.status.idle":"2022-04-26T10:39:21.749345Z","shell.execute_reply.started":"2022-04-26T10:39:21.742103Z","shell.execute_reply":"2022-04-26T10:39:21.747115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Snorkel uses a series of integer class labels, starting at zero. The negative label signifies no decision.","metadata":{}},{"cell_type":"code","source":"# Class labels to apply\nABSTAIN = -1\nBEDROOM = 0\nBATHROOM = 1\nOTHER = 2","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:39:21.751092Z","iopub.execute_input":"2022-04-26T10:39:21.751550Z","iopub.status.idle":"2022-04-26T10:39:21.756980Z","shell.execute_reply.started":"2022-04-26T10:39:21.751511Z","shell.execute_reply":"2022-04-26T10:39:21.756271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We find all the training images, and collate them by Hotel ID (the directory name).","metadata":{}},{"cell_type":"code","source":"data = {\"full_path\": [], \"image\": [], \"hotel_id\": []}\nfor subdir in os.listdir(DATA_DIR + '/train_images/'):\n    hotel_id = int(subdir)\n    for image in os.listdir(f'{DATA_DIR}/train_images/{hotel_id}'):\n        path = f'{DATA_DIR}/train_images/{hotel_id}/{image}'\n        data['image'].append(image)\n        data['full_path'].append(path)\n        data['hotel_id'].append(hotel_id)\n\ndf = pd.DataFrame(data)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:39:21.758378Z","iopub.execute_input":"2022-04-26T10:39:21.758844Z","iopub.status.idle":"2022-04-26T10:39:37.319304Z","shell.execute_reply.started":"2022-04-26T10:39:21.758762Z","shell.execute_reply":"2022-04-26T10:39:37.318637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below line, if uncommented, reduces the amount of data used for testing.","metadata":{}},{"cell_type":"code","source":"#df = df.sample(1000, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:39:37.320657Z","iopub.execute_input":"2022-04-26T10:39:37.320936Z","iopub.status.idle":"2022-04-26T10:39:37.329918Z","shell.execute_reply.started":"2022-04-26T10:39:37.320900Z","shell.execute_reply":"2022-04-26T10:39:37.329201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We use YOLOv3 to construct a Snorkel \"preprocessor\". This will process each sample, and extract computer-readable information. In this case, it's the objects detected by YOLO.","metadata":{}},{"cell_type":"code","source":"from pytorchyolo import models, detect\nyolo_model = models.load_model(YOLO_DIR + 'config/yolov3.cfg', YOLO_DIR + 'weights/yolov3.weights')\n\nlabels = []\nwith open(YOLO_DIR + '/data/coco.names', 'r') as f:\n    for line in f:\n        labels.append(line.strip())\n\n@preprocessor(memoize=True)\ndef object_detection(x):\n    # Load the image as a numpy array\n    img = cv2.imread(x.full_path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    # Runs the YOLO model on the image \n    boxes = detect.detect_image(yolo_model, img)\n    \n    # Re-map the integer class label in to a string\n    objects = []\n    for x1, y1, x2, y2, confidence, c in boxes:\n        label = labels[int(c)]\n        objects.append((x1, y1, x2, y2, confidence, label))\n    x.object_boxes = objects\n    return x","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:39:37.331588Z","iopub.execute_input":"2022-04-26T10:39:37.332246Z","iopub.status.idle":"2022-04-26T10:39:46.646545Z","shell.execute_reply.started":"2022-04-26T10:39:37.332203Z","shell.execute_reply":"2022-04-26T10:39:46.642729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each of the following is a \"labeling function\" for Snorkel. They are expected to be heuristic, somewhat noisy classifiers, which are allowed to return \"don't know\" or abstain. Each looks for one or more features with reasonable confidence, and will declare it to be a particular type of room if found. Snorkel will later combine these noisy labels together.","metadata":{}},{"cell_type":"code","source":"@labeling_function(pre=[object_detection])\ndef toilet_in_bathroom(x):\n    for x1, y1, x2, y2, confidence, label in x.object_boxes:\n        if label == 'toilet' and confidence > 0.5:\n            return BATHROOM\n    return ABSTAIN\n\n@labeling_function(pre=[object_detection])\ndef sink_in_bathroom(x):\n    for x1, y1, x2, y2, confidence, label in x.object_boxes:\n        if label == 'sink' and confidence > 0.5:\n            return BATHROOM\n    return ABSTAIN\n\n@labeling_function(pre=[object_detection])\ndef bed_in_bedroom(x):\n    for x1, y1, x2, y2, confidence, label in x.object_boxes:\n        if label == 'bed' and confidence > 0.5:\n            return BEDROOM\n    return ABSTAIN\n\n@labeling_function(pre=[object_detection])\ndef oven_in_kitchen(x):\n    for x1, y1, x2, y2, confidence, label in x.object_boxes:\n        if label == 'oven' and confidence > 0.5:\n            return OTHER\n    return ABSTAIN\n\n@labeling_function(pre=[object_detection])\ndef microwave_in_kitchen(x):\n    for x1, y1, x2, y2, confidence, label in x.object_boxes:\n        if label == 'microwave' and confidence > 0.5:\n            return OTHER\n    return ABSTAIN\n\n@labeling_function(pre=[object_detection])\ndef fridge_in_kitchen(x):\n    for x1, y1, x2, y2, confidence, label in x.object_boxes:\n        if label == 'refrigerator' and confidence > 0.5:\n            return OTHER\n    return ABSTAIN\n\n@labeling_function(pre=[object_detection])\ndef default_other(x):\n    reliable_classifiers = [toilet_in_bathroom, sink_in_bathroom, bed_in_bedroom]\n    if all([c(x) == ABSTAIN for c in reliable_classifiers]):\n        return OTHER\n    return ABSTAIN","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:39:46.647624Z","iopub.execute_input":"2022-04-26T10:39:46.647933Z","iopub.status.idle":"2022-04-26T10:39:46.674188Z","shell.execute_reply.started":"2022-04-26T10:39:46.647895Z","shell.execute_reply":"2022-04-26T10:39:46.670819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We apply our label functions to our data, to extract the prediction that each function produced for each output as a 2D array.","metadata":{}},{"cell_type":"code","source":"lfs = [toilet_in_bathroom, sink_in_bathroom, bed_in_bedroom, oven_in_kitchen, microwave_in_kitchen, fridge_in_kitchen, default_other]\n\napplier = PandasLFApplier(lfs=lfs)\nL = applier.apply(df=df)\nL","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:39:46.680044Z","iopub.execute_input":"2022-04-26T10:39:46.680245Z","iopub.status.idle":"2022-04-26T10:43:13.589372Z","shell.execute_reply.started":"2022-04-26T10:39:46.680221Z","shell.execute_reply":"2022-04-26T10:43:13.588698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can view coverage, overlap, and conflict statistics for each function, showing how valuable they are for labelling, or whether they're redundant.","metadata":{}},{"cell_type":"code","source":"LFAnalysis(L=L, lfs=lfs).lf_summary()","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:43:13.590578Z","iopub.execute_input":"2022-04-26T10:43:13.591950Z","iopub.status.idle":"2022-04-26T10:43:13.619195Z","shell.execute_reply.started":"2022-04-26T10:43:13.591908Z","shell.execute_reply":"2022-04-26T10:43:13.618393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We show a sample of what each rule is picking up and labelling, to see whether these are reasonable. It's interesting to note that some of the \"confused\" images we see (e.g. the bed it thinks is a fridge) are rotated wrong - correcting that may be a path to improvement.","metadata":{}},{"cell_type":"code","source":"NUM_IMAGES = 5\n\nfig, ax = plt.subplots(len(lfs), NUM_IMAGES, figsize=(22,8 * len(lfs)))\nfor i, lf in enumerate(lfs):\n    ax[i][0].set_ylabel(lf.name)\n    \n    matched = df.iloc[np.not_equal(L[:, i], ABSTAIN)]\n    sample = matched.sample(min(len(matched), NUM_IMAGES), random_state=1000+i)\n    for j, path in enumerate(sample['full_path']):\n        ax[i][j].imshow(mpimg.imread(path))","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:43:13.620574Z","iopub.execute_input":"2022-04-26T10:43:13.620880Z","iopub.status.idle":"2022-04-26T10:43:41.307156Z","shell.execute_reply.started":"2022-04-26T10:43:13.620840Z","shell.execute_reply":"2022-04-26T10:43:41.306295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Snorkel can train a model based on the provided labels. My understanding of this step is limited, but I believe it will end up evaluating functions by their conflicts, weighting the most reliably agreeing ones higher.","metadata":{}},{"cell_type":"code","source":"label_model = LabelModel(cardinality=3, verbose=True)\nlabel_model.fit(L_train=L, n_epochs=500, log_freq=100, seed=123)","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:46:48.774135Z","iopub.execute_input":"2022-04-26T10:46:48.774419Z","iopub.status.idle":"2022-04-26T10:46:49.162947Z","shell.execute_reply.started":"2022-04-26T10:46:48.774388Z","shell.execute_reply":"2022-04-26T10:46:49.162056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now produce our noisy labels. Note that this is always a best-effort process, and so a confidence measure is produced alongside the peak prediction.","metadata":{}},{"cell_type":"code","source":"P, C = label_model.predict(L, return_probs=True)\nLABELS = ['unknown', 'bedroom', 'bathroom', 'other']\n\nldf = df.copy()[['hotel_id', 'image']]\nldf['room_type'] = [LABELS[p+1] for p in P]\nfor i, l in enumerate(LABELS[1:]):\n    ldf['p_' + l] = C[:, i]\nldf.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:46:49.164812Z","iopub.execute_input":"2022-04-26T10:46:49.165070Z","iopub.status.idle":"2022-04-26T10:46:49.207092Z","shell.execute_reply.started":"2022-04-26T10:46:49.165033Z","shell.execute_reply":"2022-04-26T10:46:49.206365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Summarise how many rooms ended up labeled as each type:","metadata":{}},{"cell_type":"code","source":"for l in LABELS:\n    c = len(ldf[ldf['room_type'] == l])\n    print(f'{l}: {c}')","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:46:49.208345Z","iopub.execute_input":"2022-04-26T10:46:49.208566Z","iopub.status.idle":"2022-04-26T10:46:49.225038Z","shell.execute_reply.started":"2022-04-26T10:46:49.208535Z","shell.execute_reply":"2022-04-26T10:46:49.224305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And save the results.","metadata":{}},{"cell_type":"code","source":"ldf.to_csv('room-types.csv')","metadata":{"execution":{"iopub.status.busy":"2022-04-26T10:46:49.229306Z","iopub.execute_input":"2022-04-26T10:46:49.229963Z","iopub.status.idle":"2022-04-26T10:46:49.256733Z","shell.execute_reply.started":"2022-04-26T10:46:49.229915Z","shell.execute_reply":"2022-04-26T10:46:49.255997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Snorkel models produce a \"noisy\" label set, don't generalise, and are rarely used directly. The normal path from here would be to train a classifier using the noisy labels, which gains the ability to generalise to other examples. A trained classifier is also likely to fill in the \"unknown\" elements, and provided it's not overfitted too much, may have a smoothing effect that lets it express reduced confidence in anything we labelled wrongly with these heuristics.\n\nFor this competition, these heuristics may well be \"good enough\" to use directly for some purposes.","metadata":{}}]}