{
  "id": 64651,
  "title": "brief summary of 2nd place",
  "url": "/competitions/google-ai-open-images-visual-relationship-track/writeups/tito-brief-summary-of-2nd-place",
  "author_name": "",
  "post_date": "2018-09-03T10:36:22.297Z",
  "votes": 24,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Congrats to all the winners! and I'd like to thank Google and Kaggle for this interesting competition. </p>\n\n<p>This is My brief summary.</p>\n\n<h1>Model 1: object detection (yolo)</h1>\n\n<p>I used yolo for object detection with following modification.</p>\n\n<ul>\n<li>Removed confidence term from loss function</li>\n<li>Added Class weight</li>\n<li>Class masking to forces only on labeled classes <br>\nBut, did not mask the area where BB exists (BBs does not exist on same area) <br>\n  But, masked for child class (BBs can exist on same area for child classes)   </li>\n</ul>\n\n<h1>Model 2: visual relationship (InceptionResNetV 2)</h1>\n\n<p>Relation data(challenge-2018-train-vrd.csv) is mapped to all combinations of human labeled BBs(challenge-2018-train-vrd-bbox.csv) and the data is used as training data of model2.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379292/10248/example.png\" alt=\"example\"></p>\n\n<h3>2-1: relation 'is'</h3>\n\n<p>I made training data by mapping the ground truth BBOX(challenge-2018-train-vrd-bbox.csv)  to ground truth relational data(challenge-2018-train-vrd.csv), giving a target.  </p>\n\n<ul>\n<li>Objective: Material (wooden, plastic, ..., None). </li>\n<li>Features: Cropped image, BBOX class, BB position / size etc.</li>\n</ul>\n\n<p>Example:    </p>\n\n<p>bbox data:</p>\n\n<pre>      image1,bbox1,Man\n      image1,bbox2,Guiter\n      image1,bbox3,Chair\n</pre>\n\n<p>Relationship data:</p>\n\n<pre>      image1, bbox2, bbox2, Guitar, Wooden, is\n      image1, bbox3, bbox3, Chair, Wooden, is\n      image1, bbox1, bbox2, Man, Guitar, hold\n</pre>\n\n<p>then, training data would be</p>\n\n<pre>      image1,bbox1,bbox1, Man,None,is  \n      image1,bbox2,bbox2, Guiter,Wooden,is\n      image1,bbox3,bbox3, Chair ,Wooden,is  \n</pre>\n\n<p>　    here, target is None if the data is not in Relationship data.</p>\n\n<h3>2-2: Triplet Relationships</h3>\n\n<p>First, I created all pairs of BBOX in the same image with some filter.\nThen, I made training data by mapping the pairs to ground truth relational data, giving a target.</p>\n\n<ul>\n<li>Objective: Relationship (at, on, ..., None).   </li>\n<li>Features: Cropped image including two BBOX with box line, LabelName1, LabelName2, XCenter1, YCenter1, XCenter2, YCenter2, Size1, Size2, Aspect1, Aspect2, XCenterDiff, YCenterDiff, CenterDiff, XCenter, YCenter, IOU</li>\n</ul>\n\n<p>Example:\nbbox data :</p>\n\n<pre>      image1,bbox1,Man\n      image1,bbox2,Guiter\n      image1,bbox3,Chair\n</pre>\n\n<p>Relationship data:</p>\n\n<pre>      image1,bbox1,bbox2,Man,Guiter,hold\n</pre>\n\n<p>Then training data would be  </p>\n\n<pre>      image1,bbox1,bbox2,Man,Guiter,hold\n      image1,bbox1,bbox3,Man,Chair,None\n      image1,bbox2,bbox1,Guiter,Man,None\n      image1,bbox2,bbox3,Guiter,Chair,None\n      image1,bbox3,bbox1,Chair,Man,None\n      image1,bbox3,bbox2,Chair,Guiter,None\n</pre>\n\n<p>　  here, target is None if the data is not in Relationship data.</p>\n\n<h1>Model 3: Score Prediction (Light GBM)</h1>\n\n<p>Relation data(challenge-2018-train-vrd.csv) is mapped to all combinations of predicted BBs(output of model1), and the data is used as training data of model3.</p>\n\n<ul>\n<li>Features: model1 output(BBOX classes, BBOX scores, BB positions / sizes etc.), and model2 output(Relationship and its probability)</li>\n<li>Objective:  whether the prediction of model2 is true or not</li>\n</ul>\n\n<p>Model2 and model3 can be merged to one NN model.</p>\n\n<h1>About yolo object function</h1>\n\n<p>The data set of this competition has the following three characteristics.</p>\n\n<h2>(1), every images are not checked for all classes.</h2>\n\n<p>For example, there are cases where cats are not checked even if there is a cat in the image.\nThis causes false penalty if the model detects a cat during learning.</p>\n\n<h2>(2), each object has a parent-child relationship.</h2>\n\n<p>For example, cat class is a child class of the Animal class, so a BBOX of cat is also a BBOX of Animal.</p>\n\n<h2>(3), the 500 classes are not balanced.</h2>\n\n<p>To handle these better, I first deleted the confidence term from YOLO's objective function.\nBy deleting this term, class probability will be responsible for confidence.\nOnce class probability also have confidence role, this make it possible to mask unchecked classes.\nBy masking the cat class in the example above it becomes possible to eliminate the penalty when cat is detected by the model.</p>\n\n<p>However, It is unlikely that different classes of boxes will be in the same area(same place, and same size, and same aspect).\nSo, I did not mask for the same area where some bbox exists.</p>\n\n<p>But, there is an exception to this, as bbox of child class can exist in the same area(Animal can be a cat as well).\nSo I masked the child class of the corresponding BBOX.</p>\n\n<p>In order to cope with an imbalance of (3), I removed same images.\nBut sampling images is not enough to balance the classes, so I introduced class weight.</p>",
  "messages": [
    {
      "id": "379292",
      "postDate": "08/31/2018 05:59:41",
      "content": "<p>Congrats to all the winners! and I'd like to thank Google and Kaggle for this interesting competition. </p>\n\n<p>This is My brief summary.</p>\n\n<h1>Model 1: object detection (yolo)</h1>\n\n<p>I used yolo for object detection with following modification.</p>\n\n<ul>\n<li>Removed confidence term from loss function</li>\n<li>Added Class weight</li>\n<li>Class masking to forces only on labeled classes <br>\nBut, did not mask the area where BB exists (BBs does not exist on same area) <br>\n  But, masked for child class (BBs can exist on same area for child classes)   </li>\n</ul>\n\n<h1>Model 2: visual relationship (InceptionResNetV 2)</h1>\n\n<p>Relation data(challenge-2018-train-vrd.csv) is mapped to all combinations of human labeled BBs(challenge-2018-train-vrd-bbox.csv) and the data is used as training data of model2.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379292/10248/example.png\" alt=\"example\"></p>\n\n<h3>2-1: relation 'is'</h3>\n\n<p>I made training data by mapping the ground truth BBOX(challenge-2018-train-vrd-bbox.csv)  to ground truth relational data(challenge-2018-train-vrd.csv), giving a target.  </p>\n\n<ul>\n<li>Objective: Material (wooden, plastic, ..., None). </li>\n<li>Features: Cropped image, BBOX class, BB position / size etc.</li>\n</ul>\n\n<p>Example:    </p>\n\n<p>bbox data:</p>\n\n<pre>      image1,bbox1,Man\n      image1,bbox2,Guiter\n      image1,bbox3,Chair\n</pre>\n\n<p>Relationship data:</p>\n\n<pre>      image1, bbox2, bbox2, Guitar, Wooden, is\n      image1, bbox3, bbox3, Chair, Wooden, is\n      image1, bbox1, bbox2, Man, Guitar, hold\n</pre>\n\n<p>then, training data would be</p>\n\n<pre>      image1,bbox1,bbox1, Man,None,is  \n      image1,bbox2,bbox2, Guiter,Wooden,is\n      image1,bbox3,bbox3, Chair ,Wooden,is  \n</pre>\n\n<p>　    here, target is None if the data is not in Relationship data.</p>\n\n<h3>2-2: Triplet Relationships</h3>\n\n<p>First, I created all pairs of BBOX in the same image with some filter.\nThen, I made training data by mapping the pairs to ground truth relational data, giving a target.</p>\n\n<ul>\n<li>Objective: Relationship (at, on, ..., None).   </li>\n<li>Features: Cropped image including two BBOX with box line, LabelName1, LabelName2, XCenter1, YCenter1, XCenter2, YCenter2, Size1, Size2, Aspect1, Aspect2, XCenterDiff, YCenterDiff, CenterDiff, XCenter, YCenter, IOU</li>\n</ul>\n\n<p>Example:\nbbox data :</p>\n\n<pre>      image1,bbox1,Man\n      image1,bbox2,Guiter\n      image1,bbox3,Chair\n</pre>\n\n<p>Relationship data:</p>\n\n<pre>      image1,bbox1,bbox2,Man,Guiter,hold\n</pre>\n\n<p>Then training data would be  </p>\n\n<pre>      image1,bbox1,bbox2,Man,Guiter,hold\n      image1,bbox1,bbox3,Man,Chair,None\n      image1,bbox2,bbox1,Guiter,Man,None\n      image1,bbox2,bbox3,Guiter,Chair,None\n      image1,bbox3,bbox1,Chair,Man,None\n      image1,bbox3,bbox2,Chair,Guiter,None\n</pre>\n\n<p>　  here, target is None if the data is not in Relationship data.</p>\n\n<h1>Model 3: Score Prediction (Light GBM)</h1>\n\n<p>Relation data(challenge-2018-train-vrd.csv) is mapped to all combinations of predicted BBs(output of model1), and the data is used as training data of model3.</p>\n\n<ul>\n<li>Features: model1 output(BBOX classes, BBOX scores, BB positions / sizes etc.), and model2 output(Relationship and its probability)</li>\n<li>Objective:  whether the prediction of model2 is true or not</li>\n</ul>\n\n<p>Model2 and model3 can be merged to one NN model.</p>\n\n<h1>About yolo object function</h1>\n\n<p>The data set of this competition has the following three characteristics.</p>\n\n<h2>(1), every images are not checked for all classes.</h2>\n\n<p>For example, there are cases where cats are not checked even if there is a cat in the image.\nThis causes false penalty if the model detects a cat during learning.</p>\n\n<h2>(2), each object has a parent-child relationship.</h2>\n\n<p>For example, cat class is a child class of the Animal class, so a BBOX of cat is also a BBOX of Animal.</p>\n\n<h2>(3), the 500 classes are not balanced.</h2>\n\n<p>To handle these better, I first deleted the confidence term from YOLO's objective function.\nBy deleting this term, class probability will be responsible for confidence.\nOnce class probability also have confidence role, this make it possible to mask unchecked classes.\nBy masking the cat class in the example above it becomes possible to eliminate the penalty when cat is detected by the model.</p>\n\n<p>However, It is unlikely that different classes of boxes will be in the same area(same place, and same size, and same aspect).\nSo, I did not mask for the same area where some bbox exists.</p>\n\n<p>But, there is an exception to this, as bbox of child class can exist in the same area(Animal can be a cat as well).\nSo I masked the child class of the corresponding BBOX.</p>\n\n<p>In order to cope with an imbalance of (3), I removed same images.\nBut sampling images is not enough to balance the classes, so I introduced class weight.</p>",
      "rawMarkdown": "Congrats to all the winners! and I'd like to thank Google and Kaggle for this interesting competition. \n\nThis is My brief summary.\n\n# Model 1: object detection (yolo)\nI used yolo for object detection with following modification.\n\n- Removed confidence term from loss function\n- Added Class weight\n- Class masking to forces only on labeled classes  \n    But, did not mask the area where BB exists (BBs does not exist on same area)  \n      But, masked for child class (BBs can exist on same area for child classes)   \n\n# Model 2: visual relationship (InceptionResNetV 2)\n     \nRelation data(challenge-2018-train-vrd.csv) is mapped to all combinations of human labeled BBs(challenge-2018-train-vrd-bbox.csv) and the data is used as training data of model2.\n\n![example](https://storage.googleapis.com/kaggle-forum-message-attachments/379292/10248/example.png)\n\n\n### 2-1: relation 'is'\nI made training data by mapping the ground truth BBOX(challenge-2018-train-vrd-bbox.csv)  to ground truth relational data(challenge-2018-train-vrd.csv), giving a target.  \n\n- Objective: Material (wooden, plastic, ..., None). \n- Features: Cropped image, BBOX class, BB position / size etc.\n\n\nExample:\t\n\nbbox data:\n\n<pre>      image1,bbox1,Man\n      image1,bbox2,Guiter\n      image1,bbox3,Chair\n</pre>\n\n\nRelationship data:\n\n<pre>      image1, bbox2, bbox2, Guitar, Wooden, is\n      image1, bbox3, bbox3, Chair, Wooden, is\n      image1, bbox1, bbox2, Man, Guitar, hold\n</pre>\n\nthen, training data would be\n\n<pre>      image1,bbox1,bbox1, Man,None,is  \n      image1,bbox2,bbox2, Guiter,Wooden,is\n      image1,bbox3,bbox3, Chair ,Wooden,is  \n</pre>\n　    here, target is None if the data is not in Relationship data.\n\n\n### 2-2: Triplet Relationships  \n\nFirst, I created all pairs of BBOX in the same image with some filter.\nThen, I made training data by mapping the pairs to ground truth relational data, giving a target.\n\n- Objective: Relationship (at, on, ..., None).   \n- Features: Cropped image including two BBOX with box line, LabelName1, LabelName2, XCenter1, YCenter1, XCenter2, YCenter2, Size1, Size2, Aspect1, Aspect2, XCenterDiff, YCenterDiff, CenterDiff, XCenter, YCenter, IOU\n\nExample:\nbbox data :\n\n<pre>      image1,bbox1,Man\n      image1,bbox2,Guiter\n      image1,bbox3,Chair\n</pre>\n\nRelationship data:\n\n<pre>      image1,bbox1,bbox2,Man,Guiter,hold\n</pre>\n\n\nThen training data would be  \n\n<pre>      image1,bbox1,bbox2,Man,Guiter,hold\n      image1,bbox1,bbox3,Man,Chair,None\n      image1,bbox2,bbox1,Guiter,Man,None\n      image1,bbox2,bbox3,Guiter,Chair,None\n      image1,bbox3,bbox1,Chair,Man,None\n      image1,bbox3,bbox2,Chair,Guiter,None\n</pre>\n　  here, target is None if the data is not in Relationship data.\n\n\n\n\n\n\n# Model 3: Score Prediction (Light GBM)\nRelation data(challenge-2018-train-vrd.csv) is mapped to all combinations of predicted BBs(output of model1), and the data is used as training data of model3.\n\n- Features: model1 output(BBOX classes, BBOX scores, BB positions / sizes etc.), and model2 output(Relationship and its probability)\n- Objective:  whether the prediction of model2 is true or not\n\nModel2 and model3 can be merged to one NN model.\n\n# About yolo object function\n\nThe data set of this competition has the following three characteristics.\n\n## (1), every images are not checked for all classes.\nFor example, there are cases where cats are not checked even if there is a cat in the image.\nThis causes false penalty if the model detects a cat during learning.\n\n## (2), each object has a parent-child relationship.\nFor example, cat class is a child class of the Animal class, so a BBOX of cat is also a BBOX of Animal.\n\n## (3), the 500 classes are not balanced.\n\nTo handle these better, I first deleted the confidence term from YOLO's objective function.\nBy deleting this term, class probability will be responsible for confidence.\nOnce class probability also have confidence role, this make it possible to mask unchecked classes.\nBy masking the cat class in the example above it becomes possible to eliminate the penalty when cat is detected by the model.\n\nHowever, It is unlikely that different classes of boxes will be in the same area(same place, and same size, and same aspect).\nSo, I did not mask for the same area where some bbox exists.\n\nBut, there is an exception to this, as bbox of child class can exist in the same area(Animal can be a cat as well).\nSo I masked the child class of the corresponding BBOX.\n\nIn order to cope with an imbalance of (3), I removed same images.\nBut sampling images is not enough to balance the classes, so I introduced class weight.",
      "votes": null
    },
    {
      "id": "379323",
      "postDate": "08/31/2018 07:01:18",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "379328",
      "postDate": "08/31/2018 07:11:11",
      "content": "<p>Thank you, Ahmet!</p>",
      "rawMarkdown": "Thank you, Ahmet!",
      "votes": null
    },
    {
      "id": "379475",
      "postDate": "08/31/2018 12:14:10",
      "content": "<p>Congrats tito, great job! Thank you for the write up - very interesting read.</p>\n\n<p>Could I please ask you how are you using the class weight in the detector? Are you using it to increase the weight of infrequent classes or in some other way?</p>\n\n<p>Also, could I please ask you what was your reasoning behind removing the confidence score from the loss?</p>\n\n<p>Thank you very much again for the write up - much appreciated.</p>",
      "rawMarkdown": "Congrats tito, great job! Thank you for the write up - very interesting read.\n\nCould I please ask you how are you using the class weight in the detector? Are you using it to increase the weight of infrequent classes or in some other way?\n\nAlso, could I please ask you what was your reasoning behind removing the confidence score from the loss?\n\nThank you very much again for the write up - much appreciated.",
      "votes": null
    },
    {
      "id": "379485",
      "postDate": "08/31/2018 12:50:15",
      "content": "<p>Thank you radek,</p>\n\n<p>As you mentioned, I used class weight to balance the infrequent classes.</p>\n\n<p>I expected class probability takes over confidence by removing confidence term from loss function.\nThis (confidence for each class) made it possible to add class masking and class weight, etc to yolo.</p>",
      "rawMarkdown": "Thank you radek,\n\nAs you mentioned, I used class weight to balance the infrequent classes.\n\nI expected class probability takes over confidence by removing confidence term from loss function.\nThis (confidence for each class) made it possible to add class masking and class weight, etc to yolo.",
      "votes": null
    },
    {
      "id": "379981",
      "postDate": "09/01/2018 11:33:19",
      "content": "<p>Congratulations tito, and nice summary. </p>\n\n<p>Will try to implement a model with your ideas. </p>\n\n<p>And any plans to publish your code?</p>\n\n<p>All the best for the future.</p>",
      "rawMarkdown": "Congratulations tito, and nice summary. \n\nWill try to implement a model with your ideas. \n\nAnd any plans to publish your code?\n\nAll the best for the future.",
      "votes": null
    },
    {
      "id": "380066",
      "postDate": "09/01/2018 15:32:52",
      "content": "<p>Hi Nazim,</p>\n\n<p>I'm not sure whether I'll publish my code yet.\nAnyway, I'll let you know here if I publish it.</p>",
      "rawMarkdown": "Hi Nazim,\n\nI'm not sure whether I'll publish my code yet.\nAnyway, I'll let you know here if I publish it.",
      "votes": null
    },
    {
      "id": "380288",
      "postDate": "09/02/2018 08:25:37",
      "content": "<p>Congratulations tito! Thank you for the writeup. I wonder if you could further explain how you combined the cropped image plus the other non-image features (labelname1 etc) into the 2-2 model?</p>",
      "rawMarkdown": "Congratulations tito! Thank you for the writeup. I wonder if you could further explain how you combined the cropped image plus the other non-image features (labelname1 etc) into the 2-2 model?",
      "votes": null
    },
    {
      "id": "380297",
      "postDate": "09/02/2018 09:00:46",
      "content": "<p>Hi Tim H,</p>\n\n<p>I added Embeded categorical features and numerical features on the top of CNN as dens layer.\nI used following code for this.</p>\n\n<pre>BASE_MODEL='InceptionResNetV2'\ncategorical = ['name_1', 'name_2']\nnumerical = ['XCenter1','YCenter1','XCenter2','YCenter2','Size1','Size2','Aspect1','Aspect2','XCenterDiff','YCenterDiff','CenterDiff','XCenter','YCenter','IOU']\n\ndef mk_model():\n    if BASE_MODEL=='VGG16':\n        base_model =  VGG16(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='InceptionV3':\n        base_model=InceptionV3(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='Xception':\n        base_model=Xception(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='InceptionResNetV2':\n        base_model=InceptionResNetV2(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    input_list = []\n    x1=base_model.output\n    x1=GlobalAveragePooling2D()(x1)\n    input_list.append(base_model.input)\n\n    emb_list = []\n    for cat in categorical:\n        inpt = Input(shape=[1], name = cat)\n        input_list.append(inpt)\n        embd = Embedding(len(bb_values), len(bb_values))(inpt)\n        emb_list.append(embd)\n    x2 = concatenate(emb_list)\n    x2 = SpatialDropout1D(0.5)(x2)\n    x2 = Flatten()(x2)\n    x2 = Dense(512,activation='relu')(x2)\n    x2 = BatchNormalization()(x2)\n\n    inpt = Input((len(numerical),),name='numerical')\n    input_list.append(inpt)\n    x3 = Dense(512,activation='relu')(inpt)\n    x3 = Dropout(0.5)(x3)\n    x3 = BatchNormalization()(x3)\n\n    inpt = Input((len(bb_values)*2,),name='confidence')\n    input_list.append(inpt)\n    x4 = Dense(512,activation='relu')(inpt)\n    x4 = Dropout(0.5)(x4)\n    x4 = BatchNormalization()(x4)\n\n    x = concatenate([x1, x2, x3, x4])\n    x=Dense(1024,activation='relu')(x)\n    x = Dropout(0.2)(x)\n    x = BatchNormalization()(x)\n    x=Dense(1024,activation='relu')(x)\n    x = BatchNormalization()(x)\n    prediction=Dense(n_classes,activation='sigmoid')(x)\n\n    model=Model(inputs=input_list,outputs=prediction)\n    for layer in base_model.layers:\n        layer.trainable=False\n    return model, base_model\n</pre>",
      "rawMarkdown": "Hi Tim H,\n\nI added Embeded categorical features and numerical features on the top of CNN as dens layer.\nI used following code for this.\n\n<pre>BASE_MODEL='InceptionResNetV2'\ncategorical = ['name_1', 'name_2']\nnumerical = ['XCenter1','YCenter1','XCenter2','YCenter2','Size1','Size2','Aspect1','Aspect2','XCenterDiff','YCenterDiff','CenterDiff','XCenter','YCenter','IOU']\n \ndef mk_model():\n    if BASE_MODEL=='VGG16':\n        base_model =  VGG16(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='InceptionV3':\n        base_model=InceptionV3(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='Xception':\n        base_model=Xception(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='InceptionResNetV2':\n        base_model=InceptionResNetV2(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    input_list = []\n    x1=base_model.output\n    x1=GlobalAveragePooling2D()(x1)\n    input_list.append(base_model.input)\n \n    emb_list = []\n    for cat in categorical:\n        inpt = Input(shape=[1], name = cat)\n        input_list.append(inpt)\n        embd = Embedding(len(bb_values), len(bb_values))(inpt)\n        emb_list.append(embd)\n    x2 = concatenate(emb_list)\n    x2 = SpatialDropout1D(0.5)(x2)\n    x2 = Flatten()(x2)\n    x2 = Dense(512,activation='relu')(x2)\n    x2 = BatchNormalization()(x2)\n\n    inpt = Input((len(numerical),),name='numerical')\n    input_list.append(inpt)\n    x3 = Dense(512,activation='relu')(inpt)\n    x3 = Dropout(0.5)(x3)\n    x3 = BatchNormalization()(x3)\n \n    inpt = Input((len(bb_values)*2,),name='confidence')\n    input_list.append(inpt)\n    x4 = Dense(512,activation='relu')(inpt)\n    x4 = Dropout(0.5)(x4)\n    x4 = BatchNormalization()(x4)\n\n    x = concatenate([x1, x2, x3, x4])\n    x=Dense(1024,activation='relu')(x)\n    x = Dropout(0.2)(x)\n    x = BatchNormalization()(x)\n    x=Dense(1024,activation='relu')(x)\n    x = BatchNormalization()(x)\n    prediction=Dense(n_classes,activation='sigmoid')(x)\n\n    model=Model(inputs=input_list,outputs=prediction)\n    for layer in base_model.layers:\n        layer.trainable=False\n    return model, base_model\n</pre>",
      "votes": null
    },
    {
      "id": "380300",
      "postDate": "09/02/2018 09:09:33",
      "content": "<p>Thank you tito!</p>",
      "rawMarkdown": "Thank you tito!",
      "votes": null
    },
    {
      "id": "399109",
      "postDate": "10/05/2018 08:15:53",
      "content": "<p>Hi, Thank you for your reply!</p>\n\n<p>I'm trying to understand your code.\nwhat does x1 represent? x1 = base_mode.output. So is it the classification prediction from the base_model, InceptionResNet or a feature map from the model?</p>\n\n<p>x2 represents embedded representation of labels.\nx3 represents numerical inputs\nx4 are lists of bounding boxes for two objects detected.\nThen,\nyou concatenate them all to input into a model, right? </p>\n\n<p>I think I'm confused why you append base_model.input into the input_list, which makes sense to me, but base_mode.output into x1. Do you mind explaining? Thank you!</p>\n\n<p>----EDIT\nah, wait. I think I get it now. \nX1 to X4 are computed features from the layers using inputs like images, labels and bounding box annotations.</p>\n\n<p>How do you merge Model2.1 , 2.2 and 3?</p>",
      "rawMarkdown": "Hi, Thank you for your reply!\n\nI'm trying to understand your code.\nwhat does x1 represent? x1 = base_mode.output. So is it the classification prediction from the base_model, InceptionResNet or a feature map from the model?\n\nx2 represents embedded representation of labels.\nx3 represents numerical inputs\nx4 are lists of bounding boxes for two objects detected.\nThen,\nyou concatenate them all to input into a model, right? \n\nI think I'm confused why you append base_model.input into the input_list, which makes sense to me, but base_mode.output into x1. Do you mind explaining? Thank you!\n\n----EDIT\nah, wait. I think I get it now. \nX1 to X4 are computed features from the layers using inputs like images, labels and bounding box annotations.\n\nHow do you merge Model2.1 , 2.2 and 3?",
      "votes": null
    },
    {
      "id": "399291",
      "postDate": "10/05/2018 14:52:38",
      "content": "<p>Hi Kim,</p>\n\n<blockquote>\n  <p>----EDIT ah, wait. I think I get it now. X1 to X4 are computed features from the layers using inputs like images, labels and bounding box annotations.</p>\n</blockquote>\n\n<p>Yes, you understand it correctly.</p>\n\n<blockquote>\n  <p>How do you merge Model2.1 , 2.2 and 3?</p>\n</blockquote>\n\n<p>Model3 is made for  Model2.1 and Model2.2 respectively.</p>\n\n<pre><code>Model2.1 -&gt; Model3.1 -&gt; prediction for relation 'is'\nModel2.2 -&gt; Model3.2 -&gt; prediction for Triplet Relationships\n</code></pre>\n\n<p>Then I just combined the predictions to one submission.</p>",
      "rawMarkdown": "Hi Kim,\n\n&gt; ----EDIT ah, wait. I think I get it now. X1 to X4 are computed features from the layers using inputs like images, labels and bounding box annotations.\n\nYes, you understand it correctly.\n\n&gt; How do you merge Model2.1 , 2.2 and 3?\n\nModel3 is made for  Model2.1 and Model2.2 respectively.\n\n    Model2.1 -&gt; Model3.1 -&gt; prediction for relation 'is'\n    Model2.2 -&gt; Model3.2 -&gt; prediction for Triplet Relationships\n\nThen I just combined the predictions to one submission.",
      "votes": null
    },
    {
      "id": "450420",
      "postDate": "01/04/2019 22:45:26",
      "content": "<p>Could you please provide the python notebook for this competition?</p>",
      "rawMarkdown": "Could you please provide the python notebook for this competition?",
      "votes": null
    },
    {
      "id": "450525",
      "postDate": "01/05/2019 07:22:22",
      "content": "<p>Any update?</p>",
      "rawMarkdown": "Any update?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 379323,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "08/31/2018 07:01:18",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 379328,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "08/31/2018 07:11:11",
          "content": "<p>Thank you, Ahmet!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 379475,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "08/31/2018 12:14:10",
      "content": "<p>Congrats tito, great job! Thank you for the write up - very interesting read.</p>\n\n<p>Could I please ask you how are you using the class weight in the detector? Are you using it to increase the weight of infrequent classes or in some other way?</p>\n\n<p>Also, could I please ask you what was your reasoning behind removing the confidence score from the loss?</p>\n\n<p>Thank you very much again for the write up - much appreciated.</p>",
      "votes": null,
      "replies": [
        {
          "id": 379485,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "08/31/2018 12:50:15",
          "content": "<p>Thank you radek,</p>\n\n<p>As you mentioned, I used class weight to balance the infrequent classes.</p>\n\n<p>I expected class probability takes over confidence by removing confidence term from loss function.\nThis (confidence for each class) made it possible to add class masking and class weight, etc to yolo.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 379981,
      "author_name": "nazimgirach",
      "author_url": "",
      "post_date": "09/01/2018 11:33:19",
      "content": "<p>Congratulations tito, and nice summary. </p>\n\n<p>Will try to implement a model with your ideas. </p>\n\n<p>And any plans to publish your code?</p>\n\n<p>All the best for the future.</p>",
      "votes": null,
      "replies": [
        {
          "id": 380066,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "09/01/2018 15:32:52",
          "content": "<p>Hi Nazim,</p>\n\n<p>I'm not sure whether I'll publish my code yet.\nAnyway, I'll let you know here if I publish it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 450525,
          "author_name": "nazimgirach",
          "author_url": "",
          "post_date": "01/05/2019 07:22:22",
          "content": "<p>Any update?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 380288,
      "author_name": "timhartill",
      "author_url": "",
      "post_date": "09/02/2018 08:25:37",
      "content": "<p>Congratulations tito! Thank you for the writeup. I wonder if you could further explain how you combined the cropped image plus the other non-image features (labelname1 etc) into the 2-2 model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 380297,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "09/02/2018 09:00:46",
          "content": "<p>Hi Tim H,</p>\n\n<p>I added Embeded categorical features and numerical features on the top of CNN as dens layer.\nI used following code for this.</p>\n\n<pre>BASE_MODEL='InceptionResNetV2'\ncategorical = ['name_1', 'name_2']\nnumerical = ['XCenter1','YCenter1','XCenter2','YCenter2','Size1','Size2','Aspect1','Aspect2','XCenterDiff','YCenterDiff','CenterDiff','XCenter','YCenter','IOU']\n\ndef mk_model():\n    if BASE_MODEL=='VGG16':\n        base_model =  VGG16(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='InceptionV3':\n        base_model=InceptionV3(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='Xception':\n        base_model=Xception(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='InceptionResNetV2':\n        base_model=InceptionResNetV2(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    input_list = []\n    x1=base_model.output\n    x1=GlobalAveragePooling2D()(x1)\n    input_list.append(base_model.input)\n\n    emb_list = []\n    for cat in categorical:\n        inpt = Input(shape=[1], name = cat)\n        input_list.append(inpt)\n        embd = Embedding(len(bb_values), len(bb_values))(inpt)\n        emb_list.append(embd)\n    x2 = concatenate(emb_list)\n    x2 = SpatialDropout1D(0.5)(x2)\n    x2 = Flatten()(x2)\n    x2 = Dense(512,activation='relu')(x2)\n    x2 = BatchNormalization()(x2)\n\n    inpt = Input((len(numerical),),name='numerical')\n    input_list.append(inpt)\n    x3 = Dense(512,activation='relu')(inpt)\n    x3 = Dropout(0.5)(x3)\n    x3 = BatchNormalization()(x3)\n\n    inpt = Input((len(bb_values)*2,),name='confidence')\n    input_list.append(inpt)\n    x4 = Dense(512,activation='relu')(inpt)\n    x4 = Dropout(0.5)(x4)\n    x4 = BatchNormalization()(x4)\n\n    x = concatenate([x1, x2, x3, x4])\n    x=Dense(1024,activation='relu')(x)\n    x = Dropout(0.2)(x)\n    x = BatchNormalization()(x)\n    x=Dense(1024,activation='relu')(x)\n    x = BatchNormalization()(x)\n    prediction=Dense(n_classes,activation='sigmoid')(x)\n\n    model=Model(inputs=input_list,outputs=prediction)\n    for layer in base_model.layers:\n        layer.trainable=False\n    return model, base_model\n</pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 380300,
          "author_name": "timhartill",
          "author_url": "",
          "post_date": "09/02/2018 09:09:33",
          "content": "<p>Thank you tito!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 399109,
          "author_name": "lemonista",
          "author_url": "",
          "post_date": "10/05/2018 08:15:53",
          "content": "<p>Hi, Thank you for your reply!</p>\n\n<p>I'm trying to understand your code.\nwhat does x1 represent? x1 = base_mode.output. So is it the classification prediction from the base_model, InceptionResNet or a feature map from the model?</p>\n\n<p>x2 represents embedded representation of labels.\nx3 represents numerical inputs\nx4 are lists of bounding boxes for two objects detected.\nThen,\nyou concatenate them all to input into a model, right? </p>\n\n<p>I think I'm confused why you append base_model.input into the input_list, which makes sense to me, but base_mode.output into x1. Do you mind explaining? Thank you!</p>\n\n<p>----EDIT\nah, wait. I think I get it now. \nX1 to X4 are computed features from the layers using inputs like images, labels and bounding box annotations.</p>\n\n<p>How do you merge Model2.1 , 2.2 and 3?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 399291,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "10/05/2018 14:52:38",
          "content": "<p>Hi Kim,</p>\n\n<blockquote>\n  <p>----EDIT ah, wait. I think I get it now. X1 to X4 are computed features from the layers using inputs like images, labels and bounding box annotations.</p>\n</blockquote>\n\n<p>Yes, you understand it correctly.</p>\n\n<blockquote>\n  <p>How do you merge Model2.1 , 2.2 and 3?</p>\n</blockquote>\n\n<p>Model3 is made for  Model2.1 and Model2.2 respectively.</p>\n\n<pre><code>Model2.1 -&gt; Model3.1 -&gt; prediction for relation 'is'\nModel2.2 -&gt; Model3.2 -&gt; prediction for Triplet Relationships\n</code></pre>\n\n<p>Then I just combined the predictions to one submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 450420,
      "author_name": "sriramr",
      "author_url": "",
      "post_date": "01/04/2019 22:45:26",
      "content": "<p>Could you please provide the python notebook for this competition?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "379292": "Congrats to all the winners! and I'd like to thank Google and Kaggle for this interesting competition. \n\nThis is My brief summary.\n\n# Model 1: object detection (yolo)\nI used yolo for object detection with following modification.\n\n- Removed confidence term from loss function\n- Added Class weight\n- Class masking to forces only on labeled classes  \n    But, did not mask the area where BB exists (BBs does not exist on same area)  \n      But, masked for child class (BBs can exist on same area for child classes)   \n\n# Model 2: visual relationship (InceptionResNetV 2)\n     \nRelation data(challenge-2018-train-vrd.csv) is mapped to all combinations of human labeled BBs(challenge-2018-train-vrd-bbox.csv) and the data is used as training data of model2.\n\n![example](https://storage.googleapis.com/kaggle-forum-message-attachments/379292/10248/example.png)\n\n\n### 2-1: relation 'is'\nI made training data by mapping the ground truth BBOX(challenge-2018-train-vrd-bbox.csv)  to ground truth relational data(challenge-2018-train-vrd.csv), giving a target.  \n\n- Objective: Material (wooden, plastic, ..., None). \n- Features: Cropped image, BBOX class, BB position / size etc.\n\n\nExample:\t\n\nbbox data:\n\n<pre>      image1,bbox1,Man\n      image1,bbox2,Guiter\n      image1,bbox3,Chair\n</pre>\n\n\nRelationship data:\n\n<pre>      image1, bbox2, bbox2, Guitar, Wooden, is\n      image1, bbox3, bbox3, Chair, Wooden, is\n      image1, bbox1, bbox2, Man, Guitar, hold\n</pre>\n\nthen, training data would be\n\n<pre>      image1,bbox1,bbox1, Man,None,is  \n      image1,bbox2,bbox2, Guiter,Wooden,is\n      image1,bbox3,bbox3, Chair ,Wooden,is  \n</pre>\n　    here, target is None if the data is not in Relationship data.\n\n\n### 2-2: Triplet Relationships  \n\nFirst, I created all pairs of BBOX in the same image with some filter.\nThen, I made training data by mapping the pairs to ground truth relational data, giving a target.\n\n- Objective: Relationship (at, on, ..., None).   \n- Features: Cropped image including two BBOX with box line, LabelName1, LabelName2, XCenter1, YCenter1, XCenter2, YCenter2, Size1, Size2, Aspect1, Aspect2, XCenterDiff, YCenterDiff, CenterDiff, XCenter, YCenter, IOU\n\nExample:\nbbox data :\n\n<pre>      image1,bbox1,Man\n      image1,bbox2,Guiter\n      image1,bbox3,Chair\n</pre>\n\nRelationship data:\n\n<pre>      image1,bbox1,bbox2,Man,Guiter,hold\n</pre>\n\n\nThen training data would be  \n\n<pre>      image1,bbox1,bbox2,Man,Guiter,hold\n      image1,bbox1,bbox3,Man,Chair,None\n      image1,bbox2,bbox1,Guiter,Man,None\n      image1,bbox2,bbox3,Guiter,Chair,None\n      image1,bbox3,bbox1,Chair,Man,None\n      image1,bbox3,bbox2,Chair,Guiter,None\n</pre>\n　  here, target is None if the data is not in Relationship data.\n\n\n\n\n\n\n# Model 3: Score Prediction (Light GBM)\nRelation data(challenge-2018-train-vrd.csv) is mapped to all combinations of predicted BBs(output of model1), and the data is used as training data of model3.\n\n- Features: model1 output(BBOX classes, BBOX scores, BB positions / sizes etc.), and model2 output(Relationship and its probability)\n- Objective:  whether the prediction of model2 is true or not\n\nModel2 and model3 can be merged to one NN model.\n\n# About yolo object function\n\nThe data set of this competition has the following three characteristics.\n\n## (1), every images are not checked for all classes.\nFor example, there are cases where cats are not checked even if there is a cat in the image.\nThis causes false penalty if the model detects a cat during learning.\n\n## (2), each object has a parent-child relationship.\nFor example, cat class is a child class of the Animal class, so a BBOX of cat is also a BBOX of Animal.\n\n## (3), the 500 classes are not balanced.\n\nTo handle these better, I first deleted the confidence term from YOLO's objective function.\nBy deleting this term, class probability will be responsible for confidence.\nOnce class probability also have confidence role, this make it possible to mask unchecked classes.\nBy masking the cat class in the example above it becomes possible to eliminate the penalty when cat is detected by the model.\n\nHowever, It is unlikely that different classes of boxes will be in the same area(same place, and same size, and same aspect).\nSo, I did not mask for the same area where some bbox exists.\n\nBut, there is an exception to this, as bbox of child class can exist in the same area(Animal can be a cat as well).\nSo I masked the child class of the corresponding BBOX.\n\nIn order to cope with an imbalance of (3), I removed same images.\nBut sampling images is not enough to balance the classes, so I introduced class weight.",
    "379323": "Congrats!",
    "379328": "Thank you, Ahmet!",
    "379475": "Congrats tito, great job! Thank you for the write up - very interesting read.\n\nCould I please ask you how are you using the class weight in the detector? Are you using it to increase the weight of infrequent classes or in some other way?\n\nAlso, could I please ask you what was your reasoning behind removing the confidence score from the loss?\n\nThank you very much again for the write up - much appreciated.",
    "379485": "Thank you radek,\n\nAs you mentioned, I used class weight to balance the infrequent classes.\n\nI expected class probability takes over confidence by removing confidence term from loss function.\nThis (confidence for each class) made it possible to add class masking and class weight, etc to yolo.",
    "379981": "Congratulations tito, and nice summary. \n\nWill try to implement a model with your ideas. \n\nAnd any plans to publish your code?\n\nAll the best for the future.",
    "380066": "Hi Nazim,\n\nI'm not sure whether I'll publish my code yet.\nAnyway, I'll let you know here if I publish it.",
    "380288": "Congratulations tito! Thank you for the writeup. I wonder if you could further explain how you combined the cropped image plus the other non-image features (labelname1 etc) into the 2-2 model?",
    "380297": "Hi Tim H,\n\nI added Embeded categorical features and numerical features on the top of CNN as dens layer.\nI used following code for this.\n\n<pre>BASE_MODEL='InceptionResNetV2'\ncategorical = ['name_1', 'name_2']\nnumerical = ['XCenter1','YCenter1','XCenter2','YCenter2','Size1','Size2','Aspect1','Aspect2','XCenterDiff','YCenterDiff','CenterDiff','XCenter','YCenter','IOU']\n \ndef mk_model():\n    if BASE_MODEL=='VGG16':\n        base_model =  VGG16(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='InceptionV3':\n        base_model=InceptionV3(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='Xception':\n        base_model=Xception(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    elif BASE_MODEL=='InceptionResNetV2':\n        base_model=InceptionResNetV2(include_top=False, weights='imagenet', input_tensor=Input(shape=(224,224,3)))\n    input_list = []\n    x1=base_model.output\n    x1=GlobalAveragePooling2D()(x1)\n    input_list.append(base_model.input)\n \n    emb_list = []\n    for cat in categorical:\n        inpt = Input(shape=[1], name = cat)\n        input_list.append(inpt)\n        embd = Embedding(len(bb_values), len(bb_values))(inpt)\n        emb_list.append(embd)\n    x2 = concatenate(emb_list)\n    x2 = SpatialDropout1D(0.5)(x2)\n    x2 = Flatten()(x2)\n    x2 = Dense(512,activation='relu')(x2)\n    x2 = BatchNormalization()(x2)\n\n    inpt = Input((len(numerical),),name='numerical')\n    input_list.append(inpt)\n    x3 = Dense(512,activation='relu')(inpt)\n    x3 = Dropout(0.5)(x3)\n    x3 = BatchNormalization()(x3)\n \n    inpt = Input((len(bb_values)*2,),name='confidence')\n    input_list.append(inpt)\n    x4 = Dense(512,activation='relu')(inpt)\n    x4 = Dropout(0.5)(x4)\n    x4 = BatchNormalization()(x4)\n\n    x = concatenate([x1, x2, x3, x4])\n    x=Dense(1024,activation='relu')(x)\n    x = Dropout(0.2)(x)\n    x = BatchNormalization()(x)\n    x=Dense(1024,activation='relu')(x)\n    x = BatchNormalization()(x)\n    prediction=Dense(n_classes,activation='sigmoid')(x)\n\n    model=Model(inputs=input_list,outputs=prediction)\n    for layer in base_model.layers:\n        layer.trainable=False\n    return model, base_model\n</pre>",
    "380300": "Thank you tito!",
    "399109": "Hi, Thank you for your reply!\n\nI'm trying to understand your code.\nwhat does x1 represent? x1 = base_mode.output. So is it the classification prediction from the base_model, InceptionResNet or a feature map from the model?\n\nx2 represents embedded representation of labels.\nx3 represents numerical inputs\nx4 are lists of bounding boxes for two objects detected.\nThen,\nyou concatenate them all to input into a model, right? \n\nI think I'm confused why you append base_model.input into the input_list, which makes sense to me, but base_mode.output into x1. Do you mind explaining? Thank you!\n\n----EDIT\nah, wait. I think I get it now. \nX1 to X4 are computed features from the layers using inputs like images, labels and bounding box annotations.\n\nHow do you merge Model2.1 , 2.2 and 3?",
    "399291": "Hi Kim,\n\n&gt; ----EDIT ah, wait. I think I get it now. X1 to X4 are computed features from the layers using inputs like images, labels and bounding box annotations.\n\nYes, you understand it correctly.\n\n&gt; How do you merge Model2.1 , 2.2 and 3?\n\nModel3 is made for  Model2.1 and Model2.2 respectively.\n\n    Model2.1 -&gt; Model3.1 -&gt; prediction for relation 'is'\n    Model2.2 -&gt; Model3.2 -&gt; prediction for Triplet Relationships\n\nThen I just combined the predictions to one submission.",
    "450420": "Could you please provide the python notebook for this competition?",
    "450525": "Any update?"
  },
  "source": "meta"
}