{
  "id": 110329,
  "title": "Difference in local evaluation and leaderboard evaluation",
  "url": "/competitions/open-images-2019-object-detection/discussion/110329",
  "author_name": "AI Study",
  "post_date": "2019-09-27T00:30:54.974000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I don't understand what is happening. When I run an evaluation on my local validation images I get an mAP score of 40%+ but when I upload a submission I get between 5-10% on the test images.</p>\n\n<p>I really don't think I'm overfitting since I'm using the complete training set and barely get through 1 complete epoch before submitting. I also don't think the unbalanced dataset is causing this since I get similar results using a smaller balanced validation set and the complete validation set.</p>\n\n<p>I saw a post about how the xmin, ymin, xmax, ymax are in a difference sequence but I have triple checked and I think my submissions have the right order. I have added bounding boxes to some of the test images and they look reasonable.</p>\n\n<p>An example image (ignore any spacing errors from this copy/paste): \nd504e26d60244000   /m/015p6 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 /m/0jbk 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2236626%2F862648fc56cf08eec13231bf5098849b%2Fd504e26d60244000-boxes.jpg?generation=1569544119041305&amp;alt=media\" alt=\"\"></p>\n\n<p>The bounding box looks reasonable to me on the test image. Is this format or field order not correct?</p>\n\n<p>Any ideas what may cause the huge difference in my local validation scores and the leaderboard scores?</p>",
  "messages": [
    {
      "id": 636641,
      "postDate": "2019-09-30T00:52:59.673Z",
      "content": "<p>That is the whole line for the image. Only difference in the file is the comma between the imageid and the prediction string.\nd504e26d60244000,/m/015p6 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 /m/0jbk 0.9671 0.12052649 0.2865757 0.41084218 0.90251577\nI'll have to double check on the thresholds but I'm pretty sure they were the same. \nActually I'm not sure at all since the evaluation was using the script from tensorflow object detection (models\\research\\object_detection\\inference\\infer_detections.py). I'll have to look at that tomorrow. I'm seeing some predictions in the .30+ range so thats probably their threshold.</p>",
      "rawMarkdown": "That is the whole line for the image. Only difference in the file is the comma between the imageid and the prediction string.\nd504e26d60244000,/m/015p6 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 /m/0jbk 0.9671 0.12052649 0.2865757 0.41084218 0.90251577\nI'll have to double check on the thresholds but I'm pretty sure they were the same. \nActually I'm not sure at all since the evaluation was using the script from tensorflow object detection (models\\research\\object_detection\\inference\\infer_detections.py). I'll have to look at that tomorrow. I'm seeing some predictions in the .30+ range so thats probably their threshold.",
      "replies": [
        {
          "id": 636839,
          "postDate": "2019-09-30T09:01:50.110Z",
          "content": "<p>Well, my model says that it's Canary (/m/0ccs93).</p>",
          "rawMarkdown": "Well, my model says that it's Canary (/m/0ccs93)."
        },
        {
          "id": 636854,
          "postDate": "2019-09-30T09:16:42.893Z",
          "content": "<p>yep, I'll bet you 1$ it's th.... with proper th you will get at least 30 predictions for such image</p>",
          "rawMarkdown": "yep, I'll bet you 1$ it's th.... with proper th you will get at least 30 predictions for such image"
        },
        {
          "id": 637180,
          "postDate": "2019-09-30T18:22:52.223Z",
          "content": "<p>Looks like you were correct and it is a threshold issue. I created a new submission with a lower threshold from an older run that had the complete prediction list for each image and got a much better score.\nCanary, for example was under .30 but included in the new submission.\nUnfortunately not enough time left to run a full evaluation with my latest work but it helps to know what the cause was.\nThanks for the responses.</p>",
          "rawMarkdown": "Looks like you were correct and it is a threshold issue. I created a new submission with a lower threshold from an older run that had the complete prediction list for each image and got a much better score.\nCanary, for example was under .30 but included in the new submission.\nUnfortunately not enough time left to run a full evaluation with my latest work but it helps to know what the cause was.\nThanks for the responses."
        }
      ]
    },
    {
      "id": 636639,
      "postDate": "2019-09-30T00:32:47.460Z",
      "content": "<p>can you please send the full line for the image? I am sure it does not voilate anything to send 1 out of 100000 images.\nalso, did you make sure you used the same threshold for inferencing the validation and test?</p>",
      "rawMarkdown": "can you please send the full line for the image? I am sure it does not voilate anything to send 1 out of 100000 images.\nalso, did you make sure you used the same threshold for inferencing the validation and test?"
    },
    {
      "id": 636638,
      "postDate": "2019-09-30T00:00:39.197Z",
      "content": "<p>Thanks for the reply.\nThat is only part of the image. The ImageID is d504e26d60244000 so you can look at a local copy and see the whole thing if you wish.\nI previously opened the image and calculated where the pixels would be and they are in the right locations as far as I can tell.\nImage is 1024 wide, 755 high\nXmin is at pixel 123 (1024*0.12052649 calculating from upper left corner)\nXmax at 420 (1024*0.41084218)\nYmin at 216 (755*0.2865757)\nYmax at 681 (755*0.90251577)\nSo upper left corner is at 123,216 - lower right corner is at 420,681</p>\n\n<p>I checked some training images and the calculations from the given bounding boxes were from upper left so that should be right. The field order is different but I adjusted for that.</p>\n\n<p>The example I gave is in Xmin, Ymin, Xmax, Ymax order according to my calculations.</p>\n\n<p>/m/015p6 is bird which looks correct to me\n/m/0jbk is animal which looks correct to me</p>\n\n<p>I'm obviously missing something but have no idea what else to look at.</p>",
      "rawMarkdown": "Thanks for the reply.\nThat is only part of the image. The ImageID is d504e26d60244000 so you can look at a local copy and see the whole thing if you wish.\nI previously opened the image and calculated where the pixels would be and they are in the right locations as far as I can tell.\nImage is 1024 wide, 755 high\nXmin is at pixel 123 (1024*0.12052649 calculating from upper left corner)\nXmax at 420 (1024*0.41084218)\nYmin at 216 (755*0.2865757)\nYmax at 681 (755*0.90251577)\nSo upper left corner is at 123,216 - lower right corner is at 420,681\n\nI checked some training images and the calculations from the given bounding boxes were from upper left so that should be right. The field order is different but I adjusted for that.\n\nThe example I gave is in Xmin, Ymin, Xmax, Ymax order according to my calculations.\n\n/m/015p6 is bird which looks correct to me\n/m/0jbk is animal which looks correct to me\n\nI'm obviously missing something but have no idea what else to look at."
    },
    {
      "id": 636163,
      "postDate": "2019-09-28T22:26:36.097Z",
      "content": "<p>Wait, what are your coordinates for this image? Y range doesn't seem to be 0.28 - 0.90. Maybe your coordinates are flipped or something like this?</p>",
      "rawMarkdown": "Wait, what are your coordinates for this image? Y range doesn't seem to be 0.28 - 0.90. Maybe your coordinates are flipped or something like this?"
    },
    {
      "id": 634921,
      "postDate": "2019-09-27T00:30:54.973Z",
      "content": "<p>I don't understand what is happening. When I run an evaluation on my local validation images I get an mAP score of 40%+ but when I upload a submission I get between 5-10% on the test images.</p>\n\n<p>I really don't think I'm overfitting since I'm using the complete training set and barely get through 1 complete epoch before submitting. I also don't think the unbalanced dataset is causing this since I get similar results using a smaller balanced validation set and the complete validation set.</p>\n\n<p>I saw a post about how the xmin, ymin, xmax, ymax are in a difference sequence but I have triple checked and I think my submissions have the right order. I have added bounding boxes to some of the test images and they look reasonable.</p>\n\n<p>An example image (ignore any spacing errors from this copy/paste): \nd504e26d60244000   /m/015p6 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 /m/0jbk 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2236626%2F862648fc56cf08eec13231bf5098849b%2Fd504e26d60244000-boxes.jpg?generation=1569544119041305&amp;alt=media\" alt=\"\"></p>\n\n<p>The bounding box looks reasonable to me on the test image. Is this format or field order not correct?</p>\n\n<p>Any ideas what may cause the huge difference in my local validation scores and the leaderboard scores?</p>",
      "rawMarkdown": "I don't understand what is happening. When I run an evaluation on my local validation images I get an mAP score of 40%+ but when I upload a submission I get between 5-10% on the test images.\n\nI really don't think I'm overfitting since I'm using the complete training set and barely get through 1 complete epoch before submitting. I also don't think the unbalanced dataset is causing this since I get similar results using a smaller balanced validation set and the complete validation set.\n\nI saw a post about how the xmin, ymin, xmax, ymax are in a difference sequence but I have triple checked and I think my submissions have the right order. I have added bounding boxes to some of the test images and they look reasonable.\n\nAn example image (ignore any spacing errors from this copy/paste): \nd504e26d60244000   /m/015p6 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 /m/0jbk 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2236626%2F862648fc56cf08eec13231bf5098849b%2Fd504e26d60244000-boxes.jpg?generation=1569544119041305&amp;alt=media)\n\nThe bounding box looks reasonable to me on the test image. Is this format or field order not correct?\n\nAny ideas what may cause the huge difference in my local validation scores and the leaderboard scores?\n"
    },
    {
      "id": 636162,
      "postDate": "2019-09-28T22:22:24.657Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 636641,
      "author_name": "AI Study",
      "author_url": "",
      "post_date": "2019-09-30T00:52:59.673000",
      "content": "<p>That is the whole line for the image. Only difference in the file is the comma between the imageid and the prediction string.\nd504e26d60244000,/m/015p6 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 /m/0jbk 0.9671 0.12052649 0.2865757 0.41084218 0.90251577\nI'll have to double check on the thresholds but I'm pretty sure they were the same. \nActually I'm not sure at all since the evaluation was using the script from tensorflow object detection (models\\research\\object_detection\\inference\\infer_detections.py). I'll have to look at that tomorrow. I'm seeing some predictions in the .30+ range so thats probably their threshold.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 636839,
          "author_name": "Artyom Palvelev",
          "author_url": "",
          "post_date": "2019-09-30T09:01:50.110000",
          "content": "<p>Well, my model says that it's Canary (/m/0ccs93).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 636854,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2019-09-30T09:16:42.893000",
          "content": "<p>yep, I'll bet you 1$ it's th.... with proper th you will get at least 30 predictions for such image</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 637180,
          "author_name": "AI Study",
          "author_url": "",
          "post_date": "2019-09-30T18:22:52.223000",
          "content": "<p>Looks like you were correct and it is a threshold issue. I created a new submission with a lower threshold from an older run that had the complete prediction list for each image and got a much better score.\nCanary, for example was under .30 but included in the new submission.\nUnfortunately not enough time left to run a full evaluation with my latest work but it helps to know what the cause was.\nThanks for the responses.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 636639,
      "author_name": "Moshel",
      "author_url": "",
      "post_date": "2019-09-30T00:32:47.460000",
      "content": "<p>can you please send the full line for the image? I am sure it does not voilate anything to send 1 out of 100000 images.\nalso, did you make sure you used the same threshold for inferencing the validation and test?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636638,
      "author_name": "AI Study",
      "author_url": "",
      "post_date": "2019-09-30T00:00:39.197000",
      "content": "<p>Thanks for the reply.\nThat is only part of the image. The ImageID is d504e26d60244000 so you can look at a local copy and see the whole thing if you wish.\nI previously opened the image and calculated where the pixels would be and they are in the right locations as far as I can tell.\nImage is 1024 wide, 755 high\nXmin is at pixel 123 (1024*0.12052649 calculating from upper left corner)\nXmax at 420 (1024*0.41084218)\nYmin at 216 (755*0.2865757)\nYmax at 681 (755*0.90251577)\nSo upper left corner is at 123,216 - lower right corner is at 420,681</p>\n\n<p>I checked some training images and the calculations from the given bounding boxes were from upper left so that should be right. The field order is different but I adjusted for that.</p>\n\n<p>The example I gave is in Xmin, Ymin, Xmax, Ymax order according to my calculations.</p>\n\n<p>/m/015p6 is bird which looks correct to me\n/m/0jbk is animal which looks correct to me</p>\n\n<p>I'm obviously missing something but have no idea what else to look at.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636163,
      "author_name": "Artyom Palvelev",
      "author_url": "",
      "post_date": "2019-09-28T22:26:36.097000",
      "content": "<p>Wait, what are your coordinates for this image? Y range doesn't seem to be 0.28 - 0.90. Maybe your coordinates are flipped or something like this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636162,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-28T22:22:24.657000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "636641": "That is the whole line for the image. Only difference in the file is the comma between the imageid and the prediction string.\nd504e26d60244000,/m/015p6 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 /m/0jbk 0.9671 0.12052649 0.2865757 0.41084218 0.90251577\nI'll have to double check on the thresholds but I'm pretty sure they were the same. \nActually I'm not sure at all since the evaluation was using the script from tensorflow object detection (models\\research\\object_detection\\inference\\infer_detections.py). I'll have to look at that tomorrow. I'm seeing some predictions in the .30+ range so thats probably their threshold.",
    "636639": "can you please send the full line for the image? I am sure it does not voilate anything to send 1 out of 100000 images.\nalso, did you make sure you used the same threshold for inferencing the validation and test?",
    "636638": "Thanks for the reply.\nThat is only part of the image. The ImageID is d504e26d60244000 so you can look at a local copy and see the whole thing if you wish.\nI previously opened the image and calculated where the pixels would be and they are in the right locations as far as I can tell.\nImage is 1024 wide, 755 high\nXmin is at pixel 123 (1024*0.12052649 calculating from upper left corner)\nXmax at 420 (1024*0.41084218)\nYmin at 216 (755*0.2865757)\nYmax at 681 (755*0.90251577)\nSo upper left corner is at 123,216 - lower right corner is at 420,681\n\nI checked some training images and the calculations from the given bounding boxes were from upper left so that should be right. The field order is different but I adjusted for that.\n\nThe example I gave is in Xmin, Ymin, Xmax, Ymax order according to my calculations.\n\n/m/015p6 is bird which looks correct to me\n/m/0jbk is animal which looks correct to me\n\nI'm obviously missing something but have no idea what else to look at.",
    "636163": "Wait, what are your coordinates for this image? Y range doesn't seem to be 0.28 - 0.90. Maybe your coordinates are flipped or something like this?",
    "634921": "I don't understand what is happening. When I run an evaluation on my local validation images I get an mAP score of 40%+ but when I upload a submission I get between 5-10% on the test images.\n\nI really don't think I'm overfitting since I'm using the complete training set and barely get through 1 complete epoch before submitting. I also don't think the unbalanced dataset is causing this since I get similar results using a smaller balanced validation set and the complete validation set.\n\nI saw a post about how the xmin, ymin, xmax, ymax are in a difference sequence but I have triple checked and I think my submissions have the right order. I have added bounding boxes to some of the test images and they look reasonable.\n\nAn example image (ignore any spacing errors from this copy/paste): \nd504e26d60244000   /m/015p6 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 /m/0jbk 0.9671 0.12052649 0.2865757 0.41084218 0.90251577 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2236626%2F862648fc56cf08eec13231bf5098849b%2Fd504e26d60244000-boxes.jpg?generation=1569544119041305&amp;alt=media)\n\nThe bounding box looks reasonable to me on the test image. Is this format or field order not correct?\n\nAny ideas what may cause the huge difference in my local validation scores and the leaderboard scores?\n",
    "636162": ""
  }
}