{
  "id": 123653,
  "title": "Crop and Resize? And flipped.",
  "url": "/competitions/pku-autonomous-driving/discussion/123653",
  "author_name": "",
  "post_date": "2019-12-29T10:47:04.498810Z",
  "votes": 29,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Dear Competition organisor,\n     May I ask \n(1)  if the image is cropped and resized, are we still predicting the same \"X,Y,Z\" translation in world space?  Because if the image is cropped and resized, then the camera intrinsic will not be applicable here. To be more specific: the cars rendered on the image will not be overlapping with the actual RGB image.\n (2)  Moreover, if the image is flipped, since the camera intrinsic doesn't strictly overlap with the image centre, the flipped image will not simply be X=-X. But more like X = 2*delta_x /Z - x_camera\n     I wish the organisor take my above two comments seriously and give a more detailed answer. Because in the training set, there is no such augmentation, hence no ground truth annotation is given. So the participants of this challenge can not just \"guess\" what is the right answer. (I think @tito asked the same question). \n   The above two issues are quite technical, given the first round of the leaderboard didn't even take the False Negative into consideration, I presumptuously  assume that the challenge organisor didn't think these two issues thoroughly.  Or at least: provide the participants with enough knowledge about how final test set's ground truth is annotated if the image is cropped or flipped.</p>\n\n<p>Therefore, if the organisor can give more knowledge whether the test annotation is done in which of the following manner:</p>\n\n<p><em><strong>(1)  annotate the test image and then crop.</strong></em></p>\n\n<p><em><strong>(2) crop the image and then annotate the cropped the image.</strong></em></p>\n\n<p><em><strong>They are two different tasks!!!!!</strong></em></p>\n\n<p>For example, with the cropped image, the current camera intrinsic will not predict the mesh align with the RGB image:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F16463%2F3681cffe067ba30473b6122d6cafc457%2FID_2c40ef6e2.jpg?generation=1577669853089412&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "705715",
      "postDate": "12/29/2019 10:47:04",
      "content": "<p>Dear Competition organisor,\n     May I ask \n(1)  if the image is cropped and resized, are we still predicting the same \"X,Y,Z\" translation in world space?  Because if the image is cropped and resized, then the camera intrinsic will not be applicable here. To be more specific: the cars rendered on the image will not be overlapping with the actual RGB image.\n (2)  Moreover, if the image is flipped, since the camera intrinsic doesn't strictly overlap with the image centre, the flipped image will not simply be X=-X. But more like X = 2*delta_x /Z - x_camera\n     I wish the organisor take my above two comments seriously and give a more detailed answer. Because in the training set, there is no such augmentation, hence no ground truth annotation is given. So the participants of this challenge can not just \"guess\" what is the right answer. (I think @tito asked the same question). \n   The above two issues are quite technical, given the first round of the leaderboard didn't even take the False Negative into consideration, I presumptuously  assume that the challenge organisor didn't think these two issues thoroughly.  Or at least: provide the participants with enough knowledge about how final test set's ground truth is annotated if the image is cropped or flipped.</p>\n\n<p>Therefore, if the organisor can give more knowledge whether the test annotation is done in which of the following manner:</p>\n\n<p><em><strong>(1)  annotate the test image and then crop.</strong></em></p>\n\n<p><em><strong>(2) crop the image and then annotate the cropped the image.</strong></em></p>\n\n<p><em><strong>They are two different tasks!!!!!</strong></em></p>\n\n<p>For example, with the cropped image, the current camera intrinsic will not predict the mesh align with the RGB image:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F16463%2F3681cffe067ba30473b6122d6cafc457%2FID_2c40ef6e2.jpg?generation=1577669853089412&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Dear Competition organisor,\n     May I ask \n(1)  if the image is cropped and resized, are we still predicting the same \"X,Y,Z\" translation in world space?  Because if the image is cropped and resized, then the camera intrinsic will not be applicable here. To be more specific: the cars rendered on the image will not be overlapping with the actual RGB image.\n (2)  Moreover, if the image is flipped, since the camera intrinsic doesn't strictly overlap with the image centre, the flipped image will not simply be X=-X. But more like X = 2*delta_x /Z - x_camera\n     I wish the organisor take my above two comments seriously and give a more detailed answer. Because in the training set, there is no such augmentation, hence no ground truth annotation is given. So the participants of this challenge can not just \"guess\" what is the right answer. (I think @tito asked the same question). \n   The above two issues are quite technical, given the first round of the leaderboard didn't even take the False Negative into consideration, I presumptuously  assume that the challenge organisor didn't think these two issues thoroughly.  Or at least: provide the participants with enough knowledge about how final test set's ground truth is annotated if the image is cropped or flipped.\n\nTherefore, if the organisor can give more knowledge whether the test annotation is done in which of the following manner:\n\n***(1)  annotate the test image and then crop.***\n\n***(2) crop the image and then annotate the cropped the image.***\n\n***They are two different tasks!!!!!***\n\nFor example, with the cropped image, the current camera intrinsic will not predict the mesh align with the RGB image:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F16463%2F3681cffe067ba30473b6122d6cafc457%2FID_2c40ef6e2.jpg?generation=1577669853089412&amp;alt=media)",
      "votes": null
    },
    {
      "id": "706016",
      "postDate": "12/29/2019 20:19:34",
      "content": "<p>Valid questions <a href=\"/stevenwudi\">@stevenwudi</a> , sadly I am not expecting answers from the hosts. They seems to be somewhat absent besides the very obvious FN fix in their metric.</p>",
      "rawMarkdown": "Valid questions @stevenwudi , sadly I am not expecting answers from the hosts. They seems to be somewhat absent besides the very obvious FN fix in their metric.",
      "votes": null
    },
    {
      "id": "706582",
      "postDate": "12/30/2019 15:40:46",
      "content": "<p>i observed similar attitude from hosts in past lyft competition :(</p>",
      "rawMarkdown": "i observed similar attitude from hosts in past lyft competition :(",
      "votes": null
    },
    {
      "id": "712276",
      "postDate": "01/07/2020 02:56:49",
      "content": "<p>What is the value of deltax and xcamera? I can't understand this formula yet. <a href=\"/stevenwudi\">@stevenwudi</a> </p>",
      "rawMarkdown": "What is the value of deltax and xcamera? I can't understand this formula yet. @stevenwudi",
      "votes": null
    },
    {
      "id": "712282",
      "postDate": "01/07/2020 02:59:31",
      "content": "<p>I directly use -X as the gt of the flipped image, and the LB do not improve compared to the previous result.</p>",
      "rawMarkdown": "I directly use -X as the gt of the flipped image, and the LB do not improve compared to the previous result.",
      "votes": null
    },
    {
      "id": "721795",
      "postDate": "01/17/2020 17:45:07",
      "content": "<p>deltax I understand as the centre offset of the camera intrinsic from the image centre.</p>",
      "rawMarkdown": "deltax I understand as the centre offset of the camera intrinsic from the image centre.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 706016,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "12/29/2019 20:19:34",
      "content": "<p>Valid questions <a href=\"/stevenwudi\">@stevenwudi</a> , sadly I am not expecting answers from the hosts. They seems to be somewhat absent besides the very obvious FN fix in their metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 706582,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "12/30/2019 15:40:46",
          "content": "<p>i observed similar attitude from hosts in past lyft competition :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 712276,
      "author_name": "lida0372",
      "author_url": "",
      "post_date": "01/07/2020 02:56:49",
      "content": "<p>What is the value of deltax and xcamera? I can't understand this formula yet. <a href=\"/stevenwudi\">@stevenwudi</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 721795,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "01/17/2020 17:45:07",
          "content": "<p>deltax I understand as the centre offset of the camera intrinsic from the image centre.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 712282,
      "author_name": "lida0372",
      "author_url": "",
      "post_date": "01/07/2020 02:59:31",
      "content": "<p>I directly use -X as the gt of the flipped image, and the LB do not improve compared to the previous result.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "705715": "Dear Competition organisor,\n     May I ask \n(1)  if the image is cropped and resized, are we still predicting the same \"X,Y,Z\" translation in world space?  Because if the image is cropped and resized, then the camera intrinsic will not be applicable here. To be more specific: the cars rendered on the image will not be overlapping with the actual RGB image.\n (2)  Moreover, if the image is flipped, since the camera intrinsic doesn't strictly overlap with the image centre, the flipped image will not simply be X=-X. But more like X = 2*delta_x /Z - x_camera\n     I wish the organisor take my above two comments seriously and give a more detailed answer. Because in the training set, there is no such augmentation, hence no ground truth annotation is given. So the participants of this challenge can not just \"guess\" what is the right answer. (I think @tito asked the same question). \n   The above two issues are quite technical, given the first round of the leaderboard didn't even take the False Negative into consideration, I presumptuously  assume that the challenge organisor didn't think these two issues thoroughly.  Or at least: provide the participants with enough knowledge about how final test set's ground truth is annotated if the image is cropped or flipped.\n\nTherefore, if the organisor can give more knowledge whether the test annotation is done in which of the following manner:\n\n***(1)  annotate the test image and then crop.***\n\n***(2) crop the image and then annotate the cropped the image.***\n\n***They are two different tasks!!!!!***\n\nFor example, with the cropped image, the current camera intrinsic will not predict the mesh align with the RGB image:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F16463%2F3681cffe067ba30473b6122d6cafc457%2FID_2c40ef6e2.jpg?generation=1577669853089412&amp;alt=media)",
    "706016": "Valid questions @stevenwudi , sadly I am not expecting answers from the hosts. They seems to be somewhat absent besides the very obvious FN fix in their metric.",
    "706582": "i observed similar attitude from hosts in past lyft competition :(",
    "712276": "What is the value of deltax and xcamera? I can't understand this formula yet. @stevenwudi",
    "712282": "I directly use -X as the gt of the flipped image, and the LB do not improve compared to the previous result.",
    "721795": "deltax I understand as the centre offset of the camera intrinsic from the image centre."
  },
  "source": "meta"
}