{
  "id": 114139,
  "title": "How do we transform a a point from an image to 3D world space(x,y,z)?",
  "url": "/competitions/pku-autonomous-driving/discussion/114139",
  "author_name": "",
  "post_date": "2019-10-24T14:15:29.869872600Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We have the camera intrinsic parameters, but the images are just RGB, without the depth information is there a mathematical way to transform points from images to 3D coordinates?</p>",
  "messages": [
    {
      "id": "656672",
      "postDate": "10/24/2019 14:15:29",
      "content": "<p>We have the camera intrinsic parameters, but the images are just RGB, without the depth information is there a mathematical way to transform points from images to 3D coordinates?</p>",
      "rawMarkdown": "We have the camera intrinsic parameters, but the images are just RGB, without the depth information is there a mathematical way to transform points from images to 3D coordinates?",
      "votes": null
    },
    {
      "id": "657086",
      "postDate": "10/24/2019 23:08:22",
      "content": "<p>I would imagine that would be difficult, without knowing more about the camera and the circumstances in which the photograph was taken.</p>\n\n<p><strong>EXAMPLE</strong>\nIf you have ever played with a camera that has interchangeable lenses or a built-in zoom or something similar, you know that you can remain at a fixed point in space and render the \"scene in front of you\" in a variety of perspective widths, depths of field, etc.</p>\n\n<p>In other words, given that one point in physical space can, if you are altering the settings on the camera, produce multiple \"measuring sticks\", I would think it would be hard without knowing more, as I mentioned before.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3788151%2F6a6cb7ba8c66c6e7568d8910009a77a6%2Fred-barn-sequence.jpg?generation=1571958738361462&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://www.nikonusa.com/en/learn-and-explore/a/tips-and-techniques/understanding-focal-length.html\">Nikon</a></p>\n\n<p>In the picture above, how \"far\" is the red barn?</p>\n\n<p>If the camera uses a fixed-lens, maybe there is a way to extrapolate, but even then I imagine its not trivial.</p>\n\n<p><strong>DISCLOSURE</strong>\nI am not formally trained in anything related to the above. Just a (fairly serious) hobbyist photographer. Put differently, while I <em>think</em> I am right, I absolutely do <em>not</em> want to steer you wrong, so you should definitely get a second opinion.</p>\n\n<p>Cheers, and hope this helps. If this idea works, its brilliant!</p>",
      "rawMarkdown": "I would imagine that would be difficult, without knowing more about the camera and the circumstances in which the photograph was taken.\n\n**EXAMPLE**\nIf you have ever played with a camera that has interchangeable lenses or a built-in zoom or something similar, you know that you can remain at a fixed point in space and render the \"scene in front of you\" in a variety of perspective widths, depths of field, etc.\n\nIn other words, given that one point in physical space can, if you are altering the settings on the camera, produce multiple \"measuring sticks\", I would think it would be hard without knowing more, as I mentioned before.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3788151%2F6a6cb7ba8c66c6e7568d8910009a77a6%2Fred-barn-sequence.jpg?generation=1571958738361462&amp;alt=media)\n\n[Nikon](https://www.nikonusa.com/en/learn-and-explore/a/tips-and-techniques/understanding-focal-length.html)\n\nIn the picture above, how \"far\" is the red barn?\n\nIf the camera uses a fixed-lens, maybe there is a way to extrapolate, but even then I imagine its not trivial.\n\n**DISCLOSURE**\nI am not formally trained in anything related to the above. Just a (fairly serious) hobbyist photographer. Put differently, while I *think* I am right, I absolutely do *not* want to steer you wrong, so you should definitely get a second opinion.\n\nCheers, and hope this helps. If this idea works, its brilliant!",
      "votes": null
    },
    {
      "id": "657319",
      "postDate": "10/25/2019 03:22:28",
      "content": "<p>I make a notebookt to transform a point from an image to 3D world space, if it is wrong, please tell me.😄 \n<a href=\"https://www.kaggle.com/zstusnoopy/visualize-the-location-and-3d-bounding-box-of-car?scriptVersionId=22518286\">Visualize the location and 3d bounding box of car</a></p>",
      "rawMarkdown": "I make a notebookt to transform a point from an image to 3D world space, if it is wrong, please tell me.😄 \n[Visualize the location and 3d bounding box of car](https://www.kaggle.com/zstusnoopy/visualize-the-location-and-3d-bounding-box-of-car?scriptVersionId=22518286)",
      "votes": null
    },
    {
      "id": "657515",
      "postDate": "10/25/2019 07:47:54",
      "content": "<p>Hi Li! Very nice kernel, but in this line you used the Z coordinate from the prediction string. I replaced it for a dummy value, but it didn't work :(\n```</p>\n\n<h1>image coordinate to world coordinate</h1>\n\n<p>def img_cor_2_world_cor():\n    x_img, y_img, z_img = img_cor_points[0]\n```</p>",
      "rawMarkdown": "Hi Li! Very nice kernel, but in this line you used the Z coordinate from the prediction string. I replaced it for a dummy value, but it didn't work :(\n```\n# image coordinate to world coordinate\ndef img_cor_2_world_cor():\n    x_img, y_img, z_img = img_cor_points[0]\n```",
      "votes": null
    },
    {
      "id": "657631",
      "postDate": "10/25/2019 10:08:16",
      "content": "<p>The Z coordinate should be the distance between camera and car. In my kernel, I first convert the (x, y, z) from world coordinate to image coordinate, then I convert image coordinate back to world coordinate. The value can't be changed.</p>",
      "rawMarkdown": "The Z coordinate should be the distance between camera and car. In my kernel, I first convert the (x, y, z) from world coordinate to image coordinate, then I convert image coordinate back to world coordinate. The value can't be changed.",
      "votes": null
    },
    {
      "id": "657805",
      "postDate": "10/25/2019 12:40:02",
      "content": "<p>Hmm. <a href=\"/zstusnoopy\">@zstusnoopy</a> Let me take a deeper look? As I mentioned above, if I am wrong, I will happily admit it, as I think this path could prove useful, if possible.</p>\n\n<p>In other words, more than happy to be wrong. In either event, a brief look at the notebook shows some really nice/interesting work. Congrats on that!</p>\n\n<p>🤘 </p>",
      "rawMarkdown": "Hmm. @zstusnoopy Let me take a deeper look? As I mentioned above, if I am wrong, I will happily admit it, as I think this path could prove useful, if possible.\n\nIn other words, more than happy to be wrong. In either event, a brief look at the notebook shows some really nice/interesting work. Congrats on that!\n\n🤘",
      "votes": null
    },
    {
      "id": "657825",
      "postDate": "10/25/2019 13:02:07",
      "content": "<p>Happy tohelp!</p>",
      "rawMarkdown": "Happy tohelp!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 657086,
      "author_name": "zer0state",
      "author_url": "",
      "post_date": "10/24/2019 23:08:22",
      "content": "<p>I would imagine that would be difficult, without knowing more about the camera and the circumstances in which the photograph was taken.</p>\n\n<p><strong>EXAMPLE</strong>\nIf you have ever played with a camera that has interchangeable lenses or a built-in zoom or something similar, you know that you can remain at a fixed point in space and render the \"scene in front of you\" in a variety of perspective widths, depths of field, etc.</p>\n\n<p>In other words, given that one point in physical space can, if you are altering the settings on the camera, produce multiple \"measuring sticks\", I would think it would be hard without knowing more, as I mentioned before.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3788151%2F6a6cb7ba8c66c6e7568d8910009a77a6%2Fred-barn-sequence.jpg?generation=1571958738361462&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://www.nikonusa.com/en/learn-and-explore/a/tips-and-techniques/understanding-focal-length.html\">Nikon</a></p>\n\n<p>In the picture above, how \"far\" is the red barn?</p>\n\n<p>If the camera uses a fixed-lens, maybe there is a way to extrapolate, but even then I imagine its not trivial.</p>\n\n<p><strong>DISCLOSURE</strong>\nI am not formally trained in anything related to the above. Just a (fairly serious) hobbyist photographer. Put differently, while I <em>think</em> I am right, I absolutely do <em>not</em> want to steer you wrong, so you should definitely get a second opinion.</p>\n\n<p>Cheers, and hope this helps. If this idea works, its brilliant!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 657319,
      "author_name": "zstusnoopy",
      "author_url": "",
      "post_date": "10/25/2019 03:22:28",
      "content": "<p>I make a notebookt to transform a point from an image to 3D world space, if it is wrong, please tell me.😄 \n<a href=\"https://www.kaggle.com/zstusnoopy/visualize-the-location-and-3d-bounding-box-of-car?scriptVersionId=22518286\">Visualize the location and 3d bounding box of car</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 657515,
          "author_name": "doncalculator",
          "author_url": "",
          "post_date": "10/25/2019 07:47:54",
          "content": "<p>Hi Li! Very nice kernel, but in this line you used the Z coordinate from the prediction string. I replaced it for a dummy value, but it didn't work :(\n```</p>\n\n<h1>image coordinate to world coordinate</h1>\n\n<p>def img_cor_2_world_cor():\n    x_img, y_img, z_img = img_cor_points[0]\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 657631,
          "author_name": "zstusnoopy",
          "author_url": "",
          "post_date": "10/25/2019 10:08:16",
          "content": "<p>The Z coordinate should be the distance between camera and car. In my kernel, I first convert the (x, y, z) from world coordinate to image coordinate, then I convert image coordinate back to world coordinate. The value can't be changed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 657805,
          "author_name": "zer0state",
          "author_url": "",
          "post_date": "10/25/2019 12:40:02",
          "content": "<p>Hmm. <a href=\"/zstusnoopy\">@zstusnoopy</a> Let me take a deeper look? As I mentioned above, if I am wrong, I will happily admit it, as I think this path could prove useful, if possible.</p>\n\n<p>In other words, more than happy to be wrong. In either event, a brief look at the notebook shows some really nice/interesting work. Congrats on that!</p>\n\n<p>🤘 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 657825,
          "author_name": "zstusnoopy",
          "author_url": "",
          "post_date": "10/25/2019 13:02:07",
          "content": "<p>Happy tohelp!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "656672": "We have the camera intrinsic parameters, but the images are just RGB, without the depth information is there a mathematical way to transform points from images to 3D coordinates?",
    "657086": "I would imagine that would be difficult, without knowing more about the camera and the circumstances in which the photograph was taken.\n\n**EXAMPLE**\nIf you have ever played with a camera that has interchangeable lenses or a built-in zoom or something similar, you know that you can remain at a fixed point in space and render the \"scene in front of you\" in a variety of perspective widths, depths of field, etc.\n\nIn other words, given that one point in physical space can, if you are altering the settings on the camera, produce multiple \"measuring sticks\", I would think it would be hard without knowing more, as I mentioned before.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3788151%2F6a6cb7ba8c66c6e7568d8910009a77a6%2Fred-barn-sequence.jpg?generation=1571958738361462&amp;alt=media)\n\n[Nikon](https://www.nikonusa.com/en/learn-and-explore/a/tips-and-techniques/understanding-focal-length.html)\n\nIn the picture above, how \"far\" is the red barn?\n\nIf the camera uses a fixed-lens, maybe there is a way to extrapolate, but even then I imagine its not trivial.\n\n**DISCLOSURE**\nI am not formally trained in anything related to the above. Just a (fairly serious) hobbyist photographer. Put differently, while I *think* I am right, I absolutely do *not* want to steer you wrong, so you should definitely get a second opinion.\n\nCheers, and hope this helps. If this idea works, its brilliant!",
    "657319": "I make a notebookt to transform a point from an image to 3D world space, if it is wrong, please tell me.😄 \n[Visualize the location and 3d bounding box of car](https://www.kaggle.com/zstusnoopy/visualize-the-location-and-3d-bounding-box-of-car?scriptVersionId=22518286)",
    "657515": "Hi Li! Very nice kernel, but in this line you used the Z coordinate from the prediction string. I replaced it for a dummy value, but it didn't work :(\n```\n# image coordinate to world coordinate\ndef img_cor_2_world_cor():\n    x_img, y_img, z_img = img_cor_points[0]\n```",
    "657631": "The Z coordinate should be the distance between camera and car. In my kernel, I first convert the (x, y, z) from world coordinate to image coordinate, then I convert image coordinate back to world coordinate. The value can't be changed.",
    "657805": "Hmm. @zstusnoopy Let me take a deeper look? As I mentioned above, if I am wrong, I will happily admit it, as I think this path could prove useful, if possible.\n\nIn other words, more than happy to be wrong. In either event, a brief look at the notebook shows some really nice/interesting work. Congrats on that!\n\n🤘",
    "657825": "Happy tohelp!"
  },
  "source": "meta"
}