{
  "id": 120083,
  "title": "Understand Camera Intrinsic Parameters and X, Y, Z",
  "url": "/competitions/pku-autonomous-driving/discussion/120083",
  "author_name": "",
  "post_date": "2019-12-03T14:33:04.427469900Z",
  "votes": 59,
  "comment_count": 30,
  "views": 0,
  "content": "<p><img src=\"http://image71.360doc.com/DownloadImg/2014/04/1014/40652696_5.jpg\" alt=\"1\"></p>\n\n<p><strong>There are acutally four coordinate the world coordinate (Xw, Yw, Zw)、camera coordinate (Xc, Yc, Zc)、image coordinates (x, y) and pixel coordinates (u, v).</strong></p>\n\n<p><strong>In this competition the world coordinate (Xw, Yw, Zw) is same with camera coordinate (Xc, Yc, Zc), the camera is origin.</strong></p>\n\n<p><strong>Our goal is transforming the world coordinate (Xw, Yw, Zw) to pixel coordinates (u, v).</strong></p>\n\n<h2>1 -</h2>\n\n<p><strong>As I said before, (Xc, Yc, Zc) = (Xw, Yw, Zw), so first step is transforming camera coordinate (Xc, Yc, Zc)(meter) to image coordinates (x, y)(millimeter) and f is focal length.</strong></p>\n\n<p><img src=\"http://www.pianshen.com/images/647/fd1c1f4cfa17e5bb5218d80a425c3a57.png\" alt=\"4\"></p>\n\n<h2>2 -</h2>\n\n<p><strong>Then transforming image coordinates (x, y)(millimeter) to pixel coordinates (u, v)(pixel), and dx, dy are the actual size of pixels on the sensitive chip</strong></p>\n\n<p><img src=\"http://www.pianshen.com/images/592/6a06910d057cbd436d710b9dd978bf68.png\" alt=\"3\"></p>\n\n<h2>3 -</h2>\n\n<p><strong>Actually, we use matrix of Camera Intrinsic Parameters to compute. fx = f / dx and fy = f / dy.</strong>\n*<em>In this competition, the Camera Intrinsic is fx = 2304.5479; fy = 2305.8757; u0 = 1686.2379; v0 = 1354.9849;</em>*</p>\n\n<p><img src=\"http://www.pianshen.com/images/654/e166052ae19a43a7a3e628f1722373d6.png\" alt=\"4\"></p>\n\n<h2>Please give me some upvotes, if you think it's useful, it's a support for my work, thanks😎</h2>\n\n<p>These are my other topics.\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/123385\">The coordinate system of this competition</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120076\">The algorithm that baidu apollo chooses!</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120015\">Algorithm Selection for Beginner!</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120381\">Algorithm Selection in 6D Pose Estimation</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120443\">How the Competition Organizer get train.csv !</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120582\">Source of 6D Pose Estimation</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653\">A way to improve the accuracy of model</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120710\">Object Detection in 20 Years</a></p>\n\n<p>My notebook:\n<a href=\"https://www.kaggle.com/diegojohnson/centernet-objects-as-points\">Centernet - Objects as Points</a>\n<a href=\"https://www.kaggle.com/diegojohnson/a-way-to-regress-translation-and-rotation?scriptVersionId=24677131\">A Way to Regress Translation and Rotation</a>\n<a href=\"https://www.kaggle.com/diegojohnson/best-algorithm-so-far-google-s-new-paper-in-nov\">Best Algorithm so far - Google's new paper in Nov</a>\n<a href=\"https://www.kaggle.com/diegojohnson/a-clear-view-of-car-pose\">A clear view of car pose</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/124480\">Dataset with preprocess</a></p>",
  "messages": [
    {
      "id": "686785",
      "postDate": "12/03/2019 14:33:04",
      "content": "<p><img src=\"http://image71.360doc.com/DownloadImg/2014/04/1014/40652696_5.jpg\" alt=\"1\"></p>\n\n<p><strong>There are acutally four coordinate the world coordinate (Xw, Yw, Zw)、camera coordinate (Xc, Yc, Zc)、image coordinates (x, y) and pixel coordinates (u, v).</strong></p>\n\n<p><strong>In this competition the world coordinate (Xw, Yw, Zw) is same with camera coordinate (Xc, Yc, Zc), the camera is origin.</strong></p>\n\n<p><strong>Our goal is transforming the world coordinate (Xw, Yw, Zw) to pixel coordinates (u, v).</strong></p>\n\n<h2>1 -</h2>\n\n<p><strong>As I said before, (Xc, Yc, Zc) = (Xw, Yw, Zw), so first step is transforming camera coordinate (Xc, Yc, Zc)(meter) to image coordinates (x, y)(millimeter) and f is focal length.</strong></p>\n\n<p><img src=\"http://www.pianshen.com/images/647/fd1c1f4cfa17e5bb5218d80a425c3a57.png\" alt=\"4\"></p>\n\n<h2>2 -</h2>\n\n<p><strong>Then transforming image coordinates (x, y)(millimeter) to pixel coordinates (u, v)(pixel), and dx, dy are the actual size of pixels on the sensitive chip</strong></p>\n\n<p><img src=\"http://www.pianshen.com/images/592/6a06910d057cbd436d710b9dd978bf68.png\" alt=\"3\"></p>\n\n<h2>3 -</h2>\n\n<p><strong>Actually, we use matrix of Camera Intrinsic Parameters to compute. fx = f / dx and fy = f / dy.</strong>\n*<em>In this competition, the Camera Intrinsic is fx = 2304.5479; fy = 2305.8757; u0 = 1686.2379; v0 = 1354.9849;</em>*</p>\n\n<p><img src=\"http://www.pianshen.com/images/654/e166052ae19a43a7a3e628f1722373d6.png\" alt=\"4\"></p>\n\n<h2>Please give me some upvotes, if you think it's useful, it's a support for my work, thanks😎</h2>\n\n<p>These are my other topics.\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/123385\">The coordinate system of this competition</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120076\">The algorithm that baidu apollo chooses!</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120015\">Algorithm Selection for Beginner!</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120381\">Algorithm Selection in 6D Pose Estimation</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120443\">How the Competition Organizer get train.csv !</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120582\">Source of 6D Pose Estimation</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653\">A way to improve the accuracy of model</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120710\">Object Detection in 20 Years</a></p>\n\n<p>My notebook:\n<a href=\"https://www.kaggle.com/diegojohnson/centernet-objects-as-points\">Centernet - Objects as Points</a>\n<a href=\"https://www.kaggle.com/diegojohnson/a-way-to-regress-translation-and-rotation?scriptVersionId=24677131\">A Way to Regress Translation and Rotation</a>\n<a href=\"https://www.kaggle.com/diegojohnson/best-algorithm-so-far-google-s-new-paper-in-nov\">Best Algorithm so far - Google's new paper in Nov</a>\n<a href=\"https://www.kaggle.com/diegojohnson/a-clear-view-of-car-pose\">A clear view of car pose</a>\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/124480\">Dataset with preprocess</a></p>",
      "rawMarkdown": "![1](http://image71.360doc.com/DownloadImg/2014/04/1014/40652696_5.jpg)\n\n**There are acutally four coordinate the world coordinate (Xw, Yw, Zw)、camera coordinate (Xc, Yc, Zc)、image coordinates (x, y) and pixel coordinates (u, v).**\n\n**In this competition the world coordinate (Xw, Yw, Zw) is same with camera coordinate (Xc, Yc, Zc), the camera is origin.**\n\n**Our goal is transforming the world coordinate (Xw, Yw, Zw) to pixel coordinates (u, v).**\n\n## 1 - \n**As I said before, (Xc, Yc, Zc) = (Xw, Yw, Zw), so first step is transforming camera coordinate (Xc, Yc, Zc)(meter) to image coordinates (x, y)(millimeter) and f is focal length.**\n\n![4](http://www.pianshen.com/images/647/fd1c1f4cfa17e5bb5218d80a425c3a57.png)\n\n## 2 - \n**Then transforming image coordinates (x, y)(millimeter) to pixel coordinates (u, v)(pixel), and dx, dy are the actual size of pixels on the sensitive chip**\n\n![3](http://www.pianshen.com/images/592/6a06910d057cbd436d710b9dd978bf68.png)\n\n## 3 -\n**Actually, we use matrix of Camera Intrinsic Parameters to compute. fx = f / dx and fy = f / dy.**\n**In this competition, the Camera Intrinsic is fx = 2304.5479; fy = 2305.8757; u0 = 1686.2379; v0 = 1354.9849;**\n\n![4](http://www.pianshen.com/images/654/e166052ae19a43a7a3e628f1722373d6.png)\n\n## Please give me some upvotes, if you think it's useful, it's a support for my work, thanks😎 \n\nThese are my other topics.\n[The coordinate system of this competition](https://www.kaggle.com/c/pku-autonomous-driving/discussion/123385)\n[The algorithm that baidu apollo chooses!](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120076)\n[Algorithm Selection for Beginner!](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120015)\n[Algorithm Selection in 6D Pose Estimation](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120381)\n[How the Competition Organizer get train.csv !](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120443)\n[Source of 6D Pose Estimation](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120582)\n[A way to improve the accuracy of model](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653)\n[Object Detection in 20 Years](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120710)\n\n\nMy notebook:\n[Centernet - Objects as Points](https://www.kaggle.com/diegojohnson/centernet-objects-as-points)\n[A Way to Regress Translation and Rotation](https://www.kaggle.com/diegojohnson/a-way-to-regress-translation-and-rotation?scriptVersionId=24677131)\n[Best Algorithm so far - Google's new paper in Nov](https://www.kaggle.com/diegojohnson/best-algorithm-so-far-google-s-new-paper-in-nov)\n[A clear view of car pose](https://www.kaggle.com/diegojohnson/a-clear-view-of-car-pose)\n[Dataset with preprocess](https://www.kaggle.com/c/pku-autonomous-driving/discussion/124480)",
      "votes": null
    },
    {
      "id": "686826",
      "postDate": "12/03/2019 15:40:25",
      "content": "<p>I see many notebooks use a function to complete 3d-to-2d, but I'd like to know the principle inside.</p>\n\n<p>so I google it and write this topic, hope you like it.😎 </p>",
      "rawMarkdown": "I see many notebooks use a function to complete 3d-to-2d, but I'd like to know the principle inside.\n\nso I google it and write this topic, hope you like it.😎",
      "votes": null
    },
    {
      "id": "687345",
      "postDate": "12/04/2019 09:18:43",
      "content": "<p>Very Informative\nThanks <a href=\"/diegojohnson\">@diegojohnson</a> </p>",
      "rawMarkdown": "Very Informative\nThanks @diegojohnson",
      "votes": null
    },
    {
      "id": "687359",
      "postDate": "12/04/2019 09:35:58",
      "content": "<p>My pleasure😏 </p>",
      "rawMarkdown": "My pleasure😏",
      "votes": null
    },
    {
      "id": "689381",
      "postDate": "12/06/2019 21:31:47",
      "content": "<p>What are R and T in the last equation?</p>",
      "rawMarkdown": "What are R and T in the last equation?",
      "votes": null
    },
    {
      "id": "689392",
      "postDate": "12/06/2019 22:00:16",
      "content": "<p>Rotation matrix and Translation vector.</p>",
      "rawMarkdown": "Rotation matrix and Translation vector.",
      "votes": null
    },
    {
      "id": "689519",
      "postDate": "12/07/2019 02:08:45",
      "content": "<p>Yeah👍 </p>",
      "rawMarkdown": "Yeah👍",
      "votes": null
    },
    {
      "id": "691468",
      "postDate": "12/10/2019 07:10:23",
      "content": "<p>Thank You!</p>",
      "rawMarkdown": "Thank You!",
      "votes": null
    },
    {
      "id": "691536",
      "postDate": "12/10/2019 08:38:16",
      "content": "<p>my pleasure</p>",
      "rawMarkdown": "my pleasure",
      "votes": null
    },
    {
      "id": "695574",
      "postDate": "12/15/2019 11:20:57",
      "content": "<p>hi, are you sure <code>In this competition the world coordinate (Xw, Yw, Zw) is same with camera coordinate (Xc, Yc, Zc), the camera is origin.</code>，if world coordinate is same with camera coordinate ,which means we don't need matmul the matrix [R, T; 0, 1] when convert to pixel coordinates?</p>",
      "rawMarkdown": "hi, are you sure `In this competition the world coordinate (Xw, Yw, Zw) is same with camera coordinate (Xc, Yc, Zc), the camera is origin.`，if world coordinate is same with camera coordinate ,which means we don't need matmul the matrix [R, T; 0, 1] when convert to pixel coordinates?",
      "votes": null
    },
    {
      "id": "695676",
      "postDate": "12/15/2019 12:58:57",
      "content": "<p>'Extrinsic parameters is the position of the origin of the world coordinate system expressed in coordinates of the camera-centered coordinate system. '</p>\n\n<p>In this competition, the origin of the world coordinate is camera. <a href=\"/hustkevin1037\">@hustkevin1037</a> </p>",
      "rawMarkdown": "'Extrinsic parameters is the position of the origin of the world coordinate system expressed in coordinates of the camera-centered coordinate system. '\n\nIn this competition, the origin of the world coordinate is camera. @hustkevin1037",
      "votes": null
    },
    {
      "id": "696215",
      "postDate": "12/16/2019 09:41:24",
      "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> this gives me some more better picture..\n1) What is difference between Image coordinates and pixel coordinates . I suppose one of them could be image screen for  camera\n2) For model training we use u,v ? just want to understand rationale behind is it because of mask ?</p>",
      "rawMarkdown": "diegojohnson this gives me some more better picture..\n1) What is difference between Image coordinates and pixel coordinates . I suppose one of them could be image screen for  camera\n2) For model training we use u,v ? just want to understand rationale behind is it because of mask ?",
      "votes": null
    },
    {
      "id": "696229",
      "postDate": "12/16/2019 10:15:57",
      "content": "<p>Actually, the article has told you answer.\nAn image consists of a lot of pixels, and pixel coordinates can tell you the location of pixel.</p>",
      "rawMarkdown": "Actually, the article has told you answer.\nAn image consists of a lot of pixels, and pixel coordinates can tell you the location of pixel.",
      "votes": null
    },
    {
      "id": "696373",
      "postDate": "12/16/2019 14:28:31",
      "content": "<p>Thanks..i have taken the pos from your notebook\n1)How are we building the labels for it ,below is 128 is max limit for no of cars that are of interest ?\n2) if there are more than one prediction strings for an image ,how is that handled, coords is list of dicts\n<code>[{'id': 16, 'yaw': 0.254839, 'pitch': -2.57534, 'roll': -3.10256, 'x': 7.96539, 'y': 3.20066, 'z': 11.0225}, {'id': 56, 'yaw': 0.181647, 'pitch': -1.46947, 'roll': -3.12159, 'x': 9.60332, 'y': 4.66632, 'z': 19.339}, {'id': 70, 'yaw': 0.163072, 'pitch': -1.56865, 'roll': -3.11754, 'x': 10.39, 'y': 11.2219, 'z': 59.7825}, {'id': 70, 'yaw': 0.141942, 'pitch': -3.1395, 'roll': 3.11969, 'x': -9.59236, 'y': 5.13662, 'z': 24.7337}, {'id': 46, 'yaw': 0.163068, 'pitch': -2.08578, 'roll': -3.11754, 'x': 9.83335, 'y': 13.2689, 'z': 72.9323}]</code>\nwhat would be the labels for id 70 image\n3) purpose of Heatmap\n<code>\ndef pose(s, u, v):\n    regr = np.zeros([128, 128, 6], dtype='float32')\n    coords = str_to_coords(s)\n    for p_x, p_y, regr_dict in zip(u, v, coords):\n        if p_x &amp;gt;= 0 and p_x &amp;lt; 128 and p_y &amp;gt;= 0 and p_y &amp;lt; 128:\n            regr_dict.pop('id')\n            regr[floor(p_y), floor(p_x)] = [regr_dict[n] for n in regr_dict] <br>\n    return regr\n</code></p>",
      "rawMarkdown": "Thanks..i have taken the pos from your notebook\n1)How are we building the labels for it ,below is 128 is max limit for no of cars that are of interest ?\n2) if there are more than one prediction strings for an image ,how is that handled, coords is list of dicts\n`[{'id': 16, 'yaw': 0.254839, 'pitch': -2.57534, 'roll': -3.10256, 'x': 7.96539, 'y': 3.20066, 'z': 11.0225}, {'id': 56, 'yaw': 0.181647, 'pitch': -1.46947, 'roll': -3.12159, 'x': 9.60332, 'y': 4.66632, 'z': 19.339}, {'id': 70, 'yaw': 0.163072, 'pitch': -1.56865, 'roll': -3.11754, 'x': 10.39, 'y': 11.2219, 'z': 59.7825}, {'id': 70, 'yaw': 0.141942, 'pitch': -3.1395, 'roll': 3.11969, 'x': -9.59236, 'y': 5.13662, 'z': 24.7337}, {'id': 46, 'yaw': 0.163068, 'pitch': -2.08578, 'roll': -3.11754, 'x': 9.83335, 'y': 13.2689, 'z': 72.9323}]`\nwhat would be the labels for id 70 image\n3) purpose of Heatmap\n```\ndef pose(s, u, v):\n    regr = np.zeros([128, 128, 6], dtype='float32')\n    coords = str_to_coords(s)\n    for p_x, p_y, regr_dict in zip(u, v, coords):\n        if p_x &gt;= 0 and p_x &lt; 128 and p_y &gt;= 0 and p_y &lt; 128:\n            regr_dict.pop('id')\n            regr[floor(p_y), floor(p_x)] = [regr_dict[n] for n in regr_dict]    \n    return regr\n```",
      "votes": null
    },
    {
      "id": "696415",
      "postDate": "12/16/2019 15:25:08",
      "content": "<p>the regression label's shape is 128 * 128 * 6\npose function is for putting (yaw, pith, roll, x, y, z) in right position of 128 * 128\nI suggest you look at other runable notebook, acutally there's something wrong in my notebook😅 , and I haven't found it.\nand origin paper is the best place you learn this model, because all public notebook is from there. <a href=\"/jaideepvalani\">@jaideepvalani</a> </p>",
      "rawMarkdown": "the regression label's shape is 128 * 128 * 6\npose function is for putting (yaw, pith, roll, x, y, z) in right position of 128 * 128\nI suggest you look at other runable notebook, acutally there's something wrong in my notebook😅 , and I haven't found it.\nand origin paper is the best place you learn this model, because all public notebook is from there. @jaideepvalani",
      "votes": null
    },
    {
      "id": "699376",
      "postDate": "12/20/2019 11:15:41",
      "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> \n1 ) one thing i noticed while labeling after now m gaining more info, is that we use x,y 2d points just to assign these points their corresponding world coordinates and it is actually those we are doing regression for .Please correct me if wrong.</p>\n\n<p>2) How will these go during the inference. is overall process is like you are teaching model through the points on the 2d plane that these points position of car looking object is actually these coordinates of same car in real world??</p>\n\n<p><a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a></p>",
      "rawMarkdown": "diegojohnson \n1 ) one thing i noticed while labeling after now m gaining more info, is that we use x,y 2d points just to assign these points their corresponding world coordinates and it is actually those we are doing regression for .Please correct me if wrong.\n\n2) How will these go during the inference. is overall process is like you are teaching model through the points on the 2d plane that these points position of car looking object is actually these coordinates of same car in real world??\n\nhttps://www.kaggle.com/hocop1/centernet-baseline",
      "votes": null
    },
    {
      "id": "699516",
      "postDate": "12/20/2019 14:35:09",
      "content": "<p>we regress the values(x,y,z,yaw,pitch,roll whatever you want) in the car's postion.\nI think this picture is very clear. \n<img src=\"https://www.kaggleusercontent.com/kf/25098336/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..WjysAO3-6H-Z2s64qukatQ.E81-tzdiF0SR4VR3aFrzxMG-lDRTiVpw7U-WGxLHPTmOLj2JSP6Zn6cSmbF2fFO48Tk3vB6-p5n1opkydcQrl_bz3Z1URDEpcWnHtM-zBWWZ0r68nL8qKN2lfQvKZkjfbIzF8C77JEu6zgB5R0OFUyB3UzmvXSJafcORWYAGvBymsJznoK3QFsZvNRY8HpmJ.jXABv1dwhWPRTZyt98ziQA/__results___files/__results___23_0.png\" alt=\"1\"></p>",
      "rawMarkdown": "we regress the values(x,y,z,yaw,pitch,roll whatever you want) in the car's postion.\nI think this picture is very clear. \n![1](https://www.kaggleusercontent.com/kf/25098336/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..WjysAO3-6H-Z2s64qukatQ.E81-tzdiF0SR4VR3aFrzxMG-lDRTiVpw7U-WGxLHPTmOLj2JSP6Zn6cSmbF2fFO48Tk3vB6-p5n1opkydcQrl_bz3Z1URDEpcWnHtM-zBWWZ0r68nL8qKN2lfQvKZkjfbIzF8C77JEu6zgB5R0OFUyB3UzmvXSJafcORWYAGvBymsJznoK3QFsZvNRY8HpmJ.jXABv1dwhWPRTZyt98ziQA/__results___files/__results___23_0.png)",
      "votes": null
    },
    {
      "id": "699615",
      "postDate": "12/20/2019 16:32:42",
      "content": "<p>sorry about it.. i did wanted to ..\nyour picture is not loaded</p>",
      "rawMarkdown": "sorry about it.. i did wanted to ..\nyour picture is not loaded",
      "votes": null
    },
    {
      "id": "699840",
      "postDate": "12/21/2019 02:34:26",
      "content": "<p>It just the picture posted in my notebook. I think it's very clear. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F4afa7e282fbb6a913db336539bfd5a85%2F2019-12-21_103355.png?generation=1576895663407777&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "It just the picture posted in my notebook. I think it's very clear. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F4afa7e282fbb6a913db336539bfd5a85%2F2019-12-21_103355.png?generation=1576895663407777&amp;alt=media)",
      "votes": null
    },
    {
      "id": "699856",
      "postDate": "12/21/2019 03:13:12",
      "content": "<p>Thanku ..\n<code>mask_loss = mask * torch.log(pred_mask + 1e-12) + (1 - mask) * torch.log(1 - pred_mask + 1e-12)</code>\nI use this loss but unable to determine why it ends in nan \nCompared pytorch standard f.bce with logits  .\nDo you see any trouble ? I use mixed fp </p>",
      "rawMarkdown": "Thanku ..\n`mask_loss = mask * torch.log(pred_mask + 1e-12) + (1 - mask) * torch.log(1 - pred_mask + 1e-12)`\nI use this loss but unable to determine why it ends in nan \nCompared pytorch standard f.bce with logits  .\nDo you see any trouble ? I use mixed fp",
      "votes": null
    },
    {
      "id": "700445",
      "postDate": "12/22/2019 03:11:54",
      "content": "<p>Did you use <strong>sigmoid</strong>?\npred_mask = torch.sigmoid(prediction[:, 0])</p>",
      "rawMarkdown": "Did you use **sigmoid**?\npred_mask = torch.sigmoid(prediction[:, 0])",
      "votes": null
    },
    {
      "id": "700453",
      "postDate": "12/22/2019 03:45:42",
      "content": "<p>Yes . I figured out the reason for nan ,it was because to  accommodate higher bs I used mixed fp  . ,1e-12 was rounding to zero it was reason loss ending in nans ,I changed to 1e-7\nAnd dint dont get nan any more .</p>\n\n<p>2  Second issue is why results appear different from std fct  I dont know \n.\n3. M now on lb :) one thing is sure all public kernel based on hiccups kernel wont go beyond .08 like that .Loss is extremely fluctuating so it may not  let model generalize well reason I think of it is scales of regression and confidence are different . Few different things to be tried up.Would u like to team up ? </p>",
      "rawMarkdown": "Yes . I figured out the reason for nan ,it was because to  accommodate higher bs I used mixed fp  . ,1e-12 was rounding to zero it was reason loss ending in nans ,I changed to 1e-7\nAnd dint dont get nan any more .\n \n2  Second issue is why results appear different from std fct  I dont know \n.\n3. M now on lb :) one thing is sure all public kernel based on hiccups kernel wont go beyond .08 like that .Loss is extremely fluctuating so it may not  let model generalize well reason I think of it is scales of regression and confidence are different . Few different things to be tried up.Would u like to team up ?",
      "votes": null
    },
    {
      "id": "700478",
      "postDate": "12/22/2019 04:36:49",
      "content": "<p>sorry, I've teamed up with others.</p>",
      "rawMarkdown": "sorry, I've teamed up with others.",
      "votes": null
    },
    {
      "id": "702814",
      "postDate": "12/25/2019 07:54:10",
      "content": "<p>X,Y coordinates we need to predict are the image coordinates, right?</p>",
      "rawMarkdown": "X,Y coordinates we need to predict are the image coordinates, right?",
      "votes": null
    },
    {
      "id": "702818",
      "postDate": "12/25/2019 08:01:14",
      "content": "<p>World coordinate <a href=\"/prayagmadhu\">@prayagmadhu</a> </p>",
      "rawMarkdown": "World coordinate @prayagmadhu",
      "votes": null
    },
    {
      "id": "702819",
      "postDate": "12/25/2019 08:02:14",
      "content": "<p>is the coordinates given train.csv camera coordinates ?\nSo, we are given camera coordinates of cars in images, and task is to predict yaw, pitch roll and xyz image coordinates of unmasked cars, is this right ?</p>",
      "rawMarkdown": "is the coordinates given train.csv camera coordinates ?\nSo, we are given camera coordinates of cars in images, and task is to predict yaw, pitch roll and xyz image coordinates of unmasked cars, is this right ?",
      "votes": null
    },
    {
      "id": "702826",
      "postDate": "12/25/2019 08:19:03",
      "content": "<p>task is also to predict camera coordinates of unmasked cars.</p>",
      "rawMarkdown": "task is also to predict camera coordinates of unmasked cars.",
      "votes": null
    },
    {
      "id": "704377",
      "postDate": "12/27/2019 12:02:57",
      "content": "<p>In the last equation all intrinsic camera parameters are given except Zc, camera Z coordinate.\nMay I ask stupid question - how to get it?</p>",
      "rawMarkdown": "In the last equation all intrinsic camera parameters are given except Zc, camera Z coordinate.\nMay I ask stupid question - how to get it?",
      "votes": null
    },
    {
      "id": "704392",
      "postDate": "12/27/2019 12:28:31",
      "content": "<p>You will get Xc, Yc, Zc together. You can look at public notebook to get more compution detail.</p>",
      "rawMarkdown": "You will get Xc, Yc, Zc together. You can look at public notebook to get more compution detail.",
      "votes": null
    },
    {
      "id": "704399",
      "postDate": "12/27/2019 12:37:58",
      "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> , \nthanks for a guess!</p>\n\n<p>Which notebook you mean? Could you please drop a link?</p>",
      "rawMarkdown": "diegojohnson , \nthanks for a guess!\n\nWhich notebook you mean? Could you please drop a link?",
      "votes": null
    },
    {
      "id": "704404",
      "postDate": "12/27/2019 12:48:20",
      "content": "<p>Anyone is right, the computation is same.</p>",
      "rawMarkdown": "Anyone is right, the computation is same.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 686826,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/03/2019 15:40:25",
      "content": "<p>I see many notebooks use a function to complete 3d-to-2d, but I'd like to know the principle inside.</p>\n\n<p>so I google it and write this topic, hope you like it.😎 </p>",
      "votes": null,
      "replies": [
        {
          "id": 695574,
          "author_name": "hustkevin1037",
          "author_url": "",
          "post_date": "12/15/2019 11:20:57",
          "content": "<p>hi, are you sure <code>In this competition the world coordinate (Xw, Yw, Zw) is same with camera coordinate (Xc, Yc, Zc), the camera is origin.</code>，if world coordinate is same with camera coordinate ,which means we don't need matmul the matrix [R, T; 0, 1] when convert to pixel coordinates?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 695676,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/15/2019 12:58:57",
          "content": "<p>'Extrinsic parameters is the position of the origin of the world coordinate system expressed in coordinates of the camera-centered coordinate system. '</p>\n\n<p>In this competition, the origin of the world coordinate is camera. <a href=\"/hustkevin1037\">@hustkevin1037</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699376,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/20/2019 11:15:41",
          "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> \n1 ) one thing i noticed while labeling after now m gaining more info, is that we use x,y 2d points just to assign these points their corresponding world coordinates and it is actually those we are doing regression for .Please correct me if wrong.</p>\n\n<p>2) How will these go during the inference. is overall process is like you are teaching model through the points on the 2d plane that these points position of car looking object is actually these coordinates of same car in real world??</p>\n\n<p><a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699516,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/20/2019 14:35:09",
          "content": "<p>we regress the values(x,y,z,yaw,pitch,roll whatever you want) in the car's postion.\nI think this picture is very clear. \n<img src=\"https://www.kaggleusercontent.com/kf/25098336/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..WjysAO3-6H-Z2s64qukatQ.E81-tzdiF0SR4VR3aFrzxMG-lDRTiVpw7U-WGxLHPTmOLj2JSP6Zn6cSmbF2fFO48Tk3vB6-p5n1opkydcQrl_bz3Z1URDEpcWnHtM-zBWWZ0r68nL8qKN2lfQvKZkjfbIzF8C77JEu6zgB5R0OFUyB3UzmvXSJafcORWYAGvBymsJznoK3QFsZvNRY8HpmJ.jXABv1dwhWPRTZyt98ziQA/__results___files/__results___23_0.png\" alt=\"1\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699615,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/20/2019 16:32:42",
          "content": "<p>sorry about it.. i did wanted to ..\nyour picture is not loaded</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699840,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/21/2019 02:34:26",
          "content": "<p>It just the picture posted in my notebook. I think it's very clear. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F4afa7e282fbb6a913db336539bfd5a85%2F2019-12-21_103355.png?generation=1576895663407777&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699856,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/21/2019 03:13:12",
          "content": "<p>Thanku ..\n<code>mask_loss = mask * torch.log(pred_mask + 1e-12) + (1 - mask) * torch.log(1 - pred_mask + 1e-12)</code>\nI use this loss but unable to determine why it ends in nan \nCompared pytorch standard f.bce with logits  .\nDo you see any trouble ? I use mixed fp </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 700445,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/22/2019 03:11:54",
          "content": "<p>Did you use <strong>sigmoid</strong>?\npred_mask = torch.sigmoid(prediction[:, 0])</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 700453,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/22/2019 03:45:42",
          "content": "<p>Yes . I figured out the reason for nan ,it was because to  accommodate higher bs I used mixed fp  . ,1e-12 was rounding to zero it was reason loss ending in nans ,I changed to 1e-7\nAnd dint dont get nan any more .</p>\n\n<p>2  Second issue is why results appear different from std fct  I dont know \n.\n3. M now on lb :) one thing is sure all public kernel based on hiccups kernel wont go beyond .08 like that .Loss is extremely fluctuating so it may not  let model generalize well reason I think of it is scales of regression and confidence are different . Few different things to be tried up.Would u like to team up ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 700478,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/22/2019 04:36:49",
          "content": "<p>sorry, I've teamed up with others.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 687345,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "12/04/2019 09:18:43",
      "content": "<p>Very Informative\nThanks <a href=\"/diegojohnson\">@diegojohnson</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 687359,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/04/2019 09:35:58",
          "content": "<p>My pleasure😏 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 689381,
      "author_name": "ericfowler",
      "author_url": "",
      "post_date": "12/06/2019 21:31:47",
      "content": "<p>What are R and T in the last equation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 689392,
          "author_name": "mxbonn",
          "author_url": "",
          "post_date": "12/06/2019 22:00:16",
          "content": "<p>Rotation matrix and Translation vector.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 689519,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/07/2019 02:08:45",
          "content": "<p>Yeah👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 702814,
          "author_name": "prayagmadhu",
          "author_url": "",
          "post_date": "12/25/2019 07:54:10",
          "content": "<p>X,Y coordinates we need to predict are the image coordinates, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 702818,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/25/2019 08:01:14",
          "content": "<p>World coordinate <a href=\"/prayagmadhu\">@prayagmadhu</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 691468,
      "author_name": "markcrass",
      "author_url": "",
      "post_date": "12/10/2019 07:10:23",
      "content": "<p>Thank You!</p>",
      "votes": null,
      "replies": [
        {
          "id": 691536,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/10/2019 08:38:16",
          "content": "<p>my pleasure</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 696215,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "12/16/2019 09:41:24",
      "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> this gives me some more better picture..\n1) What is difference between Image coordinates and pixel coordinates . I suppose one of them could be image screen for  camera\n2) For model training we use u,v ? just want to understand rationale behind is it because of mask ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 696229,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/16/2019 10:15:57",
          "content": "<p>Actually, the article has told you answer.\nAn image consists of a lot of pixels, and pixel coordinates can tell you the location of pixel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 696373,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/16/2019 14:28:31",
          "content": "<p>Thanks..i have taken the pos from your notebook\n1)How are we building the labels for it ,below is 128 is max limit for no of cars that are of interest ?\n2) if there are more than one prediction strings for an image ,how is that handled, coords is list of dicts\n<code>[{'id': 16, 'yaw': 0.254839, 'pitch': -2.57534, 'roll': -3.10256, 'x': 7.96539, 'y': 3.20066, 'z': 11.0225}, {'id': 56, 'yaw': 0.181647, 'pitch': -1.46947, 'roll': -3.12159, 'x': 9.60332, 'y': 4.66632, 'z': 19.339}, {'id': 70, 'yaw': 0.163072, 'pitch': -1.56865, 'roll': -3.11754, 'x': 10.39, 'y': 11.2219, 'z': 59.7825}, {'id': 70, 'yaw': 0.141942, 'pitch': -3.1395, 'roll': 3.11969, 'x': -9.59236, 'y': 5.13662, 'z': 24.7337}, {'id': 46, 'yaw': 0.163068, 'pitch': -2.08578, 'roll': -3.11754, 'x': 9.83335, 'y': 13.2689, 'z': 72.9323}]</code>\nwhat would be the labels for id 70 image\n3) purpose of Heatmap\n<code>\ndef pose(s, u, v):\n    regr = np.zeros([128, 128, 6], dtype='float32')\n    coords = str_to_coords(s)\n    for p_x, p_y, regr_dict in zip(u, v, coords):\n        if p_x &amp;gt;= 0 and p_x &amp;lt; 128 and p_y &amp;gt;= 0 and p_y &amp;lt; 128:\n            regr_dict.pop('id')\n            regr[floor(p_y), floor(p_x)] = [regr_dict[n] for n in regr_dict] <br>\n    return regr\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 696415,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/16/2019 15:25:08",
          "content": "<p>the regression label's shape is 128 * 128 * 6\npose function is for putting (yaw, pith, roll, x, y, z) in right position of 128 * 128\nI suggest you look at other runable notebook, acutally there's something wrong in my notebook😅 , and I haven't found it.\nand origin paper is the best place you learn this model, because all public notebook is from there. <a href=\"/jaideepvalani\">@jaideepvalani</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 702819,
      "author_name": "prayagmadhu",
      "author_url": "",
      "post_date": "12/25/2019 08:02:14",
      "content": "<p>is the coordinates given train.csv camera coordinates ?\nSo, we are given camera coordinates of cars in images, and task is to predict yaw, pitch roll and xyz image coordinates of unmasked cars, is this right ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 702826,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/25/2019 08:19:03",
          "content": "<p>task is also to predict camera coordinates of unmasked cars.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 704377,
      "author_name": "polishch",
      "author_url": "",
      "post_date": "12/27/2019 12:02:57",
      "content": "<p>In the last equation all intrinsic camera parameters are given except Zc, camera Z coordinate.\nMay I ask stupid question - how to get it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 704392,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/27/2019 12:28:31",
          "content": "<p>You will get Xc, Yc, Zc together. You can look at public notebook to get more compution detail.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 704399,
          "author_name": "polishch",
          "author_url": "",
          "post_date": "12/27/2019 12:37:58",
          "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> , \nthanks for a guess!</p>\n\n<p>Which notebook you mean? Could you please drop a link?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 704404,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/27/2019 12:48:20",
          "content": "<p>Anyone is right, the computation is same.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "686785": "![1](http://image71.360doc.com/DownloadImg/2014/04/1014/40652696_5.jpg)\n\n**There are acutally four coordinate the world coordinate (Xw, Yw, Zw)、camera coordinate (Xc, Yc, Zc)、image coordinates (x, y) and pixel coordinates (u, v).**\n\n**In this competition the world coordinate (Xw, Yw, Zw) is same with camera coordinate (Xc, Yc, Zc), the camera is origin.**\n\n**Our goal is transforming the world coordinate (Xw, Yw, Zw) to pixel coordinates (u, v).**\n\n## 1 - \n**As I said before, (Xc, Yc, Zc) = (Xw, Yw, Zw), so first step is transforming camera coordinate (Xc, Yc, Zc)(meter) to image coordinates (x, y)(millimeter) and f is focal length.**\n\n![4](http://www.pianshen.com/images/647/fd1c1f4cfa17e5bb5218d80a425c3a57.png)\n\n## 2 - \n**Then transforming image coordinates (x, y)(millimeter) to pixel coordinates (u, v)(pixel), and dx, dy are the actual size of pixels on the sensitive chip**\n\n![3](http://www.pianshen.com/images/592/6a06910d057cbd436d710b9dd978bf68.png)\n\n## 3 -\n**Actually, we use matrix of Camera Intrinsic Parameters to compute. fx = f / dx and fy = f / dy.**\n**In this competition, the Camera Intrinsic is fx = 2304.5479; fy = 2305.8757; u0 = 1686.2379; v0 = 1354.9849;**\n\n![4](http://www.pianshen.com/images/654/e166052ae19a43a7a3e628f1722373d6.png)\n\n## Please give me some upvotes, if you think it's useful, it's a support for my work, thanks😎 \n\nThese are my other topics.\n[The coordinate system of this competition](https://www.kaggle.com/c/pku-autonomous-driving/discussion/123385)\n[The algorithm that baidu apollo chooses!](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120076)\n[Algorithm Selection for Beginner!](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120015)\n[Algorithm Selection in 6D Pose Estimation](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120381)\n[How the Competition Organizer get train.csv !](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120443)\n[Source of 6D Pose Estimation](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120582)\n[A way to improve the accuracy of model](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653)\n[Object Detection in 20 Years](https://www.kaggle.com/c/pku-autonomous-driving/discussion/120710)\n\n\nMy notebook:\n[Centernet - Objects as Points](https://www.kaggle.com/diegojohnson/centernet-objects-as-points)\n[A Way to Regress Translation and Rotation](https://www.kaggle.com/diegojohnson/a-way-to-regress-translation-and-rotation?scriptVersionId=24677131)\n[Best Algorithm so far - Google's new paper in Nov](https://www.kaggle.com/diegojohnson/best-algorithm-so-far-google-s-new-paper-in-nov)\n[A clear view of car pose](https://www.kaggle.com/diegojohnson/a-clear-view-of-car-pose)\n[Dataset with preprocess](https://www.kaggle.com/c/pku-autonomous-driving/discussion/124480)",
    "686826": "I see many notebooks use a function to complete 3d-to-2d, but I'd like to know the principle inside.\n\nso I google it and write this topic, hope you like it.😎",
    "687345": "Very Informative\nThanks @diegojohnson",
    "687359": "My pleasure😏",
    "689381": "What are R and T in the last equation?",
    "689392": "Rotation matrix and Translation vector.",
    "689519": "Yeah👍",
    "691468": "Thank You!",
    "691536": "my pleasure",
    "695574": "hi, are you sure `In this competition the world coordinate (Xw, Yw, Zw) is same with camera coordinate (Xc, Yc, Zc), the camera is origin.`，if world coordinate is same with camera coordinate ,which means we don't need matmul the matrix [R, T; 0, 1] when convert to pixel coordinates?",
    "695676": "'Extrinsic parameters is the position of the origin of the world coordinate system expressed in coordinates of the camera-centered coordinate system. '\n\nIn this competition, the origin of the world coordinate is camera. @hustkevin1037",
    "696215": "diegojohnson this gives me some more better picture..\n1) What is difference between Image coordinates and pixel coordinates . I suppose one of them could be image screen for  camera\n2) For model training we use u,v ? just want to understand rationale behind is it because of mask ?",
    "696229": "Actually, the article has told you answer.\nAn image consists of a lot of pixels, and pixel coordinates can tell you the location of pixel.",
    "696373": "Thanks..i have taken the pos from your notebook\n1)How are we building the labels for it ,below is 128 is max limit for no of cars that are of interest ?\n2) if there are more than one prediction strings for an image ,how is that handled, coords is list of dicts\n`[{'id': 16, 'yaw': 0.254839, 'pitch': -2.57534, 'roll': -3.10256, 'x': 7.96539, 'y': 3.20066, 'z': 11.0225}, {'id': 56, 'yaw': 0.181647, 'pitch': -1.46947, 'roll': -3.12159, 'x': 9.60332, 'y': 4.66632, 'z': 19.339}, {'id': 70, 'yaw': 0.163072, 'pitch': -1.56865, 'roll': -3.11754, 'x': 10.39, 'y': 11.2219, 'z': 59.7825}, {'id': 70, 'yaw': 0.141942, 'pitch': -3.1395, 'roll': 3.11969, 'x': -9.59236, 'y': 5.13662, 'z': 24.7337}, {'id': 46, 'yaw': 0.163068, 'pitch': -2.08578, 'roll': -3.11754, 'x': 9.83335, 'y': 13.2689, 'z': 72.9323}]`\nwhat would be the labels for id 70 image\n3) purpose of Heatmap\n```\ndef pose(s, u, v):\n    regr = np.zeros([128, 128, 6], dtype='float32')\n    coords = str_to_coords(s)\n    for p_x, p_y, regr_dict in zip(u, v, coords):\n        if p_x &gt;= 0 and p_x &lt; 128 and p_y &gt;= 0 and p_y &lt; 128:\n            regr_dict.pop('id')\n            regr[floor(p_y), floor(p_x)] = [regr_dict[n] for n in regr_dict]    \n    return regr\n```",
    "696415": "the regression label's shape is 128 * 128 * 6\npose function is for putting (yaw, pith, roll, x, y, z) in right position of 128 * 128\nI suggest you look at other runable notebook, acutally there's something wrong in my notebook😅 , and I haven't found it.\nand origin paper is the best place you learn this model, because all public notebook is from there. @jaideepvalani",
    "699376": "diegojohnson \n1 ) one thing i noticed while labeling after now m gaining more info, is that we use x,y 2d points just to assign these points their corresponding world coordinates and it is actually those we are doing regression for .Please correct me if wrong.\n\n2) How will these go during the inference. is overall process is like you are teaching model through the points on the 2d plane that these points position of car looking object is actually these coordinates of same car in real world??\n\nhttps://www.kaggle.com/hocop1/centernet-baseline",
    "699516": "we regress the values(x,y,z,yaw,pitch,roll whatever you want) in the car's postion.\nI think this picture is very clear. \n![1](https://www.kaggleusercontent.com/kf/25098336/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..WjysAO3-6H-Z2s64qukatQ.E81-tzdiF0SR4VR3aFrzxMG-lDRTiVpw7U-WGxLHPTmOLj2JSP6Zn6cSmbF2fFO48Tk3vB6-p5n1opkydcQrl_bz3Z1URDEpcWnHtM-zBWWZ0r68nL8qKN2lfQvKZkjfbIzF8C77JEu6zgB5R0OFUyB3UzmvXSJafcORWYAGvBymsJznoK3QFsZvNRY8HpmJ.jXABv1dwhWPRTZyt98ziQA/__results___files/__results___23_0.png)",
    "699615": "sorry about it.. i did wanted to ..\nyour picture is not loaded",
    "699840": "It just the picture posted in my notebook. I think it's very clear. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3876174%2F4afa7e282fbb6a913db336539bfd5a85%2F2019-12-21_103355.png?generation=1576895663407777&amp;alt=media)",
    "699856": "Thanku ..\n`mask_loss = mask * torch.log(pred_mask + 1e-12) + (1 - mask) * torch.log(1 - pred_mask + 1e-12)`\nI use this loss but unable to determine why it ends in nan \nCompared pytorch standard f.bce with logits  .\nDo you see any trouble ? I use mixed fp",
    "700445": "Did you use **sigmoid**?\npred_mask = torch.sigmoid(prediction[:, 0])",
    "700453": "Yes . I figured out the reason for nan ,it was because to  accommodate higher bs I used mixed fp  . ,1e-12 was rounding to zero it was reason loss ending in nans ,I changed to 1e-7\nAnd dint dont get nan any more .\n \n2  Second issue is why results appear different from std fct  I dont know \n.\n3. M now on lb :) one thing is sure all public kernel based on hiccups kernel wont go beyond .08 like that .Loss is extremely fluctuating so it may not  let model generalize well reason I think of it is scales of regression and confidence are different . Few different things to be tried up.Would u like to team up ?",
    "700478": "sorry, I've teamed up with others.",
    "702814": "X,Y coordinates we need to predict are the image coordinates, right?",
    "702818": "World coordinate @prayagmadhu",
    "702819": "is the coordinates given train.csv camera coordinates ?\nSo, we are given camera coordinates of cars in images, and task is to predict yaw, pitch roll and xyz image coordinates of unmasked cars, is this right ?",
    "702826": "task is also to predict camera coordinates of unmasked cars.",
    "704377": "In the last equation all intrinsic camera parameters are given except Zc, camera Z coordinate.\nMay I ask stupid question - how to get it?",
    "704392": "You will get Xc, Yc, Zc together. You can look at public notebook to get more compution detail.",
    "704399": "diegojohnson , \nthanks for a guess!\n\nWhich notebook you mean? Could you please drop a link?",
    "704404": "Anyone is right, the computation is same."
  },
  "source": "meta"
}