{
  "id": 127034,
  "title": "7th place solution",
  "url": "/competitions/pku-autonomous-driving/discussion/127034",
  "author_name": "phalanx",
  "post_date": "2020-01-22T00:10:49.192000",
  "votes": 39,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Congrats everyone with excellent result!\nI summarize and write down the part of my solution and our post process.\nFor other part:\n<a href=\"https://www.kaggle.com/bamps53\">camaro</a> part: <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127056\">(part of) 7th place solution with code</a>\n<a href=\"https://www.kaggle.com/hesene\">Jhui He</a> and <a href=\"https://www.kaggle.com/lanjunyelan\">yelan</a> part: <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034#726362\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034#726362</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2F6062aa1f543d25daab45fa475adb223d%2Fdrive_pipeline%20(6\" alt=\"\">.jpg?generation=1579652111438826&amp;alt=media)</p>\n\n<h2>Model1: detection and pose estimation</h2>\n\n<h3>common setting</h3>\n\n<p><strong>detection model</strong>\n- Mask RCNN(mask head removed)\n- backbone: resnext101-32x4d\n- lvis pretrained</p>\n\n<p><strong>pose estimation model</strong>\n- HRNet-w18c, efficientnetb0/b3\n- imagenet pretrained</p>\n\n<p><strong>Loss</strong>\n- classification: BCE\n- detectoin: Focal Loss\n- pose regression: L1 Loss</p>\n\n<p><strong>detection: Optimizer and scheduler</strong>\n- optimizer: SGD(lr=0.01, momentum=0.9, weight_decay=1e-4, nesterov=True)\n- scheduler: CosineAnnealingWarmRestarts</p>\n\n<p><strong>pose estimation: Optimizer and scheduler</strong>\n- optimizer: Adam(lr=0.0001)\n- scheduler: None</p>\n\n<p><strong>Augmentation</strong>\n- detection: horizontal flip\n- pose estimation: horizontal flip, shift, rotate, random blightness/contrast</p>\n\n<h3>1. pretrain on the boxy-vehicle-dataset</h3>\n\n<p>At first, I train my model on the <a href=\"https://boxy-dataset.com/boxy/\">boxy-vehicle-dataset</a>.\nThis dataset include axis aligned bounding box and 3d cuboids, but I use only 2d bbox.<br>\n<strong>training setting</strong>:\n- image resolution: 1232x1028\n- epochs: 10\n- batch_size: 4</p>\n\n<h3>2. finetune on competition dataset</h3>\n\n<p><strong>Model</strong>\n- add depth head on top of model</p>\n\n<p><strong>preprocess</strong>\n- split train vs val = 9 vs 1\n- create 3d bbox using label, then create axis aligned bbox\n- depth -&gt; 1 / sigmoid(depth) - 1</p>\n\n<p><strong>training setting</strong>\n- image resolution: 800 x 2800, 1400x3300\n- depth loss: L1 Loss\n- epochs: 50\n- batch_size: 4</p>\n\n<h3>3. pose estimation(yaw_sin, yaw_cos, pitch)</h3>\n\n<p><strong>preprocess</strong>\n- crop image by bbox and resize</p>\n\n<p><strong>training setting</strong>\n- image resolution: 320x480\n- epochs: 30\n- batch_size: 128</p>\n\n<p>public LB/private LB\n- 800x2800, single fold: 0.119/0.106\n- 1400x3300, single fold: 0.106/0.111</p>\n\n<h1>Model2: centernet</h1>\n\n<p><strong>model</strong>\n- <a href=\"https://github.com/xingyizhou/CenterNet\">pytorch dla centernet</a>.\n- regression of yaw_sin, yaw_cos, pitch, depth, 2d bbox size(w, h), 3d bbox size(w, h, l)\n- classification of object centerness(heat map)</p>\n\n<p><strong>Loss</strong>\n- regression: L1 Loss\n- classification: Focal Loss</p>\n\n<p><strong>optimizer and scheduler</strong>\n- optimizer: Adam(lr=5e-4)\n- scheduler: CosineAnnealingLR(lr=5e-5)</p>\n\n<p><strong>Augmentation</strong>\n- pose estimation: horizontal flip, shift, random blightness/contrast</p>\n\n<p><strong>preprocess</strong>\n- depth -&gt; 1 / sigmoid(depth) - 1</p>\n\n<p><strong>training setting</strong>\n- split train vs val = 8 vs 2\n- epochs: 30\n- batch_size: 12</p>\n\n<p>public LB/ private LB\n- single fold: 0.100/0.096</p>\n\n<h1>Ensemble: Linear assignment</h1>\n\n<p>We use different model(faster rcnn, centernet), so it is difficult to ensemble predicitons.\nSo we decided to ensemble nearset points between predictions.\nWe use hungalian algorithm for linear assignment.\nPlease refere to below code.\n```\nfrom scipy.optimize import linear_sum_assignment\ndistance_th = 30\nyaw_th = 10</p>\n\n<p>sub1 = pd.read_csv('sub1.csv')\ny1 = sub1['PredictionString'].str.split(' ').values\nX1 = sub1['ImageId'].values</p>\n\n<p>sub2 = pd.read_csv('sub2.csv')\ny2 = sub2['PredictionString'].str.split(' ').values\nX2 = sub2['ImageId'].values\n​\nfor idx in tqdm(range(len(sub1)), position=0):\n    if str(np.nan) != str(y1[idx]) and str(np.nan) != str(y2[idx]):\n        label1 = np.array(y1[idx]).reshape(-1, 7).astype(float)\n        label2 = np.array(y2[idx]).reshape(-1, 7).astype(float)\n        center_points1 = get_imgcoords(label1) # [N, 3], (img_x, img_y, img_z)\n        center_points2 = get_imgcoords(label2)</p>\n\n<pre><code>    cost_matrix = np.zeros([len(center_points1), len(center_points1)])\n    for idx1, i in enumerate(center_points1):\n        for idx2, j in enumerate(center_points2):\n            cost_matrix[idx1, idx2] = np.linalg.norm(i - j)\n\n    match1, match2 = linear_sum_assignment(cost_matrix)\n    for i, j in zip(match1, match2):\n        if cost_matrix[i, j] &amp;lt; distance_th:\n            if np.abs(label1[i][1] - label2[j][1]) &amp;lt; yaw_th:\n                label1[j] = (label1[i] + label2[j]) / 2\n    label1[idx] = np.concatenate(tmp1).astype(str)\n</code></pre>\n\n<h1>use label1 for submission</h1>\n\n<p>```</p>",
  "messages": [
    {
      "id": 725256,
      "postDate": "2020-01-22T00:10:49.193Z",
      "content": "<p>Congrats everyone with excellent result!\nI summarize and write down the part of my solution and our post process.\nFor other part:\n<a href=\"https://www.kaggle.com/bamps53\">camaro</a> part: <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127056\">(part of) 7th place solution with code</a>\n<a href=\"https://www.kaggle.com/hesene\">Jhui He</a> and <a href=\"https://www.kaggle.com/lanjunyelan\">yelan</a> part: <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034#726362\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034#726362</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2F6062aa1f543d25daab45fa475adb223d%2Fdrive_pipeline%20(6\" alt=\"\">.jpg?generation=1579652111438826&amp;alt=media)</p>\n\n<h2>Model1: detection and pose estimation</h2>\n\n<h3>common setting</h3>\n\n<p><strong>detection model</strong>\n- Mask RCNN(mask head removed)\n- backbone: resnext101-32x4d\n- lvis pretrained</p>\n\n<p><strong>pose estimation model</strong>\n- HRNet-w18c, efficientnetb0/b3\n- imagenet pretrained</p>\n\n<p><strong>Loss</strong>\n- classification: BCE\n- detectoin: Focal Loss\n- pose regression: L1 Loss</p>\n\n<p><strong>detection: Optimizer and scheduler</strong>\n- optimizer: SGD(lr=0.01, momentum=0.9, weight_decay=1e-4, nesterov=True)\n- scheduler: CosineAnnealingWarmRestarts</p>\n\n<p><strong>pose estimation: Optimizer and scheduler</strong>\n- optimizer: Adam(lr=0.0001)\n- scheduler: None</p>\n\n<p><strong>Augmentation</strong>\n- detection: horizontal flip\n- pose estimation: horizontal flip, shift, rotate, random blightness/contrast</p>\n\n<h3>1. pretrain on the boxy-vehicle-dataset</h3>\n\n<p>At first, I train my model on the <a href=\"https://boxy-dataset.com/boxy/\">boxy-vehicle-dataset</a>.\nThis dataset include axis aligned bounding box and 3d cuboids, but I use only 2d bbox.<br>\n<strong>training setting</strong>:\n- image resolution: 1232x1028\n- epochs: 10\n- batch_size: 4</p>\n\n<h3>2. finetune on competition dataset</h3>\n\n<p><strong>Model</strong>\n- add depth head on top of model</p>\n\n<p><strong>preprocess</strong>\n- split train vs val = 9 vs 1\n- create 3d bbox using label, then create axis aligned bbox\n- depth -&gt; 1 / sigmoid(depth) - 1</p>\n\n<p><strong>training setting</strong>\n- image resolution: 800 x 2800, 1400x3300\n- depth loss: L1 Loss\n- epochs: 50\n- batch_size: 4</p>\n\n<h3>3. pose estimation(yaw_sin, yaw_cos, pitch)</h3>\n\n<p><strong>preprocess</strong>\n- crop image by bbox and resize</p>\n\n<p><strong>training setting</strong>\n- image resolution: 320x480\n- epochs: 30\n- batch_size: 128</p>\n\n<p>public LB/private LB\n- 800x2800, single fold: 0.119/0.106\n- 1400x3300, single fold: 0.106/0.111</p>\n\n<h1>Model2: centernet</h1>\n\n<p><strong>model</strong>\n- <a href=\"https://github.com/xingyizhou/CenterNet\">pytorch dla centernet</a>.\n- regression of yaw_sin, yaw_cos, pitch, depth, 2d bbox size(w, h), 3d bbox size(w, h, l)\n- classification of object centerness(heat map)</p>\n\n<p><strong>Loss</strong>\n- regression: L1 Loss\n- classification: Focal Loss</p>\n\n<p><strong>optimizer and scheduler</strong>\n- optimizer: Adam(lr=5e-4)\n- scheduler: CosineAnnealingLR(lr=5e-5)</p>\n\n<p><strong>Augmentation</strong>\n- pose estimation: horizontal flip, shift, random blightness/contrast</p>\n\n<p><strong>preprocess</strong>\n- depth -&gt; 1 / sigmoid(depth) - 1</p>\n\n<p><strong>training setting</strong>\n- split train vs val = 8 vs 2\n- epochs: 30\n- batch_size: 12</p>\n\n<p>public LB/ private LB\n- single fold: 0.100/0.096</p>\n\n<h1>Ensemble: Linear assignment</h1>\n\n<p>We use different model(faster rcnn, centernet), so it is difficult to ensemble predicitons.\nSo we decided to ensemble nearset points between predictions.\nWe use hungalian algorithm for linear assignment.\nPlease refere to below code.\n```\nfrom scipy.optimize import linear_sum_assignment\ndistance_th = 30\nyaw_th = 10</p>\n\n<p>sub1 = pd.read_csv('sub1.csv')\ny1 = sub1['PredictionString'].str.split(' ').values\nX1 = sub1['ImageId'].values</p>\n\n<p>sub2 = pd.read_csv('sub2.csv')\ny2 = sub2['PredictionString'].str.split(' ').values\nX2 = sub2['ImageId'].values\n​\nfor idx in tqdm(range(len(sub1)), position=0):\n    if str(np.nan) != str(y1[idx]) and str(np.nan) != str(y2[idx]):\n        label1 = np.array(y1[idx]).reshape(-1, 7).astype(float)\n        label2 = np.array(y2[idx]).reshape(-1, 7).astype(float)\n        center_points1 = get_imgcoords(label1) # [N, 3], (img_x, img_y, img_z)\n        center_points2 = get_imgcoords(label2)</p>\n\n<pre><code>    cost_matrix = np.zeros([len(center_points1), len(center_points1)])\n    for idx1, i in enumerate(center_points1):\n        for idx2, j in enumerate(center_points2):\n            cost_matrix[idx1, idx2] = np.linalg.norm(i - j)\n\n    match1, match2 = linear_sum_assignment(cost_matrix)\n    for i, j in zip(match1, match2):\n        if cost_matrix[i, j] &amp;lt; distance_th:\n            if np.abs(label1[i][1] - label2[j][1]) &amp;lt; yaw_th:\n                label1[j] = (label1[i] + label2[j]) / 2\n    label1[idx] = np.concatenate(tmp1).astype(str)\n</code></pre>\n\n<h1>use label1 for submission</h1>\n\n<p>```</p>",
      "rawMarkdown": "Congrats everyone with excellent result!\nI summarize and write down the part of my solution and our post process.\nFor other part:\n[camaro](https://www.kaggle.com/bamps53) part: [(part of) 7th place solution with code](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127056)\n[Jhui He](https://www.kaggle.com/hesene) and [yelan](https://www.kaggle.com/lanjunyelan) part: [https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034#726362](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034#726362)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2F6062aa1f543d25daab45fa475adb223d%2Fdrive_pipeline%20(6).jpg?generation=1579652111438826&amp;alt=media)\n\n## Model1: detection and pose estimation\n### common setting\n<strong>detection model</strong>\n- Mask RCNN(mask head removed)\n- backbone: resnext101-32x4d\n- lvis pretrained\n\n<strong>pose estimation model</strong>\n- HRNet-w18c, efficientnetb0/b3\n- imagenet pretrained\n\n<strong>Loss</strong>\n- classification: BCE\n- detectoin: Focal Loss\n- pose regression: L1 Loss\n\n<strong>detection: Optimizer and scheduler</strong>\n- optimizer: SGD(lr=0.01, momentum=0.9, weight_decay=1e-4, nesterov=True)\n- scheduler: CosineAnnealingWarmRestarts\n\n<strong>pose estimation: Optimizer and scheduler</strong>\n- optimizer: Adam(lr=0.0001)\n- scheduler: None\n\n<strong>Augmentation</strong>\n- detection: horizontal flip\n- pose estimation: horizontal flip, shift, rotate, random blightness/contrast\n\n### 1. pretrain on the boxy-vehicle-dataset\nAt first, I train my model on the [boxy-vehicle-dataset](https://boxy-dataset.com/boxy/).\nThis dataset include axis aligned bounding box and 3d cuboids, but I use only 2d bbox.<br>\n<strong>training setting</strong>:\n- image resolution: 1232x1028\n- epochs: 10\n- batch_size: 4\n### 2. finetune on competition dataset\n<strong>Model</strong>\n- add depth head on top of model\n\n<strong>preprocess</strong>\n- split train vs val = 9 vs 1\n- create 3d bbox using label, then create axis aligned bbox\n- depth -&gt; 1 / sigmoid(depth) - 1\n\n<strong>training setting</strong>\n- image resolution: 800 x 2800, 1400x3300\n- depth loss: L1 Loss\n- epochs: 50\n- batch_size: 4\n\n### 3. pose estimation(yaw_sin, yaw_cos, pitch)\n<strong>preprocess</strong>\n- crop image by bbox and resize\n\n<strong>training setting</strong>\n- image resolution: 320x480\n- epochs: 30\n- batch_size: 128\n\npublic LB/private LB\n- 800x2800, single fold: 0.119/0.106\n- 1400x3300, single fold: 0.106/0.111\n\n# Model2: centernet\n<strong>model</strong>\n- [pytorch dla centernet](https://github.com/xingyizhou/CenterNet).\n- regression of yaw_sin, yaw_cos, pitch, depth, 2d bbox size(w, h), 3d bbox size(w, h, l)\n- classification of object centerness(heat map)\n\n<strong>Loss</strong>\n- regression: L1 Loss\n- classification: Focal Loss\n\n<strong>optimizer and scheduler</strong>\n- optimizer: Adam(lr=5e-4)\n- scheduler: CosineAnnealingLR(lr=5e-5)\n\n<strong>Augmentation</strong>\n- pose estimation: horizontal flip, shift, random blightness/contrast\n\n<strong>preprocess</strong>\n- depth -&gt; 1 / sigmoid(depth) - 1\n\n<strong>training setting</strong>\n- split train vs val = 8 vs 2\n- epochs: 30\n- batch_size: 12\n\npublic LB/ private LB\n- single fold: 0.100/0.096\n\n# Ensemble: Linear assignment\nWe use different model(faster rcnn, centernet), so it is difficult to ensemble predicitons.\nSo we decided to ensemble nearset points between predictions.\nWe use hungalian algorithm for linear assignment.\nPlease refere to below code.\n```\nfrom scipy.optimize import linear_sum_assignment\ndistance_th = 30\nyaw_th = 10\n\nsub1 = pd.read_csv('sub1.csv')\ny1 = sub1['PredictionString'].str.split(' ').values\nX1 = sub1['ImageId'].values\n\nsub2 = pd.read_csv('sub2.csv')\ny2 = sub2['PredictionString'].str.split(' ').values\nX2 = sub2['ImageId'].values\n​\nfor idx in tqdm(range(len(sub1)), position=0):\n    if str(np.nan) != str(y1[idx]) and str(np.nan) != str(y2[idx]):\n        label1 = np.array(y1[idx]).reshape(-1, 7).astype(float)\n        label2 = np.array(y2[idx]).reshape(-1, 7).astype(float)\n        center_points1 = get_imgcoords(label1) # [N, 3], (img_x, img_y, img_z)\n        center_points2 = get_imgcoords(label2)\n        \n        cost_matrix = np.zeros([len(center_points1), len(center_points1)])\n        for idx1, i in enumerate(center_points1):\n            for idx2, j in enumerate(center_points2):\n                cost_matrix[idx1, idx2] = np.linalg.norm(i - j)\n        \n        match1, match2 = linear_sum_assignment(cost_matrix)\n        for i, j in zip(match1, match2):\n            if cost_matrix[i, j] &lt; distance_th:\n                if np.abs(label1[i][1] - label2[j][1]) &lt; yaw_th:\n                    label1[j] = (label1[i] + label2[j]) / 2\n        label1[idx] = np.concatenate(tmp1).astype(str)\n\n# use label1 for submission\n```",
      "votes": 39
    },
    {
      "id": 726362,
      "postDate": "2020-01-23T01:27:09.163Z",
      "content": "<p><strong>Part of Centernet solution:</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2241982%2Ffc1e876999b387f06bc7edb3d0e92e04%2Fpipline.png?generation=1579742255199041&amp;alt=media\" alt=\"\">\n- Scale-&gt;4096*1024 and Filter ignored targets by test mask: Public: 0.101, Private:0.106\n- Pseudo-label and Filter ignored targets by test mask, Public: 0.106, Private:0.110\n<strong><em>Aug</em></strong>\n<code>\ntrain_transform = albumentations.Compose([\n    albumentations.JpegCompression(quality_lower=99, quality_upper=100,p=0.5),\n    albumentations.OneOf([\n        albumentations.RandomGamma(gamma_limit=(60, 120), p=0.5),\n        albumentations.RandomBrightnessContrast(brightness_limit=0.2, contrast_limit=0.2, p=0.5),\n    ]),\n    albumentations.GaussNoise(var_limit=(10, 30), p=0.5),\n    albumentations.Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225), max_pixel_value=255.0, p=1.0)\n])\n</code>\n<strong><em>Depth Loss</em></strong>\n<code>\npred_depth = (1. / sigmoid_ddd(prediction[:, -1])) - 1\ngt_depth = regr[:, -1]\npos_inds = gt_depth.gt(0).float()\ndepth_loss = (torch.abs(pred_depth - gt_depth) * pos_inds).sum(1).sum(1) / pos_inds.sum(1).sum(1)\ndepth_loss = depth_loss.mean(0)\n</code>\n<strong><em>Learning rate and optimizer</em></strong>\n- Learning rate is 0.0005 in first 15 epoch, Then cycle cosine learning rate trains per 5 epochs, total epoch is 40~60. At the large scale, I use the small scale model to fine-tune with a small learning rate(1e-5)\n- optimizer: Adam</p>\n\n<p><strong>Key point of this solution:</strong>\n1.  Heatmap(umich_gaussian) and Focal loss(same as <a href=\"https://github.com/xingyizhou/CenterNet\">https://github.com/xingyizhou/CenterNet</a>)\n2. Depth loss\n3. Bigger scale</p>\n\n<p>Thanks for my teammate's great work and Thanks hocop1's kernel(<a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a>)</p>",
      "rawMarkdown": "**Part of Centernet solution:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2241982%2Ffc1e876999b387f06bc7edb3d0e92e04%2Fpipline.png?generation=1579742255199041&amp;alt=media)\n- Scale-&gt;4096*1024 and Filter ignored targets by test mask: Public: 0.101, Private:0.106\n- Pseudo-label and Filter ignored targets by test mask, Public: 0.106, Private:0.110\n***Aug***\n```\ntrain_transform = albumentations.Compose([\n    albumentations.JpegCompression(quality_lower=99, quality_upper=100,p=0.5),\n    albumentations.OneOf([\n        albumentations.RandomGamma(gamma_limit=(60, 120), p=0.5),\n        albumentations.RandomBrightnessContrast(brightness_limit=0.2, contrast_limit=0.2, p=0.5),\n    ]),\n    albumentations.GaussNoise(var_limit=(10, 30), p=0.5),\n    albumentations.Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225), max_pixel_value=255.0, p=1.0)\n])\n```\n***Depth Loss***\n```\npred_depth = (1. / sigmoid_ddd(prediction[:, -1])) - 1\ngt_depth = regr[:, -1]\npos_inds = gt_depth.gt(0).float()\ndepth_loss = (torch.abs(pred_depth - gt_depth) * pos_inds).sum(1).sum(1) / pos_inds.sum(1).sum(1)\ndepth_loss = depth_loss.mean(0)\n```\n***Learning rate and optimizer***\n- Learning rate is 0.0005 in first 15 epoch, Then cycle cosine learning rate trains per 5 epochs, total epoch is 40~60. At the large scale, I use the small scale model to fine-tune with a small learning rate(1e-5)\n- optimizer: Adam\n\n**Key point of this solution:**\n1.  Heatmap(umich_gaussian) and Focal loss(same as https://github.com/xingyizhou/CenterNet)\n2. Depth loss\n3. Bigger scale\n\nThanks for my teammate's great work and Thanks hocop1's kernel(https://www.kaggle.com/hocop1/centernet-baseline)",
      "votes": 5
    },
    {
      "id": 728020,
      "postDate": "2020-01-24T10:40:52.320Z",
      "content": "<p>Congrats !! Amazing! Thanks for sharing !</p>",
      "rawMarkdown": "Congrats !! Amazing! Thanks for sharing !",
      "votes": 4
    },
    {
      "id": 727485,
      "postDate": "2020-01-23T18:40:46.833Z",
      "content": "<p>Great solution, thank you so much for sharing!\nMay I ask what do you do if the number of detected cars is not the same for each model?\nI'd say my approach is 1/3 of this one and when I tried to ensemble, many times I had to face that not all the predictions include the same amount of detected cars.\nThanks!</p>",
      "rawMarkdown": "Great solution, thank you so much for sharing!\nMay I ask what do you do if the number of detected cars is not the same for each model?\nI'd say my approach is 1/3 of this one and when I tried to ensemble, many times I had to face that not all the predictions include the same amount of detected cars.\nThanks!",
      "votes": -1,
      "replies": [
        {
          "id": 727742,
          "postDate": "2020-01-24T02:15:48.327Z",
          "content": "<p>Thanks Nanashi!\nWe use hungarian algorithm(cost: euclidian distance) for ensemble, so some predictions is assigned based on cost matrix and the others is ignored.\nIn our pipeline image, some centernet prediction(red points) is not assigned and ignored.</p>",
          "rawMarkdown": "Thanks Nanashi!\nWe use hungarian algorithm(cost: euclidian distance) for ensemble, so some predictions is assigned based on cost matrix and the others is ignored.\nIn our pipeline image, some centernet prediction(red points) is not assigned and ignored.",
          "votes": 1
        }
      ]
    },
    {
      "id": 726397,
      "postDate": "2020-01-23T01:53:42.073Z",
      "content": "<p>Thanks for sharing about the DLA solution! By the way, do you use Adam as the optimizer? And how do you set the learning rate? It seems I used a similar network but didn't know why the net couldn't converge. Thanks for your help!</p>",
      "rawMarkdown": "Thanks for sharing about the DLA solution! By the way, do you use Adam as the optimizer? And how do you set the learning rate? It seems I used a similar network but didn't know why the net couldn't converge. Thanks for your help!",
      "replies": [
        {
          "id": 726411,
          "postDate": "2020-01-23T02:05:46.257Z",
          "content": "<p>I update it</p>",
          "rawMarkdown": "I update it",
          "votes": 1
        },
        {
          "id": 726412,
          "postDate": "2020-01-23T02:06:57.680Z",
          "content": "<p>Many thanks! I'll try this setting! 👍 </p>",
          "rawMarkdown": "Many thanks! I'll try this setting! 👍 "
        }
      ]
    },
    {
      "id": 725300,
      "postDate": "2020-01-22T01:30:25.287Z",
      "content": "<p>Congratulations! Nice game</p>",
      "rawMarkdown": "Congratulations! Nice game"
    },
    {
      "id": 725279,
      "postDate": "2020-01-22T00:52:11.893Z",
      "content": "<p>Congratulations! Thanks for sharing! May I ask the optimizer and the learning rate you use to train centerNet-dla34? I've tried a few times, it never converges. 😔  </p>",
      "rawMarkdown": "Congratulations! Thanks for sharing! May I ask the optimizer and the learning rate you use to train centerNet-dla34? I've tried a few times, it never converges. 😔  "
    },
    {
      "id": 725345,
      "postDate": "2020-01-22T02:29:15.913Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 732537,
      "postDate": "2020-01-29T23:26:24.077Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 726362,
      "author_name": "He",
      "author_url": "",
      "post_date": "2020-01-23T01:27:09.163000",
      "content": "<p><strong>Part of Centernet solution:</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2241982%2Ffc1e876999b387f06bc7edb3d0e92e04%2Fpipline.png?generation=1579742255199041&amp;alt=media\" alt=\"\">\n- Scale-&gt;4096*1024 and Filter ignored targets by test mask: Public: 0.101, Private:0.106\n- Pseudo-label and Filter ignored targets by test mask, Public: 0.106, Private:0.110\n<strong><em>Aug</em></strong>\n<code>\ntrain_transform = albumentations.Compose([\n    albumentations.JpegCompression(quality_lower=99, quality_upper=100,p=0.5),\n    albumentations.OneOf([\n        albumentations.RandomGamma(gamma_limit=(60, 120), p=0.5),\n        albumentations.RandomBrightnessContrast(brightness_limit=0.2, contrast_limit=0.2, p=0.5),\n    ]),\n    albumentations.GaussNoise(var_limit=(10, 30), p=0.5),\n    albumentations.Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225), max_pixel_value=255.0, p=1.0)\n])\n</code>\n<strong><em>Depth Loss</em></strong>\n<code>\npred_depth = (1. / sigmoid_ddd(prediction[:, -1])) - 1\ngt_depth = regr[:, -1]\npos_inds = gt_depth.gt(0).float()\ndepth_loss = (torch.abs(pred_depth - gt_depth) * pos_inds).sum(1).sum(1) / pos_inds.sum(1).sum(1)\ndepth_loss = depth_loss.mean(0)\n</code>\n<strong><em>Learning rate and optimizer</em></strong>\n- Learning rate is 0.0005 in first 15 epoch, Then cycle cosine learning rate trains per 5 epochs, total epoch is 40~60. At the large scale, I use the small scale model to fine-tune with a small learning rate(1e-5)\n- optimizer: Adam</p>\n\n<p><strong>Key point of this solution:</strong>\n1.  Heatmap(umich_gaussian) and Focal loss(same as <a href=\"https://github.com/xingyizhou/CenterNet\">https://github.com/xingyizhou/CenterNet</a>)\n2. Depth loss\n3. Bigger scale</p>\n\n<p>Thanks for my teammate's great work and Thanks hocop1's kernel(<a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a>)</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 728020,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-24T10:40:52.320000",
      "content": "<p>Congrats !! Amazing! Thanks for sharing !</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 727485,
      "author_name": "Nanashi",
      "author_url": "",
      "post_date": "2020-01-23T18:40:46.833000",
      "content": "<p>Great solution, thank you so much for sharing!\nMay I ask what do you do if the number of detected cars is not the same for each model?\nI'd say my approach is 1/3 of this one and when I tried to ensemble, many times I had to face that not all the predictions include the same amount of detected cars.\nThanks!</p>",
      "votes": -1,
      "replies": [
        {
          "id": 727742,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2020-01-24T02:15:48.327000",
          "content": "<p>Thanks Nanashi!\nWe use hungarian algorithm(cost: euclidian distance) for ensemble, so some predictions is assigned based on cost matrix and the others is ignored.\nIn our pipeline image, some centernet prediction(red points) is not assigned and ignored.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 726397,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-01-23T01:53:42.073000",
      "content": "<p>Thanks for sharing about the DLA solution! By the way, do you use Adam as the optimizer? And how do you set the learning rate? It seems I used a similar network but didn't know why the net couldn't converge. Thanks for your help!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 726411,
          "author_name": "He",
          "author_url": "",
          "post_date": "2020-01-23T02:05:46.257000",
          "content": "<p>I update it</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 726412,
          "author_name": "Helen",
          "author_url": "",
          "post_date": "2020-01-23T02:06:57.680000",
          "content": "<p>Many thanks! I'll try this setting! 👍 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 725300,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2020-01-22T01:30:25.287000",
      "content": "<p>Congratulations! Nice game</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 725279,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-01-22T00:52:11.893000",
      "content": "<p>Congratulations! Thanks for sharing! May I ask the optimizer and the learning rate you use to train centerNet-dla34? I've tried a few times, it never converges. 😔  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 725345,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-22T02:29:15.913000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732537,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-29T23:26:24.077000",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "725256": "Congrats everyone with excellent result!\nI summarize and write down the part of my solution and our post process.\nFor other part:\n[camaro](https://www.kaggle.com/bamps53) part: [(part of) 7th place solution with code](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127056)\n[Jhui He](https://www.kaggle.com/hesene) and [yelan](https://www.kaggle.com/lanjunyelan) part: [https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034#726362](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034#726362)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2F6062aa1f543d25daab45fa475adb223d%2Fdrive_pipeline%20(6).jpg?generation=1579652111438826&amp;alt=media)\n\n## Model1: detection and pose estimation\n### common setting\n<strong>detection model</strong>\n- Mask RCNN(mask head removed)\n- backbone: resnext101-32x4d\n- lvis pretrained\n\n<strong>pose estimation model</strong>\n- HRNet-w18c, efficientnetb0/b3\n- imagenet pretrained\n\n<strong>Loss</strong>\n- classification: BCE\n- detectoin: Focal Loss\n- pose regression: L1 Loss\n\n<strong>detection: Optimizer and scheduler</strong>\n- optimizer: SGD(lr=0.01, momentum=0.9, weight_decay=1e-4, nesterov=True)\n- scheduler: CosineAnnealingWarmRestarts\n\n<strong>pose estimation: Optimizer and scheduler</strong>\n- optimizer: Adam(lr=0.0001)\n- scheduler: None\n\n<strong>Augmentation</strong>\n- detection: horizontal flip\n- pose estimation: horizontal flip, shift, rotate, random blightness/contrast\n\n### 1. pretrain on the boxy-vehicle-dataset\nAt first, I train my model on the [boxy-vehicle-dataset](https://boxy-dataset.com/boxy/).\nThis dataset include axis aligned bounding box and 3d cuboids, but I use only 2d bbox.<br>\n<strong>training setting</strong>:\n- image resolution: 1232x1028\n- epochs: 10\n- batch_size: 4\n### 2. finetune on competition dataset\n<strong>Model</strong>\n- add depth head on top of model\n\n<strong>preprocess</strong>\n- split train vs val = 9 vs 1\n- create 3d bbox using label, then create axis aligned bbox\n- depth -&gt; 1 / sigmoid(depth) - 1\n\n<strong>training setting</strong>\n- image resolution: 800 x 2800, 1400x3300\n- depth loss: L1 Loss\n- epochs: 50\n- batch_size: 4\n\n### 3. pose estimation(yaw_sin, yaw_cos, pitch)\n<strong>preprocess</strong>\n- crop image by bbox and resize\n\n<strong>training setting</strong>\n- image resolution: 320x480\n- epochs: 30\n- batch_size: 128\n\npublic LB/private LB\n- 800x2800, single fold: 0.119/0.106\n- 1400x3300, single fold: 0.106/0.111\n\n# Model2: centernet\n<strong>model</strong>\n- [pytorch dla centernet](https://github.com/xingyizhou/CenterNet).\n- regression of yaw_sin, yaw_cos, pitch, depth, 2d bbox size(w, h), 3d bbox size(w, h, l)\n- classification of object centerness(heat map)\n\n<strong>Loss</strong>\n- regression: L1 Loss\n- classification: Focal Loss\n\n<strong>optimizer and scheduler</strong>\n- optimizer: Adam(lr=5e-4)\n- scheduler: CosineAnnealingLR(lr=5e-5)\n\n<strong>Augmentation</strong>\n- pose estimation: horizontal flip, shift, random blightness/contrast\n\n<strong>preprocess</strong>\n- depth -&gt; 1 / sigmoid(depth) - 1\n\n<strong>training setting</strong>\n- split train vs val = 8 vs 2\n- epochs: 30\n- batch_size: 12\n\npublic LB/ private LB\n- single fold: 0.100/0.096\n\n# Ensemble: Linear assignment\nWe use different model(faster rcnn, centernet), so it is difficult to ensemble predicitons.\nSo we decided to ensemble nearset points between predictions.\nWe use hungalian algorithm for linear assignment.\nPlease refere to below code.\n```\nfrom scipy.optimize import linear_sum_assignment\ndistance_th = 30\nyaw_th = 10\n\nsub1 = pd.read_csv('sub1.csv')\ny1 = sub1['PredictionString'].str.split(' ').values\nX1 = sub1['ImageId'].values\n\nsub2 = pd.read_csv('sub2.csv')\ny2 = sub2['PredictionString'].str.split(' ').values\nX2 = sub2['ImageId'].values\n​\nfor idx in tqdm(range(len(sub1)), position=0):\n    if str(np.nan) != str(y1[idx]) and str(np.nan) != str(y2[idx]):\n        label1 = np.array(y1[idx]).reshape(-1, 7).astype(float)\n        label2 = np.array(y2[idx]).reshape(-1, 7).astype(float)\n        center_points1 = get_imgcoords(label1) # [N, 3], (img_x, img_y, img_z)\n        center_points2 = get_imgcoords(label2)\n        \n        cost_matrix = np.zeros([len(center_points1), len(center_points1)])\n        for idx1, i in enumerate(center_points1):\n            for idx2, j in enumerate(center_points2):\n                cost_matrix[idx1, idx2] = np.linalg.norm(i - j)\n        \n        match1, match2 = linear_sum_assignment(cost_matrix)\n        for i, j in zip(match1, match2):\n            if cost_matrix[i, j] &lt; distance_th:\n                if np.abs(label1[i][1] - label2[j][1]) &lt; yaw_th:\n                    label1[j] = (label1[i] + label2[j]) / 2\n        label1[idx] = np.concatenate(tmp1).astype(str)\n\n# use label1 for submission\n```",
    "726362": "**Part of Centernet solution:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2241982%2Ffc1e876999b387f06bc7edb3d0e92e04%2Fpipline.png?generation=1579742255199041&amp;alt=media)\n- Scale-&gt;4096*1024 and Filter ignored targets by test mask: Public: 0.101, Private:0.106\n- Pseudo-label and Filter ignored targets by test mask, Public: 0.106, Private:0.110\n***Aug***\n```\ntrain_transform = albumentations.Compose([\n    albumentations.JpegCompression(quality_lower=99, quality_upper=100,p=0.5),\n    albumentations.OneOf([\n        albumentations.RandomGamma(gamma_limit=(60, 120), p=0.5),\n        albumentations.RandomBrightnessContrast(brightness_limit=0.2, contrast_limit=0.2, p=0.5),\n    ]),\n    albumentations.GaussNoise(var_limit=(10, 30), p=0.5),\n    albumentations.Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225), max_pixel_value=255.0, p=1.0)\n])\n```\n***Depth Loss***\n```\npred_depth = (1. / sigmoid_ddd(prediction[:, -1])) - 1\ngt_depth = regr[:, -1]\npos_inds = gt_depth.gt(0).float()\ndepth_loss = (torch.abs(pred_depth - gt_depth) * pos_inds).sum(1).sum(1) / pos_inds.sum(1).sum(1)\ndepth_loss = depth_loss.mean(0)\n```\n***Learning rate and optimizer***\n- Learning rate is 0.0005 in first 15 epoch, Then cycle cosine learning rate trains per 5 epochs, total epoch is 40~60. At the large scale, I use the small scale model to fine-tune with a small learning rate(1e-5)\n- optimizer: Adam\n\n**Key point of this solution:**\n1.  Heatmap(umich_gaussian) and Focal loss(same as https://github.com/xingyizhou/CenterNet)\n2. Depth loss\n3. Bigger scale\n\nThanks for my teammate's great work and Thanks hocop1's kernel(https://www.kaggle.com/hocop1/centernet-baseline)",
    "728020": "Congrats !! Amazing! Thanks for sharing !",
    "727485": "Great solution, thank you so much for sharing!\nMay I ask what do you do if the number of detected cars is not the same for each model?\nI'd say my approach is 1/3 of this one and when I tried to ensemble, many times I had to face that not all the predictions include the same amount of detected cars.\nThanks!",
    "726397": "Thanks for sharing about the DLA solution! By the way, do you use Adam as the optimizer? And how do you set the learning rate? It seems I used a similar network but didn't know why the net couldn't converge. Thanks for your help!",
    "725300": "Congratulations! Nice game",
    "725279": "Congratulations! Thanks for sharing! May I ask the optimizer and the learning rate you use to train centerNet-dla34? I've tried a few times, it never converges. 😔  ",
    "725345": "",
    "732537": "Congrats and thanks for sharing!"
  }
}