{
  "id": 199540,
  "title": "25nd place Summary",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/chung-hsien-tsai-25nd-place-summary",
  "author_name": "",
  "post_date": "2020-11-26T06:20:18.587Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<h2>Tools I used</h2>\n<ul>\n<li>Tensorflow/Keras</li>\n<li>Colab (with TPU, free version)</li>\n<li>Google Cloud Storage<br>\n I rendered the data offline, packed the data into .tfrec format, and uploaded them to Google Cloud Storage (~2TB), To ensure that the size would not be exploded, I compress the image into PNG format</li>\n</ul>\n<h2>Dataset</h2>\n<ul>\n<li>Used training (not full version*) and validation data</li>\n<li>Used default setting (from sample code) for rendering except min_future_frame ( = 10)</li>\n</ul>\n<h2>Network backbone</h2>\n<ul>\n<li>Xception<br>\n I had tried a lot of backbone provided from Keras and different version of Resnet (101, 50, 34, 18), and found Xception's quality was the best (through validation)</li>\n<li>Batch size = 256</li>\n</ul>\n<h2>Other tricks</h2>\n<ul>\n<li>Validation<br>\nValidation data and part of training data were used, and I separated both training and validation data into five parts by y-position of agent's center points. I made sure that the same (or similar) maps would not appear in both training and validation data simultaneously</li>\n<li>Multimode<br>\n I predicted 4 modes and picked the top-3 (determined by confidences) modes. Besides, I had tried more modes (&gt; 4) but found only 4 modes had larger confidence values while the other modes' never be picked. The unbalanced issue seems common in multimode trick</li>\n<li>Label smoothing  <br>\n I \"multiply\" a small gaussion noise (mean = 1, std = 0.00333) to each ground truth.</li>\n<li>Custom loss function<br>\n I used the evaluation metric for multimode (provided from official) as loss function. To gave more penalty to the corner cases, the loss was multiply by a weight, which was proportional to  the \"Angle\" of each agent<br>\n Angle = acos(x/sqrt(x**2 + y ** 2)), (x, y) is the coordinate of the final position of each agent </li>\n<li>Discard History<br>\n I used 10 history frames for training, but I found that around 5% testing data had history frames less than 10. So, I discard the history of target history frame randomly</li>\n</ul>\n<p>*Due to the limitation of memory in colab, I could not use preprocess full training data. So sad.</p>",
  "messages": [
    {
      "id": "1091578",
      "postDate": "11/26/2020 05:41:40",
      "content": "<h2>Tools I used</h2>\n<ul>\n<li>Tensorflow/Keras</li>\n<li>Colab (with TPU, free version)</li>\n<li>Google Cloud Storage<br>\n I rendered the data offline, packed the data into .tfrec format, and uploaded them to Google Cloud Storage (~2TB), To ensure that the size would not be exploded, I compress the image into PNG format</li>\n</ul>\n<h2>Dataset</h2>\n<ul>\n<li>Used training (not full version*) and validation data</li>\n<li>Used default setting (from sample code) for rendering except min_future_frame ( = 10)</li>\n</ul>\n<h2>Network backbone</h2>\n<ul>\n<li>Xception<br>\n I had tried a lot of backbone provided from Keras and different version of Resnet (101, 50, 34, 18), and found Xception's quality was the best (through validation)</li>\n<li>Batch size = 256</li>\n</ul>\n<h2>Other tricks</h2>\n<ul>\n<li>Validation<br>\nValidation data and part of training data were used, and I separated both training and validation data into five parts by y-position of agent's center points. I made sure that the same (or similar) maps would not appear in both training and validation data simultaneously</li>\n<li>Multimode<br>\n I predicted 4 modes and picked the top-3 (determined by confidences) modes. Besides, I had tried more modes (&gt; 4) but found only 4 modes had larger confidence values while the other modes' never be picked. The unbalanced issue seems common in multimode trick</li>\n<li>Label smoothing  <br>\n I \"multiply\" a small gaussion noise (mean = 1, std = 0.00333) to each ground truth.</li>\n<li>Custom loss function<br>\n I used the evaluation metric for multimode (provided from official) as loss function. To gave more penalty to the corner cases, the loss was multiply by a weight, which was proportional to  the \"Angle\" of each agent<br>\n Angle = acos(x/sqrt(x**2 + y ** 2)), (x, y) is the coordinate of the final position of each agent </li>\n<li>Discard History<br>\n I used 10 history frames for training, but I found that around 5% testing data had history frames less than 10. So, I discard the history of target history frame randomly</li>\n</ul>\n<p>*Due to the limitation of memory in colab, I could not use preprocess full training data. So sad.</p>",
      "rawMarkdown": "## Tools I used\n- Tensorflow/Keras\n- Colab (with TPU, free version)\n- Google Cloud Storage\n     I rendered the data offline, packed the data into .tfrec format, and uploaded them to Google Cloud Storage (~2TB), To ensure that the size would not be exploded, I compress the image into PNG format\n\n## Dataset\n- Used training (not full version*) and validation data\n- Used default setting (from sample code) for rendering except min_future_frame ( = 10)\n\n## Network backbone\n- Xception\n     I had tried a lot of backbone provided from Keras and different version of Resnet (101, 50, 34, 18), and found Xception's quality was the best (through validation)\n- Batch size = 256\n\n## Other tricks\n- Validation\n    Validation data and part of training data were used, and I separated both training and validation data into five parts by y-position of agent's center points. I made sure that the same (or similar) maps would not appear in both training and validation data simultaneously\n- Multimode\n     I predicted 4 modes and picked the top-3 (determined by confidences) modes. Besides, I had tried more modes (> 4) but found only 4 modes had larger confidence values while the other modes' never be picked. The unbalanced issue seems common in multimode trick\n- Label smoothing  \n     I \"multiply\" a small gaussion noise (mean = 1, std = 0.00333) to each ground truth.\n- Custom loss function\n     I used the evaluation metric for multimode (provided from official) as loss function. To gave more penalty to the corner cases, the loss was multiply by a weight, which was proportional to  the \"Angle\" of each agent\n     Angle = acos(x/sqrt(x**2 + y ** 2)), (x, y) is the coordinate of the final position of each agent \n- Discard History\n     I used 10 history frames for training, but I found that around 5% testing data had history frames less than 10. So, I discard the history of target history frame randomly\n\n*Due to the limitation of memory in colab, I could not use preprocess full training data. So sad.",
      "votes": null
    },
    {
      "id": "1091616",
      "postDate": "11/26/2020 06:22:47",
      "content": "<p>Thanks for sharing. How much did turning data into tfrecord speed up training for you? I also ended up having some success with a 4-mode model and then choosing 3-modes from it. I tried similar things with 10-mode but found that only about 5 would ever get picked. Of the bottom 5 they were only selected less than 100 times on the validation set and even then with low confidence. I tried augmenting with dropout on the confidence output so it would enforce other branches to get some attention, but ultimately seemed to just slow convergence</p>",
      "rawMarkdown": "Thanks for sharing. How much did turning data into tfrecord speed up training for you? I also ended up having some success with a 4-mode model and then choosing 3-modes from it. I tried similar things with 10-mode but found that only about 5 would ever get picked. Of the bottom 5 they were only selected less than 100 times on the validation set and even then with low confidence. I tried augmenting with dropout on the confidence output so it would enforce other branches to get some attention, but ultimately seemed to just slow convergence",
      "votes": null
    },
    {
      "id": "1091927",
      "postDate": "11/26/2020 11:40:47",
      "content": "<p>Well done! How much improvement did your custom loss function give over the regular one?</p>",
      "rawMarkdown": "Well done! How much improvement did your custom loss function give over the regular one?",
      "votes": null
    },
    {
      "id": "1092015",
      "postDate": "11/26/2020 13:28:34",
      "content": "<blockquote>\n  <blockquote>\n    <p>How much did turning data into tfrecord speed up training for you?<br>\n    Around 2~3 days. Actually I used a lot of machines to do this.</p>\n  </blockquote>\n</blockquote>\n<p>To enhance the weak modes, I had also tried a loss function which formulate the confidence estimation as a classification problem. The confidence of the mode with the least MSE would be \"1\" while the others are \"0\", and I calculated the cross entropy of the confidences as part of the loss function. However it did not work either.</p>",
      "rawMarkdown": "> > How much did turning data into tfrecord speed up training for you?\nAround 2~3 days. Actually I used a lot of machines to do this.\n\nTo enhance the weak modes, I had also tried a loss function which formulate the confidence estimation as a classification problem. The confidence of the mode with the least MSE would be \"1\" while the others are \"0\", and I calculated the cross entropy of the confidences as part of the loss function. However it did not work either.",
      "votes": null
    },
    {
      "id": "1092035",
      "postDate": "11/26/2020 13:43:55",
      "content": "<p>I only compared the loss function when doing the cross validation (using little parts of the dataset). I remembered the improvement was around \"-1\" for the evaluation metric. However, I am not sure that there would be so much improvement when using all of my training data</p>",
      "rawMarkdown": "I only compared the loss function when doing the cross validation (using little parts of the dataset). I remembered the improvement was around \"-1\" for the evaluation metric. However, I am not sure that there would be so much improvement when using all of my training data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091616,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "11/26/2020 06:22:47",
      "content": "<p>Thanks for sharing. How much did turning data into tfrecord speed up training for you? I also ended up having some success with a 4-mode model and then choosing 3-modes from it. I tried similar things with 10-mode but found that only about 5 would ever get picked. Of the bottom 5 they were only selected less than 100 times on the validation set and even then with low confidence. I tried augmenting with dropout on the confidence output so it would enforce other branches to get some attention, but ultimately seemed to just slow convergence</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092015,
          "author_name": "hardworkingkaggler",
          "author_url": "",
          "post_date": "11/26/2020 13:28:34",
          "content": "<blockquote>\n  <blockquote>\n    <p>How much did turning data into tfrecord speed up training for you?<br>\n    Around 2~3 days. Actually I used a lot of machines to do this.</p>\n  </blockquote>\n</blockquote>\n<p>To enhance the weak modes, I had also tried a loss function which formulate the confidence estimation as a classification problem. The confidence of the mode with the least MSE would be \"1\" while the others are \"0\", and I calculated the cross entropy of the confidences as part of the loss function. However it did not work either.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091927,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "11/26/2020 11:40:47",
      "content": "<p>Well done! How much improvement did your custom loss function give over the regular one?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1092035,
          "author_name": "hardworkingkaggler",
          "author_url": "",
          "post_date": "11/26/2020 13:43:55",
          "content": "<p>I only compared the loss function when doing the cross validation (using little parts of the dataset). I remembered the improvement was around \"-1\" for the evaluation metric. However, I am not sure that there would be so much improvement when using all of my training data</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1091578": "## Tools I used\n- Tensorflow/Keras\n- Colab (with TPU, free version)\n- Google Cloud Storage\n     I rendered the data offline, packed the data into .tfrec format, and uploaded them to Google Cloud Storage (~2TB), To ensure that the size would not be exploded, I compress the image into PNG format\n\n## Dataset\n- Used training (not full version*) and validation data\n- Used default setting (from sample code) for rendering except min_future_frame ( = 10)\n\n## Network backbone\n- Xception\n     I had tried a lot of backbone provided from Keras and different version of Resnet (101, 50, 34, 18), and found Xception's quality was the best (through validation)\n- Batch size = 256\n\n## Other tricks\n- Validation\n    Validation data and part of training data were used, and I separated both training and validation data into five parts by y-position of agent's center points. I made sure that the same (or similar) maps would not appear in both training and validation data simultaneously\n- Multimode\n     I predicted 4 modes and picked the top-3 (determined by confidences) modes. Besides, I had tried more modes (> 4) but found only 4 modes had larger confidence values while the other modes' never be picked. The unbalanced issue seems common in multimode trick\n- Label smoothing  \n     I \"multiply\" a small gaussion noise (mean = 1, std = 0.00333) to each ground truth.\n- Custom loss function\n     I used the evaluation metric for multimode (provided from official) as loss function. To gave more penalty to the corner cases, the loss was multiply by a weight, which was proportional to  the \"Angle\" of each agent\n     Angle = acos(x/sqrt(x**2 + y ** 2)), (x, y) is the coordinate of the final position of each agent \n- Discard History\n     I used 10 history frames for training, but I found that around 5% testing data had history frames less than 10. So, I discard the history of target history frame randomly\n\n*Due to the limitation of memory in colab, I could not use preprocess full training data. So sad.",
    "1091616": "Thanks for sharing. How much did turning data into tfrecord speed up training for you? I also ended up having some success with a 4-mode model and then choosing 3-modes from it. I tried similar things with 10-mode but found that only about 5 would ever get picked. Of the bottom 5 they were only selected less than 100 times on the validation set and even then with low confidence. I tried augmenting with dropout on the confidence output so it would enforce other branches to get some attention, but ultimately seemed to just slow convergence",
    "1091927": "Well done! How much improvement did your custom loss function give over the regular one?",
    "1092015": "> > How much did turning data into tfrecord speed up training for you?\nAround 2~3 days. Actually I used a lot of machines to do this.\n\nTo enhance the weak modes, I had also tried a loss function which formulate the confidence estimation as a classification problem. The confidence of the mode with the least MSE would be \"1\" while the others are \"0\", and I calculated the cross entropy of the confidences as part of the loss function. However it did not work either.",
    "1092035": "I only compared the loss function when doing the cross validation (using little parts of the dataset). I remembered the improvement was around \"-1\" for the evaluation metric. However, I am not sure that there would be so much improvement when using all of my training data"
  },
  "source": "meta"
}