{
  "id": 199406,
  "title": "Ideas that looked good but didn't work",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199406",
  "author_name": "A_Elsheikh",
  "post_date": "2020-11-25T16:25:39.919000",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Here is a thread to share some ideas that didn't work. I will start with sharing three ideas that I wished could provide me with an edge but didn't really work and didn't endup in my submission:</p>\n<p>1- Hard sample mining by target position. The distribution of <code>log(distance traveled)</code> for something available for all samples <code>min_frame_future</code> shows a bimodal distribution. I used two heads like the paper by Uber and this informaiton might have already been embeded there (BTW, distance traveled is highly correlated with history position, velocity)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2Fe4452384e0d1979a3a8218f06fbdf773%2Flog_distance_traveled.png?generation=1606320889799403&amp;alt=media\" alt=\"\"><br>\n2- Stratification by the target position after rotation in the image space (signed and <code>log(1+p)</code> transformed) as well. See the picture for the clustering as a stratification target.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2F6978227fdadbeee5a57bdf219edc01a7%2Fcluster_2.png?generation=1606321127011260&amp;alt=media\" alt=\"\"></p>\n<p>3- Boosting by embedding the outputs from the first layer model in some of the frames for the second level. For example the three modes could be embedded in three different layers.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2Ff67833dbdf2b5f14a44f3fbc2d8da5cb%2Fboosting_small.png?generation=1606321472949678&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1090961,
      "postDate": "2020-11-25T17:36:38.883Z",
      "content": "<p>I also tried using weighted sampler approach (which should be standard in classification problems), but the result was terrible. Specifically I used R square of the trajectory and the inverse of its density as sampler weights.</p>",
      "rawMarkdown": "I also tried using weighted sampler approach (which should be standard in classification problems), but the result was terrible. Specifically I used R square of the trajectory and the inverse of its density as sampler weights.",
      "votes": 3,
      "replies": [
        {
          "id": 1091503,
          "postDate": "2020-11-26T04:28:32.697Z",
          "content": "<p>Here is the gist. My idea was since the training data as in trajectories is so imbalanced, I tried to use inverse of densities of R-square to produce a \"balanced\" training set when sampling.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2F77fcb47cb390f9c6767bb3ab4d5bf25b%2FPicture1.png?generation=1606364756126527&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Here is the gist. My idea was since the training data as in trajectories is so imbalanced, I tried to use inverse of densities of R-square to produce a \"balanced\" training set when sampling.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2F77fcb47cb390f9c6767bb3ab4d5bf25b%2FPicture1.png?generation=1606364756126527&alt=media)",
          "votes": 2
        }
      ]
    },
    {
      "id": 1090871,
      "postDate": "2020-11-25T16:25:39.920Z",
      "content": "<p>Here is a thread to share some ideas that didn't work. I will start with sharing three ideas that I wished could provide me with an edge but didn't really work and didn't endup in my submission:</p>\n<p>1- Hard sample mining by target position. The distribution of <code>log(distance traveled)</code> for something available for all samples <code>min_frame_future</code> shows a bimodal distribution. I used two heads like the paper by Uber and this informaiton might have already been embeded there (BTW, distance traveled is highly correlated with history position, velocity)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2Fe4452384e0d1979a3a8218f06fbdf773%2Flog_distance_traveled.png?generation=1606320889799403&amp;alt=media\" alt=\"\"><br>\n2- Stratification by the target position after rotation in the image space (signed and <code>log(1+p)</code> transformed) as well. See the picture for the clustering as a stratification target.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2F6978227fdadbeee5a57bdf219edc01a7%2Fcluster_2.png?generation=1606321127011260&amp;alt=media\" alt=\"\"></p>\n<p>3- Boosting by embedding the outputs from the first layer model in some of the frames for the second level. For example the three modes could be embedded in three different layers.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2Ff67833dbdf2b5f14a44f3fbc2d8da5cb%2Fboosting_small.png?generation=1606321472949678&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here is a thread to share some ideas that didn't work. I will start with sharing three ideas that I wished could provide me with an edge but didn't really work and didn't endup in my submission:\n\n1- Hard sample mining by target position. The distribution of `log(distance traveled)` for something available for all samples `min_frame_future` shows a bimodal distribution. I used two heads like the paper by Uber and this informaiton might have already been embeded there (BTW, distance traveled is highly correlated with history position, velocity)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2Fe4452384e0d1979a3a8218f06fbdf773%2Flog_distance_traveled.png?generation=1606320889799403&alt=media)\n2- Stratification by the target position after rotation in the image space (signed and `log(1+p)` transformed) as well. See the picture for the clustering as a stratification target.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2F6978227fdadbeee5a57bdf219edc01a7%2Fcluster_2.png?generation=1606321127011260&alt=media)\n\n3- Boosting by embedding the outputs from the first layer model in some of the frames for the second level. For example the three modes could be embedded in three different layers.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2Ff67833dbdf2b5f14a44f3fbc2d8da5cb%2Fboosting_small.png?generation=1606321472949678&alt=media)",
      "votes": 4
    },
    {
      "id": 1091045,
      "postDate": "2020-11-25T18:35:04.007Z",
      "content": "<p>One trick that I tried but not very successful is that you can estimate the expected future positions using agent's history velocity, then add this expected positions to the final prediction layer. So your model only need to learn the difference between this expected trajectory and the true trajectory, and it doesn't need to output value like ~100, which is slightly difficult for neural network to do.<br>\nHowever, this seems to only work at the beginning, where you will start with smaller loss. Later this added positions seem to make the training less effective.</p>",
      "rawMarkdown": "One trick that I tried but not very successful is that you can estimate the expected future positions using agent's history velocity, then add this expected positions to the final prediction layer. So your model only need to learn the difference between this expected trajectory and the true trajectory, and it doesn't need to output value like ~100, which is slightly difficult for neural network to do.\nHowever, this seems to only work at the beginning, where you will start with smaller loss. Later this added positions seem to make the training less effective.",
      "votes": 1,
      "replies": [
        {
          "id": 1091048,
          "postDate": "2020-11-25T18:40:45.400Z",
          "content": "<p>Did try that as well -- same conclusion. The only thing that ended up in my submission is a <code>cumsum</code> step on the outputs (i.e. NN only outputs the differences similar to most timeseries models) but I didn't have enough time to test if this had a positive/negative impact on the good working models.</p>",
          "rawMarkdown": "Did try that as well -- same conclusion. The only thing that ended up in my submission is a `cumsum` step on the outputs (i.e. NN only outputs the differences similar to most timeseries models) but I didn't have enough time to test if this had a positive/negative impact on the good working models.\n",
          "votes": 1
        },
        {
          "id": 1091052,
          "postDate": "2020-11-25T18:42:14.740Z",
          "content": "<p>As I am writing this, I should have tested making the NN outputs as the differences from the base model (constant velocity) !!</p>",
          "rawMarkdown": "As I am writing this, I should have tested making the NN outputs as the differences from the base model (constant velocity) !!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1090906,
      "postDate": "2020-11-25T17:02:11.643Z",
      "content": "<p>Ha, I was trying the 3rd method.</p>",
      "rawMarkdown": "Ha, I was trying the 3rd method.",
      "votes": 1,
      "replies": [
        {
          "id": 1090954,
          "postDate": "2020-11-25T17:31:12.570Z",
          "content": "<p>Theoretically, it should work. See the recent work \"Gradient Boosting Neural Networks: GrowNet\"  <a href=\"https://arxiv.org/abs/2002.07971\" target=\"_blank\">https://arxiv.org/abs/2002.07971</a> and it is a general framewrok for stacking, combining heterogeneous models both in image size, different backbone models, etc.</p>",
          "rawMarkdown": "Theoretically, it should work. See the recent work \"Gradient Boosting Neural Networks: GrowNet\"  https://arxiv.org/abs/2002.07971 and it is a general framewrok for stacking, combining heterogeneous models both in image size, different backbone models, etc.",
          "votes": 1
        },
        {
          "id": 1091032,
          "postDate": "2020-11-25T18:29:01.237Z",
          "content": "<p>I guess I probably didn't have enough time to train it since I develop it like yesterday. Also, I added some other random trick which isn't best idea :D</p>",
          "rawMarkdown": "I guess I probably didn't have enough time to train it since I develop it like yesterday. Also, I added some other random trick which isn't best idea :D"
        }
      ]
    },
    {
      "id": 1091566,
      "postDate": "2020-11-26T05:35:04.297Z",
      "content": "<p>In a short training, OHEM helped me (about 1.0 score improvement). However, it didn't help me for long training of ~full epoch, and I lost a lot of time because I stuck with it :(</p>",
      "rawMarkdown": "In a short training, OHEM helped me (about 1.0 score improvement). However, it didn't help me for long training of ~full epoch, and I lost a lot of time because I stuck with it :(",
      "votes": 2
    },
    {
      "id": 1091202,
      "postDate": "2020-11-25T21:11:28.900Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1091239,
          "postDate": "2020-11-25T22:13:23.813Z",
          "content": "<p>I only tried the LSTM head described in this paper <a href=\"https://arxiv.org/abs/1808.05819\" target=\"_blank\">https://arxiv.org/abs/1808.05819</a>. It was much harder to optimize so I moved on and didn't test it later when things started working better.</p>",
          "rawMarkdown": "I only tried the LSTM head described in this paper https://arxiv.org/abs/1808.05819. It was much harder to optimize so I moved on and didn't test it later when things started working better."
        },
        {
          "id": 1091245,
          "postDate": "2020-11-25T22:18:01.947Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1091246,
          "postDate": "2020-11-25T22:20:40.747Z",
          "content": "<p>Thanks for sharing. The map data provides a lot of information that you can't afford to not use.</p>",
          "rawMarkdown": "Thanks for sharing. The map data provides a lot of information that you can't afford to not use."
        },
        {
          "id": 1091252,
          "postDate": "2020-11-25T22:25:48.397Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1090961,
      "author_name": "ryan",
      "author_url": "",
      "post_date": "2020-11-25T17:36:38.883000",
      "content": "<p>I also tried using weighted sampler approach (which should be standard in classification problems), but the result was terrible. Specifically I used R square of the trajectory and the inverse of its density as sampler weights.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1091503,
          "author_name": "ryan",
          "author_url": "",
          "post_date": "2020-11-26T04:28:32.697000",
          "content": "<p>Here is the gist. My idea was since the training data as in trajectories is so imbalanced, I tried to use inverse of densities of R-square to produce a \"balanced\" training set when sampling.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2168668%2F77fcb47cb390f9c6767bb3ab4d5bf25b%2FPicture1.png?generation=1606364756126527&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1091045,
      "author_name": "Louis Yang",
      "author_url": "",
      "post_date": "2020-11-25T18:35:04.007000",
      "content": "<p>One trick that I tried but not very successful is that you can estimate the expected future positions using agent's history velocity, then add this expected positions to the final prediction layer. So your model only need to learn the difference between this expected trajectory and the true trajectory, and it doesn't need to output value like ~100, which is slightly difficult for neural network to do.<br>\nHowever, this seems to only work at the beginning, where you will start with smaller loss. Later this added positions seem to make the training less effective.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1091048,
          "author_name": "A_Elsheikh",
          "author_url": "",
          "post_date": "2020-11-25T18:40:45.400000",
          "content": "<p>Did try that as well -- same conclusion. The only thing that ended up in my submission is a <code>cumsum</code> step on the outputs (i.e. NN only outputs the differences similar to most timeseries models) but I didn't have enough time to test if this had a positive/negative impact on the good working models.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091052,
          "author_name": "A_Elsheikh",
          "author_url": "",
          "post_date": "2020-11-25T18:42:14.740000",
          "content": "<p>As I am writing this, I should have tested making the NN outputs as the differences from the base model (constant velocity) !!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1090906,
      "author_name": "Louis Yang",
      "author_url": "",
      "post_date": "2020-11-25T17:02:11.643000",
      "content": "<p>Ha, I was trying the 3rd method.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1090954,
          "author_name": "A_Elsheikh",
          "author_url": "",
          "post_date": "2020-11-25T17:31:12.570000",
          "content": "<p>Theoretically, it should work. See the recent work \"Gradient Boosting Neural Networks: GrowNet\"  <a href=\"https://arxiv.org/abs/2002.07971\" target=\"_blank\">https://arxiv.org/abs/2002.07971</a> and it is a general framewrok for stacking, combining heterogeneous models both in image size, different backbone models, etc.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091032,
          "author_name": "Louis Yang",
          "author_url": "",
          "post_date": "2020-11-25T18:29:01.237000",
          "content": "<p>I guess I probably didn't have enough time to train it since I develop it like yesterday. Also, I added some other random trick which isn't best idea :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091566,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2020-11-26T05:35:04.297000",
      "content": "<p>In a short training, OHEM helped me (about 1.0 score improvement). However, it didn't help me for long training of ~full epoch, and I lost a lot of time because I stuck with it :(</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1091202,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-25T21:11:28.900000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1091239,
          "author_name": "A_Elsheikh",
          "author_url": "",
          "post_date": "2020-11-25T22:13:23.813000",
          "content": "<p>I only tried the LSTM head described in this paper <a href=\"https://arxiv.org/abs/1808.05819\" target=\"_blank\">https://arxiv.org/abs/1808.05819</a>. It was much harder to optimize so I moved on and didn't test it later when things started working better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1091245,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-25T22:18:01.947000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1091246,
          "author_name": "A_Elsheikh",
          "author_url": "",
          "post_date": "2020-11-25T22:20:40.747000",
          "content": "<p>Thanks for sharing. The map data provides a lot of information that you can't afford to not use.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1091252,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-25T22:25:48.397000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1090961": "I also tried using weighted sampler approach (which should be standard in classification problems), but the result was terrible. Specifically I used R square of the trajectory and the inverse of its density as sampler weights.",
    "1090871": "Here is a thread to share some ideas that didn't work. I will start with sharing three ideas that I wished could provide me with an edge but didn't really work and didn't endup in my submission:\n\n1- Hard sample mining by target position. The distribution of `log(distance traveled)` for something available for all samples `min_frame_future` shows a bimodal distribution. I used two heads like the paper by Uber and this informaiton might have already been embeded there (BTW, distance traveled is highly correlated with history position, velocity)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2Fe4452384e0d1979a3a8218f06fbdf773%2Flog_distance_traveled.png?generation=1606320889799403&alt=media)\n2- Stratification by the target position after rotation in the image space (signed and `log(1+p)` transformed) as well. See the picture for the clustering as a stratification target.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2F6978227fdadbeee5a57bdf219edc01a7%2Fcluster_2.png?generation=1606321127011260&alt=media)\n\n3- Boosting by embedding the outputs from the first layer model in some of the frames for the second level. For example the three modes could be embedded in three different layers.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F257647%2Ff67833dbdf2b5f14a44f3fbc2d8da5cb%2Fboosting_small.png?generation=1606321472949678&alt=media)",
    "1091045": "One trick that I tried but not very successful is that you can estimate the expected future positions using agent's history velocity, then add this expected positions to the final prediction layer. So your model only need to learn the difference between this expected trajectory and the true trajectory, and it doesn't need to output value like ~100, which is slightly difficult for neural network to do.\nHowever, this seems to only work at the beginning, where you will start with smaller loss. Later this added positions seem to make the training less effective.",
    "1090906": "Ha, I was trying the 3rd method.",
    "1091566": "In a short training, OHEM helped me (about 1.0 score improvement). However, it didn't help me for long training of ~full epoch, and I lost a lot of time because I stuck with it :(",
    "1091202": ""
  }
}