{
  "id": 187920,
  "title": "What can a CNN do?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/187920",
  "author_name": "ryches",
  "post_date": "2020-09-30T22:21:48.505000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I've been spending a bit of my idle training time on the side trying to do some benchmarking of toy scenarios that I've always been curious about with CNN's. The notebook is still in progress but figured I would share it and query the community for additional scenarios that they would like to be tested. <a href=\"https://www.kaggle.com/ryches/what-can-a-cnn-do\" target=\"_blank\">https://www.kaggle.com/ryches/what-can-a-cnn-do</a></p>\n<p>This kind of benchmarking is primarily of interest to me in the context of this competition because I think estimating distance to other vehicles is likely very important to this task. It is interesting to see how the CNN's do when this is their sole goal and seeing how they break down when resolution is changed. </p>\n<p>I encountered massive degradation in performance when trying to transfer from 128x128 images to 300x300 images and in past tasks it hasn't seemed to be too terrible of an issue. It would be highly beneficial if we could train on lower resolution and then do some fine-tuning on higher resolution with more context, but my preliminary experiments show this might not be hugely beneficial. </p>\n<p>One idea that I am considering is hard coding the final layers of the network to calculate the euclidean distance and have the network explicitly try to learn the x1, y1, x2, y2 coordinates to plug into those frozen computations. Making the network more explicitly learn the objective and see if that can generalize to more resolutions. </p>\n<p>Another idea is to use a U-net architecture and instead of predicting the distance between points have it try to plot the connecting points. </p>",
  "messages": [
    {
      "id": 1033348,
      "postDate": "2020-09-30T22:21:48.507Z",
      "content": "<p>I've been spending a bit of my idle training time on the side trying to do some benchmarking of toy scenarios that I've always been curious about with CNN's. The notebook is still in progress but figured I would share it and query the community for additional scenarios that they would like to be tested. <a href=\"https://www.kaggle.com/ryches/what-can-a-cnn-do\" target=\"_blank\">https://www.kaggle.com/ryches/what-can-a-cnn-do</a></p>\n<p>This kind of benchmarking is primarily of interest to me in the context of this competition because I think estimating distance to other vehicles is likely very important to this task. It is interesting to see how the CNN's do when this is their sole goal and seeing how they break down when resolution is changed. </p>\n<p>I encountered massive degradation in performance when trying to transfer from 128x128 images to 300x300 images and in past tasks it hasn't seemed to be too terrible of an issue. It would be highly beneficial if we could train on lower resolution and then do some fine-tuning on higher resolution with more context, but my preliminary experiments show this might not be hugely beneficial. </p>\n<p>One idea that I am considering is hard coding the final layers of the network to calculate the euclidean distance and have the network explicitly try to learn the x1, y1, x2, y2 coordinates to plug into those frozen computations. Making the network more explicitly learn the objective and see if that can generalize to more resolutions. </p>\n<p>Another idea is to use a U-net architecture and instead of predicting the distance between points have it try to plot the connecting points. </p>",
      "rawMarkdown": "I've been spending a bit of my idle training time on the side trying to do some benchmarking of toy scenarios that I've always been curious about with CNN's. The notebook is still in progress but figured I would share it and query the community for additional scenarios that they would like to be tested. https://www.kaggle.com/ryches/what-can-a-cnn-do\n\nThis kind of benchmarking is primarily of interest to me in the context of this competition because I think estimating distance to other vehicles is likely very important to this task. It is interesting to see how the CNN's do when this is their sole goal and seeing how they break down when resolution is changed. \n\nI encountered massive degradation in performance when trying to transfer from 128x128 images to 300x300 images and in past tasks it hasn't seemed to be too terrible of an issue. It would be highly beneficial if we could train on lower resolution and then do some fine-tuning on higher resolution with more context, but my preliminary experiments show this might not be hugely beneficial. \n\nOne idea that I am considering is hard coding the final layers of the network to calculate the euclidean distance and have the network explicitly try to learn the x1, y1, x2, y2 coordinates to plug into those frozen computations. Making the network more explicitly learn the objective and see if that can generalize to more resolutions. \n\nAnother idea is to use a U-net architecture and instead of predicting the distance between points have it try to plot the connecting points. \n",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1033348": "I've been spending a bit of my idle training time on the side trying to do some benchmarking of toy scenarios that I've always been curious about with CNN's. The notebook is still in progress but figured I would share it and query the community for additional scenarios that they would like to be tested. https://www.kaggle.com/ryches/what-can-a-cnn-do\n\nThis kind of benchmarking is primarily of interest to me in the context of this competition because I think estimating distance to other vehicles is likely very important to this task. It is interesting to see how the CNN's do when this is their sole goal and seeing how they break down when resolution is changed. \n\nI encountered massive degradation in performance when trying to transfer from 128x128 images to 300x300 images and in past tasks it hasn't seemed to be too terrible of an issue. It would be highly beneficial if we could train on lower resolution and then do some fine-tuning on higher resolution with more context, but my preliminary experiments show this might not be hugely beneficial. \n\nOne idea that I am considering is hard coding the final layers of the network to calculate the euclidean distance and have the network explicitly try to learn the x1, y1, x2, y2 coordinates to plug into those frozen computations. Making the network more explicitly learn the objective and see if that can generalize to more resolutions. \n\nAnother idea is to use a U-net architecture and instead of predicting the distance between points have it try to plot the connecting points. \n"
  }
}