{
  "id": 199890,
  "title": "Why concating metadata with deep neural network output doesn't work?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199890",
  "author_name": "",
  "post_date": "2020-11-27T19:51:32.702877200Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>One simple idea that I have was to concate additional information like <code>history_positions</code>, speed, history_yaws to the final dense layers with the resnet output. And, throw a few dense layer after that. However, I found that doing so can lead to overfit pretty easily (huge gap between train and validation score), and give you worse result. <br>\nI was wondering if anyone has tried similar thing and got similar problem? And, why does this happen?</p>\n<p>Also, is it a general pattern that adding dense layers after these deep neural net in image recognition will lead to worse result?</p>",
  "messages": [
    {
      "id": "1093535",
      "postDate": "11/27/2020 19:51:32",
      "content": "<p>One simple idea that I have was to concate additional information like <code>history_positions</code>, speed, history_yaws to the final dense layers with the resnet output. And, throw a few dense layer after that. However, I found that doing so can lead to overfit pretty easily (huge gap between train and validation score), and give you worse result. <br>\nI was wondering if anyone has tried similar thing and got similar problem? And, why does this happen?</p>\n<p>Also, is it a general pattern that adding dense layers after these deep neural net in image recognition will lead to worse result?</p>",
      "rawMarkdown": "One simple idea that I have was to concate additional information like `history_positions`, speed, history_yaws to the final dense layers with the resnet output. And, throw a few dense layer after that. However, I found that doing so can lead to overfit pretty easily (huge gap between train and validation score), and give you worse result. \nI was wondering if anyone has tried similar thing and got similar problem? And, why does this happen?\n\nAlso, is it a general pattern that adding dense layers after these deep neural net in image recognition will lead to worse result?",
      "votes": null
    },
    {
      "id": "1093547",
      "postDate": "11/27/2020 20:15:56",
      "content": "<p>Yes. I tried the same and got the same result. I think it is because the network might have already picked up on this information from the raster? Not too sure though.</p>",
      "rawMarkdown": "Yes. I tried the same and got the same result. I think it is because the network might have already picked up on this information from the raster? Not too sure though.",
      "votes": null
    },
    {
      "id": "1093698",
      "postDate": "11/28/2020 00:03:51",
      "content": "<p>deeper heads didn't work for us (2 linear layers max)<br>\nalso, usage of additional data such as velocity only lead to a faster loss (also visible in validation) reduction but didn't yield a better result at the end of training. I guess the models were strong enough to pick all the information up from the image already and use it for the feature extraction with better results than late concatenation would give. </p>",
      "rawMarkdown": "deeper heads didn't work for us (2 linear layers max)\nalso, usage of additional data such as velocity only lead to a faster loss (also visible in validation) reduction but didn't yield a better result at the end of training. I guess the models were strong enough to pick all the information up from the image already and use it for the feature extraction with better results than late concatenation would give.",
      "votes": null
    },
    {
      "id": "1093738",
      "postDate": "11/28/2020 01:20:34",
      "content": "<p>Glad that 1st place also got the similar observation. It sounds like making the problem easier for a deep neural net is not always a good thing.<br>\nFrom your experience to other competition, do you think this apply in general to image recognition problem or deep learning? Or, is it just specific for this problem?</p>",
      "rawMarkdown": "Glad that 1st place also got the similar observation. It sounds like making the problem easier for a deep neural net is not always a good thing.\nFrom your experience to other competition, do you think this apply in general to image recognition problem or deep learning? Or, is it just specific for this problem?",
      "votes": null
    },
    {
      "id": "1093917",
      "postDate": "11/28/2020 05:54:55",
      "content": "<p>I would say with any competition you hardly know before trying if auxiliary loss/data helps. Here I was pretty sure the NN will \"learn\" velocity and acceleration just from difference of frames, we tried anyways.</p>",
      "rawMarkdown": "I would say with any competition you hardly know before trying if auxiliary loss/data helps. Here I was pretty sure the NN will \"learn\" velocity and acceleration just from difference of frames, we tried anyways.",
      "votes": null
    },
    {
      "id": "1094270",
      "postDate": "11/28/2020 12:59:32",
      "content": "<p>I had this in my solution (BTW, it is all single model) and I added date, hour, extra info. Did it contribute much to the results, it is hard to confirm given the extra long run time and the nonlinear exploration path I went through till reaching a good score.  </p>",
      "rawMarkdown": "I had this in my solution (BTW, it is all single model) and I added date, hour, extra info. Did it contribute much to the results, it is hard to confirm given the extra long run time and the nonlinear exploration path I went through till reaching a good score.",
      "votes": null
    },
    {
      "id": "1094696",
      "postDate": "11/28/2020 20:27:21",
      "content": "<p>Interesting! How did you merge them with resenet output? Did you do many regular fully connected layers and relu after that?</p>",
      "rawMarkdown": "Interesting! How did you merge them with resenet output? Did you do many regular fully connected layers and relu after that?",
      "votes": null
    },
    {
      "id": "1094758",
      "postDate": "11/28/2020 22:21:04",
      "content": "<p>Fully connected, relu, then output. Something like this:</p>\n<pre><code>x = torch.cat((x, y), axis=-1)\nx = self.backbone.fc(x)\nx = F.relu(x)\nx = self.output(x)\n</code></pre>",
      "rawMarkdown": "Fully connected, relu, then output. Something like this:\n```\nx = torch.cat((x, y), axis=-1)\nx = self.backbone.fc(x)\nx = F.relu(x)\nx = self.output(x)\n```",
      "votes": null
    },
    {
      "id": "1096022",
      "postDate": "11/30/2020 06:44:26",
      "content": "<p>We also tried add some information, like <code>agent type</code>.<br>\nWe tried to input this additional data as multiplication between image feature vectors and additional metadata feature vectors, instead of concatenation, but not worked well.</p>",
      "rawMarkdown": "We also tried add some information, like `agent type`.\nWe tried to input this additional data as multiplication between image feature vectors and additional metadata feature vectors, instead of concatenation, but not worked well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1093547,
      "author_name": "yousof9",
      "author_url": "",
      "post_date": "11/27/2020 20:15:56",
      "content": "<p>Yes. I tried the same and got the same result. I think it is because the network might have already picked up on this information from the raster? Not too sure though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1093698,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "11/28/2020 00:03:51",
      "content": "<p>deeper heads didn't work for us (2 linear layers max)<br>\nalso, usage of additional data such as velocity only lead to a faster loss (also visible in validation) reduction but didn't yield a better result at the end of training. I guess the models were strong enough to pick all the information up from the image already and use it for the feature extraction with better results than late concatenation would give. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1093738,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/28/2020 01:20:34",
          "content": "<p>Glad that 1st place also got the similar observation. It sounds like making the problem easier for a deep neural net is not always a good thing.<br>\nFrom your experience to other competition, do you think this apply in general to image recognition problem or deep learning? Or, is it just specific for this problem?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093917,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "11/28/2020 05:54:55",
          "content": "<p>I would say with any competition you hardly know before trying if auxiliary loss/data helps. Here I was pretty sure the NN will \"learn\" velocity and acceleration just from difference of frames, we tried anyways.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1094270,
      "author_name": "ggrizzly",
      "author_url": "",
      "post_date": "11/28/2020 12:59:32",
      "content": "<p>I had this in my solution (BTW, it is all single model) and I added date, hour, extra info. Did it contribute much to the results, it is hard to confirm given the extra long run time and the nonlinear exploration path I went through till reaching a good score.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1094696,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/28/2020 20:27:21",
          "content": "<p>Interesting! How did you merge them with resenet output? Did you do many regular fully connected layers and relu after that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094758,
          "author_name": "ggrizzly",
          "author_url": "",
          "post_date": "11/28/2020 22:21:04",
          "content": "<p>Fully connected, relu, then output. Something like this:</p>\n<pre><code>x = torch.cat((x, y), axis=-1)\nx = self.backbone.fc(x)\nx = F.relu(x)\nx = self.output(x)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1096022,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "11/30/2020 06:44:26",
      "content": "<p>We also tried add some information, like <code>agent type</code>.<br>\nWe tried to input this additional data as multiplication between image feature vectors and additional metadata feature vectors, instead of concatenation, but not worked well.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1093535": "One simple idea that I have was to concate additional information like `history_positions`, speed, history_yaws to the final dense layers with the resnet output. And, throw a few dense layer after that. However, I found that doing so can lead to overfit pretty easily (huge gap between train and validation score), and give you worse result. \nI was wondering if anyone has tried similar thing and got similar problem? And, why does this happen?\n\nAlso, is it a general pattern that adding dense layers after these deep neural net in image recognition will lead to worse result?",
    "1093547": "Yes. I tried the same and got the same result. I think it is because the network might have already picked up on this information from the raster? Not too sure though.",
    "1093698": "deeper heads didn't work for us (2 linear layers max)\nalso, usage of additional data such as velocity only lead to a faster loss (also visible in validation) reduction but didn't yield a better result at the end of training. I guess the models were strong enough to pick all the information up from the image already and use it for the feature extraction with better results than late concatenation would give.",
    "1093738": "Glad that 1st place also got the similar observation. It sounds like making the problem easier for a deep neural net is not always a good thing.\nFrom your experience to other competition, do you think this apply in general to image recognition problem or deep learning? Or, is it just specific for this problem?",
    "1093917": "I would say with any competition you hardly know before trying if auxiliary loss/data helps. Here I was pretty sure the NN will \"learn\" velocity and acceleration just from difference of frames, we tried anyways.",
    "1094270": "I had this in my solution (BTW, it is all single model) and I added date, hour, extra info. Did it contribute much to the results, it is hard to confirm given the extra long run time and the nonlinear exploration path I went through till reaching a good score.",
    "1094696": "Interesting! How did you merge them with resenet output? Did you do many regular fully connected layers and relu after that?",
    "1094758": "Fully connected, relu, then output. Something like this:\n```\nx = torch.cat((x, y), axis=-1)\nx = self.backbone.fc(x)\nx = F.relu(x)\nx = self.output(x)\n```",
    "1096022": "We also tried add some information, like `agent type`.\nWe tried to input this additional data as multiplication between image feature vectors and additional metadata feature vectors, instead of concatenation, but not worked well."
  },
  "source": "meta"
}