{
  "id": 199499,
  "title": "Best single model",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199499",
  "author_name": "nosound",
  "post_date": "2020-11-26T00:53:01.233000",
  "votes": 21,
  "comment_count": 42,
  "views": 0,
  "content": "<p>What is your best single model?</p>\n<p>Mine is <code>mixnet_l</code> from <a href=\"https://rwightman.github.io/pytorch-image-models/results/\" target=\"_blank\">timm</a>, with standard 3 trajectories output, trained on <code>215 (epochs) x 64 (batch size) x 2500 (iterations)=34.4m</code> samples.</p>\n<ul>\n<li>Train loss: 11.48</li>\n<li>Validation: 12.02</li>\n<li>LB public: 12.285</li>\n<li>LB private: 11.484 (13th place)</li>\n</ul>\n<p>Slightly more fitted on the training data, which indicates that the training data is not infinite. I made sure that no sample point is attended twice, but the correlations between the samples are still significant which results in this lower training loss. But the validation and the LB were very similar along the way.</p>\n<p>Note that <code>mixnet_l</code> is the heavy one in the family. It gave me a big boost to move from <code>mixnet_s</code> to <code>mixnet_m</code> and then to <code>mixnet_l</code>, approximately <code>1</code> score point per change. This is in contradiction with several discussions on the forum that smaller models are better. Looking forward to learn what other models have worked.</p>",
  "messages": [
    {
      "id": 1091365,
      "postDate": "2020-11-26T00:53:01.233Z",
      "content": "<p>What is your best single model?</p>\n<p>Mine is <code>mixnet_l</code> from <a href=\"https://rwightman.github.io/pytorch-image-models/results/\" target=\"_blank\">timm</a>, with standard 3 trajectories output, trained on <code>215 (epochs) x 64 (batch size) x 2500 (iterations)=34.4m</code> samples.</p>\n<ul>\n<li>Train loss: 11.48</li>\n<li>Validation: 12.02</li>\n<li>LB public: 12.285</li>\n<li>LB private: 11.484 (13th place)</li>\n</ul>\n<p>Slightly more fitted on the training data, which indicates that the training data is not infinite. I made sure that no sample point is attended twice, but the correlations between the samples are still significant which results in this lower training loss. But the validation and the LB were very similar along the way.</p>\n<p>Note that <code>mixnet_l</code> is the heavy one in the family. It gave me a big boost to move from <code>mixnet_s</code> to <code>mixnet_m</code> and then to <code>mixnet_l</code>, approximately <code>1</code> score point per change. This is in contradiction with several discussions on the forum that smaller models are better. Looking forward to learn what other models have worked.</p>",
      "rawMarkdown": "What is your best single model?\n\nMine is `mixnet_l` from [timm](https://rwightman.github.io/pytorch-image-models/results/), with standard 3 trajectories output, trained on `215 (epochs) x 64 (batch size) x 2500 (iterations)=34.4m` samples.\n\n- Train loss: 11.48\n- Validation: 12.02\n- LB public: 12.285\n- LB private: 11.484 (13th place)\n\nSlightly more fitted on the training data, which indicates that the training data is not infinite. I made sure that no sample point is attended twice, but the correlations between the samples are still significant which results in this lower training loss. But the validation and the LB were very similar along the way.\n\nNote that `mixnet_l` is the heavy one in the family. It gave me a big boost to move from `mixnet_s` to `mixnet_m` and then to `mixnet_l`, approximately `1` score point per change. This is in contradiction with several discussions on the forum that smaller models are better. Looking forward to learn what other models have worked.\n",
      "votes": 21
    },
    {
      "id": 1091404,
      "postDate": "2020-11-26T01:56:13.583Z",
      "content": "<p>we have two single models (b3 &amp; b6), that each would be enough to win. Will talk more about them in our summary</p>",
      "rawMarkdown": "we have two single models (b3 & b6), that each would be enough to win. Will talk more about them in our summary",
      "votes": 9,
      "replies": [
        {
          "id": 1091723,
          "postDate": "2020-11-26T08:26:22.370Z",
          "content": "<p>Really interested in seeing how you got efficientnets to work! I fell flat in that regard, and couldn't figure out why they were faring so poorly…</p>",
          "rawMarkdown": "Really interested in seeing how you got efficientnets to work! I fell flat in that regard, and couldn't figure out why they were faring so poorly...",
          "votes": 1
        },
        {
          "id": 1091738,
          "postDate": "2020-11-26T08:41:18.377Z",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> in my experience, i found efficientnets to have tendency to converge slower than resnets. could that be the reason, that you didn't train long enough. Anyone with similar experience?</p>",
          "rawMarkdown": "@fergusoci in my experience, i found efficientnets to have tendency to converge slower than resnets. could that be the reason, that you didn't train long enough. Anyone with similar experience?"
        },
        {
          "id": 1096593,
          "postDate": "2020-11-30T16:16:44.243Z",
          "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> Do you mind if I ask you - how did you adapt your efficientnets to more than 3 input channels? I'm thinking that I might have missed something there. I was using something like:</p>\n<p>`<br>\nmodel = EfficientNet.from_name('eff-b0')</p>\n<p>model._conv_stem.in_channels = num_in_channels</p>\n<p>conv_stem_weight = torch.cat([model._conv_stem.weight.clone()] * int(math.ceil(num_in_channels/3)), axis=1)</p>\n<p>model._conv_stem.weight = torch.nn.Parameter(conv_stem_weight[:, :num_in_channels, :num_in_channels, :num_in_channels], requires_grad=True)<br>\n``<br>\n<a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> Maybe you also know the answer to this, given that you opted for mixnet.</p>\n<p>Cheers!</p>",
          "rawMarkdown": "@christofhenkel Do you mind if I ask you - how did you adapt your efficientnets to more than 3 input channels? I'm thinking that I might have missed something there. I was using something like:\n\n`\nmodel = EfficientNet.from_name('eff-b0')\n\nmodel._conv_stem.in_channels = num_in_channels\n\nconv_stem_weight = torch.cat([model._conv_stem.weight.clone()] * int(math.ceil(num_in_channels/3)), axis=1)\n\nmodel._conv_stem.weight = torch.nn.Parameter(conv_stem_weight[:, :num_in_channels, :num_in_channels, :num_in_channels], requires_grad=True)\n``\n@zaharch Maybe you also know the answer to this, given that you opted for mixnet.\n\nCheers!"
        },
        {
          "id": 1096608,
          "postDate": "2020-11-30T16:32:33.360Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1096717,
          "postDate": "2020-11-30T18:09:23.963Z",
          "content": "<p>Actually, I think the issue I may have had was the grad_fn for the conv_stem was not retained from the pretrained model. So effectively the graph was broken at that point. I wonder why it even trained moderately well in that case…</p>",
          "rawMarkdown": "Actually, I think the issue I may have had was the grad_fn for the conv_stem was not retained from the pretrained model. So effectively the graph was broken at that point. I wonder why it even trained moderately well in that case..."
        },
        {
          "id": 1096732,
          "postDate": "2020-11-30T18:25:24.193Z",
          "content": "<p>we used something like this. We used the efficientnet from timm</p>\n<pre><code>        elif \"efficientnet\" in cfg[\"model_params\"][\"model_architecture\"]:\n            self.backbone.conv_stem = Conv2dSame(\n                num_in_channels,\n                self.backbone.conv_stem.out_channels,\n                kernel_size=self.backbone.conv_stem.kernel_size,\n                stride=self.backbone.conv_stem.stride,\n                padding=self.backbone.conv_stem.padding,\n                bias=False,\n            )\n</code></pre>",
          "rawMarkdown": "we used something like this. We used the efficientnet from timm\n\n```\n        elif \"efficientnet\" in cfg[\"model_params\"][\"model_architecture\"]:\n            self.backbone.conv_stem = Conv2dSame(\n                num_in_channels,\n                self.backbone.conv_stem.out_channels,\n                kernel_size=self.backbone.conv_stem.kernel_size,\n                stride=self.backbone.conv_stem.stride,\n                padding=self.backbone.conv_stem.padding,\n                bias=False,\n            )\n```",
          "votes": 1
        },
        {
          "id": 1096741,
          "postDate": "2020-11-30T18:30:53.587Z",
          "content": "<p>Thanks. Yeah, I can see where I went wrong now. Appreciate the help. 👍</p>",
          "rawMarkdown": "Thanks. Yeah, I can see where I went wrong now. Appreciate the help. 👍"
        },
        {
          "id": 1096742,
          "postDate": "2020-11-30T18:31:45.943Z",
          "content": "<p>I am expecting your final solution. what is your diamond solution</p>",
          "rawMarkdown": "I am expecting your final solution. what is your diamond solution"
        },
        {
          "id": 1096757,
          "postDate": "2020-11-30T18:50:11.203Z",
          "content": "<p>we are still waiting on finalization of the LB…</p>",
          "rawMarkdown": "we are still waiting on finalization of the LB..."
        },
        {
          "id": 1096790,
          "postDate": "2020-11-30T19:28:08.227Z",
          "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> , I am curious, why do you wait for the finalization to post your summary?</p>",
          "rawMarkdown": "@christofhenkel , I am curious, why do you wait for the finalization to post your summary?"
        },
        {
          "id": 1097140,
          "postDate": "2020-12-01T00:48:12.447Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199226\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199226</a></p>\n<p>based on this it seems like it will be quite a while before that is finalized</p>",
          "rawMarkdown": "https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199226\n\nbased on this it seems like it will be quite a while before that is finalized"
        }
      ]
    },
    {
      "id": 1091376,
      "postDate": "2020-11-26T01:07:32.150Z",
      "content": "<p>The best single model for me was the xception41 from timm, trained on 224x224 resolution on the full dataset. I have not tried to submit it but on validation it reached around 10.37 after a week of training.</p>\n<p>The few days earlier checkpoint of the same model scored 11.1 on public LB, 10.2 on private and 11.1 on validation. The validation score was still dropping, but we switched to the full dataset only a week or so before deadline. Some heavier models seems to reach the lower score with the same number of epochs, but was slower to train.</p>",
      "rawMarkdown": "The best single model for me was the xception41 from timm, trained on 224x224 resolution on the full dataset. I have not tried to submit it but on validation it reached around 10.37 after a week of training.\n\nThe few days earlier checkpoint of the same model scored 11.1 on public LB, 10.2 on private and 11.1 on validation. The validation score was still dropping, but we switched to the full dataset only a week or so before deadline. Some heavier models seems to reach the lower score with the same number of epochs, but was slower to train.",
      "votes": 7,
      "replies": [
        {
          "id": 1091385,
          "postDate": "2020-11-26T01:14:15.963Z",
          "content": "<p>Wow, awesome results with a single model. Waiting to read your more detailed summary, interesting the list of changes that you applied to that model.</p>",
          "rawMarkdown": "Wow, awesome results with a single model. Waiting to read your more detailed summary, interesting the list of changes that you applied to that model.",
          "votes": 1
        },
        {
          "id": 1091386,
          "postDate": "2020-11-26T01:15:09.113Z",
          "content": "<p>thanks!</p>\n<p>\"after a week of training.\" … that's is a long time. i trained my models only for 2 days at most.<br>\nmaybe i should have trained longer.</p>\n<p>are are your history? are you using 10? are you training on chopped data or the original data?</p>",
          "rawMarkdown": "thanks!\n\n\"after a week of training.\" ... that's is a long time. i trained my models only for 2 days at most.\nmaybe i should have trained longer.\n\nare are your history? are you using 10? are you training on chopped data or the original data?",
          "votes": 1,
          "replies": [
            {
              "id": 1091531,
              "postDate": "2020-11-26T05:05:16.147Z",
              "content": "<p>We trained with min_history=0 and didn't chop the training dataset</p>",
              "rawMarkdown": "We trained with min_history=0 and didn't chop the training dataset",
              "votes": 1
            }
          ]
        },
        {
          "id": 1091645,
          "postDate": "2020-11-26T06:55:32.417Z",
          "content": "<p>We used the reduced range of history samples at 0, 1, 2, 4, 8 frame offsets.<br>\nWe trained on the original data but used min_history_frames set to 0 to make sure we are using harder frames for training, especially since low history frames are used in the chopped dataset.</p>\n<p>We found using the full dataset helps quite a bit to improve the validation score, but switched to it quite late in the competition.</p>",
          "rawMarkdown": "We used the reduced range of history samples at 0, 1, 2, 4, 8 frame offsets.\nWe trained on the original data but used min_history_frames set to 0 to make sure we are using harder frames for training, especially since low history frames are used in the chopped dataset.\n\nWe found using the full dataset helps quite a bit to improve the validation score, but switched to it quite late in the competition.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1091787,
      "postDate": "2020-11-26T09:31:00.780Z",
      "content": "<p>Best single model was resnet18 224x224 trained on chopped version of train_full:</p>\n<ul>\n<li>Valid loss: 12.321</li>\n<li>LB Public: 12.807</li>\n<li>LB Private: 11.626</li>\n</ul>\n<p>Final submission was just ensemble of checkpoints from this model.</p>",
      "rawMarkdown": "Best single model was resnet18 224x224 trained on chopped version of train_full:\n\n- Valid loss: 12.321\n- LB Public: 12.807\n- LB Private: 11.626\n\nFinal submission was just ensemble of checkpoints from this model.",
      "votes": 5,
      "replies": [
        {
          "id": 1091869,
          "postDate": "2020-11-26T10:52:18.357Z",
          "content": "<p>Impressive result for simple <code>resnet18</code>. Looks like simple models are very competitive after all.</p>",
          "rawMarkdown": "Impressive result for simple `resnet18`. Looks like simple models are very competitive after all.",
          "votes": 2
        },
        {
          "id": 1091875,
          "postDate": "2020-11-26T10:54:37.143Z",
          "content": "<p>I think the second winning model after EF is ResNet18 alone</p>",
          "rawMarkdown": "I think the second winning model after EF is ResNet18 alone"
        },
        {
          "id": 1091954,
          "postDate": "2020-11-26T12:18:08.087Z",
          "content": "<p>Well done! How many times did you iterate through your chopped dataset? And what history_num_frames did you opt for?</p>",
          "rawMarkdown": "Well done! How many times did you iterate through your chopped dataset? And what history_num_frames did you opt for?"
        },
        {
          "id": 1091974,
          "postDate": "2020-11-26T12:39:47.950Z",
          "content": "<p>The final model was trained using 10 chopped versions of train_full, with history set to 10. I pre generated the rasters and left it training for over a week on a single 2080ti, which was around 15 repeated passes of all chopped versions. I'll perhaps add a summary later today :)</p>",
          "rawMarkdown": "The final model was trained using 10 chopped versions of train_full, with history set to 10. I pre generated the rasters and left it training for over a week on a single 2080ti, which was around 15 repeated passes of all chopped versions. I'll perhaps add a summary later today :)",
          "votes": 1
        },
        {
          "id": 1091989,
          "postDate": "2020-11-26T12:56:56.423Z",
          "content": "<p>Interesting, I had found that the models stopped learning after about 12 views of the same image. Did you find any useful augmentation techniques?</p>",
          "rawMarkdown": "Interesting, I had found that the models stopped learning after about 12 views of the same image. Did you find any useful augmentation techniques?"
        }
      ]
    },
    {
      "id": 1091396,
      "postDate": "2020-11-26T01:29:43.917Z",
      "content": "<p>Mine is the baseline model Resnet34, batch_size 64, train on 24M samples</p>\n<ul>\n<li>train loss: 16.961 (complete training avg), 9.201 (last 200 batches avg)</li>\n<li>val loss: 18.310</li>\n<li>LB public: 17.549</li>\n<li>LB private: 16.981<br>\nLooks like this model is having a hard time on just the validation set :)</li>\n</ul>",
      "rawMarkdown": "Mine is the baseline model Resnet34, batch_size 64, train on 24M samples\n- train loss: 16.961 (complete training avg), 9.201 (last 200 batches avg)\n- val loss: 18.310\n- LB public: 17.549\n- LB private: 16.981\nLooks like this model is having a hard time on just the validation set :)",
      "votes": 5,
      "replies": [
        {
          "id": 1091878,
          "postDate": "2020-11-26T10:55:09.183Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1091389,
      "postDate": "2020-11-26T01:16:09.393Z",
      "content": "<p>Interesting. when you say no sample was attended to twice what exactly do you mean by that? How did you implement that? Like a sampler without replacement? Was there anything else special you did to get such low scores?  I stuck mostly with resnet18 and tried upping to resnet50, but saw virtually no difference. Maybe down to the sampling</p>",
      "rawMarkdown": "Interesting. when you say no sample was attended to twice what exactly do you mean by that? How did you implement that? Like a sampler without replacement? Was there anything else special you did to get such low scores?  I stuck mostly with resnet18 and tried upping to resnet50, but saw virtually no difference. Maybe down to the sampling",
      "votes": 3,
      "replies": [
        {
          "id": 1091905,
          "postDate": "2020-11-26T11:15:08.547Z",
          "content": "<p>Yes, it's like a sampler without replacement. I generated a permutation with a seed and picked indices from it sequentially. With a seed, - so I can stop and rerun it. </p>",
          "rawMarkdown": "Yes, it's like a sampler without replacement. I generated a permutation with a seed and picked indices from it sequentially. With a seed, - so I can stop and rerun it. "
        }
      ]
    },
    {
      "id": 1091879,
      "postDate": "2020-11-26T10:55:53.067Z",
      "content": "<p>Have Anyone tried with GeM Pooling i got a reasonable score in it better than before i had</p>",
      "rawMarkdown": "Have Anyone tried with GeM Pooling i got a reasonable score in it better than before i had",
      "votes": 1,
      "replies": [
        {
          "id": 1091907,
          "postDate": "2020-11-26T11:16:10.117Z",
          "content": "<p>Yes, our best single model uses GeM</p>",
          "rawMarkdown": "Yes, our best single model uses GeM",
          "votes": 7
        },
        {
          "id": 1098643,
          "postDate": "2020-12-01T18:49:46.220Z",
          "content": "<p>I wonder what is the p that you use for GeM pooling?</p>",
          "rawMarkdown": "I wonder what is the p that you use for GeM pooling?"
        },
        {
          "id": 1098779,
          "postDate": "2020-12-01T21:01:44.470Z",
          "content": "<p>Automatically trained, never checked how it ended up</p>",
          "rawMarkdown": "Automatically trained, never checked how it ended up"
        }
      ]
    },
    {
      "id": 1091537,
      "postDate": "2020-11-26T05:12:26.817Z",
      "content": "<p>I tried ResNeSt , ResNeXt, MixNet, B0,B3, MobileNetV3, ResNet34,18,50,101  , RegNet, ReXNet, SeResNext,and lot more and totally we had 24 models, <br>\nBut we didn't have any use of them since we cannot train any of these into a better one</p>",
      "rawMarkdown": "I tried ResNeSt , ResNeXt, MixNet, B0,B3, MobileNetV3, ResNet34,18,50,101  , RegNet, ReXNet, SeResNext,and lot more and totally we had 24 models, \nBut we didn't have any use of them since we cannot train any of these into a better one",
      "votes": 1
    },
    {
      "id": 1091453,
      "postDate": "2020-11-26T03:17:22.673Z",
      "content": "<p>Congrats! How did you select 3.44m sample? Is it coming from train zarr or full train zarr?</p>",
      "rawMarkdown": "Congrats! How did you select 3.44m sample? Is it coming from train zarr or full train zarr?",
      "votes": 1,
      "replies": [
        {
          "id": 1091855,
          "postDate": "2020-11-26T10:43:59.373Z",
          "content": "<p>It is a full train, I will give more info in my writeup, working on it</p>",
          "rawMarkdown": "It is a full train, I will give more info in my writeup, working on it"
        }
      ]
    },
    {
      "id": 1091372,
      "postDate": "2020-11-26T01:00:42.490Z",
      "content": "<p>good work!</p>\n<p>\"215x64x2500\", how do i interprete the dims?<br>\nwhat is the image size and history you used?</p>",
      "rawMarkdown": "good work!\n\n\"215x64x2500\", how do i interprete the dims?\nwhat is the image size and history you used?",
      "votes": 2,
      "replies": [
        {
          "id": 1091374,
          "postDate": "2020-11-26T01:03:19.823Z",
          "content": "<p>I added the dims description to the original post. Regarding the rest of the details, I will write a summary tomorrow! Need to sleep.</p>",
          "rawMarkdown": "I added the dims description to the original post. Regarding the rest of the details, I will write a summary tomorrow! Need to sleep.",
          "votes": 1
        },
        {
          "id": 1091382,
          "postDate": "2020-11-26T01:13:10.377Z",
          "content": "<p>Hmm… I thought people usually refer one epoch to one pass of all the training examples.</p>",
          "rawMarkdown": "Hmm... I thought people usually refer one epoch to one pass of all the training examples.",
          "votes": 2
        },
        {
          "id": 1091387,
          "postDate": "2020-11-26T01:15:28.907Z",
          "content": "<p>I was always bad with definitions, please excuse. My epoch is 2500 iterations in this case.</p>",
          "rawMarkdown": "I was always bad with definitions, please excuse. My epoch is 2500 iterations in this case.",
          "votes": 1
        },
        {
          "id": 1091392,
          "postDate": "2020-11-26T01:20:01.160Z",
          "content": "<p>Nevertheless, it is a very interesting idea! Glad the Mixnet works! <br>\nNow I got to learn some new models other than old fashion Resnet!</p>",
          "rawMarkdown": "Nevertheless, it is a very interesting idea! Glad the Mixnet works! \nNow I got to learn some new models other than old fashion Resnet!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1096982,
      "postDate": "2020-11-30T23:05:44.913Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> excellent work. Well done</p>",
      "rawMarkdown": "Hi @zaharch excellent work. Well done"
    },
    {
      "id": 1091910,
      "postDate": "2020-11-26T11:19:16.987Z",
      "content": "<p>The main solution for this competition i think need more CPU and High computing</p>",
      "rawMarkdown": "The main solution for this competition i think need more CPU and High computing"
    },
    {
      "id": 1091535,
      "postDate": "2020-11-26T05:09:37.610Z",
      "content": "<p>Look like we get some better score if we subtract <code>0.009</code> from transformed points in inference part<br>\nlike </p>\n<pre><code># convert into world coordinates and compute offsets\n        for idx in range(len(preds)):\n            for mode in range(3):\n                preds[idx, mode, :, :] = transform_points(preds[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]-0.009\n</code></pre>",
      "rawMarkdown": "Look like we get some better score if we subtract `0.009` from transformed points in inference part\nlike \n```        \n# convert into world coordinates and compute offsets\n        for idx in range(len(preds)):\n            for mode in range(3):\n                preds[idx, mode, :, :] = transform_points(preds[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]-0.009\n```"
    }
  ],
  "comments": [
    {
      "id": 1091404,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2020-11-26T01:56:13.583000",
      "content": "<p>we have two single models (b3 &amp; b6), that each would be enough to win. Will talk more about them in our summary</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1091723,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-11-26T08:26:22.370000",
          "content": "<p>Really interested in seeing how you got efficientnets to work! I fell flat in that regard, and couldn't figure out why they were faring so poorly…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091738,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-11-26T08:41:18.377000",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> in my experience, i found efficientnets to have tendency to converge slower than resnets. could that be the reason, that you didn't train long enough. Anyone with similar experience?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096593,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-11-30T16:16:44.243000",
          "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> Do you mind if I ask you - how did you adapt your efficientnets to more than 3 input channels? I'm thinking that I might have missed something there. I was using something like:</p>\n<p>`<br>\nmodel = EfficientNet.from_name('eff-b0')</p>\n<p>model._conv_stem.in_channels = num_in_channels</p>\n<p>conv_stem_weight = torch.cat([model._conv_stem.weight.clone()] * int(math.ceil(num_in_channels/3)), axis=1)</p>\n<p>model._conv_stem.weight = torch.nn.Parameter(conv_stem_weight[:, :num_in_channels, :num_in_channels, :num_in_channels], requires_grad=True)<br>\n``<br>\n<a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> Maybe you also know the answer to this, given that you opted for mixnet.</p>\n<p>Cheers!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096608,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-30T16:32:33.360000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096717,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-11-30T18:09:23.963000",
          "content": "<p>Actually, I think the issue I may have had was the grad_fn for the conv_stem was not retained from the pretrained model. So effectively the graph was broken at that point. I wonder why it even trained moderately well in that case…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096732,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-11-30T18:25:24.193000",
          "content": "<p>we used something like this. We used the efficientnet from timm</p>\n<pre><code>        elif \"efficientnet\" in cfg[\"model_params\"][\"model_architecture\"]:\n            self.backbone.conv_stem = Conv2dSame(\n                num_in_channels,\n                self.backbone.conv_stem.out_channels,\n                kernel_size=self.backbone.conv_stem.kernel_size,\n                stride=self.backbone.conv_stem.stride,\n                padding=self.backbone.conv_stem.padding,\n                bias=False,\n            )\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1096741,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-11-30T18:30:53.587000",
          "content": "<p>Thanks. Yeah, I can see where I went wrong now. Appreciate the help. 👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096742,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-11-30T18:31:45.943000",
          "content": "<p>I am expecting your final solution. what is your diamond solution</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096757,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-11-30T18:50:11.203000",
          "content": "<p>we are still waiting on finalization of the LB…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096790,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-11-30T19:28:08.227000",
          "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> , I am curious, why do you wait for the finalization to post your summary?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1097140,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-12-01T00:48:12.447000",
          "content": "<p><a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199226\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/199226</a></p>\n<p>based on this it seems like it will be quite a while before that is finalized</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091376,
      "author_name": "Dmytro Poplavskiy",
      "author_url": "",
      "post_date": "2020-11-26T01:07:32.150000",
      "content": "<p>The best single model for me was the xception41 from timm, trained on 224x224 resolution on the full dataset. I have not tried to submit it but on validation it reached around 10.37 after a week of training.</p>\n<p>The few days earlier checkpoint of the same model scored 11.1 on public LB, 10.2 on private and 11.1 on validation. The validation score was still dropping, but we switched to the full dataset only a week or so before deadline. Some heavier models seems to reach the lower score with the same number of epochs, but was slower to train.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1091385,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-11-26T01:14:15.963000",
          "content": "<p>Wow, awesome results with a single model. Waiting to read your more detailed summary, interesting the list of changes that you applied to that model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091386,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-11-26T01:15:09.113000",
          "content": "<p>thanks!</p>\n<p>\"after a week of training.\" … that's is a long time. i trained my models only for 2 days at most.<br>\nmaybe i should have trained longer.</p>\n<p>are are your history? are you using 10? are you training on chopped data or the original data?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1091531,
              "author_name": "Artem.Sanakoev",
              "author_url": "",
              "post_date": "2020-11-26T05:05:16.147000",
              "content": "<p>We trained with min_history=0 and didn't chop the training dataset</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1091645,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2020-11-26T06:55:32.417000",
          "content": "<p>We used the reduced range of history samples at 0, 1, 2, 4, 8 frame offsets.<br>\nWe trained on the original data but used min_history_frames set to 0 to make sure we are using harder frames for training, especially since low history frames are used in the chopped dataset.</p>\n<p>We found using the full dataset helps quite a bit to improve the validation score, but switched to it quite late in the competition.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1091787,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2020-11-26T09:31:00.780000",
      "content": "<p>Best single model was resnet18 224x224 trained on chopped version of train_full:</p>\n<ul>\n<li>Valid loss: 12.321</li>\n<li>LB Public: 12.807</li>\n<li>LB Private: 11.626</li>\n</ul>\n<p>Final submission was just ensemble of checkpoints from this model.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1091869,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-11-26T10:52:18.357000",
          "content": "<p>Impressive result for simple <code>resnet18</code>. Looks like simple models are very competitive after all.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1091875,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-11-26T10:54:37.143000",
          "content": "<p>I think the second winning model after EF is ResNet18 alone</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1091954,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-11-26T12:18:08.087000",
          "content": "<p>Well done! How many times did you iterate through your chopped dataset? And what history_num_frames did you opt for?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1091974,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2020-11-26T12:39:47.950000",
          "content": "<p>The final model was trained using 10 chopped versions of train_full, with history set to 10. I pre generated the rasters and left it training for over a week on a single 2080ti, which was around 15 repeated passes of all chopped versions. I'll perhaps add a summary later today :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091989,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-11-26T12:56:56.423000",
          "content": "<p>Interesting, I had found that the models stopped learning after about 12 views of the same image. Did you find any useful augmentation techniques?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091396,
      "author_name": "Louis Yang",
      "author_url": "",
      "post_date": "2020-11-26T01:29:43.917000",
      "content": "<p>Mine is the baseline model Resnet34, batch_size 64, train on 24M samples</p>\n<ul>\n<li>train loss: 16.961 (complete training avg), 9.201 (last 200 batches avg)</li>\n<li>val loss: 18.310</li>\n<li>LB public: 17.549</li>\n<li>LB private: 16.981<br>\nLooks like this model is having a hard time on just the validation set :)</li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 1091878,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-26T10:55:09.183000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091389,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-11-26T01:16:09.393000",
      "content": "<p>Interesting. when you say no sample was attended to twice what exactly do you mean by that? How did you implement that? Like a sampler without replacement? Was there anything else special you did to get such low scores?  I stuck mostly with resnet18 and tried upping to resnet50, but saw virtually no difference. Maybe down to the sampling</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1091905,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-11-26T11:15:08.547000",
          "content": "<p>Yes, it's like a sampler without replacement. I generated a permutation with a seed and picked indices from it sequentially. With a seed, - so I can stop and rerun it. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091879,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-11-26T10:55:53.067000",
      "content": "<p>Have Anyone tried with GeM Pooling i got a reasonable score in it better than before i had</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1091907,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-11-26T11:16:10.117000",
          "content": "<p>Yes, our best single model uses GeM</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1098643,
          "author_name": "Louis Yang",
          "author_url": "",
          "post_date": "2020-12-01T18:49:46.220000",
          "content": "<p>I wonder what is the p that you use for GeM pooling?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1098779,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-12-01T21:01:44.470000",
          "content": "<p>Automatically trained, never checked how it ended up</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091537,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-11-26T05:12:26.817000",
      "content": "<p>I tried ResNeSt , ResNeXt, MixNet, B0,B3, MobileNetV3, ResNet34,18,50,101  , RegNet, ReXNet, SeResNext,and lot more and totally we had 24 models, <br>\nBut we didn't have any use of them since we cannot train any of these into a better one</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1091453,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-11-26T03:17:22.673000",
      "content": "<p>Congrats! How did you select 3.44m sample? Is it coming from train zarr or full train zarr?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1091855,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-11-26T10:43:59.373000",
          "content": "<p>It is a full train, I will give more info in my writeup, working on it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091372,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-11-26T01:00:42.490000",
      "content": "<p>good work!</p>\n<p>\"215x64x2500\", how do i interprete the dims?<br>\nwhat is the image size and history you used?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1091374,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-11-26T01:03:19.823000",
          "content": "<p>I added the dims description to the original post. Regarding the rest of the details, I will write a summary tomorrow! Need to sleep.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091382,
          "author_name": "Louis Yang",
          "author_url": "",
          "post_date": "2020-11-26T01:13:10.377000",
          "content": "<p>Hmm… I thought people usually refer one epoch to one pass of all the training examples.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1091387,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-11-26T01:15:28.907000",
          "content": "<p>I was always bad with definitions, please excuse. My epoch is 2500 iterations in this case.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091392,
          "author_name": "Louis Yang",
          "author_url": "",
          "post_date": "2020-11-26T01:20:01.160000",
          "content": "<p>Nevertheless, it is a very interesting idea! Glad the Mixnet works! <br>\nNow I got to learn some new models other than old fashion Resnet!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1096982,
      "author_name": "Patrick Uzuwe",
      "author_url": "",
      "post_date": "2020-11-30T23:05:44.913000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> excellent work. Well done</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1091910,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-11-26T11:19:16.987000",
      "content": "<p>The main solution for this competition i think need more CPU and High computing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1091535,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-11-26T05:09:37.610000",
      "content": "<p>Look like we get some better score if we subtract <code>0.009</code> from transformed points in inference part<br>\nlike </p>\n<pre><code># convert into world coordinates and compute offsets\n        for idx in range(len(preds)):\n            for mode in range(3):\n                preds[idx, mode, :, :] = transform_points(preds[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]-0.009\n</code></pre>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1091365": "What is your best single model?\n\nMine is `mixnet_l` from [timm](https://rwightman.github.io/pytorch-image-models/results/), with standard 3 trajectories output, trained on `215 (epochs) x 64 (batch size) x 2500 (iterations)=34.4m` samples.\n\n- Train loss: 11.48\n- Validation: 12.02\n- LB public: 12.285\n- LB private: 11.484 (13th place)\n\nSlightly more fitted on the training data, which indicates that the training data is not infinite. I made sure that no sample point is attended twice, but the correlations between the samples are still significant which results in this lower training loss. But the validation and the LB were very similar along the way.\n\nNote that `mixnet_l` is the heavy one in the family. It gave me a big boost to move from `mixnet_s` to `mixnet_m` and then to `mixnet_l`, approximately `1` score point per change. This is in contradiction with several discussions on the forum that smaller models are better. Looking forward to learn what other models have worked.\n",
    "1091404": "we have two single models (b3 & b6), that each would be enough to win. Will talk more about them in our summary",
    "1091376": "The best single model for me was the xception41 from timm, trained on 224x224 resolution on the full dataset. I have not tried to submit it but on validation it reached around 10.37 after a week of training.\n\nThe few days earlier checkpoint of the same model scored 11.1 on public LB, 10.2 on private and 11.1 on validation. The validation score was still dropping, but we switched to the full dataset only a week or so before deadline. Some heavier models seems to reach the lower score with the same number of epochs, but was slower to train.",
    "1091787": "Best single model was resnet18 224x224 trained on chopped version of train_full:\n\n- Valid loss: 12.321\n- LB Public: 12.807\n- LB Private: 11.626\n\nFinal submission was just ensemble of checkpoints from this model.",
    "1091396": "Mine is the baseline model Resnet34, batch_size 64, train on 24M samples\n- train loss: 16.961 (complete training avg), 9.201 (last 200 batches avg)\n- val loss: 18.310\n- LB public: 17.549\n- LB private: 16.981\nLooks like this model is having a hard time on just the validation set :)",
    "1091389": "Interesting. when you say no sample was attended to twice what exactly do you mean by that? How did you implement that? Like a sampler without replacement? Was there anything else special you did to get such low scores?  I stuck mostly with resnet18 and tried upping to resnet50, but saw virtually no difference. Maybe down to the sampling",
    "1091879": "Have Anyone tried with GeM Pooling i got a reasonable score in it better than before i had",
    "1091537": "I tried ResNeSt , ResNeXt, MixNet, B0,B3, MobileNetV3, ResNet34,18,50,101  , RegNet, ReXNet, SeResNext,and lot more and totally we had 24 models, \nBut we didn't have any use of them since we cannot train any of these into a better one",
    "1091453": "Congrats! How did you select 3.44m sample? Is it coming from train zarr or full train zarr?",
    "1091372": "good work!\n\n\"215x64x2500\", how do i interprete the dims?\nwhat is the image size and history you used?",
    "1096982": "Hi @zaharch excellent work. Well done",
    "1091910": "The main solution for this competition i think need more CPU and High computing",
    "1091535": "Look like we get some better score if we subtract `0.009` from transformed points in inference part\nlike \n```        \n# convert into world coordinates and compute offsets\n        for idx in range(len(preds)):\n            for mode in range(3):\n                preds[idx, mode, :, :] = transform_points(preds[idx, mode, :, :], world_from_agents[idx]) - centroids[idx][:2]-0.009\n```"
  }
}