{
  "id": 199523,
  "title": "Random undersampling by cumulative distance - maybe it doesn't seem to work",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199523",
  "author_name": "",
  "post_date": "2020-11-26T04:13:11.655634900Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>After I read discussion of <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>, Improving convergence and sample efficiency<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814</a></p>\n<p>I tried to calculate cumulative distance of target_positions.</p>\n<p>train.zarr with min_frame_future=10<br>\nlength of AgentDataset: 17,003,687</p>\n<p>This plot is sample distribution about non-zero-end of target_positions. number of samples are 850,184 (1/20 of AgentDatast).  Most sample has 50 target_positions data.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F95718930bd30d9c7b54044db4449359c%2Fplot1.png?generation=1606362496690194&amp;alt=media\" alt=\"\"></p>\n<p>And histogram of cumulative distance is as follows. (number of samples are 17,003,687)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F4618fd472d51e0cda5ab12a44fab159e%2Fplot2.png?generation=1606362515038305&amp;alt=media\" alt=\"\"></p>\n<p>And this is cumulative graph.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F422cd1396d43aee8a97963a6713fb017%2Fplot3.png?generation=1606362532269209&amp;alt=media\" alt=\"\"></p>\n<p>About 60% of data have less than 3 of cumulative distance.</p>\n<table>\n<thead>\n<tr>\n<th>distance</th>\n<th>percentile</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 ~ 1</td>\n<td>0.3108957002248925</td>\n</tr>\n<tr>\n<td>1 ~ 2</td>\n<td>0.15160976917937763</td>\n</tr>\n<tr>\n<td>2 ~ 3</td>\n<td>0.08709361738164917</td>\n</tr>\n<tr>\n<td>3 ~ 4</td>\n<td>0.057570243617852124</td>\n</tr>\n<tr>\n<td>4 ~ 5</td>\n<td>0.030621018508934505</td>\n</tr>\n<tr>\n<td>5 ~ 6</td>\n<td>0.01739799855090196</td>\n</tr>\n<tr>\n<td>6 ~ 7</td>\n<td>0.011909304338825422</td>\n</tr>\n<tr>\n<td>7 ~ 8</td>\n<td>0.009271287156662589</td>\n</tr>\n<tr>\n<td>8 ~ 9</td>\n<td>0.007785491140741341</td>\n</tr>\n<tr>\n<td>9 ~ 10</td>\n<td>0.006836990580862512</td>\n</tr>\n</tbody>\n</table>\n<p>So i undersampled and made mask of data which of cumulative distance is less than 10 randomly by the inverse of percentile.</p>\n<table>\n<thead>\n<tr>\n<th>distance</th>\n<th>ratio</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 ~ 1</td>\n<td>0.021991267733573804</td>\n</tr>\n<tr>\n<td>1 ~ 2</td>\n<td>0.04509597645237031</td>\n</tr>\n<tr>\n<td>2 ~ 3</td>\n<td>0.0785016260250442</td>\n</tr>\n<tr>\n<td>3 ~ 4</td>\n<td>0.1187591045514077</td>\n</tr>\n<tr>\n<td>4 ~ 5</td>\n<td>0.22327769988668522</td>\n</tr>\n<tr>\n<td>5 ~ 6</td>\n<td>0.3929756955007945</td>\n</tr>\n<tr>\n<td>6 ~ 7</td>\n<td>0.5740881571540075</td>\n</tr>\n<tr>\n<td>7 ~ 8</td>\n<td>0.7374370424875981</td>\n</tr>\n<tr>\n<td>8 ~ 9</td>\n<td>0.8781707482890396</td>\n</tr>\n<tr>\n<td>9 ~ 10</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<p>This is histogram of undersampled data.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F77659f08f6f965d19c7edceddf6ed204%2Fplot4.png?generation=1606362629250668&amp;alt=media\" alt=\"\"></p>\n<p>The total length of undersampled data is 6,415,723. It is 37.7% of original data.</p>\n<p>And I trained with this 6M data on resnet18 for 4 epochs. <br>\ntrain loss is average of 2k iter(2k x 16 = 32k samples)</p>\n<table>\n<thead>\n<tr>\n<th>epoch</th>\n<th>tr_loss</th>\n<th>val_loss</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>25.582</td>\n<td>23.814</td>\n<td>22.676</td>\n</tr>\n<tr>\n<td>2</td>\n<td>20.947</td>\n<td>21.242</td>\n<td>20.529</td>\n</tr>\n<tr>\n<td>3</td>\n<td>18.935</td>\n<td>21.900</td>\n<td>20.585</td>\n</tr>\n<tr>\n<td>4</td>\n<td>17.359</td>\n<td>21.878</td>\n<td>20.484</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>parameters</th>\n<th>values</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>raster_size</td>\n<td>267, 267</td>\n</tr>\n<tr>\n<td>pixel_size</td>\n<td>0.5, 0.5</td>\n</tr>\n<tr>\n<td>history_num_frames</td>\n<td>1</td>\n</tr>\n<tr>\n<td>history_delta_time</td>\n<td>1</td>\n</tr>\n<tr>\n<td>batch_size</td>\n<td>16</td>\n</tr>\n<tr>\n<td>AdamW</td>\n<td>lr=1.0e-4, weight_decay=0.01</td>\n</tr>\n<tr>\n<td>CosineAnnealingWarmRestarts</td>\n<td>T_0=80k, T_mult=1, eta_min=1.0e-5</td>\n</tr>\n</tbody>\n</table>\n<p>I think I may have trained in wrong way. I used quite large weight_decay. Anyway training one epoch took ~17 hours on GCE with vCPU:16, GPU: T4 and ~13 hours with vCPU:24, GPU: T4.</p>\n<p>I guess random undersampling by cumulative distance is not good idea. The information losses may be quite significant. Samples of small cumulative distance may have various trajectories. I think reducing overlapped sample is more appropriate. I wanted to do more experiment but I couldn't because of lack of time.</p>\n<p>And my final submission is just resnet18 with train.zarr, no undersampling. I also tried mobilenetV2 but resnet18 is better.</p>\n<table>\n<thead>\n<tr>\n<th>epoch</th>\n<th>private LB</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>19.101</td>\n<td>19.886</td>\n</tr>\n<tr>\n<td>2</td>\n<td>18.185</td>\n<td>18.675</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>parameters</th>\n<th>values</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>raster_size</td>\n<td>267, 267</td>\n</tr>\n<tr>\n<td>pixel_size</td>\n<td>0.5, 0.5</td>\n</tr>\n<tr>\n<td>history_num_frames</td>\n<td>1</td>\n</tr>\n<tr>\n<td>history_delta_time</td>\n<td>1</td>\n</tr>\n<tr>\n<td>batch_size</td>\n<td>16</td>\n</tr>\n<tr>\n<td>AdamW</td>\n<td>weight_decay=2.43e-5</td>\n</tr>\n<tr>\n<td>first epoch with OneCycleLR</td>\n<td>max_lr=1.0e-4, div_factor=10</td>\n</tr>\n<tr>\n<td>second epoch with CosineAnnealing</td>\n<td>lr=1.0e-5, eta_min=1.0e-6</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "1091489",
      "postDate": "11/26/2020 04:13:11",
      "content": "<p>After I read discussion of <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>, Improving convergence and sample efficiency<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814</a></p>\n<p>I tried to calculate cumulative distance of target_positions.</p>\n<p>train.zarr with min_frame_future=10<br>\nlength of AgentDataset: 17,003,687</p>\n<p>This plot is sample distribution about non-zero-end of target_positions. number of samples are 850,184 (1/20 of AgentDatast).  Most sample has 50 target_positions data.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F95718930bd30d9c7b54044db4449359c%2Fplot1.png?generation=1606362496690194&amp;alt=media\" alt=\"\"></p>\n<p>And histogram of cumulative distance is as follows. (number of samples are 17,003,687)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F4618fd472d51e0cda5ab12a44fab159e%2Fplot2.png?generation=1606362515038305&amp;alt=media\" alt=\"\"></p>\n<p>And this is cumulative graph.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F422cd1396d43aee8a97963a6713fb017%2Fplot3.png?generation=1606362532269209&amp;alt=media\" alt=\"\"></p>\n<p>About 60% of data have less than 3 of cumulative distance.</p>\n<table>\n<thead>\n<tr>\n<th>distance</th>\n<th>percentile</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 ~ 1</td>\n<td>0.3108957002248925</td>\n</tr>\n<tr>\n<td>1 ~ 2</td>\n<td>0.15160976917937763</td>\n</tr>\n<tr>\n<td>2 ~ 3</td>\n<td>0.08709361738164917</td>\n</tr>\n<tr>\n<td>3 ~ 4</td>\n<td>0.057570243617852124</td>\n</tr>\n<tr>\n<td>4 ~ 5</td>\n<td>0.030621018508934505</td>\n</tr>\n<tr>\n<td>5 ~ 6</td>\n<td>0.01739799855090196</td>\n</tr>\n<tr>\n<td>6 ~ 7</td>\n<td>0.011909304338825422</td>\n</tr>\n<tr>\n<td>7 ~ 8</td>\n<td>0.009271287156662589</td>\n</tr>\n<tr>\n<td>8 ~ 9</td>\n<td>0.007785491140741341</td>\n</tr>\n<tr>\n<td>9 ~ 10</td>\n<td>0.006836990580862512</td>\n</tr>\n</tbody>\n</table>\n<p>So i undersampled and made mask of data which of cumulative distance is less than 10 randomly by the inverse of percentile.</p>\n<table>\n<thead>\n<tr>\n<th>distance</th>\n<th>ratio</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 ~ 1</td>\n<td>0.021991267733573804</td>\n</tr>\n<tr>\n<td>1 ~ 2</td>\n<td>0.04509597645237031</td>\n</tr>\n<tr>\n<td>2 ~ 3</td>\n<td>0.0785016260250442</td>\n</tr>\n<tr>\n<td>3 ~ 4</td>\n<td>0.1187591045514077</td>\n</tr>\n<tr>\n<td>4 ~ 5</td>\n<td>0.22327769988668522</td>\n</tr>\n<tr>\n<td>5 ~ 6</td>\n<td>0.3929756955007945</td>\n</tr>\n<tr>\n<td>6 ~ 7</td>\n<td>0.5740881571540075</td>\n</tr>\n<tr>\n<td>7 ~ 8</td>\n<td>0.7374370424875981</td>\n</tr>\n<tr>\n<td>8 ~ 9</td>\n<td>0.8781707482890396</td>\n</tr>\n<tr>\n<td>9 ~ 10</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<p>This is histogram of undersampled data.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F77659f08f6f965d19c7edceddf6ed204%2Fplot4.png?generation=1606362629250668&amp;alt=media\" alt=\"\"></p>\n<p>The total length of undersampled data is 6,415,723. It is 37.7% of original data.</p>\n<p>And I trained with this 6M data on resnet18 for 4 epochs. <br>\ntrain loss is average of 2k iter(2k x 16 = 32k samples)</p>\n<table>\n<thead>\n<tr>\n<th>epoch</th>\n<th>tr_loss</th>\n<th>val_loss</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>25.582</td>\n<td>23.814</td>\n<td>22.676</td>\n</tr>\n<tr>\n<td>2</td>\n<td>20.947</td>\n<td>21.242</td>\n<td>20.529</td>\n</tr>\n<tr>\n<td>3</td>\n<td>18.935</td>\n<td>21.900</td>\n<td>20.585</td>\n</tr>\n<tr>\n<td>4</td>\n<td>17.359</td>\n<td>21.878</td>\n<td>20.484</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>parameters</th>\n<th>values</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>raster_size</td>\n<td>267, 267</td>\n</tr>\n<tr>\n<td>pixel_size</td>\n<td>0.5, 0.5</td>\n</tr>\n<tr>\n<td>history_num_frames</td>\n<td>1</td>\n</tr>\n<tr>\n<td>history_delta_time</td>\n<td>1</td>\n</tr>\n<tr>\n<td>batch_size</td>\n<td>16</td>\n</tr>\n<tr>\n<td>AdamW</td>\n<td>lr=1.0e-4, weight_decay=0.01</td>\n</tr>\n<tr>\n<td>CosineAnnealingWarmRestarts</td>\n<td>T_0=80k, T_mult=1, eta_min=1.0e-5</td>\n</tr>\n</tbody>\n</table>\n<p>I think I may have trained in wrong way. I used quite large weight_decay. Anyway training one epoch took ~17 hours on GCE with vCPU:16, GPU: T4 and ~13 hours with vCPU:24, GPU: T4.</p>\n<p>I guess random undersampling by cumulative distance is not good idea. The information losses may be quite significant. Samples of small cumulative distance may have various trajectories. I think reducing overlapped sample is more appropriate. I wanted to do more experiment but I couldn't because of lack of time.</p>\n<p>And my final submission is just resnet18 with train.zarr, no undersampling. I also tried mobilenetV2 but resnet18 is better.</p>\n<table>\n<thead>\n<tr>\n<th>epoch</th>\n<th>private LB</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>19.101</td>\n<td>19.886</td>\n</tr>\n<tr>\n<td>2</td>\n<td>18.185</td>\n<td>18.675</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>parameters</th>\n<th>values</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>raster_size</td>\n<td>267, 267</td>\n</tr>\n<tr>\n<td>pixel_size</td>\n<td>0.5, 0.5</td>\n</tr>\n<tr>\n<td>history_num_frames</td>\n<td>1</td>\n</tr>\n<tr>\n<td>history_delta_time</td>\n<td>1</td>\n</tr>\n<tr>\n<td>batch_size</td>\n<td>16</td>\n</tr>\n<tr>\n<td>AdamW</td>\n<td>weight_decay=2.43e-5</td>\n</tr>\n<tr>\n<td>first epoch with OneCycleLR</td>\n<td>max_lr=1.0e-4, div_factor=10</td>\n</tr>\n<tr>\n<td>second epoch with CosineAnnealing</td>\n<td>lr=1.0e-5, eta_min=1.0e-6</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "After I read discussion of @ryches, Improving convergence and sample efficiency\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\n\nI tried to calculate cumulative distance of target_positions.\n\ntrain.zarr with min_frame_future=10\nlength of AgentDataset: 17,003,687\n\nThis plot is sample distribution about non-zero-end of target_positions. number of samples are 850,184 (1/20 of AgentDatast).  Most sample has 50 target_positions data.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F95718930bd30d9c7b54044db4449359c%2Fplot1.png?generation=1606362496690194&alt=media)\n\n\nAnd histogram of cumulative distance is as follows. (number of samples are 17,003,687)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F4618fd472d51e0cda5ab12a44fab159e%2Fplot2.png?generation=1606362515038305&alt=media)\n\nAnd this is cumulative graph.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F422cd1396d43aee8a97963a6713fb017%2Fplot3.png?generation=1606362532269209&alt=media)\n\nAbout 60% of data have less than 3 of cumulative distance.\n\n|distance|percentile|\n|---|---|\n|0 ~ 1|0.3108957002248925|\n|1 ~ 2|0.15160976917937763|\n|2 ~ 3|0.08709361738164917|\n|3 ~ 4|0.057570243617852124|\n|4 ~ 5|0.030621018508934505|\n|5 ~ 6|0.01739799855090196|\n|6 ~ 7|0.011909304338825422|\n|7 ~ 8|0.009271287156662589|\n|8 ~ 9|0.007785491140741341|\n|9 ~ 10|0.006836990580862512|\n\n\nSo i undersampled and made mask of data which of cumulative distance is less than 10 randomly by the inverse of percentile.\n\n|distance|ratio|\n|---|---|\n|0 ~ 1|0.021991267733573804|\n|1 ~ 2|0.04509597645237031|\n|2 ~ 3|0.0785016260250442|\n|3 ~ 4|0.1187591045514077|\n|4 ~ 5|0.22327769988668522|\n|5 ~ 6|0.3929756955007945|\n|6 ~ 7|0.5740881571540075|\n|7 ~ 8|0.7374370424875981|\n|8 ~ 9|0.8781707482890396|\n|9 ~ 10|1.0|\n\n\nThis is histogram of undersampled data.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F77659f08f6f965d19c7edceddf6ed204%2Fplot4.png?generation=1606362629250668&alt=media)\n\nThe total length of undersampled data is 6,415,723. It is 37.7% of original data.\n\nAnd I trained with this 6M data on resnet18 for 4 epochs. \ntrain loss is average of 2k iter(2k x 16 = 32k samples)\n\n|epoch|tr_loss|val_loss|LB|\n| --- | --- | --- | --- |\n|1|25.582|23.814|22.676|\n|2|20.947|21.242|20.529|\n|3|18.935|21.900|20.585|\n|4|17.359|21.878|20.484|\n\n|parameters|values|\n|---|---|\n|raster_size|267, 267|\n|pixel_size|0.5, 0.5|\n|history_num_frames|1|\n|history_delta_time|1|\n|batch_size|16|\n|AdamW|lr=1.0e-4, weight_decay=0.01|\n|CosineAnnealingWarmRestarts|T_0=80k, T_mult=1, eta_min=1.0e-5|\n\n\nI think I may have trained in wrong way. I used quite large weight_decay. Anyway training one epoch took ~17 hours on GCE with vCPU:16, GPU: T4 and ~13 hours with vCPU:24, GPU: T4.\n\nI guess random undersampling by cumulative distance is not good idea. The information losses may be quite significant. Samples of small cumulative distance may have various trajectories. I think reducing overlapped sample is more appropriate. I wanted to do more experiment but I couldn't because of lack of time.\n\nAnd my final submission is just resnet18 with train.zarr, no undersampling. I also tried mobilenetV2 but resnet18 is better.\n\n|epoch|private LB|public LB\n|---|---|---|\n|1|19.101|19.886|\n|2|18.185|18.675|\n\n|parameters|values|\n|---|---|\n|raster_size|267, 267|\n|pixel_size|0.5, 0.5|\n|history_num_frames|1|\n|history_delta_time|1|\n|batch_size|16|\n|AdamW|weight_decay=2.43e-5|\n|first epoch with OneCycleLR|max_lr=1.0e-4, div_factor=10|\n|second epoch with CosineAnnealing|lr=1.0e-5, eta_min=1.0e-6|",
      "votes": null
    },
    {
      "id": "1091493",
      "postDate": "11/26/2020 04:19:14",
      "content": "<p>I had similar results. Any kind of resampling of the data caused worse performance. </p>",
      "rawMarkdown": "I had similar results. Any kind of resampling of the data caused worse performance.",
      "votes": null
    },
    {
      "id": "1091501",
      "postDate": "11/26/2020 04:25:44",
      "content": "<p>I agree. Undersampling doesn't guarantee better result.</p>",
      "rawMarkdown": "I agree. Undersampling doesn't guarantee better result.",
      "votes": null
    },
    {
      "id": "1091597",
      "postDate": "11/26/2020 06:07:30",
      "content": "<p>I have also try oversample which also don't work too well. I guess the reason could be that to learn those outlier, you will impact the learning on other examples too much. Or, I simply didn't train the entire data set. Perhaps, if you train the entire epoch, it will be OK.</p>",
      "rawMarkdown": "I have also try oversample which also don't work too well. I guess the reason could be that to learn those outlier, you will impact the learning on other examples too much. Or, I simply didn't train the entire data set. Perhaps, if you train the entire epoch, it will be OK.",
      "votes": null
    },
    {
      "id": "1091599",
      "postDate": "11/26/2020 06:09:23",
      "content": "<p>I think batch size also matter a lot. I was using batch_size = 16 and couldn't get pass the 23 threshold. Switching to 64 can bring you lower loss when train more data.</p>",
      "rawMarkdown": "I think batch size also matter a lot. I was using batch_size = 16 and couldn't get pass the 23 threshold. Switching to 64 can bring you lower loss when train more data.",
      "votes": null
    },
    {
      "id": "1091624",
      "postDate": "11/26/2020 06:25:33",
      "content": "<p>I inspected the high loss samples quite a bit and it didn't seem like there was a great deal of systemic problems. So training on one outlier didn't seem to be particularly likely to reduce loss on other outliers. I tried an augmentation where I shifted a single channel out of the full history or surrounding vehicles by a couple pixels in a random direction and often found that that could shift a prediction from 1k+ loss down to under 50</p>",
      "rawMarkdown": "I inspected the high loss samples quite a bit and it didn't seem like there was a great deal of systemic problems. So training on one outlier didn't seem to be particularly likely to reduce loss on other outliers. I tried an augmentation where I shifted a single channel out of the full history or surrounding vehicles by a couple pixels in a random direction and often found that that could shift a prediction from 1k+ loss down to under 50",
      "votes": null
    },
    {
      "id": "1091626",
      "postDate": "11/26/2020 06:26:54",
      "content": "<p>It was interesting to see the system so fragile to this augmentation, I presume the error of the perception system could cause these small single pixel shifts so making it robust to this would probably be important. </p>",
      "rawMarkdown": "It was interesting to see the system so fragile to this augmentation, I presume the error of the perception system could cause these small single pixel shifts so making it robust to this would probably be important.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091493,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "11/26/2020 04:19:14",
      "content": "<p>I had similar results. Any kind of resampling of the data caused worse performance. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1091501,
          "author_name": "sunghyunjun",
          "author_url": "",
          "post_date": "11/26/2020 04:25:44",
          "content": "<p>I agree. Undersampling doesn't guarantee better result.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091597,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/26/2020 06:07:30",
          "content": "<p>I have also try oversample which also don't work too well. I guess the reason could be that to learn those outlier, you will impact the learning on other examples too much. Or, I simply didn't train the entire data set. Perhaps, if you train the entire epoch, it will be OK.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091624,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/26/2020 06:25:33",
          "content": "<p>I inspected the high loss samples quite a bit and it didn't seem like there was a great deal of systemic problems. So training on one outlier didn't seem to be particularly likely to reduce loss on other outliers. I tried an augmentation where I shifted a single channel out of the full history or surrounding vehicles by a couple pixels in a random direction and often found that that could shift a prediction from 1k+ loss down to under 50</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091626,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/26/2020 06:26:54",
          "content": "<p>It was interesting to see the system so fragile to this augmentation, I presume the error of the perception system could cause these small single pixel shifts so making it robust to this would probably be important. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091599,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/26/2020 06:09:23",
      "content": "<p>I think batch size also matter a lot. I was using batch_size = 16 and couldn't get pass the 23 threshold. Switching to 64 can bring you lower loss when train more data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1091489": "After I read discussion of @ryches, Improving convergence and sample efficiency\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\n\nI tried to calculate cumulative distance of target_positions.\n\ntrain.zarr with min_frame_future=10\nlength of AgentDataset: 17,003,687\n\nThis plot is sample distribution about non-zero-end of target_positions. number of samples are 850,184 (1/20 of AgentDatast).  Most sample has 50 target_positions data.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F95718930bd30d9c7b54044db4449359c%2Fplot1.png?generation=1606362496690194&alt=media)\n\n\nAnd histogram of cumulative distance is as follows. (number of samples are 17,003,687)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F4618fd472d51e0cda5ab12a44fab159e%2Fplot2.png?generation=1606362515038305&alt=media)\n\nAnd this is cumulative graph.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F422cd1396d43aee8a97963a6713fb017%2Fplot3.png?generation=1606362532269209&alt=media)\n\nAbout 60% of data have less than 3 of cumulative distance.\n\n|distance|percentile|\n|---|---|\n|0 ~ 1|0.3108957002248925|\n|1 ~ 2|0.15160976917937763|\n|2 ~ 3|0.08709361738164917|\n|3 ~ 4|0.057570243617852124|\n|4 ~ 5|0.030621018508934505|\n|5 ~ 6|0.01739799855090196|\n|6 ~ 7|0.011909304338825422|\n|7 ~ 8|0.009271287156662589|\n|8 ~ 9|0.007785491140741341|\n|9 ~ 10|0.006836990580862512|\n\n\nSo i undersampled and made mask of data which of cumulative distance is less than 10 randomly by the inverse of percentile.\n\n|distance|ratio|\n|---|---|\n|0 ~ 1|0.021991267733573804|\n|1 ~ 2|0.04509597645237031|\n|2 ~ 3|0.0785016260250442|\n|3 ~ 4|0.1187591045514077|\n|4 ~ 5|0.22327769988668522|\n|5 ~ 6|0.3929756955007945|\n|6 ~ 7|0.5740881571540075|\n|7 ~ 8|0.7374370424875981|\n|8 ~ 9|0.8781707482890396|\n|9 ~ 10|1.0|\n\n\nThis is histogram of undersampled data.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4706215%2F77659f08f6f965d19c7edceddf6ed204%2Fplot4.png?generation=1606362629250668&alt=media)\n\nThe total length of undersampled data is 6,415,723. It is 37.7% of original data.\n\nAnd I trained with this 6M data on resnet18 for 4 epochs. \ntrain loss is average of 2k iter(2k x 16 = 32k samples)\n\n|epoch|tr_loss|val_loss|LB|\n| --- | --- | --- | --- |\n|1|25.582|23.814|22.676|\n|2|20.947|21.242|20.529|\n|3|18.935|21.900|20.585|\n|4|17.359|21.878|20.484|\n\n|parameters|values|\n|---|---|\n|raster_size|267, 267|\n|pixel_size|0.5, 0.5|\n|history_num_frames|1|\n|history_delta_time|1|\n|batch_size|16|\n|AdamW|lr=1.0e-4, weight_decay=0.01|\n|CosineAnnealingWarmRestarts|T_0=80k, T_mult=1, eta_min=1.0e-5|\n\n\nI think I may have trained in wrong way. I used quite large weight_decay. Anyway training one epoch took ~17 hours on GCE with vCPU:16, GPU: T4 and ~13 hours with vCPU:24, GPU: T4.\n\nI guess random undersampling by cumulative distance is not good idea. The information losses may be quite significant. Samples of small cumulative distance may have various trajectories. I think reducing overlapped sample is more appropriate. I wanted to do more experiment but I couldn't because of lack of time.\n\nAnd my final submission is just resnet18 with train.zarr, no undersampling. I also tried mobilenetV2 but resnet18 is better.\n\n|epoch|private LB|public LB\n|---|---|---|\n|1|19.101|19.886|\n|2|18.185|18.675|\n\n|parameters|values|\n|---|---|\n|raster_size|267, 267|\n|pixel_size|0.5, 0.5|\n|history_num_frames|1|\n|history_delta_time|1|\n|batch_size|16|\n|AdamW|weight_decay=2.43e-5|\n|first epoch with OneCycleLR|max_lr=1.0e-4, div_factor=10|\n|second epoch with CosineAnnealing|lr=1.0e-5, eta_min=1.0e-6|",
    "1091493": "I had similar results. Any kind of resampling of the data caused worse performance.",
    "1091501": "I agree. Undersampling doesn't guarantee better result.",
    "1091597": "I have also try oversample which also don't work too well. I guess the reason could be that to learn those outlier, you will impact the learning on other examples too much. Or, I simply didn't train the entire data set. Perhaps, if you train the entire epoch, it will be OK.",
    "1091599": "I think batch size also matter a lot. I was using batch_size = 16 and couldn't get pass the 23 threshold. Switching to 64 can bring you lower loss when train more data.",
    "1091624": "I inspected the high loss samples quite a bit and it didn't seem like there was a great deal of systemic problems. So training on one outlier didn't seem to be particularly likely to reduce loss on other outliers. I tried an augmentation where I shifted a single channel out of the full history or surrounding vehicles by a couple pixels in a random direction and often found that that could shift a prediction from 1k+ loss down to under 50",
    "1091626": "It was interesting to see the system so fragile to this augmentation, I presume the error of the perception system could cause these small single pixel shifts so making it robust to this would probably be important."
  },
  "source": "meta"
}