{
  "id": 416935,
  "title": "19th place solution: 1D Unet + Transformer",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/416935",
  "author_name": "",
  "post_date": "2023-06-13T14:03:14.333607200Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<h3>Summary:</h3>\n<p>My solution was based on big deep learning models trained end to end with heavy data augmentations. Key elements of my solution:<br>\n5 folds CV with GroupKFold on subject.<br>\nEnsemble of 1D Unet built on top of Resnet with attention and a Vision transformer like model.<br>\nStrong data augmentation to counter overfitting - cutmix helped a lot.</p>\n<h3>Validation and preprocessing</h3>\n<p>Both DeFOG and tDCS FOG datasets are concatenated. tDCS FOG is divided by g to normalize units between datasets. Later every channel is normalized using mean and std across the entire dataset.</p>\n<p>Data is split into 5 folds using GroupKFold on subject. This is repeated multiple times to find split that have a consistent number of patients and labeled timesteps across folds.</p>\n<p>Every model is trained on windows of 8192 timesteps. During training windows are randomly sampled from available sequences. For the inference Unet processed the whole input in one step, while the transformer model used a sliding window approach.</p>\n<h3>Data augmentation</h3>\n<p>Strong data augmentation coupled with big neural networks was a key part of my solution. I used 3 augmentation techniques randomly sampled out of 6 available:</p>\n<ul>\n<li>Jitter</li>\n<li>Scaling</li>\n<li>Flip_left_right</li>\n<li>Permutation</li>\n<li>Magnitude_warp</li>\n<li>Window_warp</li>\n</ul>\n<p>and cutmix on top of everything</p>\n<h3>Models</h3>\n<p>1D Unet built on top of a Resnet model that uses both channel and spatial attention in decoder and encoder layers. Unet models were trained using AdamW optimizer with using <br>\nCosineAnnealingWarmRestarts scheduler. Hyperparameters:<br>\nepochs 150<br>\nlr = 2e-3<br>\nweight_decay = 1e-2<br>\nresnet_blocks = (3, 3, 27, 3)<br>\nactivation = nn.GELU<br>\nstem_kernel_size = 17<br>\nkernel_size = 9</p>\n<p>Vision transformer like models but working on 1D sequence instead of 2D image. Those were trained with AdamW optimizer using OneCycleLR scheduler. Hyperparameters:<br>\nepochs = 2400<br>\nlr = 2e-4<br>\nweight_decay = 2e-1<br>\nembed_dim = 768<br>\nnlayers = 12<br>\ndropout = 0.25<br>\nnheads = 12<br>\npatch_size = 64</p>\n<p>Transformer models seem to perform better on my local CV and public leaderboard, but convnet models work better on private leaderboard. Probably my Transformer models overfit to the training data since I trained them for a lot longer than my convnet models, and my validation scheme failed to catch that. Submission using only my convnet models would place me in the gold range, but I didn’t pick it for scoring. My final submission was an ensemble of all folds of convnet and transformer models with 0.4 weight for the convnet and 0.6 for transformer.</p>\n<h3>Mistakes</h3>\n<p>I think I focused too much on modeling and not enough on the data itself. I spent a lot of time on pseudo-labeling that I didn’t use in the end because it did not help my CV and public score. Results on one fold were significantly worse than on the others and I should investigate that, instead I focused on pseudo-labeling too much. I didn’t think of training separate models on DeFOG and tDCS FOG.</p>",
  "messages": [
    {
      "id": "2300895",
      "postDate": "06/13/2023 14:03:14",
      "content": "<h3>Summary:</h3>\n<p>My solution was based on big deep learning models trained end to end with heavy data augmentations. Key elements of my solution:<br>\n5 folds CV with GroupKFold on subject.<br>\nEnsemble of 1D Unet built on top of Resnet with attention and a Vision transformer like model.<br>\nStrong data augmentation to counter overfitting - cutmix helped a lot.</p>\n<h3>Validation and preprocessing</h3>\n<p>Both DeFOG and tDCS FOG datasets are concatenated. tDCS FOG is divided by g to normalize units between datasets. Later every channel is normalized using mean and std across the entire dataset.</p>\n<p>Data is split into 5 folds using GroupKFold on subject. This is repeated multiple times to find split that have a consistent number of patients and labeled timesteps across folds.</p>\n<p>Every model is trained on windows of 8192 timesteps. During training windows are randomly sampled from available sequences. For the inference Unet processed the whole input in one step, while the transformer model used a sliding window approach.</p>\n<h3>Data augmentation</h3>\n<p>Strong data augmentation coupled with big neural networks was a key part of my solution. I used 3 augmentation techniques randomly sampled out of 6 available:</p>\n<ul>\n<li>Jitter</li>\n<li>Scaling</li>\n<li>Flip_left_right</li>\n<li>Permutation</li>\n<li>Magnitude_warp</li>\n<li>Window_warp</li>\n</ul>\n<p>and cutmix on top of everything</p>\n<h3>Models</h3>\n<p>1D Unet built on top of a Resnet model that uses both channel and spatial attention in decoder and encoder layers. Unet models were trained using AdamW optimizer with using <br>\nCosineAnnealingWarmRestarts scheduler. Hyperparameters:<br>\nepochs 150<br>\nlr = 2e-3<br>\nweight_decay = 1e-2<br>\nresnet_blocks = (3, 3, 27, 3)<br>\nactivation = nn.GELU<br>\nstem_kernel_size = 17<br>\nkernel_size = 9</p>\n<p>Vision transformer like models but working on 1D sequence instead of 2D image. Those were trained with AdamW optimizer using OneCycleLR scheduler. Hyperparameters:<br>\nepochs = 2400<br>\nlr = 2e-4<br>\nweight_decay = 2e-1<br>\nembed_dim = 768<br>\nnlayers = 12<br>\ndropout = 0.25<br>\nnheads = 12<br>\npatch_size = 64</p>\n<p>Transformer models seem to perform better on my local CV and public leaderboard, but convnet models work better on private leaderboard. Probably my Transformer models overfit to the training data since I trained them for a lot longer than my convnet models, and my validation scheme failed to catch that. Submission using only my convnet models would place me in the gold range, but I didn’t pick it for scoring. My final submission was an ensemble of all folds of convnet and transformer models with 0.4 weight for the convnet and 0.6 for transformer.</p>\n<h3>Mistakes</h3>\n<p>I think I focused too much on modeling and not enough on the data itself. I spent a lot of time on pseudo-labeling that I didn’t use in the end because it did not help my CV and public score. Results on one fold were significantly worse than on the others and I should investigate that, instead I focused on pseudo-labeling too much. I didn’t think of training separate models on DeFOG and tDCS FOG.</p>",
      "rawMarkdown": "### Summary:\n\nMy solution was based on big deep learning models trained end to end with heavy data augmentations. Key elements of my solution:\n5 folds CV with GroupKFold on subject.\nEnsemble of 1D Unet built on top of Resnet with attention and a Vision transformer like model.\nStrong data augmentation to counter overfitting - cutmix helped a lot.\n\n### Validation and preprocessing\n\nBoth DeFOG and tDCS FOG datasets are concatenated. tDCS FOG is divided by g to normalize units between datasets. Later every channel is normalized using mean and std across the entire dataset.\n\nData is split into 5 folds using GroupKFold on subject. This is repeated multiple times to find split that have a consistent number of patients and labeled timesteps across folds.\n\nEvery model is trained on windows of 8192 timesteps. During training windows are randomly sampled from available sequences. For the inference Unet processed the whole input in one step, while the transformer model used a sliding window approach.\n\n### Data augmentation\n\nStrong data augmentation coupled with big neural networks was a key part of my solution. I used 3 augmentation techniques randomly sampled out of 6 available:\n- Jitter\n- Scaling\n- Flip_left_right\n- Permutation\n- Magnitude_warp\n- Window_warp\n\nand cutmix on top of everything\n\n### Models\n\n1D Unet built on top of a Resnet model that uses both channel and spatial attention in decoder and encoder layers. Unet models were trained using AdamW optimizer with using \nCosineAnnealingWarmRestarts scheduler. Hyperparameters:\nepochs 150\nlr = 2e-3\nweight_decay = 1e-2\nresnet_blocks = (3, 3, 27, 3)\nactivation = nn.GELU\nstem_kernel_size = 17\nkernel_size = 9\n\nVision transformer like models but working on 1D sequence instead of 2D image. Those were trained with AdamW optimizer using OneCycleLR scheduler. Hyperparameters:\nepochs = 2400\nlr = 2e-4\nweight_decay = 2e-1\nembed_dim = 768\nnlayers = 12\ndropout = 0.25\nnheads = 12\npatch_size = 64\n\nTransformer models seem to perform better on my local CV and public leaderboard, but convnet models work better on private leaderboard. Probably my Transformer models overfit to the training data since I trained them for a lot longer than my convnet models, and my validation scheme failed to catch that. Submission using only my convnet models would place me in the gold range, but I didn’t pick it for scoring. My final submission was an ensemble of all folds of convnet and transformer models with 0.4 weight for the convnet and 0.6 for transformer.\n\n### Mistakes\n\nI think I focused too much on modeling and not enough on the data itself. I spent a lot of time on pseudo-labeling that I didn’t use in the end because it did not help my CV and public score. Results on one fold were significantly worse than on the others and I should investigate that, instead I focused on pseudo-labeling too much. I didn’t think of training separate models on DeFOG and tDCS FOG.",
      "votes": null
    },
    {
      "id": "2303307",
      "postDate": "06/15/2023 08:23:59",
      "content": "<p>What did you  mean by \"Data is split into 5 folds using GroupKFold on subject. This is repeated multiple times to find split that have a consistent number of patients and labeled timesteps across folds.\" Do you mean that each fold has the same number of different patients? Similiar ratios of three events?</p>",
      "rawMarkdown": "What did you  mean by \"Data is split into 5 folds using GroupKFold on subject. This is repeated multiple times to find split that have a consistent number of patients and labeled timesteps across folds.\" Do you mean that each fold has the same number of different patients? Similiar ratios of three events?",
      "votes": null
    },
    {
      "id": "2309101",
      "postDate": "06/19/2023 11:58:34",
      "content": "<p>Yes, I wanted my folds to have similar number of patients and similar numbers of labeled timesteps across three event types.</p>",
      "rawMarkdown": "Yes, I wanted my folds to have similar number of patients and similar numbers of labeled timesteps across three event types.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2303307,
      "author_name": "exjustice",
      "author_url": "",
      "post_date": "06/15/2023 08:23:59",
      "content": "<p>What did you  mean by \"Data is split into 5 folds using GroupKFold on subject. This is repeated multiple times to find split that have a consistent number of patients and labeled timesteps across folds.\" Do you mean that each fold has the same number of different patients? Similiar ratios of three events?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2309101,
          "author_name": "nordberdt",
          "author_url": "",
          "post_date": "06/19/2023 11:58:34",
          "content": "<p>Yes, I wanted my folds to have similar number of patients and similar numbers of labeled timesteps across three event types.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2300895": "### Summary:\n\nMy solution was based on big deep learning models trained end to end with heavy data augmentations. Key elements of my solution:\n5 folds CV with GroupKFold on subject.\nEnsemble of 1D Unet built on top of Resnet with attention and a Vision transformer like model.\nStrong data augmentation to counter overfitting - cutmix helped a lot.\n\n### Validation and preprocessing\n\nBoth DeFOG and tDCS FOG datasets are concatenated. tDCS FOG is divided by g to normalize units between datasets. Later every channel is normalized using mean and std across the entire dataset.\n\nData is split into 5 folds using GroupKFold on subject. This is repeated multiple times to find split that have a consistent number of patients and labeled timesteps across folds.\n\nEvery model is trained on windows of 8192 timesteps. During training windows are randomly sampled from available sequences. For the inference Unet processed the whole input in one step, while the transformer model used a sliding window approach.\n\n### Data augmentation\n\nStrong data augmentation coupled with big neural networks was a key part of my solution. I used 3 augmentation techniques randomly sampled out of 6 available:\n- Jitter\n- Scaling\n- Flip_left_right\n- Permutation\n- Magnitude_warp\n- Window_warp\n\nand cutmix on top of everything\n\n### Models\n\n1D Unet built on top of a Resnet model that uses both channel and spatial attention in decoder and encoder layers. Unet models were trained using AdamW optimizer with using \nCosineAnnealingWarmRestarts scheduler. Hyperparameters:\nepochs 150\nlr = 2e-3\nweight_decay = 1e-2\nresnet_blocks = (3, 3, 27, 3)\nactivation = nn.GELU\nstem_kernel_size = 17\nkernel_size = 9\n\nVision transformer like models but working on 1D sequence instead of 2D image. Those were trained with AdamW optimizer using OneCycleLR scheduler. Hyperparameters:\nepochs = 2400\nlr = 2e-4\nweight_decay = 2e-1\nembed_dim = 768\nnlayers = 12\ndropout = 0.25\nnheads = 12\npatch_size = 64\n\nTransformer models seem to perform better on my local CV and public leaderboard, but convnet models work better on private leaderboard. Probably my Transformer models overfit to the training data since I trained them for a lot longer than my convnet models, and my validation scheme failed to catch that. Submission using only my convnet models would place me in the gold range, but I didn’t pick it for scoring. My final submission was an ensemble of all folds of convnet and transformer models with 0.4 weight for the convnet and 0.6 for transformer.\n\n### Mistakes\n\nI think I focused too much on modeling and not enough on the data itself. I spent a lot of time on pseudo-labeling that I didn’t use in the end because it did not help my CV and public score. Results on one fold were significantly worse than on the others and I should investigate that, instead I focused on pseudo-labeling too much. I didn’t think of training separate models on DeFOG and tDCS FOG.",
    "2303307": "What did you  mean by \"Data is split into 5 folds using GroupKFold on subject. This is repeated multiple times to find split that have a consistent number of patients and labeled timesteps across folds.\" Do you mean that each fold has the same number of different patients? Similiar ratios of three events?",
    "2309101": "Yes, I wanted my folds to have similar number of patients and similar numbers of labeled timesteps across three event types."
  },
  "source": "meta"
}