{
  "id": 416513,
  "title": "10th place solution: U-Net with squeeze & excitation",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/writeups/james-day-10th-place-solution-u-net-with-squeeze-e",
  "author_name": "",
  "post_date": "2023-06-12T02:48:18.037Z",
  "votes": 20,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Congrats to the winners! Thanks to the competition organizers for putting together an interesting challenge! Here's my solution.</p>\n<h1>Model architecture</h1>\n<p>I used a 1D convolutional U-Net with squeeze-and-excitation and 5 encoder/decoder pairs.</p>\n<p>Squeeze-and-excitation seemed to be very beneficial, presumably because it allows the model to take global context into consideration while classifying each sample. I processed the data in extremely long context windows (10240 samples).</p>\n<h1>Features</h1>\n<ul>\n<li><strong>Raw acleration values:</strong> AccV, AccML, AccAP<ul>\n<li>I did not normalize these in any way. </li></ul></li>\n<li><strong>Time features:</strong> <br>\n<code>\ndf['NormalizedTime'] = df['Time'] / df['Time'].max()\n</code><br>\n<code>\ndf['SinNormalizedTime'] = np.sin(df['NormalizedTime'] * np.pi)\n</code></li>\n</ul>\n<p>I also experimented with adding a variety of frequency domain features that were calculated using wavelet transforms but that didn't help.</p>\n<h1>Training data augmentation</h1>\n<ul>\n<li><strong>Random low pass filtering:</strong><ul>\n<li>Frequency cutoff was 5% - 37.5% the sample rate</li>\n<li>Applied to half the training sequences</li></ul></li>\n<li><strong>Random time warp:</strong><ul>\n<li>Used linear interpolation to change the sequence length by +/- 10% (or any value in between; the scale was sampled from a uniform distribution)</li>\n<li>Applied to half the training sequences</li></ul></li>\n<li><strong>Random flip:</strong><ul>\n<li>Multiplied AccML by -1 to reverse right &amp; left</li>\n<li>Applied to half the training sequences</li></ul></li>\n<li><strong>Random magnitude warping:</strong><ul>\n<li>The difference between each acceleration feature's value and its mean value was multiplied by a coefficient randomly sampled from a gaussian distribution with a mean of 0 and a standard deviation of 0.1</li>\n<li>Applied to half the training sequences</li></ul></li>\n<li><strong>Noisy time features:</strong><ul>\n<li>Normalized times within each context window shifted by value sampled from gaussian distribution with mean of 0 and standard deviation of 0.05</li>\n<li>Applied before calculating SinNormalizedTime (so the same noise impacts both features).</li>\n<li>Applied to ALL the training sequences</li></ul></li>\n</ul>\n<h1>Inference time data augmentation</h1>\n<p>Each sample was classified 16 times by each model.</p>\n<ul>\n<li>With and without multiplying AccML by -1 to reverse right &amp; left</li>\n<li>Sequences were classified in overlapping context windows with a stride equal to 1/8 the window length. Similar to random crop data augmentation.</li>\n</ul>\n<p>The values saved to the submission file were the simple mean of all predictions from all models.</p>\n<h1>Handling defog vs tdcsfog</h1>\n<p>I used the same models for both datasets. I did not do anything to normalize the sample rates or feature values. I did not even convert the features to have the same units. Normalization seemed to be harmful.</p>\n<h1>Ensembling / random seed hacking</h1>\n<p>I used 2 near-identical models that were trained with identical hyperparameters from the same cross-validation fold, but with different random seeds for weight initialization &amp; shuffling the training data. They were filtered to have mAP scores in the top 20% of my local cross validation and top 50% of my LB scores. This probably improved my score by around 0.01 - 0.02 vs. just using 2 random models.</p>\n<p><strong>Inference notebook:</strong> <a href=\"https://www.kaggle.com/jsday96/parkinsons-overlapping-se-unet-frequency-domain\" target=\"_blank\">https://www.kaggle.com/jsday96/parkinsons-overlapping-se-unet-frequency-domain</a></p>",
  "messages": [
    {
      "id": "2296407",
      "postDate": "06/11/2023 19:30:18",
      "content": "<p>Congrats to the winners! Thanks to the competition organizers for putting together an interesting challenge! Here's my solution.</p>\n<h1>Model architecture</h1>\n<p>I used a 1D convolutional U-Net with squeeze-and-excitation and 5 encoder/decoder pairs.</p>\n<p>Squeeze-and-excitation seemed to be very beneficial, presumably because it allows the model to take global context into consideration while classifying each sample. I processed the data in extremely long context windows (10240 samples).</p>\n<h1>Features</h1>\n<ul>\n<li><strong>Raw acleration values:</strong> AccV, AccML, AccAP<ul>\n<li>I did not normalize these in any way. </li></ul></li>\n<li><strong>Time features:</strong> <br>\n<code>\ndf['NormalizedTime'] = df['Time'] / df['Time'].max()\n</code><br>\n<code>\ndf['SinNormalizedTime'] = np.sin(df['NormalizedTime'] * np.pi)\n</code></li>\n</ul>\n<p>I also experimented with adding a variety of frequency domain features that were calculated using wavelet transforms but that didn't help.</p>\n<h1>Training data augmentation</h1>\n<ul>\n<li><strong>Random low pass filtering:</strong><ul>\n<li>Frequency cutoff was 5% - 37.5% the sample rate</li>\n<li>Applied to half the training sequences</li></ul></li>\n<li><strong>Random time warp:</strong><ul>\n<li>Used linear interpolation to change the sequence length by +/- 10% (or any value in between; the scale was sampled from a uniform distribution)</li>\n<li>Applied to half the training sequences</li></ul></li>\n<li><strong>Random flip:</strong><ul>\n<li>Multiplied AccML by -1 to reverse right &amp; left</li>\n<li>Applied to half the training sequences</li></ul></li>\n<li><strong>Random magnitude warping:</strong><ul>\n<li>The difference between each acceleration feature's value and its mean value was multiplied by a coefficient randomly sampled from a gaussian distribution with a mean of 0 and a standard deviation of 0.1</li>\n<li>Applied to half the training sequences</li></ul></li>\n<li><strong>Noisy time features:</strong><ul>\n<li>Normalized times within each context window shifted by value sampled from gaussian distribution with mean of 0 and standard deviation of 0.05</li>\n<li>Applied before calculating SinNormalizedTime (so the same noise impacts both features).</li>\n<li>Applied to ALL the training sequences</li></ul></li>\n</ul>\n<h1>Inference time data augmentation</h1>\n<p>Each sample was classified 16 times by each model.</p>\n<ul>\n<li>With and without multiplying AccML by -1 to reverse right &amp; left</li>\n<li>Sequences were classified in overlapping context windows with a stride equal to 1/8 the window length. Similar to random crop data augmentation.</li>\n</ul>\n<p>The values saved to the submission file were the simple mean of all predictions from all models.</p>\n<h1>Handling defog vs tdcsfog</h1>\n<p>I used the same models for both datasets. I did not do anything to normalize the sample rates or feature values. I did not even convert the features to have the same units. Normalization seemed to be harmful.</p>\n<h1>Ensembling / random seed hacking</h1>\n<p>I used 2 near-identical models that were trained with identical hyperparameters from the same cross-validation fold, but with different random seeds for weight initialization &amp; shuffling the training data. They were filtered to have mAP scores in the top 20% of my local cross validation and top 50% of my LB scores. This probably improved my score by around 0.01 - 0.02 vs. just using 2 random models.</p>\n<p><strong>Inference notebook:</strong> <a href=\"https://www.kaggle.com/jsday96/parkinsons-overlapping-se-unet-frequency-domain\" target=\"_blank\">https://www.kaggle.com/jsday96/parkinsons-overlapping-se-unet-frequency-domain</a></p>",
      "rawMarkdown": "Congrats to the winners! Thanks to the competition organizers for putting together an interesting challenge! Here's my solution.\n\n# Model architecture\n\nI used a 1D convolutional U-Net with squeeze-and-excitation and 5 encoder/decoder pairs.\n\nSqueeze-and-excitation seemed to be very beneficial, presumably because it allows the model to take global context into consideration while classifying each sample. I processed the data in extremely long context windows (10240 samples).\n\n# Features\n\n- **Raw acleration values:** AccV, AccML, AccAP\n - I did not normalize these in any way. \n- **Time features:** \n`\ndf['NormalizedTime'] = df['Time'] / df['Time'].max()\n`\n`\ndf['SinNormalizedTime'] = np.sin(df['NormalizedTime'] * np.pi)\n`\n\nI also experimented with adding a variety of frequency domain features that were calculated using wavelet transforms but that didn't help.\n\n# Training data augmentation\n\n- **Random low pass filtering:**\n - Frequency cutoff was 5% - 37.5% the sample rate\n - Applied to half the training sequences\n- **Random time warp:**\n - Used linear interpolation to change the sequence length by +/- 10% (or any value in between; the scale was sampled from a uniform distribution)\n - Applied to half the training sequences\n- **Random flip:**\n - Multiplied AccML by -1 to reverse right & left\n - Applied to half the training sequences\n- **Random magnitude warping:**\n - The difference between each acceleration feature's value and its mean value was multiplied by a coefficient randomly sampled from a gaussian distribution with a mean of 0 and a standard deviation of 0.1\n - Applied to half the training sequences\n- **Noisy time features:**\n - Normalized times within each context window shifted by value sampled from gaussian distribution with mean of 0 and standard deviation of 0.05\n - Applied before calculating SinNormalizedTime (so the same noise impacts both features).\n - Applied to ALL the training sequences\n\n# Inference time data augmentation\n\nEach sample was classified 16 times by each model.\n- With and without multiplying AccML by -1 to reverse right & left\n- Sequences were classified in overlapping context windows with a stride equal to 1/8 the window length. Similar to random crop data augmentation.\n\nThe values saved to the submission file were the simple mean of all predictions from all models.\n\n# Handling defog vs tdcsfog\n\nI used the same models for both datasets. I did not do anything to normalize the sample rates or feature values. I did not even convert the features to have the same units. Normalization seemed to be harmful.\n\n# Ensembling / random seed hacking\n\nI used 2 near-identical models that were trained with identical hyperparameters from the same cross-validation fold, but with different random seeds for weight initialization & shuffling the training data. They were filtered to have mAP scores in the top 20% of my local cross validation and top 50% of my LB scores. This probably improved my score by around 0.01 - 0.02 vs. just using 2 random models.\n\n\n\n**Inference notebook:** https://www.kaggle.com/jsday96/parkinsons-overlapping-se-unet-frequency-domain",
      "votes": null
    },
    {
      "id": "2296655",
      "postDate": "06/12/2023 03:40:02",
      "content": "<blockquote>\n  <p>Multiplied AccML by -1 to reverse right &amp; left</p>\n</blockquote>\n<p>In hindsight, this augmentation seems so obvious. Left and right are symmetric so flipping the accelerometer data here makes perfect sense.<br>\nGreat work 👍.</p>",
      "rawMarkdown": ">Multiplied AccML by -1 to reverse right & left\n\nIn hindsight, this augmentation seems so obvious. Left and right are symmetric so flipping the accelerometer data here makes perfect sense.\nGreat work 👍.",
      "votes": null
    },
    {
      "id": "2297549",
      "postDate": "06/12/2023 16:47:24",
      "content": "<p><code>I used the same models for both datasets. I did not do anything to normalize the sample rates or feature values. I did not even convert the features to have the same units. Normalization seemed to be harmful.</code><br>\nI noticed this as well. Any thoughts on why that might be the case?</p>",
      "rawMarkdown": "`I used the same models for both datasets. I did not do anything to normalize the sample rates or feature values. I did not even convert the features to have the same units. Normalization seemed to be harmful.`\nI noticed this as well. Any thoughts on why that might be the case?",
      "votes": null
    },
    {
      "id": "2299883",
      "postDate": "06/12/2023 23:12:06",
      "content": "<p>I have a suspicion my models may have been using the scale of the features as an indication of which dataset the sequences came from.</p>\n<p>Squeeze &amp; excitation allows the model to conditionally multiply certain channels in the convolutional layer outputs by 0, effectively toggling certain convolutional filters on and off based on global context information. So the model could have dataset-specific components. I would guess leaving the features unnormalized with inconsistent units made it easier for it to do that or forced it into an approach which generalizes relatively well. </p>\n<p>Looking back at my old notes, it appears normalizing to consistent units and adding a dataset type indicator as a feature works <em>almost</em> as well as using unscaled data. Scores from 10-fold cross validation with different subjects in each fold included below.</p>\n<table>\n<thead>\n<tr>\n<th>Consistent units?</th>\n<th>Dataset type as feature?</th>\n<th>Local CV score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Yes</td>\n<td>No</td>\n<td>0.341</td>\n</tr>\n<tr>\n<td>No</td>\n<td>Yes</td>\n<td>0.355</td>\n</tr>\n<tr>\n<td>No</td>\n<td>No</td>\n<td>0.367</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I have a suspicion my models may have been using the scale of the features as an indication of which dataset the sequences came from.\n\nSqueeze & excitation allows the model to conditionally multiply certain channels in the convolutional layer outputs by 0, effectively toggling certain convolutional filters on and off based on global context information. So the model could have dataset-specific components. I would guess leaving the features unnormalized with inconsistent units made it easier for it to do that or forced it into an approach which generalizes relatively well. \n\nLooking back at my old notes, it appears normalizing to consistent units and adding a dataset type indicator as a feature works *almost* as well as using unscaled data. Scores from 10-fold cross validation with different subjects in each fold included below.\n\n| Consistent units? | Dataset type as feature? | Local CV score |\n| --- | --- | --- |\n| Yes | No | 0.341 |\n| No | Yes | 0.355 |\n| No | No | 0.367 |",
      "votes": null
    },
    {
      "id": "2301086",
      "postDate": "06/13/2023 15:56:48",
      "content": "<p>Interesting, thanks for following up! Your explanation seems reasonable to me, though I'll need to read up on \"squeeze and excitation\" since this is the first time I've heard of it.</p>\n<p>Really creative solution btw, especially all the data augmentation you did. </p>",
      "rawMarkdown": "Interesting, thanks for following up! Your explanation seems reasonable to me, though I'll need to read up on \"squeeze and excitation\" since this is the first time I've heard of it.\n\nReally creative solution btw, especially all the data augmentation you did.",
      "votes": null
    },
    {
      "id": "2307231",
      "postDate": "06/18/2023 02:57:23",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> — congrats on the gold! </p>\n<p>Can I ask for more explanation on the random seed hacking bit? Are you saying you essentially fine-tuned the random seed like a hyper parameter? Or do you mean something else?  If the former, it seems like you’d be tuning the random seed to your validation set with no generalization, right?</p>",
      "rawMarkdown": "Hey @jsday96 — congrats on the gold! \n\nCan I ask for more explanation on the random seed hacking bit? Are you saying you essentially fine-tuned the random seed like a hyper parameter? Or do you mean something else?  If the former, it seems like you’d be tuning the random seed to your validation set with no generalization, right?",
      "votes": null
    },
    {
      "id": "2308251",
      "postDate": "06/18/2023 19:01:52",
      "content": "<p>It means I trained multiple models and used the CV scores and LB scores to filter them. Essentially…</p>\n<ol>\n<li>Train 20 models with identical hyperparameters &amp; the same cross validation fold, but different initial weights &amp; different data shuffling.</li>\n<li>Pick the 4 models that had the highest CV score during #1's training runs.</li>\n<li>Check the public LB score each of the models from #2 achieves individually.</li>\n<li>Submit ensemble that uses the top 2 models from #3.</li>\n</ol>\n<p>This strategy stems from an issue I had early on in the competition where certain training runs seemed to \"get lucky\" and produced models that generalized a lot better than others. The data augmentation described in my original post helped to reduce the amount of random variation in the LB scores for my individual models, so cherry picking models didn't matter as much with all the data augmentation in place, but I still did it for the final ensemble regardless.</p>",
      "rawMarkdown": "It means I trained multiple models and used the CV scores and LB scores to filter them. Essentially...\n1. Train 20 models with identical hyperparameters & the same cross validation fold, but different initial weights & different data shuffling.\n2. Pick the 4 models that had the highest CV score during #1's training runs.\n3. Check the public LB score each of the models from #2 achieves individually.\n4. Submit ensemble that uses the top 2 models from #3.\n\nThis strategy stems from an issue I had early on in the competition where certain training runs seemed to \"get lucky\" and produced models that generalized a lot better than others. The data augmentation described in my original post helped to reduce the amount of random variation in the LB scores for my individual models, so cherry picking models didn't matter as much with all the data augmentation in place, but I still did it for the final ensemble regardless.",
      "votes": null
    },
    {
      "id": "2444527",
      "postDate": "09/18/2023 10:12:41",
      "content": "<p>So you optimized your initialization and data shuffling, based on your CV first and then on the public leaderboard. But this would not generalize to the private leaderboard, right? Since it is a completely different left-out set, you cannot say if your seed would do well on the unseen data.<br>\nOr is there a correlation between good public /private leaderboard based on your seed selection, I wouldn't say so though!</p>",
      "rawMarkdown": "So you optimized your initialization and data shuffling, based on your CV first and then on the public leaderboard. But this would not generalize to the private leaderboard, right? Since it is a completely different left-out set, you cannot say if your seed would do well on the unseen data.\nOr is there a correlation between good public /private leaderboard based on your seed selection, I wouldn't say so though!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2296655,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/12/2023 03:40:02",
      "content": "<blockquote>\n  <p>Multiplied AccML by -1 to reverse right &amp; left</p>\n</blockquote>\n<p>In hindsight, this augmentation seems so obvious. Left and right are symmetric so flipping the accelerometer data here makes perfect sense.<br>\nGreat work 👍.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2297549,
      "author_name": "abandura",
      "author_url": "",
      "post_date": "06/12/2023 16:47:24",
      "content": "<p><code>I used the same models for both datasets. I did not do anything to normalize the sample rates or feature values. I did not even convert the features to have the same units. Normalization seemed to be harmful.</code><br>\nI noticed this as well. Any thoughts on why that might be the case?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2299883,
          "author_name": "jsday96",
          "author_url": "",
          "post_date": "06/12/2023 23:12:06",
          "content": "<p>I have a suspicion my models may have been using the scale of the features as an indication of which dataset the sequences came from.</p>\n<p>Squeeze &amp; excitation allows the model to conditionally multiply certain channels in the convolutional layer outputs by 0, effectively toggling certain convolutional filters on and off based on global context information. So the model could have dataset-specific components. I would guess leaving the features unnormalized with inconsistent units made it easier for it to do that or forced it into an approach which generalizes relatively well. </p>\n<p>Looking back at my old notes, it appears normalizing to consistent units and adding a dataset type indicator as a feature works <em>almost</em> as well as using unscaled data. Scores from 10-fold cross validation with different subjects in each fold included below.</p>\n<table>\n<thead>\n<tr>\n<th>Consistent units?</th>\n<th>Dataset type as feature?</th>\n<th>Local CV score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Yes</td>\n<td>No</td>\n<td>0.341</td>\n</tr>\n<tr>\n<td>No</td>\n<td>Yes</td>\n<td>0.355</td>\n</tr>\n<tr>\n<td>No</td>\n<td>No</td>\n<td>0.367</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": [
            {
              "id": 2301086,
              "author_name": "abandura",
              "author_url": "",
              "post_date": "06/13/2023 15:56:48",
              "content": "<p>Interesting, thanks for following up! Your explanation seems reasonable to me, though I'll need to read up on \"squeeze and excitation\" since this is the first time I've heard of it.</p>\n<p>Really creative solution btw, especially all the data augmentation you did. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2307231,
      "author_name": "austinhinkel",
      "author_url": "",
      "post_date": "06/18/2023 02:57:23",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> — congrats on the gold! </p>\n<p>Can I ask for more explanation on the random seed hacking bit? Are you saying you essentially fine-tuned the random seed like a hyper parameter? Or do you mean something else?  If the former, it seems like you’d be tuning the random seed to your validation set with no generalization, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2308251,
          "author_name": "jsday96",
          "author_url": "",
          "post_date": "06/18/2023 19:01:52",
          "content": "<p>It means I trained multiple models and used the CV scores and LB scores to filter them. Essentially…</p>\n<ol>\n<li>Train 20 models with identical hyperparameters &amp; the same cross validation fold, but different initial weights &amp; different data shuffling.</li>\n<li>Pick the 4 models that had the highest CV score during #1's training runs.</li>\n<li>Check the public LB score each of the models from #2 achieves individually.</li>\n<li>Submit ensemble that uses the top 2 models from #3.</li>\n</ol>\n<p>This strategy stems from an issue I had early on in the competition where certain training runs seemed to \"get lucky\" and produced models that generalized a lot better than others. The data augmentation described in my original post helped to reduce the amount of random variation in the LB scores for my individual models, so cherry picking models didn't matter as much with all the data augmentation in place, but I still did it for the final ensemble regardless.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2444527,
              "author_name": "hugodeheer",
              "author_url": "",
              "post_date": "09/18/2023 10:12:41",
              "content": "<p>So you optimized your initialization and data shuffling, based on your CV first and then on the public leaderboard. But this would not generalize to the private leaderboard, right? Since it is a completely different left-out set, you cannot say if your seed would do well on the unseen data.<br>\nOr is there a correlation between good public /private leaderboard based on your seed selection, I wouldn't say so though!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2296407": "Congrats to the winners! Thanks to the competition organizers for putting together an interesting challenge! Here's my solution.\n\n# Model architecture\n\nI used a 1D convolutional U-Net with squeeze-and-excitation and 5 encoder/decoder pairs.\n\nSqueeze-and-excitation seemed to be very beneficial, presumably because it allows the model to take global context into consideration while classifying each sample. I processed the data in extremely long context windows (10240 samples).\n\n# Features\n\n- **Raw acleration values:** AccV, AccML, AccAP\n - I did not normalize these in any way. \n- **Time features:** \n`\ndf['NormalizedTime'] = df['Time'] / df['Time'].max()\n`\n`\ndf['SinNormalizedTime'] = np.sin(df['NormalizedTime'] * np.pi)\n`\n\nI also experimented with adding a variety of frequency domain features that were calculated using wavelet transforms but that didn't help.\n\n# Training data augmentation\n\n- **Random low pass filtering:**\n - Frequency cutoff was 5% - 37.5% the sample rate\n - Applied to half the training sequences\n- **Random time warp:**\n - Used linear interpolation to change the sequence length by +/- 10% (or any value in between; the scale was sampled from a uniform distribution)\n - Applied to half the training sequences\n- **Random flip:**\n - Multiplied AccML by -1 to reverse right & left\n - Applied to half the training sequences\n- **Random magnitude warping:**\n - The difference between each acceleration feature's value and its mean value was multiplied by a coefficient randomly sampled from a gaussian distribution with a mean of 0 and a standard deviation of 0.1\n - Applied to half the training sequences\n- **Noisy time features:**\n - Normalized times within each context window shifted by value sampled from gaussian distribution with mean of 0 and standard deviation of 0.05\n - Applied before calculating SinNormalizedTime (so the same noise impacts both features).\n - Applied to ALL the training sequences\n\n# Inference time data augmentation\n\nEach sample was classified 16 times by each model.\n- With and without multiplying AccML by -1 to reverse right & left\n- Sequences were classified in overlapping context windows with a stride equal to 1/8 the window length. Similar to random crop data augmentation.\n\nThe values saved to the submission file were the simple mean of all predictions from all models.\n\n# Handling defog vs tdcsfog\n\nI used the same models for both datasets. I did not do anything to normalize the sample rates or feature values. I did not even convert the features to have the same units. Normalization seemed to be harmful.\n\n# Ensembling / random seed hacking\n\nI used 2 near-identical models that were trained with identical hyperparameters from the same cross-validation fold, but with different random seeds for weight initialization & shuffling the training data. They were filtered to have mAP scores in the top 20% of my local cross validation and top 50% of my LB scores. This probably improved my score by around 0.01 - 0.02 vs. just using 2 random models.\n\n\n\n**Inference notebook:** https://www.kaggle.com/jsday96/parkinsons-overlapping-se-unet-frequency-domain",
    "2296655": ">Multiplied AccML by -1 to reverse right & left\n\nIn hindsight, this augmentation seems so obvious. Left and right are symmetric so flipping the accelerometer data here makes perfect sense.\nGreat work 👍.",
    "2297549": "`I used the same models for both datasets. I did not do anything to normalize the sample rates or feature values. I did not even convert the features to have the same units. Normalization seemed to be harmful.`\nI noticed this as well. Any thoughts on why that might be the case?",
    "2299883": "I have a suspicion my models may have been using the scale of the features as an indication of which dataset the sequences came from.\n\nSqueeze & excitation allows the model to conditionally multiply certain channels in the convolutional layer outputs by 0, effectively toggling certain convolutional filters on and off based on global context information. So the model could have dataset-specific components. I would guess leaving the features unnormalized with inconsistent units made it easier for it to do that or forced it into an approach which generalizes relatively well. \n\nLooking back at my old notes, it appears normalizing to consistent units and adding a dataset type indicator as a feature works *almost* as well as using unscaled data. Scores from 10-fold cross validation with different subjects in each fold included below.\n\n| Consistent units? | Dataset type as feature? | Local CV score |\n| --- | --- | --- |\n| Yes | No | 0.341 |\n| No | Yes | 0.355 |\n| No | No | 0.367 |",
    "2301086": "Interesting, thanks for following up! Your explanation seems reasonable to me, though I'll need to read up on \"squeeze and excitation\" since this is the first time I've heard of it.\n\nReally creative solution btw, especially all the data augmentation you did.",
    "2307231": "Hey @jsday96 — congrats on the gold! \n\nCan I ask for more explanation on the random seed hacking bit? Are you saying you essentially fine-tuned the random seed like a hyper parameter? Or do you mean something else?  If the former, it seems like you’d be tuning the random seed to your validation set with no generalization, right?",
    "2308251": "It means I trained multiple models and used the CV scores and LB scores to filter them. Essentially...\n1. Train 20 models with identical hyperparameters & the same cross validation fold, but different initial weights & different data shuffling.\n2. Pick the 4 models that had the highest CV score during #1's training runs.\n3. Check the public LB score each of the models from #2 achieves individually.\n4. Submit ensemble that uses the top 2 models from #3.\n\nThis strategy stems from an issue I had early on in the competition where certain training runs seemed to \"get lucky\" and produced models that generalized a lot better than others. The data augmentation described in my original post helped to reduce the amount of random variation in the LB scores for my individual models, so cherry picking models didn't matter as much with all the data augmentation in place, but I still did it for the final ensemble regardless.",
    "2444527": "So you optimized your initialization and data shuffling, based on your CV first and then on the public leaderboard. But this would not generalize to the private leaderboard, right? Since it is a completely different left-out set, you cannot say if your seed would do well on the unseen data.\nOr is there a correlation between good public /private leaderboard based on your seed selection, I wouldn't say so though!"
  },
  "source": "meta"
}