{
  "id": 415992,
  "title": "6th place solution: spectrograms, wavelets, convnets, unets, and transformers",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/writeups/shujun-ahmet-6th-place-solution-spectrograms-wavel",
  "author_name": "",
  "post_date": "2023-06-09T08:07:24.680Z",
  "votes": 66,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First, thanks to the competition hosts for a meaningful competition with interesting data. Second, congrats to the winners! My team was able to shake up a bit but ended up one place shy of the prize zone. All in all, I'm still happy since we did well to survive the shakeup. I'm also excited to see what top 5 teams did to create a significant gap between us. </p>\n<p>Our final ensemble is a combination of spectrogram models, wavelet models, and 1D conv models, which scores 0.369/0.462. Below I will discuss them as well as other important technical details. See below for an overview <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F12e8af388e67c174b467596a85afebdf%2FFOG_overview.png?generation=1686277343043841&amp;alt=media\" alt=\"\"></p>\n<h1>Validation setup</h1>\n<p>Validation setup is important due to the noisy data. I ended up with a nested CV setup. The procedure is as follows:</p>\n<ol>\n<li>split data into 4 folds stratified by data type and grouped by subject</li>\n<li>set aside the validation fold (i call this the outer fold)</li>\n<li>resplit the 3 training folds into 4 folds and do cross validation (I call these inner folds)</li>\n<li>take last epoch or epoch with best validation score</li>\n<li>evaluate on outer fold with 4 inner fold models averaged for each outer fold <br>\nWith this setup, we can more accurately simulate a situation where we have 250 sequences in the test set and we avg the fold model predictions. Later on, I switched to training on full inner fold data for 4 times without validation and evaluate with last epoch models on the outer fold set. </li>\n</ol>\n<h1>Input features</h1>\n<p>All of our models use the 3 waves and the pct time feature. In addition, Ahmet uses some metadata features in his models.</p>\n<h1>Spectrogram Models</h1>\n<p>When I first saw the data I thought it looked like some sort of waveform, like audio data, so I thought it might work well to use spectrograms to model it. The 3 dimension waves are transformed into 2D spectrograms with STFT. Importantly, transforming the data in spectrograms significantly downscaled the data in the time dimension, so since I use a hop length of 64/50, each frame represents a 0.5 secs window and I'm basically making predictions for 0.5 sec windows. During training, labels are resized with torchvision's resize to fit the size of the time dimension of the spectrograms and during inference the model output is resized back to full dimensionality. Sequences are cut into chunks of 128 secs (256 spectrogram frames) to generate spectrograms. </p>\n<p>Another important thing with using spectrograms is that if we use a regular type 2D conv model like resnet18, it wouldn't preserve the full dimensionality of the spectrogram (e.g a 256x256 becomes 8x8 after resnet18). In order to circumvent that, I thought to use a UNet to upsample the small feature map after the conv network. Following that, the spectrograms are pooled along the frequency dimension so I have a 1D sequence, which is then inputted into a transformer network before outputting predictions. </p>\n<p>Best submitted single spectrogram model scores 0.432/0.372. Spectrogram models are good at predicting StartHesitation and Turn but bad at Walking. </p>\n<h1>Wavelet Models</h1>\n<p>Wavelets are similar to spectrograms but also different because wavelets have different frequency/time resolutions at different frequencies. Transforming a wave into a wavelet also does not reduce the dimensionality of the scaleogram (I think this is the term for the image you get after wavelet transform). Since there's no downsampling in the time dimension, Unet is no longer needed and I simply use a resnet18/34, which downsample the scaleogram to the same time resolution as spectrogram models after Unet. In turn, I'm also classifying 0.5 sec windows. Similarly, sequences are cut into chunks of 128 secs (256 spectrogram frames) to generate spectrograms.</p>\n<p>Best submitted single spectrogram model scores 0.386/0.345. Wavelet models are good at predicting Walking but bad at StartHesitation and Turn, so it complements spectrogram models nicely.  </p>\n<h1>Transformer modeling</h1>\n<p>Just plugging in a transformer actually does not work so well because it overfits, so I bias the transformer self-attention with a float bias mask to condition each prediction on its adjacent time points, which helps the model predict long events well. </p>\n<p>What the float bias mask looks like (see above) <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F3db704367c0eb1a6f7e46f82684e3f31%2FFOG_attention_mask.png?generation=1686277393310664&amp;alt=media\" alt=\"\"></p>\n<pre><code>torch.((L,L))\n\n\n    j:\n                m[i,j]=(-(i-j)/L)**\n\n             i==j:\n                m[i,i]=\n     m\n</code></pre>\n<h1>Data augmentation</h1>\n<p>I used data augmentation for spectrogram/wavelet models including the following (mostly from audiomentations): </p>\n<ol>\n<li>time stretch </li>\n<li>gaussian noise</li>\n<li>pitch shift </li>\n<li>wave scale aug (randomly multiply wave by 0.75~1.5 to get scale invariance)</li>\n<li>time feature shift aug.  </li>\n</ol>\n<pre><code>        self.augment = Compose([\n            AddGaussianNoise(=0.001, =0.015, =0.5),\n            TimeStretch(=0.8, =1.25, =0.5,leave_length_unchanged=False,n_fft=n_fft,hop_length=hop_length),\n            PitchShift(=-4, =4, =0.5,n_fft=n_fft,hop_length=hop_length),\n        ])\n</code></pre>\n<pre><code>\n np.random.uniform()&gt;.:\n     ['wave']*=np.random.uniform(.,.)\n</code></pre>\n<pre><code> feature shift aug\n self and np()&gt;:\n    data=data+np(-,)\n</code></pre>\n<h1>Frequency encoding and range</h1>\n<p>It's important to not use the high frequency bins of fourier transform, so I simply discard them and only keep the first 64 bins, corresponding to 0-15 hz. For spectrogram models, I also encode the frequency bin with torch.linspace(0,15,n_bins) expanded in the time and channel dimension and concatted so the input to the 2D conv network has 4 channels (3 directions of spectrograms + frequency encoding). It was also useful to resample the waves to a lower frequency, which I think reduces the level of noise. I used 32, 64, and 128 hz for spectrogram models and 64 hz for wavelet models. Defog waves are resampled to match the sample rate of tdcsfog waves.</p>\n<pre><code> self==:\n    data=FA(data,,self.sample_rate)\n:\n    data=FA(data,,self.sample_rate)\n</code></pre>\n<h1>1D conv Models</h1>\n<p>1D conv models are Ahmet's solution. Please see below for details:</p>\n<ol>\n<li>First align defog and tdcs on time axis (downsampled by 32 and 25, but kept their std as a feature)</li>\n<li>pct_time, total_len, Test are used as independent features. Their prediction is summed with the prediction from the 1D CNN.</li>\n<li>Because the input was only around 7 seconds long, cumsum features are also fed into 1D CNN.</li>\n<li>Outlier dominant subject is downweighted.</li>\n<li>Used snapshot ensembling.</li>\n<li>Used notype data by applying max on the predictions.</li>\n</ol>\n<p>1D conv models are weaker compared to the other 2, scoring 0.373/0.293, but are still a nice addition to the ensemble. Interestingly, 1D conv models and spectrogram models have a similar gap of 0.09 between public and private, whereas wavelet models have only a gap of 0.04. We think this is due to a change in class balance between private/public where public has more start hesitation and private has more walking. </p>\n<h1>Ensemble Weight Tuning</h1>\n<p>For our big ensemble, the weights are first hand tuned as a starting point and then I used GP_minimize to maximize CV score. We used 2 weight tuning setups at the end 1. map of 4 folds + map of full data excluding Subject 2d57c2, 2.  map of 3 folds excluding fold with Subject 2d57c2 + map of full data excluding Subject 2d57c2. We do this because we consider Subject 2d57c2 to an outlier. </p>\n<pre><code>results=gp\n</code></pre>\n<p>The weights for our models are (we downweight loss of 2d57c2 to 0.2 in some of them)</p>\n<p>[0.1821, 0.2792, 0.1052] Ahemt model<br>\n[0.2153, 0.0, 0.0] test257 (32 hz spectrogram)<br>\n[0.6026, 0.1734, 0.0] test262 (64 hz spectrogram)<br>\n[0.0, 0.0287, 0.2579] test264 wavelet<br>\n[0.0, 0.2168, 0.0] test265 (32 hz spectrogram) downweight 2d57c2<br>\n[0.0, 0.1734, 0.2997] test266 wavelet downweight 2d57c2<br>\n[0.0, 0.1284, 0.1124] test263 128 hz spec<br>\n[0.0, 0.0, 0.2248] test271 wavelet double freq scales</p>\n<p>Let me know if you have questions, and I wouldn't be surprised if I forgot to mention some details . The code is a bit messy atm but i will clean up and release it soon. </p>",
  "messages": [
    {
      "id": "2293189",
      "postDate": "06/09/2023 02:37:04",
      "content": "<p>First, thanks to the competition hosts for a meaningful competition with interesting data. Second, congrats to the winners! My team was able to shake up a bit but ended up one place shy of the prize zone. All in all, I'm still happy since we did well to survive the shakeup. I'm also excited to see what top 5 teams did to create a significant gap between us. </p>\n<p>Our final ensemble is a combination of spectrogram models, wavelet models, and 1D conv models, which scores 0.369/0.462. Below I will discuss them as well as other important technical details. See below for an overview <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F12e8af388e67c174b467596a85afebdf%2FFOG_overview.png?generation=1686277343043841&amp;alt=media\" alt=\"\"></p>\n<h1>Validation setup</h1>\n<p>Validation setup is important due to the noisy data. I ended up with a nested CV setup. The procedure is as follows:</p>\n<ol>\n<li>split data into 4 folds stratified by data type and grouped by subject</li>\n<li>set aside the validation fold (i call this the outer fold)</li>\n<li>resplit the 3 training folds into 4 folds and do cross validation (I call these inner folds)</li>\n<li>take last epoch or epoch with best validation score</li>\n<li>evaluate on outer fold with 4 inner fold models averaged for each outer fold <br>\nWith this setup, we can more accurately simulate a situation where we have 250 sequences in the test set and we avg the fold model predictions. Later on, I switched to training on full inner fold data for 4 times without validation and evaluate with last epoch models on the outer fold set. </li>\n</ol>\n<h1>Input features</h1>\n<p>All of our models use the 3 waves and the pct time feature. In addition, Ahmet uses some metadata features in his models.</p>\n<h1>Spectrogram Models</h1>\n<p>When I first saw the data I thought it looked like some sort of waveform, like audio data, so I thought it might work well to use spectrograms to model it. The 3 dimension waves are transformed into 2D spectrograms with STFT. Importantly, transforming the data in spectrograms significantly downscaled the data in the time dimension, so since I use a hop length of 64/50, each frame represents a 0.5 secs window and I'm basically making predictions for 0.5 sec windows. During training, labels are resized with torchvision's resize to fit the size of the time dimension of the spectrograms and during inference the model output is resized back to full dimensionality. Sequences are cut into chunks of 128 secs (256 spectrogram frames) to generate spectrograms. </p>\n<p>Another important thing with using spectrograms is that if we use a regular type 2D conv model like resnet18, it wouldn't preserve the full dimensionality of the spectrogram (e.g a 256x256 becomes 8x8 after resnet18). In order to circumvent that, I thought to use a UNet to upsample the small feature map after the conv network. Following that, the spectrograms are pooled along the frequency dimension so I have a 1D sequence, which is then inputted into a transformer network before outputting predictions. </p>\n<p>Best submitted single spectrogram model scores 0.432/0.372. Spectrogram models are good at predicting StartHesitation and Turn but bad at Walking. </p>\n<h1>Wavelet Models</h1>\n<p>Wavelets are similar to spectrograms but also different because wavelets have different frequency/time resolutions at different frequencies. Transforming a wave into a wavelet also does not reduce the dimensionality of the scaleogram (I think this is the term for the image you get after wavelet transform). Since there's no downsampling in the time dimension, Unet is no longer needed and I simply use a resnet18/34, which downsample the scaleogram to the same time resolution as spectrogram models after Unet. In turn, I'm also classifying 0.5 sec windows. Similarly, sequences are cut into chunks of 128 secs (256 spectrogram frames) to generate spectrograms.</p>\n<p>Best submitted single spectrogram model scores 0.386/0.345. Wavelet models are good at predicting Walking but bad at StartHesitation and Turn, so it complements spectrogram models nicely.  </p>\n<h1>Transformer modeling</h1>\n<p>Just plugging in a transformer actually does not work so well because it overfits, so I bias the transformer self-attention with a float bias mask to condition each prediction on its adjacent time points, which helps the model predict long events well. </p>\n<p>What the float bias mask looks like (see above) <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F3db704367c0eb1a6f7e46f82684e3f31%2FFOG_attention_mask.png?generation=1686277393310664&amp;alt=media\" alt=\"\"></p>\n<pre><code>torch.((L,L))\n\n\n    j:\n                m[i,j]=(-(i-j)/L)**\n\n             i==j:\n                m[i,i]=\n     m\n</code></pre>\n<h1>Data augmentation</h1>\n<p>I used data augmentation for spectrogram/wavelet models including the following (mostly from audiomentations): </p>\n<ol>\n<li>time stretch </li>\n<li>gaussian noise</li>\n<li>pitch shift </li>\n<li>wave scale aug (randomly multiply wave by 0.75~1.5 to get scale invariance)</li>\n<li>time feature shift aug.  </li>\n</ol>\n<pre><code>        self.augment = Compose([\n            AddGaussianNoise(=0.001, =0.015, =0.5),\n            TimeStretch(=0.8, =1.25, =0.5,leave_length_unchanged=False,n_fft=n_fft,hop_length=hop_length),\n            PitchShift(=-4, =4, =0.5,n_fft=n_fft,hop_length=hop_length),\n        ])\n</code></pre>\n<pre><code>\n np.random.uniform()&gt;.:\n     ['wave']*=np.random.uniform(.,.)\n</code></pre>\n<pre><code> feature shift aug\n self and np()&gt;:\n    data=data+np(-,)\n</code></pre>\n<h1>Frequency encoding and range</h1>\n<p>It's important to not use the high frequency bins of fourier transform, so I simply discard them and only keep the first 64 bins, corresponding to 0-15 hz. For spectrogram models, I also encode the frequency bin with torch.linspace(0,15,n_bins) expanded in the time and channel dimension and concatted so the input to the 2D conv network has 4 channels (3 directions of spectrograms + frequency encoding). It was also useful to resample the waves to a lower frequency, which I think reduces the level of noise. I used 32, 64, and 128 hz for spectrogram models and 64 hz for wavelet models. Defog waves are resampled to match the sample rate of tdcsfog waves.</p>\n<pre><code> self==:\n    data=FA(data,,self.sample_rate)\n:\n    data=FA(data,,self.sample_rate)\n</code></pre>\n<h1>1D conv Models</h1>\n<p>1D conv models are Ahmet's solution. Please see below for details:</p>\n<ol>\n<li>First align defog and tdcs on time axis (downsampled by 32 and 25, but kept their std as a feature)</li>\n<li>pct_time, total_len, Test are used as independent features. Their prediction is summed with the prediction from the 1D CNN.</li>\n<li>Because the input was only around 7 seconds long, cumsum features are also fed into 1D CNN.</li>\n<li>Outlier dominant subject is downweighted.</li>\n<li>Used snapshot ensembling.</li>\n<li>Used notype data by applying max on the predictions.</li>\n</ol>\n<p>1D conv models are weaker compared to the other 2, scoring 0.373/0.293, but are still a nice addition to the ensemble. Interestingly, 1D conv models and spectrogram models have a similar gap of 0.09 between public and private, whereas wavelet models have only a gap of 0.04. We think this is due to a change in class balance between private/public where public has more start hesitation and private has more walking. </p>\n<h1>Ensemble Weight Tuning</h1>\n<p>For our big ensemble, the weights are first hand tuned as a starting point and then I used GP_minimize to maximize CV score. We used 2 weight tuning setups at the end 1. map of 4 folds + map of full data excluding Subject 2d57c2, 2.  map of 3 folds excluding fold with Subject 2d57c2 + map of full data excluding Subject 2d57c2. We do this because we consider Subject 2d57c2 to an outlier. </p>\n<pre><code>results=gp\n</code></pre>\n<p>The weights for our models are (we downweight loss of 2d57c2 to 0.2 in some of them)</p>\n<p>[0.1821, 0.2792, 0.1052] Ahemt model<br>\n[0.2153, 0.0, 0.0] test257 (32 hz spectrogram)<br>\n[0.6026, 0.1734, 0.0] test262 (64 hz spectrogram)<br>\n[0.0, 0.0287, 0.2579] test264 wavelet<br>\n[0.0, 0.2168, 0.0] test265 (32 hz spectrogram) downweight 2d57c2<br>\n[0.0, 0.1734, 0.2997] test266 wavelet downweight 2d57c2<br>\n[0.0, 0.1284, 0.1124] test263 128 hz spec<br>\n[0.0, 0.0, 0.2248] test271 wavelet double freq scales</p>\n<p>Let me know if you have questions, and I wouldn't be surprised if I forgot to mention some details . The code is a bit messy atm but i will clean up and release it soon. </p>",
      "rawMarkdown": "First, thanks to the competition hosts for a meaningful competition with interesting data. Second, congrats to the winners! My team was able to shake up a bit but ended up one place shy of the prize zone. All in all, I'm still happy since we did well to survive the shakeup. I'm also excited to see what top 5 teams did to create a significant gap between us. \n\nOur final ensemble is a combination of spectrogram models, wavelet models, and 1D conv models, which scores 0.369/0.462. Below I will discuss them as well as other important technical details. See below for an overview ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F12e8af388e67c174b467596a85afebdf%2FFOG_overview.png?generation=1686277343043841&alt=media)\n\n# Validation setup\n\nValidation setup is important due to the noisy data. I ended up with a nested CV setup. The procedure is as follows:\n1. split data into 4 folds stratified by data type and grouped by subject\n2. set aside the validation fold (i call this the outer fold)\n3. resplit the 3 training folds into 4 folds and do cross validation (I call these inner folds)\n4. take last epoch or epoch with best validation score\n5. evaluate on outer fold with 4 inner fold models averaged for each outer fold \nWith this setup, we can more accurately simulate a situation where we have 250 sequences in the test set and we avg the fold model predictions. Later on, I switched to training on full inner fold data for 4 times without validation and evaluate with last epoch models on the outer fold set. \n\n# Input features\nAll of our models use the 3 waves and the pct time feature. In addition, Ahmet uses some metadata features in his models.\n\n# Spectrogram Models\n\nWhen I first saw the data I thought it looked like some sort of waveform, like audio data, so I thought it might work well to use spectrograms to model it. The 3 dimension waves are transformed into 2D spectrograms with STFT. Importantly, transforming the data in spectrograms significantly downscaled the data in the time dimension, so since I use a hop length of 64/50, each frame represents a 0.5 secs window and I'm basically making predictions for 0.5 sec windows. During training, labels are resized with torchvision's resize to fit the size of the time dimension of the spectrograms and during inference the model output is resized back to full dimensionality. Sequences are cut into chunks of 128 secs (256 spectrogram frames) to generate spectrograms. \n\nAnother important thing with using spectrograms is that if we use a regular type 2D conv model like resnet18, it wouldn't preserve the full dimensionality of the spectrogram (e.g a 256x256 becomes 8x8 after resnet18). In order to circumvent that, I thought to use a UNet to upsample the small feature map after the conv network. Following that, the spectrograms are pooled along the frequency dimension so I have a 1D sequence, which is then inputted into a transformer network before outputting predictions. \n\nBest submitted single spectrogram model scores 0.432/0.372. Spectrogram models are good at predicting StartHesitation and Turn but bad at Walking. \n\n# Wavelet Models\n\nWavelets are similar to spectrograms but also different because wavelets have different frequency/time resolutions at different frequencies. Transforming a wave into a wavelet also does not reduce the dimensionality of the scaleogram (I think this is the term for the image you get after wavelet transform). Since there's no downsampling in the time dimension, Unet is no longer needed and I simply use a resnet18/34, which downsample the scaleogram to the same time resolution as spectrogram models after Unet. In turn, I'm also classifying 0.5 sec windows. Similarly, sequences are cut into chunks of 128 secs (256 spectrogram frames) to generate spectrograms.\n\nBest submitted single spectrogram model scores 0.386/0.345. Wavelet models are good at predicting Walking but bad at StartHesitation and Turn, so it complements spectrogram models nicely.  \n\n# Transformer modeling\n\nJust plugging in a transformer actually does not work so well because it overfits, so I bias the transformer self-attention with a float bias mask to condition each prediction on its adjacent time points, which helps the model predict long events well. \n\nWhat the float bias mask looks like (see above) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F3db704367c0eb1a6f7e46f82684e3f31%2FFOG_attention_mask.png?generation=1686277393310664&alt=media)\n\n```\ndef get_distance_mask(L,power):\n\n    m=torch.zeros((L,L))\n\n\n    for i in range(L):\n        for j in range(L):\n            if i!=j:\n                m[i,j]=(1-abs(i-j)/L)**2\n\n            if i==j:\n                m[i,i]=1.0\n    return m\n\n```\n\n# Data augmentation\n\nI used data augmentation for spectrogram/wavelet models including the following (mostly from audiomentations): \n1. time stretch \n2. gaussian noise\n3. pitch shift \n4. wave scale aug (randomly multiply wave by 0.75~1.5 to get scale invariance)\n5. time feature shift aug.  \n\n\n```\n        self.augment = Compose([\n            AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.5),\n            TimeStretch(min_rate=0.8, max_rate=1.25, p=0.5,leave_length_unchanged=False,n_fft=n_fft,hop_length=hop_length),\n            PitchShift(min_semitones=-4, max_semitones=4, p=0.5,n_fft=n_fft,hop_length=hop_length),\n        ])\n```\n```\n#wave scale aug\nif np.random.uniform()>0.5:\n     data['wave']*=np.random.uniform(0.75,1.5)\n```\n```\n#time feature shift aug\nif self.train and np.random.uniform()>0.5:\n    data['time']=data['time']+np.random.uniform(-0.1,0.1)\n```\n\n# Frequency encoding and range\n\nIt's important to not use the high frequency bins of fourier transform, so I simply discard them and only keep the first 64 bins, corresponding to 0-15 hz. For spectrogram models, I also encode the frequency bin with torch.linspace(0,15,n_bins) expanded in the time and channel dimension and concatted so the input to the 2D conv network has 4 channels (3 directions of spectrograms + frequency encoding). It was also useful to resample the waves to a lower frequency, which I think reduces the level of noise. I used 32, 64, and 128 hz for spectrogram models and 64 hz for wavelet models. Defog waves are resampled to match the sample rate of tdcsfog waves.\n\n```\nif self.df.loc[idx,'data_type']=='defog':\n    data['wave']=FA.resample(data['wave'],100,self.sample_rate)\nelse:\n    data['wave']=FA.resample(data['wave'],128,self.sample_rate)\n```\n\n# 1D conv Models\n1D conv models are Ahmet's solution. Please see below for details:\n1. First align defog and tdcs on time axis (downsampled by 32 and 25, but kept their std as a feature)\n2. pct_time, total_len, Test are used as independent features. Their prediction is summed with the prediction from the 1D CNN.\n3. Because the input was only around 7 seconds long, cumsum features are also fed into 1D CNN.\n4. Outlier dominant subject is downweighted.\n5. Used snapshot ensembling.\n6. Used notype data by applying max on the predictions.\n\n1D conv models are weaker compared to the other 2, scoring 0.373/0.293, but are still a nice addition to the ensemble. Interestingly, 1D conv models and spectrogram models have a similar gap of 0.09 between public and private, whereas wavelet models have only a gap of 0.04. We think this is due to a change in class balance between private/public where public has more start hesitation and private has more walking. \n\n# Ensemble Weight Tuning \n\nFor our big ensemble, the weights are first hand tuned as a starting point and then I used GP_minimize to maximize CV score. We used 2 weight tuning setups at the end 1. map of 4 folds + map of full data excluding Subject 2d57c2, 2.  map of 3 folds excluding fold with Subject 2d57c2 + map of full data excluding Subject 2d57c2. We do this because we consider Subject 2d57c2 to an outlier. \n\n\n```\nresults=gp_minimize(get_score,boundaries,x0=w,verbose=1,n_jobs=48,acq_optimizer='lbfgs',random_state=0)\n```\n\nThe weights for our models are (we downweight loss of 2d57c2 to 0.2 in some of them)\n\n[0.1821, 0.2792, 0.1052] Ahemt model\n[0.2153, 0.0, 0.0] test257 (32 hz spectrogram)\n[0.6026, 0.1734, 0.0] test262 (64 hz spectrogram)\n[0.0, 0.0287, 0.2579] test264 wavelet\n[0.0, 0.2168, 0.0] test265 (32 hz spectrogram) downweight 2d57c2\n[0.0, 0.1734, 0.2997] test266 wavelet downweight 2d57c2\n[0.0, 0.1284, 0.1124] test263 128 hz spec\n[0.0, 0.0, 0.2248] test271 wavelet double freq scales\n\nLet me know if you have questions, and I wouldn't be surprised if I forgot to mention some details . The code is a bit messy atm but i will clean up and release it soon.",
      "votes": null
    },
    {
      "id": "2293202",
      "postDate": "06/09/2023 02:57:48",
      "content": "<p>It seems that spectrogram models are very powerful and really amazed me.👍 </p>",
      "rawMarkdown": "It seems that spectrogram models are very powerful and really amazed me.👍",
      "votes": null
    },
    {
      "id": "2293483",
      "postDate": "06/09/2023 08:11:25",
      "content": "<p>Thanks! Good job on the shakeup!</p>",
      "rawMarkdown": "Thanks! Good job on the shakeup!",
      "votes": null
    },
    {
      "id": "2293617",
      "postDate": "06/09/2023 10:37:39",
      "content": "<p>I also implemented something close to your wavelet model but with efficentnetB0. My model is not as performant as yours but I also got a difference of 0.04 between private and public LB.</p>\n<p>Also thanks for making me aware of <code>skopt.gp_minimize</code> (or the <code>skopt</code>package in general). Great Work!</p>",
      "rawMarkdown": "I also implemented something close to your wavelet model but with efficentnetB0. My model is not as performant as yours but I also got a difference of 0.04 between private and public LB.\n\nAlso thanks for making me aware of `skopt.gp_minimize` (or the `skopt `package in general). Great Work!",
      "votes": null
    },
    {
      "id": "2294021",
      "postDate": "06/09/2023 16:59:07",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> !</p>\n<p>Interesting to see how you used wavelets and spectrograms in order to balance the prediction power between StartHesitation, Turn and Walking….</p>\n<p>I've seen you merged TDCSFOG and DEFOG data series. Have you tried to create diferent models for each input dataset, if so what was the result?</p>\n<p>Thank you for sharing your knowledge with the community. Keep up the great work🔥💪</p>",
      "rawMarkdown": "Congratulations, @shujun717 !\n\nInteresting to see how you used wavelets and spectrograms in order to balance the prediction power between StartHesitation, Turn and Walking....\n\nI've seen you merged TDCSFOG and DEFOG data series. Have you tried to create diferent models for each input dataset, if so what was the result?\n\nThank you for sharing your knowledge with the community. Keep up the great work🔥💪",
      "votes": null
    },
    {
      "id": "2294348",
      "postDate": "06/10/2023 02:33:52",
      "content": "<p>Thanks. We tried separate models but it was worse… seeing the 1st and 2nd place solution does make me want to try again tho</p>",
      "rawMarkdown": "Thanks. We tried separate models but it was worse... seeing the 1st and 2nd place solution does make me want to try again tho",
      "votes": null
    },
    {
      "id": "2295064",
      "postDate": "06/10/2023 15:00:50",
      "content": "<p>Thanks for replying! Same here, it seem's we could have achieved greater results when considering different models…</p>\n<p>Congratulations once again🙏</p>",
      "rawMarkdown": "Thanks for replying! Same here, it seem's we could have achieved greater results when considering different models...\n\nCongratulations once again🙏",
      "votes": null
    },
    {
      "id": "2366327",
      "postDate": "07/30/2023 23:12:31",
      "content": "<p>A bit late, - have you tried to make attention bias mask learnable? [initialzed from the one showed in the post]</p>",
      "rawMarkdown": "A bit late, - have you tried to make attention bias mask learnable? [initialzed from the one showed in the post]",
      "votes": null
    },
    {
      "id": "2371304",
      "postDate": "08/03/2023 04:10:41",
      "content": "<p>Not really. We used longer sequence during inference to account for edge effects and learnable attention bias would have been bad (you can imagine if we learned 128x128 mask for 128 sequence during training and tried to do inference for 160 sequence). Additionally, with learnable attention bias mask, the model may learn absolute position information which I think wouldn't be good. </p>",
      "rawMarkdown": "Not really. We used longer sequence during inference to account for edge effects and learnable attention bias would have been bad (you can imagine if we learned 128x128 mask for 128 sequence during training and tried to do inference for 160 sequence). Additionally, with learnable attention bias mask, the model may learn absolute position information which I think wouldn't be good.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2293202,
      "author_name": "hydantess",
      "author_url": "",
      "post_date": "06/09/2023 02:57:48",
      "content": "<p>It seems that spectrogram models are very powerful and really amazed me.👍 </p>",
      "votes": null,
      "replies": [
        {
          "id": 2293483,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "06/09/2023 08:11:25",
          "content": "<p>Thanks! Good job on the shakeup!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2293617,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/09/2023 10:37:39",
      "content": "<p>I also implemented something close to your wavelet model but with efficentnetB0. My model is not as performant as yours but I also got a difference of 0.04 between private and public LB.</p>\n<p>Also thanks for making me aware of <code>skopt.gp_minimize</code> (or the <code>skopt</code>package in general). Great Work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2294021,
      "author_name": "vladiluzjr",
      "author_url": "",
      "post_date": "06/09/2023 16:59:07",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> !</p>\n<p>Interesting to see how you used wavelets and spectrograms in order to balance the prediction power between StartHesitation, Turn and Walking….</p>\n<p>I've seen you merged TDCSFOG and DEFOG data series. Have you tried to create diferent models for each input dataset, if so what was the result?</p>\n<p>Thank you for sharing your knowledge with the community. Keep up the great work🔥💪</p>",
      "votes": null,
      "replies": [
        {
          "id": 2294348,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "06/10/2023 02:33:52",
          "content": "<p>Thanks. We tried separate models but it was worse… seeing the 1st and 2nd place solution does make me want to try again tho</p>",
          "votes": null,
          "replies": [
            {
              "id": 2295064,
              "author_name": "vladiluzjr",
              "author_url": "",
              "post_date": "06/10/2023 15:00:50",
              "content": "<p>Thanks for replying! Same here, it seem's we could have achieved greater results when considering different models…</p>\n<p>Congratulations once again🙏</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2366327,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "07/30/2023 23:12:31",
      "content": "<p>A bit late, - have you tried to make attention bias mask learnable? [initialzed from the one showed in the post]</p>",
      "votes": null,
      "replies": [
        {
          "id": 2371304,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "08/03/2023 04:10:41",
          "content": "<p>Not really. We used longer sequence during inference to account for edge effects and learnable attention bias would have been bad (you can imagine if we learned 128x128 mask for 128 sequence during training and tried to do inference for 160 sequence). Additionally, with learnable attention bias mask, the model may learn absolute position information which I think wouldn't be good. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2293189": "First, thanks to the competition hosts for a meaningful competition with interesting data. Second, congrats to the winners! My team was able to shake up a bit but ended up one place shy of the prize zone. All in all, I'm still happy since we did well to survive the shakeup. I'm also excited to see what top 5 teams did to create a significant gap between us. \n\nOur final ensemble is a combination of spectrogram models, wavelet models, and 1D conv models, which scores 0.369/0.462. Below I will discuss them as well as other important technical details. See below for an overview ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F12e8af388e67c174b467596a85afebdf%2FFOG_overview.png?generation=1686277343043841&alt=media)\n\n# Validation setup\n\nValidation setup is important due to the noisy data. I ended up with a nested CV setup. The procedure is as follows:\n1. split data into 4 folds stratified by data type and grouped by subject\n2. set aside the validation fold (i call this the outer fold)\n3. resplit the 3 training folds into 4 folds and do cross validation (I call these inner folds)\n4. take last epoch or epoch with best validation score\n5. evaluate on outer fold with 4 inner fold models averaged for each outer fold \nWith this setup, we can more accurately simulate a situation where we have 250 sequences in the test set and we avg the fold model predictions. Later on, I switched to training on full inner fold data for 4 times without validation and evaluate with last epoch models on the outer fold set. \n\n# Input features\nAll of our models use the 3 waves and the pct time feature. In addition, Ahmet uses some metadata features in his models.\n\n# Spectrogram Models\n\nWhen I first saw the data I thought it looked like some sort of waveform, like audio data, so I thought it might work well to use spectrograms to model it. The 3 dimension waves are transformed into 2D spectrograms with STFT. Importantly, transforming the data in spectrograms significantly downscaled the data in the time dimension, so since I use a hop length of 64/50, each frame represents a 0.5 secs window and I'm basically making predictions for 0.5 sec windows. During training, labels are resized with torchvision's resize to fit the size of the time dimension of the spectrograms and during inference the model output is resized back to full dimensionality. Sequences are cut into chunks of 128 secs (256 spectrogram frames) to generate spectrograms. \n\nAnother important thing with using spectrograms is that if we use a regular type 2D conv model like resnet18, it wouldn't preserve the full dimensionality of the spectrogram (e.g a 256x256 becomes 8x8 after resnet18). In order to circumvent that, I thought to use a UNet to upsample the small feature map after the conv network. Following that, the spectrograms are pooled along the frequency dimension so I have a 1D sequence, which is then inputted into a transformer network before outputting predictions. \n\nBest submitted single spectrogram model scores 0.432/0.372. Spectrogram models are good at predicting StartHesitation and Turn but bad at Walking. \n\n# Wavelet Models\n\nWavelets are similar to spectrograms but also different because wavelets have different frequency/time resolutions at different frequencies. Transforming a wave into a wavelet also does not reduce the dimensionality of the scaleogram (I think this is the term for the image you get after wavelet transform). Since there's no downsampling in the time dimension, Unet is no longer needed and I simply use a resnet18/34, which downsample the scaleogram to the same time resolution as spectrogram models after Unet. In turn, I'm also classifying 0.5 sec windows. Similarly, sequences are cut into chunks of 128 secs (256 spectrogram frames) to generate spectrograms.\n\nBest submitted single spectrogram model scores 0.386/0.345. Wavelet models are good at predicting Walking but bad at StartHesitation and Turn, so it complements spectrogram models nicely.  \n\n# Transformer modeling\n\nJust plugging in a transformer actually does not work so well because it overfits, so I bias the transformer self-attention with a float bias mask to condition each prediction on its adjacent time points, which helps the model predict long events well. \n\nWhat the float bias mask looks like (see above) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F3db704367c0eb1a6f7e46f82684e3f31%2FFOG_attention_mask.png?generation=1686277393310664&alt=media)\n\n```\ndef get_distance_mask(L,power):\n\n    m=torch.zeros((L,L))\n\n\n    for i in range(L):\n        for j in range(L):\n            if i!=j:\n                m[i,j]=(1-abs(i-j)/L)**2\n\n            if i==j:\n                m[i,i]=1.0\n    return m\n\n```\n\n# Data augmentation\n\nI used data augmentation for spectrogram/wavelet models including the following (mostly from audiomentations): \n1. time stretch \n2. gaussian noise\n3. pitch shift \n4. wave scale aug (randomly multiply wave by 0.75~1.5 to get scale invariance)\n5. time feature shift aug.  \n\n\n```\n        self.augment = Compose([\n            AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.5),\n            TimeStretch(min_rate=0.8, max_rate=1.25, p=0.5,leave_length_unchanged=False,n_fft=n_fft,hop_length=hop_length),\n            PitchShift(min_semitones=-4, max_semitones=4, p=0.5,n_fft=n_fft,hop_length=hop_length),\n        ])\n```\n```\n#wave scale aug\nif np.random.uniform()>0.5:\n     data['wave']*=np.random.uniform(0.75,1.5)\n```\n```\n#time feature shift aug\nif self.train and np.random.uniform()>0.5:\n    data['time']=data['time']+np.random.uniform(-0.1,0.1)\n```\n\n# Frequency encoding and range\n\nIt's important to not use the high frequency bins of fourier transform, so I simply discard them and only keep the first 64 bins, corresponding to 0-15 hz. For spectrogram models, I also encode the frequency bin with torch.linspace(0,15,n_bins) expanded in the time and channel dimension and concatted so the input to the 2D conv network has 4 channels (3 directions of spectrograms + frequency encoding). It was also useful to resample the waves to a lower frequency, which I think reduces the level of noise. I used 32, 64, and 128 hz for spectrogram models and 64 hz for wavelet models. Defog waves are resampled to match the sample rate of tdcsfog waves.\n\n```\nif self.df.loc[idx,'data_type']=='defog':\n    data['wave']=FA.resample(data['wave'],100,self.sample_rate)\nelse:\n    data['wave']=FA.resample(data['wave'],128,self.sample_rate)\n```\n\n# 1D conv Models\n1D conv models are Ahmet's solution. Please see below for details:\n1. First align defog and tdcs on time axis (downsampled by 32 and 25, but kept their std as a feature)\n2. pct_time, total_len, Test are used as independent features. Their prediction is summed with the prediction from the 1D CNN.\n3. Because the input was only around 7 seconds long, cumsum features are also fed into 1D CNN.\n4. Outlier dominant subject is downweighted.\n5. Used snapshot ensembling.\n6. Used notype data by applying max on the predictions.\n\n1D conv models are weaker compared to the other 2, scoring 0.373/0.293, but are still a nice addition to the ensemble. Interestingly, 1D conv models and spectrogram models have a similar gap of 0.09 between public and private, whereas wavelet models have only a gap of 0.04. We think this is due to a change in class balance between private/public where public has more start hesitation and private has more walking. \n\n# Ensemble Weight Tuning \n\nFor our big ensemble, the weights are first hand tuned as a starting point and then I used GP_minimize to maximize CV score. We used 2 weight tuning setups at the end 1. map of 4 folds + map of full data excluding Subject 2d57c2, 2.  map of 3 folds excluding fold with Subject 2d57c2 + map of full data excluding Subject 2d57c2. We do this because we consider Subject 2d57c2 to an outlier. \n\n\n```\nresults=gp_minimize(get_score,boundaries,x0=w,verbose=1,n_jobs=48,acq_optimizer='lbfgs',random_state=0)\n```\n\nThe weights for our models are (we downweight loss of 2d57c2 to 0.2 in some of them)\n\n[0.1821, 0.2792, 0.1052] Ahemt model\n[0.2153, 0.0, 0.0] test257 (32 hz spectrogram)\n[0.6026, 0.1734, 0.0] test262 (64 hz spectrogram)\n[0.0, 0.0287, 0.2579] test264 wavelet\n[0.0, 0.2168, 0.0] test265 (32 hz spectrogram) downweight 2d57c2\n[0.0, 0.1734, 0.2997] test266 wavelet downweight 2d57c2\n[0.0, 0.1284, 0.1124] test263 128 hz spec\n[0.0, 0.0, 0.2248] test271 wavelet double freq scales\n\nLet me know if you have questions, and I wouldn't be surprised if I forgot to mention some details . The code is a bit messy atm but i will clean up and release it soon.",
    "2293202": "It seems that spectrogram models are very powerful and really amazed me.👍",
    "2293483": "Thanks! Good job on the shakeup!",
    "2293617": "I also implemented something close to your wavelet model but with efficentnetB0. My model is not as performant as yours but I also got a difference of 0.04 between private and public LB.\n\nAlso thanks for making me aware of `skopt.gp_minimize` (or the `skopt `package in general). Great Work!",
    "2294021": "Congratulations, @shujun717 !\n\nInteresting to see how you used wavelets and spectrograms in order to balance the prediction power between StartHesitation, Turn and Walking....\n\nI've seen you merged TDCSFOG and DEFOG data series. Have you tried to create diferent models for each input dataset, if so what was the result?\n\nThank you for sharing your knowledge with the community. Keep up the great work🔥💪",
    "2294348": "Thanks. We tried separate models but it was worse... seeing the 1st and 2nd place solution does make me want to try again tho",
    "2295064": "Thanks for replying! Same here, it seem's we could have achieved greater results when considering different models...\n\nCongratulations once again🙏",
    "2366327": "A bit late, - have you tried to make attention bias mask learnable? [initialzed from the one showed in the post]",
    "2371304": "Not really. We used longer sequence during inference to account for edge effects and learnable attention bias would have been bad (you can imagine if we learned 128x128 mask for 128 sequence during training and tried to do inference for 160 sequence). Additionally, with learnable attention bias mask, the model may learn absolute position information which I think wouldn't be good."
  },
  "source": "meta"
}