{
  "id": 275337,
  "title": "Notes on G2 Comp",
  "url": "/competitions/g2net-gravitational-wave-detection/writeups/neutron-net-notes-on-g2-comp",
  "author_name": "",
  "post_date": "2021-09-30T01:37:12.233Z",
  "votes": 25,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In total, I trained 96 full-fold models for this competition. Having nice local compute helps, but having deep insight into the problem + good exploratory prowess is more important (both of which I lack). I guess predestination / luck also plays its role :-). Some of the things I had tried off the top of my head (since I've already deleted my working directory):</p>\n<ul>\n<li>Channel shuffle</li>\n<li>Channel shifting</li>\n<li>Random noise</li>\n<li>1D Net, all my 1D nets were really shallow, just 4 layers w/o residual</li>\n<li>2D Net, couldn't get them to converge as nicely as the rest of folks</li>\n<li>Downloading actual LIGO 3 detector data for pretraining - it's very messy and has bands not represented in the synthetic data</li>\n<li>GW injection using pycbc in 1d and 2d</li>\n<li>Injected gw param regression as pretrain task</li>\n<li>Cutout</li>\n<li>Mixup</li>\n<li>Trainable CQT</li>\n<li>Very wide kernel sizes</li>\n<li>Kernel smoothing and resume training</li>\n<li>Frequency domain CQT</li>\n<li>FD 1D net, 6 channels</li>\n<li>Training on bandpassed first derivative data</li>\n<li>Training on bandpassed integrated data</li>\n<li>PCA, SVD, and FastICA transformations as channel augmentations</li>\n<li>Signal flipping (only works in 1D <code>x=-x</code>. You can train a model like this then do TTA)</li>\n<li>Signal compression and stretching - I only attempted this in 1D as TTA and for training</li>\n<li>Random Mean and STD shift TTA</li>\n<li>Whitening by creating ~6sec signals: basically I'd concat the signal with itself on the right and left, but first I'd find the location on the right that has the most similar value and slope and repeat from there. Same as the left. I thought by using longer signal, we take care of the ASD generation issue. It 'worked' but the model did not benefit</li>\n<li>Whitening using auto regression residuals</li>\n<li>Swin Transformer</li>\n<li>Galerkin transformer in 1D time domain, bucketing signal into 8hz wide chunks</li>\n<li>Probably more that I am forgetting</li>\n</ul>\n<p>I feel like instead of trying 903209230923 things, better to abandon quickly what doesn't work, focus on what does, and importantly, drive it to its logical conclusion. I was very early on the 1D train and was training nets with the full competition data in ram. But after I got to a decent level in 1D (at the time) I started exploring all these different avenues instead without focusing more in improving the net's architecture.</p>\n<p>I'd also like to express gratitude and appreciation to my team mates. I actually ran out of steam about two weeks ago, but contributions and motivation from <a href=\"https://www.kaggle.com/misaki1111\" target=\"_blank\">@misaki1111</a> and <a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> helped carry me over. Thank you.</p>",
  "messages": [
    {
      "id": "1528826",
      "postDate": "09/30/2021 01:33:53",
      "content": "<p>In total, I trained 96 full-fold models for this competition. Having nice local compute helps, but having deep insight into the problem + good exploratory prowess is more important (both of which I lack). I guess predestination / luck also plays its role :-). Some of the things I had tried off the top of my head (since I've already deleted my working directory):</p>\n<ul>\n<li>Channel shuffle</li>\n<li>Channel shifting</li>\n<li>Random noise</li>\n<li>1D Net, all my 1D nets were really shallow, just 4 layers w/o residual</li>\n<li>2D Net, couldn't get them to converge as nicely as the rest of folks</li>\n<li>Downloading actual LIGO 3 detector data for pretraining - it's very messy and has bands not represented in the synthetic data</li>\n<li>GW injection using pycbc in 1d and 2d</li>\n<li>Injected gw param regression as pretrain task</li>\n<li>Cutout</li>\n<li>Mixup</li>\n<li>Trainable CQT</li>\n<li>Very wide kernel sizes</li>\n<li>Kernel smoothing and resume training</li>\n<li>Frequency domain CQT</li>\n<li>FD 1D net, 6 channels</li>\n<li>Training on bandpassed first derivative data</li>\n<li>Training on bandpassed integrated data</li>\n<li>PCA, SVD, and FastICA transformations as channel augmentations</li>\n<li>Signal flipping (only works in 1D <code>x=-x</code>. You can train a model like this then do TTA)</li>\n<li>Signal compression and stretching - I only attempted this in 1D as TTA and for training</li>\n<li>Random Mean and STD shift TTA</li>\n<li>Whitening by creating ~6sec signals: basically I'd concat the signal with itself on the right and left, but first I'd find the location on the right that has the most similar value and slope and repeat from there. Same as the left. I thought by using longer signal, we take care of the ASD generation issue. It 'worked' but the model did not benefit</li>\n<li>Whitening using auto regression residuals</li>\n<li>Swin Transformer</li>\n<li>Galerkin transformer in 1D time domain, bucketing signal into 8hz wide chunks</li>\n<li>Probably more that I am forgetting</li>\n</ul>\n<p>I feel like instead of trying 903209230923 things, better to abandon quickly what doesn't work, focus on what does, and importantly, drive it to its logical conclusion. I was very early on the 1D train and was training nets with the full competition data in ram. But after I got to a decent level in 1D (at the time) I started exploring all these different avenues instead without focusing more in improving the net's architecture.</p>\n<p>I'd also like to express gratitude and appreciation to my team mates. I actually ran out of steam about two weeks ago, but contributions and motivation from <a href=\"https://www.kaggle.com/misaki1111\" target=\"_blank\">@misaki1111</a> and <a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> helped carry me over. Thank you.</p>",
      "rawMarkdown": "In total, I trained 96 full-fold models for this competition. Having nice local compute helps, but having deep insight into the problem + good exploratory prowess is more important (both of which I lack). I guess predestination / luck also plays its role :-). Some of the things I had tried off the top of my head (since I've already deleted my working directory):\n\n- Channel shuffle\n- Channel shifting\n- Random noise\n- 1D Net, all my 1D nets were really shallow, just 4 layers w/o residual\n- 2D Net, couldn't get them to converge as nicely as the rest of folks\n- Downloading actual LIGO 3 detector data for pretraining - it's very messy and has bands not represented in the synthetic data\n- GW injection using pycbc in 1d and 2d\n- Injected gw param regression as pretrain task\n- Cutout\n- Mixup\n- Trainable CQT\n- Very wide kernel sizes\n- Kernel smoothing and resume training\n- Frequency domain CQT\n- FD 1D net, 6 channels\n- Training on bandpassed first derivative data\n- Training on bandpassed integrated data\n- PCA, SVD, and FastICA transformations as channel augmentations\n- Signal flipping (only works in 1D `x=-x`. You can train a model like this then do TTA)\n- Signal compression and stretching - I only attempted this in 1D as TTA and for training\n- Random Mean and STD shift TTA\n- Whitening by creating ~6sec signals: basically I'd concat the signal with itself on the right and left, but first I'd find the location on the right that has the most similar value and slope and repeat from there. Same as the left. I thought by using longer signal, we take care of the ASD generation issue. It 'worked' but the model did not benefit\n- Whitening using auto regression residuals\n- Swin Transformer\n- Galerkin transformer in 1D time domain, bucketing signal into 8hz wide chunks\n- Probably more that I am forgetting\n\nI feel like instead of trying 903209230923 things, better to abandon quickly what doesn't work, focus on what does, and importantly, drive it to its logical conclusion. I was very early on the 1D train and was training nets with the full competition data in ram. But after I got to a decent level in 1D (at the time) I started exploring all these different avenues instead without focusing more in improving the net's architecture.\n\nI'd also like to express gratitude and appreciation to my team mates. I actually ran out of steam about two weeks ago, but contributions and motivation from @misaki1111 and @jpison helped carry me over. Thank you.",
      "votes": null
    },
    {
      "id": "1528835",
      "postDate": "09/30/2021 01:44:56",
      "content": "<p>I mean if you want to compare with me:</p>\n<ul>\n<li>Channel shuffle +</li>\n<li>Channel shifting + </li>\n<li>Random noise +</li>\n<li>1D Net, all my 1D nets were really shallow, just 4 layers w/o residual +</li>\n<li>2D Net, couldn't get them to converge as nicely as the rest of folks +</li>\n<li>Downloading actual LIGO 3 detector data for pretraining - it's very messy and has bands not represented in the synthetic data | omg</li>\n<li>GW injection using pycbc in 1d and 2d -</li>\n<li>Injected gw param regression as pretrain task - </li>\n<li>Cutout -</li>\n<li>Mixup +</li>\n<li>Trainable CQT +</li>\n<li>Very wide kernel sizes +</li>\n<li>Kernel smoothing and resume training -+</li>\n<li>Frequency domain CQT +</li>\n<li>FD 1D net, 6 channels -</li>\n<li>Training on bandpassed first derivative data -</li>\n<li>Training on bandpassed integrated data -</li>\n<li>PCA, SVD, and FastICA transformations as channel augmentations -</li>\n<li>Signal flipping (only works in 1D x=-x. You can train a model like this then do TTA) -</li>\n<li>Signal compression and stretching - I only attempted this in 1D as TTA and for training +</li>\n<li>Random Mean and STD shift TTA +</li>\n<li>Whitening by creating ~6sec signals: basically I'd concat the signal with itself on the right and left, but first I'd find the location on the right that has the most similar value and slope and repeat from there. Same as the left. I thought by using longer signal, we take care of the ASD generation issue. It 'worked' but the model did not benefit | -+</li>\n<li>Whitening using auto regression residuals +</li>\n<li>Swin Transformer +, but xcite was better</li>\n<li>Galerkin transformer in 1D time domain, bucketing signal into 8hz wide chunks - </li>\n</ul>\n<p>where + is I also did that and - i dont. Most of the things you did i would consider is on the money.</p>",
      "rawMarkdown": "I mean if you want to compare with me:\n\n- Channel shuffle +\n- Channel shifting + \n- Random noise +\n- 1D Net, all my 1D nets were really shallow, just 4 layers w/o residual +\n- 2D Net, couldn't get them to converge as nicely as the rest of folks +\n- Downloading actual LIGO 3 detector data for pretraining - it's very messy and has bands not represented in the synthetic data | omg\n- GW injection using pycbc in 1d and 2d -\n- Injected gw param regression as pretrain task - \n- Cutout -\n- Mixup +\n- Trainable CQT +\n- Very wide kernel sizes +\n- Kernel smoothing and resume training -+\n- Frequency domain CQT +\n- FD 1D net, 6 channels -\n- Training on bandpassed first derivative data -\n- Training on bandpassed integrated data -\n- PCA, SVD, and FastICA transformations as channel augmentations -\n- Signal flipping (only works in 1D x=-x. You can train a model like this then do TTA) -\n- Signal compression and stretching - I only attempted this in 1D as TTA and for training +\n- Random Mean and STD shift TTA +\n- Whitening by creating ~6sec signals: basically I'd concat the signal with itself on the right and left, but first I'd find the location on the right that has the most similar value and slope and repeat from there. Same as the left. I thought by using longer signal, we take care of the ASD generation issue. It 'worked' but the model did not benefit | -+\n- Whitening using auto regression residuals +\n- Swin Transformer +, but xcite was better\n- Galerkin transformer in 1D time domain, bucketing signal into 8hz wide chunks - \n\nwhere + is I also did that and - i dont. Most of the things you did i would consider is on the money.",
      "votes": null
    },
    {
      "id": "1528865",
      "postDate": "09/30/2021 02:26:59",
      "content": "<p>A couple more things I tried, forgot to mention:</p>\n<ul>\n<li>My 1D nets were shallow as stated. Instead of .flatten(), I used GAP on the higher dimen channels</li>\n<li>I did seq2seq loss. I noticed this allowed models to converge a lot quicker. Both 1D and 2D models produce reduced-sized temporal representations. For ve+ all elements = 1 and =0 for ve-. I noticed for 1D models I noticed the resulting labels identified where the GW was. For 2D models, there was no such time encoding and it looked like whatever was being produced was all close to 1 or all close to 0.</li>\n<li>Training 1D model on S2S 'faux segmentation' from above</li>\n<li>Stacking - the first ~15 models used the same folds. But when I attempted submission pipeline, stacking sucked and I made that post. For all future models, I generated random folds</li>\n<li>Difficulty + Target Iterative Stratification: Using best model, compute residual and bucket into difficulty levels. Then multi-label stratify on that and target</li>\n<li>Final experiment blend (my portion): I just use cuml sgd to select weights on all of my experiments. CV for individual models ranged from 853-876, and the best blend of these was CV=0.8798601, public=0.8818, private=0.8796.</li>\n<li>I'll continue to add here as I recall things.</li>\n</ul>",
      "rawMarkdown": "A couple more things I tried, forgot to mention:\n\n- My 1D nets were shallow as stated. Instead of .flatten(), I used GAP on the higher dimen channels\n- I did seq2seq loss. I noticed this allowed models to converge a lot quicker. Both 1D and 2D models produce reduced-sized temporal representations. For ve+ all elements = 1 and =0 for ve-. I noticed for 1D models I noticed the resulting labels identified where the GW was. For 2D models, there was no such time encoding and it looked like whatever was being produced was all close to 1 or all close to 0.\n- Training 1D model on S2S 'faux segmentation' from above\n- Stacking - the first ~15 models used the same folds. But when I attempted submission pipeline, stacking sucked and I made that post. For all future models, I generated random folds\n- Difficulty + Target Iterative Stratification: Using best model, compute residual and bucket into difficulty levels. Then multi-label stratify on that and target\n- Final experiment blend (my portion): I just use cuml sgd to select weights on all of my experiments. CV for individual models ranged from 853-876, and the best blend of these was CV=0.8798601, public=0.8818, private=0.8796.\n- I'll continue to add here as I recall things.",
      "votes": null
    },
    {
      "id": "1529013",
      "postDate": "09/30/2021 04:59:25",
      "content": "<p>This is some fabulous work <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> and I don't think what you did went on vain as we also tried most of the things that you have mentioned above , it was a tough competition and to remain in the top half solo takes a lot of heart , so I would say well done and many congratulations</p>",
      "rawMarkdown": "This is some fabulous work @authman and I don't think what you did went on vain as we also tried most of the things that you have mentioned above , it was a tough competition and to remain in the top half solo takes a lot of heart , so I would say well done and many congratulations",
      "votes": null
    },
    {
      "id": "1529282",
      "postDate": "09/30/2021 09:08:05",
      "content": "<p>sometimes it is luck and if you would encounter a \"good bug\".</p>\n<p>When I first tried learnable CQT, it didn't work. And I decided to give up on it and continue tunning on my fixed CQT. But then, I forgot to update my code (set trainable flag to False). I have 3 local machines and the codes were scattering everywhere. The codes failed to sync between my machine.</p>\n<p>Then when I tried to make a summary (in the discussion forum and my private note), I needed to visualize the cqt images. I realized that the cqt images were noisy like never before. After debugging, I realized that I was using a learnable version and it had magically worked.</p>\n<p>Lessons learned: Stop and review … maybe you would discover something?</p>\n<p><img src=\"https://i.ibb.co/s1YQ734/Selection-949.png\" alt=\"https://i.ibb.co/s1YQ734/Selection-949.png\"></p>",
      "rawMarkdown": "sometimes it is luck and if you would encounter a \"good bug\".\n\nWhen I first tried learnable CQT, it didn't work. And I decided to give up on it and continue tunning on my fixed CQT. But then, I forgot to update my code (set trainable flag to False). I have 3 local machines and the codes were scattering everywhere. The codes failed to sync between my machine.\n\nThen when I tried to make a summary (in the discussion forum and my private note), I needed to visualize the cqt images. I realized that the cqt images were noisy like never before. After debugging, I realized that I was using a learnable version and it had magically worked.\n\n\nLessons learned: Stop and review ... maybe you would discover something?\n\n![https://i.ibb.co/s1YQ734/Selection-949.png](https://i.ibb.co/s1YQ734/Selection-949.png)",
      "votes": null
    },
    {
      "id": "1529487",
      "postDate": "09/30/2021 12:51:07",
      "content": "<p>You tried a lot of things.  Even if final result may not be as good as your hopes, I am sure that the knowledge you gained from trying all this will help you in future competitions.  It is by experimenting that we learn IMHO.</p>\n<p>Well done!</p>",
      "rawMarkdown": "You tried a lot of things.  Even if final result may not be as good as your hopes, I am sure that the knowledge you gained from trying all this will help you in future competitions.  It is by experimenting that we learn IMHO.\n\nWell done!",
      "votes": null
    },
    {
      "id": "1529537",
      "postDate": "09/30/2021 13:21:28",
      "content": "<p>Hey, slightly disheartened to see your team fall down on the private leaderboard. That's a lot of experiments, I don't think even we were able to experiment that much as a team. Hope you'll get your deserving gold medal in the next competition.<br>\nbtw how did \"PCA, SVD, and FastICA transformations as channel augmentations\" performed? </p>",
      "rawMarkdown": "Hey, slightly disheartened to see your team fall down on the private leaderboard. That's a lot of experiments, I don't think even we were able to experiment that much as a team. Hope you'll get your deserving gold medal in the next competition.\nbtw how did \"PCA, SVD, and FastICA transformations as channel augmentations\" performed?",
      "votes": null
    },
    {
      "id": "1529631",
      "postDate": "09/30/2021 14:39:19",
      "content": "<p>Nothing stellar, I believe that experiment performed 'average' compared to other models I had trained. What I recall the most about it was:</p>\n<ol>\n<li>FastICA complained a lot about convergence issues, so I had to really up the number of iterations and silence warnings when fitting it.</li>\n<li>tSVD of course requires n_components &lt; input_dimensions, so for that aug, one channel was lost. Rather than deleting, just choose a random (original) one to place back in.</li>\n<li>PCA was fast and straightforward to implement.</li>\n</ol>",
      "rawMarkdown": "Nothing stellar, I believe that experiment performed 'average' compared to other models I had trained. What I recall the most about it was:\n\n1. FastICA complained a lot about convergence issues, so I had to really up the number of iterations and silence warnings when fitting it.\n2. tSVD of course requires n_components < input_dimensions, so for that aug, one channel was lost. Rather than deleting, just choose a random (original) one to place back in.\n3. PCA was fast and straightforward to implement.",
      "votes": null
    },
    {
      "id": "1559886",
      "postDate": "10/27/2021 08:06:00",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1528835,
      "author_name": "bakeryproducts",
      "author_url": "",
      "post_date": "09/30/2021 01:44:56",
      "content": "<p>I mean if you want to compare with me:</p>\n<ul>\n<li>Channel shuffle +</li>\n<li>Channel shifting + </li>\n<li>Random noise +</li>\n<li>1D Net, all my 1D nets were really shallow, just 4 layers w/o residual +</li>\n<li>2D Net, couldn't get them to converge as nicely as the rest of folks +</li>\n<li>Downloading actual LIGO 3 detector data for pretraining - it's very messy and has bands not represented in the synthetic data | omg</li>\n<li>GW injection using pycbc in 1d and 2d -</li>\n<li>Injected gw param regression as pretrain task - </li>\n<li>Cutout -</li>\n<li>Mixup +</li>\n<li>Trainable CQT +</li>\n<li>Very wide kernel sizes +</li>\n<li>Kernel smoothing and resume training -+</li>\n<li>Frequency domain CQT +</li>\n<li>FD 1D net, 6 channels -</li>\n<li>Training on bandpassed first derivative data -</li>\n<li>Training on bandpassed integrated data -</li>\n<li>PCA, SVD, and FastICA transformations as channel augmentations -</li>\n<li>Signal flipping (only works in 1D x=-x. You can train a model like this then do TTA) -</li>\n<li>Signal compression and stretching - I only attempted this in 1D as TTA and for training +</li>\n<li>Random Mean and STD shift TTA +</li>\n<li>Whitening by creating ~6sec signals: basically I'd concat the signal with itself on the right and left, but first I'd find the location on the right that has the most similar value and slope and repeat from there. Same as the left. I thought by using longer signal, we take care of the ASD generation issue. It 'worked' but the model did not benefit | -+</li>\n<li>Whitening using auto regression residuals +</li>\n<li>Swin Transformer +, but xcite was better</li>\n<li>Galerkin transformer in 1D time domain, bucketing signal into 8hz wide chunks - </li>\n</ul>\n<p>where + is I also did that and - i dont. Most of the things you did i would consider is on the money.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1528865,
      "author_name": "authman",
      "author_url": "",
      "post_date": "09/30/2021 02:26:59",
      "content": "<p>A couple more things I tried, forgot to mention:</p>\n<ul>\n<li>My 1D nets were shallow as stated. Instead of .flatten(), I used GAP on the higher dimen channels</li>\n<li>I did seq2seq loss. I noticed this allowed models to converge a lot quicker. Both 1D and 2D models produce reduced-sized temporal representations. For ve+ all elements = 1 and =0 for ve-. I noticed for 1D models I noticed the resulting labels identified where the GW was. For 2D models, there was no such time encoding and it looked like whatever was being produced was all close to 1 or all close to 0.</li>\n<li>Training 1D model on S2S 'faux segmentation' from above</li>\n<li>Stacking - the first ~15 models used the same folds. But when I attempted submission pipeline, stacking sucked and I made that post. For all future models, I generated random folds</li>\n<li>Difficulty + Target Iterative Stratification: Using best model, compute residual and bucket into difficulty levels. Then multi-label stratify on that and target</li>\n<li>Final experiment blend (my portion): I just use cuml sgd to select weights on all of my experiments. CV for individual models ranged from 853-876, and the best blend of these was CV=0.8798601, public=0.8818, private=0.8796.</li>\n<li>I'll continue to add here as I recall things.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529013,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "09/30/2021 04:59:25",
      "content": "<p>This is some fabulous work <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> and I don't think what you did went on vain as we also tried most of the things that you have mentioned above , it was a tough competition and to remain in the top half solo takes a lot of heart , so I would say well done and many congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529282,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/30/2021 09:08:05",
      "content": "<p>sometimes it is luck and if you would encounter a \"good bug\".</p>\n<p>When I first tried learnable CQT, it didn't work. And I decided to give up on it and continue tunning on my fixed CQT. But then, I forgot to update my code (set trainable flag to False). I have 3 local machines and the codes were scattering everywhere. The codes failed to sync between my machine.</p>\n<p>Then when I tried to make a summary (in the discussion forum and my private note), I needed to visualize the cqt images. I realized that the cqt images were noisy like never before. After debugging, I realized that I was using a learnable version and it had magically worked.</p>\n<p>Lessons learned: Stop and review … maybe you would discover something?</p>\n<p><img src=\"https://i.ibb.co/s1YQ734/Selection-949.png\" alt=\"https://i.ibb.co/s1YQ734/Selection-949.png\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529487,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/30/2021 12:51:07",
      "content": "<p>You tried a lot of things.  Even if final result may not be as good as your hopes, I am sure that the knowledge you gained from trying all this will help you in future competitions.  It is by experimenting that we learn IMHO.</p>\n<p>Well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1529537,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "09/30/2021 13:21:28",
      "content": "<p>Hey, slightly disheartened to see your team fall down on the private leaderboard. That's a lot of experiments, I don't think even we were able to experiment that much as a team. Hope you'll get your deserving gold medal in the next competition.<br>\nbtw how did \"PCA, SVD, and FastICA transformations as channel augmentations\" performed? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1529631,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/30/2021 14:39:19",
          "content": "<p>Nothing stellar, I believe that experiment performed 'average' compared to other models I had trained. What I recall the most about it was:</p>\n<ol>\n<li>FastICA complained a lot about convergence issues, so I had to really up the number of iterations and silence warnings when fitting it.</li>\n<li>tSVD of course requires n_components &lt; input_dimensions, so for that aug, one channel was lost. Rather than deleting, just choose a random (original) one to place back in.</li>\n<li>PCA was fast and straightforward to implement.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559886,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:06:00",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1528826": "In total, I trained 96 full-fold models for this competition. Having nice local compute helps, but having deep insight into the problem + good exploratory prowess is more important (both of which I lack). I guess predestination / luck also plays its role :-). Some of the things I had tried off the top of my head (since I've already deleted my working directory):\n\n- Channel shuffle\n- Channel shifting\n- Random noise\n- 1D Net, all my 1D nets were really shallow, just 4 layers w/o residual\n- 2D Net, couldn't get them to converge as nicely as the rest of folks\n- Downloading actual LIGO 3 detector data for pretraining - it's very messy and has bands not represented in the synthetic data\n- GW injection using pycbc in 1d and 2d\n- Injected gw param regression as pretrain task\n- Cutout\n- Mixup\n- Trainable CQT\n- Very wide kernel sizes\n- Kernel smoothing and resume training\n- Frequency domain CQT\n- FD 1D net, 6 channels\n- Training on bandpassed first derivative data\n- Training on bandpassed integrated data\n- PCA, SVD, and FastICA transformations as channel augmentations\n- Signal flipping (only works in 1D `x=-x`. You can train a model like this then do TTA)\n- Signal compression and stretching - I only attempted this in 1D as TTA and for training\n- Random Mean and STD shift TTA\n- Whitening by creating ~6sec signals: basically I'd concat the signal with itself on the right and left, but first I'd find the location on the right that has the most similar value and slope and repeat from there. Same as the left. I thought by using longer signal, we take care of the ASD generation issue. It 'worked' but the model did not benefit\n- Whitening using auto regression residuals\n- Swin Transformer\n- Galerkin transformer in 1D time domain, bucketing signal into 8hz wide chunks\n- Probably more that I am forgetting\n\nI feel like instead of trying 903209230923 things, better to abandon quickly what doesn't work, focus on what does, and importantly, drive it to its logical conclusion. I was very early on the 1D train and was training nets with the full competition data in ram. But after I got to a decent level in 1D (at the time) I started exploring all these different avenues instead without focusing more in improving the net's architecture.\n\nI'd also like to express gratitude and appreciation to my team mates. I actually ran out of steam about two weeks ago, but contributions and motivation from @misaki1111 and @jpison helped carry me over. Thank you.",
    "1528835": "I mean if you want to compare with me:\n\n- Channel shuffle +\n- Channel shifting + \n- Random noise +\n- 1D Net, all my 1D nets were really shallow, just 4 layers w/o residual +\n- 2D Net, couldn't get them to converge as nicely as the rest of folks +\n- Downloading actual LIGO 3 detector data for pretraining - it's very messy and has bands not represented in the synthetic data | omg\n- GW injection using pycbc in 1d and 2d -\n- Injected gw param regression as pretrain task - \n- Cutout -\n- Mixup +\n- Trainable CQT +\n- Very wide kernel sizes +\n- Kernel smoothing and resume training -+\n- Frequency domain CQT +\n- FD 1D net, 6 channels -\n- Training on bandpassed first derivative data -\n- Training on bandpassed integrated data -\n- PCA, SVD, and FastICA transformations as channel augmentations -\n- Signal flipping (only works in 1D x=-x. You can train a model like this then do TTA) -\n- Signal compression and stretching - I only attempted this in 1D as TTA and for training +\n- Random Mean and STD shift TTA +\n- Whitening by creating ~6sec signals: basically I'd concat the signal with itself on the right and left, but first I'd find the location on the right that has the most similar value and slope and repeat from there. Same as the left. I thought by using longer signal, we take care of the ASD generation issue. It 'worked' but the model did not benefit | -+\n- Whitening using auto regression residuals +\n- Swin Transformer +, but xcite was better\n- Galerkin transformer in 1D time domain, bucketing signal into 8hz wide chunks - \n\nwhere + is I also did that and - i dont. Most of the things you did i would consider is on the money.",
    "1528865": "A couple more things I tried, forgot to mention:\n\n- My 1D nets were shallow as stated. Instead of .flatten(), I used GAP on the higher dimen channels\n- I did seq2seq loss. I noticed this allowed models to converge a lot quicker. Both 1D and 2D models produce reduced-sized temporal representations. For ve+ all elements = 1 and =0 for ve-. I noticed for 1D models I noticed the resulting labels identified where the GW was. For 2D models, there was no such time encoding and it looked like whatever was being produced was all close to 1 or all close to 0.\n- Training 1D model on S2S 'faux segmentation' from above\n- Stacking - the first ~15 models used the same folds. But when I attempted submission pipeline, stacking sucked and I made that post. For all future models, I generated random folds\n- Difficulty + Target Iterative Stratification: Using best model, compute residual and bucket into difficulty levels. Then multi-label stratify on that and target\n- Final experiment blend (my portion): I just use cuml sgd to select weights on all of my experiments. CV for individual models ranged from 853-876, and the best blend of these was CV=0.8798601, public=0.8818, private=0.8796.\n- I'll continue to add here as I recall things.",
    "1529013": "This is some fabulous work @authman and I don't think what you did went on vain as we also tried most of the things that you have mentioned above , it was a tough competition and to remain in the top half solo takes a lot of heart , so I would say well done and many congratulations",
    "1529282": "sometimes it is luck and if you would encounter a \"good bug\".\n\nWhen I first tried learnable CQT, it didn't work. And I decided to give up on it and continue tunning on my fixed CQT. But then, I forgot to update my code (set trainable flag to False). I have 3 local machines and the codes were scattering everywhere. The codes failed to sync between my machine.\n\nThen when I tried to make a summary (in the discussion forum and my private note), I needed to visualize the cqt images. I realized that the cqt images were noisy like never before. After debugging, I realized that I was using a learnable version and it had magically worked.\n\n\nLessons learned: Stop and review ... maybe you would discover something?\n\n![https://i.ibb.co/s1YQ734/Selection-949.png](https://i.ibb.co/s1YQ734/Selection-949.png)",
    "1529487": "You tried a lot of things.  Even if final result may not be as good as your hopes, I am sure that the knowledge you gained from trying all this will help you in future competitions.  It is by experimenting that we learn IMHO.\n\nWell done!",
    "1529537": "Hey, slightly disheartened to see your team fall down on the private leaderboard. That's a lot of experiments, I don't think even we were able to experiment that much as a team. Hope you'll get your deserving gold medal in the next competition.\nbtw how did \"PCA, SVD, and FastICA transformations as channel augmentations\" performed?",
    "1529631": "Nothing stellar, I believe that experiment performed 'average' compared to other models I had trained. What I recall the most about it was:\n\n1. FastICA complained a lot about convergence issues, so I had to really up the number of iterations and silence warnings when fitting it.\n2. tSVD of course requires n_components < input_dimensions, so for that aug, one channel was lost. Rather than deleting, just choose a random (original) one to place back in.\n3. PCA was fast and straightforward to implement.",
    "1559886": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}