{
  "id": 269154,
  "title": "did you encounter a black hole in your training?",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/269154",
  "author_name": "",
  "post_date": "2021-08-30T14:28:56.595555400Z",
  "votes": 30,
  "comment_count": 29,
  "views": 0,
  "content": "<p>did you encounter a black hole in your training?</p>\n<p><img src=\"https://i.ibb.co/W0c4fDL/Selection-774.png\" alt=\"https://i.ibb.co/W0c4fDL/Selection-774.png\"></p>",
  "messages": [
    {
      "id": "1496641",
      "postDate": "08/30/2021 14:28:56",
      "content": "<p>did you encounter a black hole in your training?</p>\n<p><img src=\"https://i.ibb.co/W0c4fDL/Selection-774.png\" alt=\"https://i.ibb.co/W0c4fDL/Selection-774.png\"></p>",
      "rawMarkdown": "did you encounter a black hole in your training?\n\n![https://i.ibb.co/W0c4fDL/Selection-774.png](https://i.ibb.co/W0c4fDL/Selection-774.png)",
      "votes": null
    },
    {
      "id": "1496654",
      "postDate": "08/30/2021 14:37:57",
      "content": "<p>Yes, it happened multiple times during my trainings. <br>\nIn some cases, by reducing the learning rate I managed to get rid of it. Not always though.<br>\nFor some runs, such \"blackholes\" happened after the first epoch.<br>\nMoreover it looks lihe the weight update was so brutal that after the \"blackhole\" the learning never recovers and get stuck at AUC=0.5.</p>",
      "rawMarkdown": "Yes, it happened multiple times during my trainings. \nIn some cases, by reducing the learning rate I managed to get rid of it. Not always though.\nFor some runs, such \"blackholes\" happened after the first epoch.\nMoreover it looks lihe the weight update was so brutal that after the \"blackhole\" the learning never recovers and get stuck at AUC=0.5.",
      "votes": null
    },
    {
      "id": "1496663",
      "postDate": "08/30/2021 14:45:34",
      "content": "<p>It also happens to TUMOR competition. lol </p>",
      "rawMarkdown": "It also happens to TUMOR competition. lol",
      "votes": null
    },
    {
      "id": "1496668",
      "postDate": "08/30/2021 14:47:01",
      "content": "<p>\"weight update was so brutal that after the \"blackhole\" </p>\n<p>the probability score for the good +ve samples is almost 1.0. high score means zero loss. They have \"virtually disappeared\" from the training set</p>\n<p><img src=\"https://i.ibb.co/qWMCWFv/Selection-775.png\" alt=\"https://i.ibb.co/qWMCWFv/Selection-775.png\"></p>",
      "rawMarkdown": "\"weight update was so brutal that after the \"blackhole\" \n\nthe probability score for the good +ve samples is almost 1.0. high score means zero loss. They have \"virtually disappeared\" from the training set\n\n\n![https://i.ibb.co/qWMCWFv/Selection-775.png](https://i.ibb.co/qWMCWFv/Selection-775.png)",
      "votes": null
    },
    {
      "id": "1496731",
      "postDate": "08/30/2021 15:33:40",
      "content": "<p>It for this reason I thought focal loss would be a sword in this competition, but it ended up being a butter knife.</p>",
      "rawMarkdown": "It for this reason I thought focal loss would be a sword in this competition, but it ended up being a butter knife.",
      "votes": null
    },
    {
      "id": "1496816",
      "postDate": "08/30/2021 16:48:32",
      "content": "<p>It happens frequently when the batch size is large. Trying to survive from black hole :)</p>",
      "rawMarkdown": "It happens frequently when the batch size is large. Trying to survive from black hole :)",
      "votes": null
    },
    {
      "id": "1497096",
      "postDate": "08/30/2021 23:46:23",
      "content": "<p>For me, this has been very architecture-specific. Some architectures always converge, and some never do.</p>",
      "rawMarkdown": "For me, this has been very architecture-specific. Some architectures always converge, and some never do.",
      "votes": null
    },
    {
      "id": "1497113",
      "postDate": "08/31/2021 00:52:36",
      "content": "<p>Exactly. NFNet variants always cross the event horizon.</p>",
      "rawMarkdown": "Exactly. NFNet variants always cross the event horizon.",
      "votes": null
    },
    {
      "id": "1497118",
      "postDate": "08/31/2021 01:00:24",
      "content": "<p>Glad to hear someone else is experiencing this. I was going crazy trying to get the ECA variant of NFNet's working since it seemed to be successful in other spectrogram-type competitions.</p>",
      "rawMarkdown": "Glad to hear someone else is experiencing this. I was going crazy trying to get the ECA variant of NFNet's working since it seemed to be successful in other spectrogram-type competitions.",
      "votes": null
    },
    {
      "id": "1497130",
      "postDate": "08/31/2021 01:28:13",
      "content": "<p>I have been facing this as well. Specifically when training WaveNets</p>",
      "rawMarkdown": "I have been facing this as well. Specifically when training WaveNets",
      "votes": null
    },
    {
      "id": "1497262",
      "postDate": "08/31/2021 04:54:43",
      "content": "<p>Interesting plot! can you elaborate on what +ve label means? Thanks!</p>",
      "rawMarkdown": "Interesting plot! can you elaborate on what +ve label means? Thanks!",
      "votes": null
    },
    {
      "id": "1497349",
      "postDate": "08/31/2021 06:42:22",
      "content": "<p>could you please explain a little bit more?  I found in tumor competition, some models works, some don't.  not understanding the fundamental.</p>",
      "rawMarkdown": "could you please explain a little bit more?  I found in tumor competition, some models works, some don't.  not understanding the fundamental.",
      "votes": null
    },
    {
      "id": "1497904",
      "postDate": "08/31/2021 13:51:05",
      "content": "<p>Wow this is beyond my imagination</p>",
      "rawMarkdown": "Wow this is beyond my imagination",
      "votes": null
    },
    {
      "id": "1498225",
      "postDate": "08/31/2021 19:06:25",
      "content": "<p>good observation….</p>",
      "rawMarkdown": "good observation....",
      "votes": null
    },
    {
      "id": "1498256",
      "postDate": "08/31/2021 19:41:46",
      "content": "<p>positive and negative labels</p>",
      "rawMarkdown": "positive and negative labels",
      "votes": null
    },
    {
      "id": "1498623",
      "postDate": "09/01/2021 05:19:20",
      "content": "<p>Got it. Thanks!</p>",
      "rawMarkdown": "Got it. Thanks!",
      "votes": null
    },
    {
      "id": "1501798",
      "postDate": "09/03/2021 14:48:58",
      "content": "<p>Reduce mini batch from 32 to 16 and reduce learning rate by 1/10 works for me. Thank you for the suggestions</p>",
      "rawMarkdown": "Reduce mini batch from 32 to 16 and reduce learning rate by 1/10 works for me. Thank you for the suggestions",
      "votes": null
    },
    {
      "id": "1501809",
      "postDate": "09/03/2021 14:58:41",
      "content": "<p>Could that behavior of the networks come from an exploding gradient ? I plan to test some kind of gradient clipping to see if this would solve the issue: I wonder if somebody already checked that possibility ?</p>",
      "rawMarkdown": "Could that behavior of the networks come from an exploding gradient ? I plan to test some kind of gradient clipping to see if this would solve the issue: I wonder if somebody already checked that possibility ?",
      "votes": null
    },
    {
      "id": "1502554",
      "postDate": "09/04/2021 11:46:23",
      "content": "<p>Nice finding, considering the plots I saw in forum, there are around 120'000 (for 120'000 TPs typical model gives &gt; 0.99 probability) good +ve samples in the train set, it means that there is a possibility that more than half of samples from batch give almost zero loss, hence it means that something I'd like to call \"hidden learning rate\" is being introduced (because suppose batch size is N, but in average we have N / 2 effective samples which bring new information to the network, yet still loss is being divided by N, hence learning rate approximately reduced by half)<br>\nWe should consider it when setting schedulers.</p>\n<p>Also, should we try active learning approaches? <br>\n(exclude samples from the train set if they're already learned, i.e. have zero loss)</p>\n<p>image by <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a><br>\n<img src=\"https://ibb.co/mSWcdMq\" alt=\"predicted probabilities\"></p>",
      "rawMarkdown": "Nice finding, considering the plots I saw in forum, there are around 120'000 (for 120'000 TPs typical model gives > 0.99 probability) good +ve samples in the train set, it means that there is a possibility that more than half of samples from batch give almost zero loss, hence it means that something I'd like to call \"hidden learning rate\" is being introduced (because suppose batch size is N, but in average we have N / 2 effective samples which bring new information to the network, yet still loss is being divided by N, hence learning rate approximately reduced by half)\nWe should consider it when setting schedulers.\n\nAlso, should we try active learning approaches? \n(exclude samples from the train set if they're already learned, i.e. have zero loss)\n\nimage by @authman\n![predicted probabilities](https://ibb.co/mSWcdMq)",
      "votes": null
    },
    {
      "id": "1502616",
      "postDate": "09/04/2021 13:23:02",
      "content": "<p>Uh, why was I downvoted? Did I say something wrong? I had never even imagined that something like this could also occur</p>",
      "rawMarkdown": "Uh, why was I downvoted? Did I say something wrong? I had never even imagined that something like this could also occur",
      "votes": null
    },
    {
      "id": "1502681",
      "postDate": "09/04/2021 14:34:17",
      "content": "<p><a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> you mean nfnets are working well or poor? </p>",
      "rawMarkdown": "analokamus you mean nfnets are working well or poor?",
      "votes": null
    },
    {
      "id": "1502699",
      "postDate": "09/04/2021 14:48:25",
      "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> It does not work at all.</p>",
      "rawMarkdown": "pheadrus It does not work at all.",
      "votes": null
    },
    {
      "id": "1502880",
      "postDate": "09/04/2021 18:09:43",
      "content": "<p>Insane Observation!!</p>",
      "rawMarkdown": "Insane Observation!!",
      "votes": null
    },
    {
      "id": "1502938",
      "postDate": "09/04/2021 19:08:03",
      "content": "<p>Don't pay attention to haters, just keep modeling.</p>",
      "rawMarkdown": "Don't pay attention to haters, just keep modeling.",
      "votes": null
    },
    {
      "id": "1506183",
      "postDate": "09/08/2021 01:43:41",
      "content": "<p>update1:</p>\n<ul>\n<li>experiment is done without augmentation</li>\n<li>why do this experiment?</li>\n</ul>\n<ol>\n<li>make a baseline and compare it against training with augmentation</li>\n<li>set optimum num of epoch and lr schduler</li>\n</ol>\n<p><img src=\"https://i.ibb.co/pXcbTc8/Selection-806.png\" alt=\"https://i.ibb.co/pXcbTc8/Selection-806.png\"></p>",
      "rawMarkdown": "update1:\n- experiment is done without augmentation\n- why do this experiment?\n1.  make a baseline and compare it against training with augmentation\n2. set optimum num of epoch and lr schduler\n\n![https://i.ibb.co/pXcbTc8/Selection-806.png](https://i.ibb.co/pXcbTc8/Selection-806.png)",
      "votes": null
    },
    {
      "id": "1507108",
      "postDate": "09/08/2021 20:15:47",
      "content": "<p><code>this model</code>? What model?</p>",
      "rawMarkdown": "`this model`? What model?",
      "votes": null
    },
    {
      "id": "1507839",
      "postDate": "09/09/2021 15:32:38",
      "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> exactly the same I'm facing with wevanet/1d variants - I'd be interested to know if you did find a breakthrough </p>",
      "rawMarkdown": "tuckerarrants exactly the same I'm facing with wevanet/1d variants - I'd be interested to know if you did find a breakthrough",
      "votes": null
    },
    {
      "id": "1508069",
      "postDate": "09/09/2021 19:35:37",
      "content": "<p><a href=\"https://www.kaggle.com/kfk42kfk\" target=\"_blank\">@kfk42kfk</a> You don't think Heng shares enough?  He is kind to share his ideas, but maybe he can kep some details hidden?</p>",
      "rawMarkdown": "kfk42kfk You don't think Heng shares enough?  He is kind to share his ideas, but maybe he can kep some details hidden?",
      "votes": null
    },
    {
      "id": "1559698",
      "postDate": "10/27/2021 07:07:11",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    },
    {
      "id": "1559983",
      "postDate": "10/27/2021 08:52:43",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1496654,
      "author_name": "fabiendaniel",
      "author_url": "",
      "post_date": "08/30/2021 14:37:57",
      "content": "<p>Yes, it happened multiple times during my trainings. <br>\nIn some cases, by reducing the learning rate I managed to get rid of it. Not always though.<br>\nFor some runs, such \"blackholes\" happened after the first epoch.<br>\nMoreover it looks lihe the weight update was so brutal that after the \"blackhole\" the learning never recovers and get stuck at AUC=0.5.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1496668,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/30/2021 14:47:01",
          "content": "<p>\"weight update was so brutal that after the \"blackhole\" </p>\n<p>the probability score for the good +ve samples is almost 1.0. high score means zero loss. They have \"virtually disappeared\" from the training set</p>\n<p><img src=\"https://i.ibb.co/qWMCWFv/Selection-775.png\" alt=\"https://i.ibb.co/qWMCWFv/Selection-775.png\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1496731,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/30/2021 15:33:40",
          "content": "<p>It for this reason I thought focal loss would be a sword in this competition, but it ended up being a butter knife.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1496816,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "08/30/2021 16:48:32",
          "content": "<p>It happens frequently when the batch size is large. Trying to survive from black hole :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1497262,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "08/31/2021 04:54:43",
          "content": "<p>Interesting plot! can you elaborate on what +ve label means? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1498256,
          "author_name": "mihtw1",
          "author_url": "",
          "post_date": "08/31/2021 19:41:46",
          "content": "<p>positive and negative labels</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1498623,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "09/01/2021 05:19:20",
          "content": "<p>Got it. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1502554,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "09/04/2021 11:46:23",
          "content": "<p>Nice finding, considering the plots I saw in forum, there are around 120'000 (for 120'000 TPs typical model gives &gt; 0.99 probability) good +ve samples in the train set, it means that there is a possibility that more than half of samples from batch give almost zero loss, hence it means that something I'd like to call \"hidden learning rate\" is being introduced (because suppose batch size is N, but in average we have N / 2 effective samples which bring new information to the network, yet still loss is being divided by N, hence learning rate approximately reduced by half)<br>\nWe should consider it when setting schedulers.</p>\n<p>Also, should we try active learning approaches? <br>\n(exclude samples from the train set if they're already learned, i.e. have zero loss)</p>\n<p>image by <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a><br>\n<img src=\"https://ibb.co/mSWcdMq\" alt=\"predicted probabilities\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1496663,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "08/30/2021 14:45:34",
      "content": "<p>It also happens to TUMOR competition. lol </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1497096,
      "author_name": "marktenenholtz",
      "author_url": "",
      "post_date": "08/30/2021 23:46:23",
      "content": "<p>For me, this has been very architecture-specific. Some architectures always converge, and some never do.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1497113,
          "author_name": "analokamus",
          "author_url": "",
          "post_date": "08/31/2021 00:52:36",
          "content": "<p>Exactly. NFNet variants always cross the event horizon.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1497118,
          "author_name": "marktenenholtz",
          "author_url": "",
          "post_date": "08/31/2021 01:00:24",
          "content": "<p>Glad to hear someone else is experiencing this. I was going crazy trying to get the ECA variant of NFNet's working since it seemed to be successful in other spectrogram-type competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1497130,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "08/31/2021 01:28:13",
          "content": "<p>I have been facing this as well. Specifically when training WaveNets</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1497349,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "08/31/2021 06:42:22",
          "content": "<p>could you please explain a little bit more?  I found in tumor competition, some models works, some don't.  not understanding the fundamental.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1502681,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "09/04/2021 14:34:17",
          "content": "<p><a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> you mean nfnets are working well or poor? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1502699,
          "author_name": "analokamus",
          "author_url": "",
          "post_date": "09/04/2021 14:48:25",
          "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> It does not work at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507839,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "09/09/2021 15:32:38",
          "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> exactly the same I'm facing with wevanet/1d variants - I'd be interested to know if you did find a breakthrough </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1497904,
      "author_name": "keagle",
      "author_url": "",
      "post_date": "08/31/2021 13:51:05",
      "content": "<p>Wow this is beyond my imagination</p>",
      "votes": null,
      "replies": [
        {
          "id": 1502616,
          "author_name": "keagle",
          "author_url": "",
          "post_date": "09/04/2021 13:23:02",
          "content": "<p>Uh, why was I downvoted? Did I say something wrong? I had never even imagined that something like this could also occur</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1502938,
          "author_name": "authman",
          "author_url": "",
          "post_date": "09/04/2021 19:08:03",
          "content": "<p>Don't pay attention to haters, just keep modeling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1498225,
      "author_name": "nishanthaddagatla",
      "author_url": "",
      "post_date": "08/31/2021 19:06:25",
      "content": "<p>good observation….</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1501798,
      "author_name": "ivan3357",
      "author_url": "",
      "post_date": "09/03/2021 14:48:58",
      "content": "<p>Reduce mini batch from 32 to 16 and reduce learning rate by 1/10 works for me. Thank you for the suggestions</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1501809,
      "author_name": "fabiendaniel",
      "author_url": "",
      "post_date": "09/03/2021 14:58:41",
      "content": "<p>Could that behavior of the networks come from an exploding gradient ? I plan to test some kind of gradient clipping to see if this would solve the issue: I wonder if somebody already checked that possibility ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1502880,
      "author_name": "ayan007",
      "author_url": "",
      "post_date": "09/04/2021 18:09:43",
      "content": "<p>Insane Observation!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1506183,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/08/2021 01:43:41",
      "content": "<p>update1:</p>\n<ul>\n<li>experiment is done without augmentation</li>\n<li>why do this experiment?</li>\n</ul>\n<ol>\n<li>make a baseline and compare it against training with augmentation</li>\n<li>set optimum num of epoch and lr schduler</li>\n</ol>\n<p><img src=\"https://i.ibb.co/pXcbTc8/Selection-806.png\" alt=\"https://i.ibb.co/pXcbTc8/Selection-806.png\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1507108,
          "author_name": "kfk42kfk",
          "author_url": "",
          "post_date": "09/08/2021 20:15:47",
          "content": "<p><code>this model</code>? What model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508069,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/09/2021 19:35:37",
          "content": "<p><a href=\"https://www.kaggle.com/kfk42kfk\" target=\"_blank\">@kfk42kfk</a> You don't think Heng shares enough?  He is kind to share his ideas, but maybe he can kep some details hidden?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559698,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:07:11",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559983,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:52:43",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1496641": "did you encounter a black hole in your training?\n\n![https://i.ibb.co/W0c4fDL/Selection-774.png](https://i.ibb.co/W0c4fDL/Selection-774.png)",
    "1496654": "Yes, it happened multiple times during my trainings. \nIn some cases, by reducing the learning rate I managed to get rid of it. Not always though.\nFor some runs, such \"blackholes\" happened after the first epoch.\nMoreover it looks lihe the weight update was so brutal that after the \"blackhole\" the learning never recovers and get stuck at AUC=0.5.",
    "1496663": "It also happens to TUMOR competition. lol",
    "1496668": "\"weight update was so brutal that after the \"blackhole\" \n\nthe probability score for the good +ve samples is almost 1.0. high score means zero loss. They have \"virtually disappeared\" from the training set\n\n\n![https://i.ibb.co/qWMCWFv/Selection-775.png](https://i.ibb.co/qWMCWFv/Selection-775.png)",
    "1496731": "It for this reason I thought focal loss would be a sword in this competition, but it ended up being a butter knife.",
    "1496816": "It happens frequently when the batch size is large. Trying to survive from black hole :)",
    "1497096": "For me, this has been very architecture-specific. Some architectures always converge, and some never do.",
    "1497113": "Exactly. NFNet variants always cross the event horizon.",
    "1497118": "Glad to hear someone else is experiencing this. I was going crazy trying to get the ECA variant of NFNet's working since it seemed to be successful in other spectrogram-type competitions.",
    "1497130": "I have been facing this as well. Specifically when training WaveNets",
    "1497262": "Interesting plot! can you elaborate on what +ve label means? Thanks!",
    "1497349": "could you please explain a little bit more?  I found in tumor competition, some models works, some don't.  not understanding the fundamental.",
    "1497904": "Wow this is beyond my imagination",
    "1498225": "good observation....",
    "1498256": "positive and negative labels",
    "1498623": "Got it. Thanks!",
    "1501798": "Reduce mini batch from 32 to 16 and reduce learning rate by 1/10 works for me. Thank you for the suggestions",
    "1501809": "Could that behavior of the networks come from an exploding gradient ? I plan to test some kind of gradient clipping to see if this would solve the issue: I wonder if somebody already checked that possibility ?",
    "1502554": "Nice finding, considering the plots I saw in forum, there are around 120'000 (for 120'000 TPs typical model gives > 0.99 probability) good +ve samples in the train set, it means that there is a possibility that more than half of samples from batch give almost zero loss, hence it means that something I'd like to call \"hidden learning rate\" is being introduced (because suppose batch size is N, but in average we have N / 2 effective samples which bring new information to the network, yet still loss is being divided by N, hence learning rate approximately reduced by half)\nWe should consider it when setting schedulers.\n\nAlso, should we try active learning approaches? \n(exclude samples from the train set if they're already learned, i.e. have zero loss)\n\nimage by @authman\n![predicted probabilities](https://ibb.co/mSWcdMq)",
    "1502616": "Uh, why was I downvoted? Did I say something wrong? I had never even imagined that something like this could also occur",
    "1502681": "analokamus you mean nfnets are working well or poor?",
    "1502699": "pheadrus It does not work at all.",
    "1502880": "Insane Observation!!",
    "1502938": "Don't pay attention to haters, just keep modeling.",
    "1506183": "update1:\n- experiment is done without augmentation\n- why do this experiment?\n1.  make a baseline and compare it against training with augmentation\n2. set optimum num of epoch and lr schduler\n\n![https://i.ibb.co/pXcbTc8/Selection-806.png](https://i.ibb.co/pXcbTc8/Selection-806.png)",
    "1507108": "`this model`? What model?",
    "1507839": "tuckerarrants exactly the same I'm facing with wevanet/1d variants - I'd be interested to know if you did find a breakthrough",
    "1508069": "kfk42kfk You don't think Heng shares enough?  He is kind to share his ideas, but maybe he can kep some details hidden?",
    "1559698": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1559983": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}