{
  "id": 254086,
  "title": "Training in FP16 slightly worse ROC AUC? ",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/254086",
  "author_name": "",
  "post_date": "2021-07-20T04:07:44.294853300Z",
  "votes": 3,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi guys, I was experimenting with training on FP16 on a V100 GPU comparing the same model (4 layers of convolutions + max pool, followed by flattening and 3 layers of fully-connected layers including the final output layer) with full precision on <a href=\"https://www.kaggle.com/wabinab/g2net-as-image\" target=\"_blank\">this</a> dataset with batch size of 96, and somehow FP16 have (either slightly or far more) worse performance than full precision. It seems like, with a learning rate finder, FP16 shows much more flat \"loss against learning rate\" curve than bumpy FP32. Training on FP16, two of the three times being tested shows 0.6-0.7 RocAuc training, and tweaking the learning rate (since it has a flat region over a large range of learning rate) can gives me about same accuracy as TF32, however, training for 10 epochs the RocAuc seems to fluctuate around that value of 0.82-0.84 without clear direction of improvement. </p>\n<p>And yes, since the above has little experimentations one cannot be sure whether it is an illusion or solid. Perhaps there exists difference in improvements between different platforms (i.e. VM vs on Kaggle). </p>",
  "messages": [
    {
      "id": "1393965",
      "postDate": "07/20/2021 04:07:44",
      "content": "<p>Hi guys, I was experimenting with training on FP16 on a V100 GPU comparing the same model (4 layers of convolutions + max pool, followed by flattening and 3 layers of fully-connected layers including the final output layer) with full precision on <a href=\"https://www.kaggle.com/wabinab/g2net-as-image\" target=\"_blank\">this</a> dataset with batch size of 96, and somehow FP16 have (either slightly or far more) worse performance than full precision. It seems like, with a learning rate finder, FP16 shows much more flat \"loss against learning rate\" curve than bumpy FP32. Training on FP16, two of the three times being tested shows 0.6-0.7 RocAuc training, and tweaking the learning rate (since it has a flat region over a large range of learning rate) can gives me about same accuracy as TF32, however, training for 10 epochs the RocAuc seems to fluctuate around that value of 0.82-0.84 without clear direction of improvement. </p>\n<p>And yes, since the above has little experimentations one cannot be sure whether it is an illusion or solid. Perhaps there exists difference in improvements between different platforms (i.e. VM vs on Kaggle). </p>",
      "rawMarkdown": "Hi guys, I was experimenting with training on FP16 on a V100 GPU comparing the same model (4 layers of convolutions + max pool, followed by flattening and 3 layers of fully-connected layers including the final output layer) with full precision on [this](https://www.kaggle.com/wabinab/g2net-as-image) dataset with batch size of 96, and somehow FP16 have (either slightly or far more) worse performance than full precision. It seems like, with a learning rate finder, FP16 shows much more flat \"loss against learning rate\" curve than bumpy FP32. Training on FP16, two of the three times being tested shows 0.6-0.7 RocAuc training, and tweaking the learning rate (since it has a flat region over a large range of learning rate) can gives me about same accuracy as TF32, however, training for 10 epochs the RocAuc seems to fluctuate around that value of 0.82-0.84 without clear direction of improvement. \n\nAnd yes, since the above has little experimentations one cannot be sure whether it is an illusion or solid. Perhaps there exists difference in improvements between different platforms (i.e. VM vs on Kaggle).",
      "votes": null
    },
    {
      "id": "1394493",
      "postDate": "07/20/2021 12:10:13",
      "content": "<p>I think it is due to the effect of floating point. The value of the data in this competition is quite small (order ~ e-20), so when we change the data's dtype to FP16, they will not retain the original value. <br>\nThat can be confirmed by following simple code.</p>\n<pre><code>file_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float64)\nprint(waves)\n</code></pre>\n<pre><code>[[-5.94830548e-21 -5.84995448e-21 -5.42415169e-21 ... -6.06698987e-21\n  -5.96345722e-21 -5.75778438e-21]\n [ 9.75407048e-22  4.52586118e-22  4.58643893e-23 ... -1.09608208e-20\n  -1.09766636e-20 -1.10858129e-20]\n [-1.74871983e-21 -1.18286791e-21 -1.93223777e-21 ...  1.46502268e-21\n   2.18644864e-21  1.54085934e-21]]\n</code></pre>\n<pre><code>file_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float32)\nprint(waves)\n</code></pre>\n<pre><code>[[-5.9483055e-21 -5.8499546e-21 -5.4241516e-21 ... -6.0669897e-21\n  -5.9634572e-21 -5.7577845e-21]\n [ 9.7540710e-22  4.5258614e-22  4.5864389e-23 ... -1.0960821e-20\n  -1.0976663e-20 -1.1085813e-20]\n [-1.7487198e-21 -1.1828679e-21 -1.9322378e-21 ...  1.4650227e-21\n   2.1864486e-21  1.5408594e-21]]\n</code></pre>\n<pre><code>file_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float16)\nprint(waves)\n</code></pre>\n<pre><code>[[-0. -0. -0. ... -0. -0. -0.]\n [ 0.  0.  0. ... -0. -0. -0.]\n [-0. -0. -0. ...  0.  0.  0.]]\n</code></pre>",
      "rawMarkdown": "I think it is due to the effect of floating point. The value of the data in this competition is quite small (order ~ e-20), so when we change the data's dtype to FP16, they will not retain the original value. \nThat can be confirmed by following simple code.\n\n```\nfile_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float64)\nprint(waves)\n```\n```\n[[-5.94830548e-21 -5.84995448e-21 -5.42415169e-21 ... -6.06698987e-21\n  -5.96345722e-21 -5.75778438e-21]\n [ 9.75407048e-22  4.52586118e-22  4.58643893e-23 ... -1.09608208e-20\n  -1.09766636e-20 -1.10858129e-20]\n [-1.74871983e-21 -1.18286791e-21 -1.93223777e-21 ...  1.46502268e-21\n   2.18644864e-21  1.54085934e-21]]\n```\n\n```\nfile_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float32)\nprint(waves)\n```\n```\n[[-5.9483055e-21 -5.8499546e-21 -5.4241516e-21 ... -6.0669897e-21\n  -5.9634572e-21 -5.7577845e-21]\n [ 9.7540710e-22  4.5258614e-22  4.5864389e-23 ... -1.0960821e-20\n  -1.0976663e-20 -1.1085813e-20]\n [-1.7487198e-21 -1.1828679e-21 -1.9322378e-21 ...  1.4650227e-21\n   2.1864486e-21  1.5408594e-21]]\n```\n\n```\nfile_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float16)\nprint(waves)\n```\n```\n[[-0. -0. -0. ... -0. -0. -0.]\n [ 0.  0.  0. ... -0. -0. -0.]\n [-0. -0. -0. ...  0.  0.  0.]]\n```",
      "votes": null
    },
    {
      "id": "1394707",
      "postDate": "07/20/2021 14:36:50",
      "content": "<p>As <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> pointed out, FP16 cannot store extremely small values in the order of e-20 such as the data from this competition (this is because FP16 has only 5 bits for exponent!). <br>\nBut I don't think your case is related to the wave data because at least you were able to reach score of 0,8 or above. <br>\nI recommend you partially disable AMP autocast by adding a <code>with torch.cuda.amp.autocast(enabled=False):</code> subregion to your model, and see what happens. I noticed nnAudio package sometimes return NaN with AMP enabled. <br>\nHope this helps you!</p>",
      "rawMarkdown": "As @naoism pointed out, FP16 cannot store extremely small values in the order of e-20 such as the data from this competition (this is because FP16 has only 5 bits for exponent!). \nBut I don't think your case is related to the wave data because at least you were able to reach score of 0,8 or above. \nI recommend you partially disable AMP autocast by adding a `with torch.cuda.amp.autocast(enabled=False):` subregion to your model, and see what happens. I noticed nnAudio package sometimes return NaN with AMP enabled. \nHope this helps you!",
      "votes": null
    },
    {
      "id": "1394746",
      "postDate": "07/20/2021 15:10:43",
      "content": "<p>Thank you for your comment. Indeed, it seems that wave data may be not relevant if the score of 0.8 was obtained.</p>",
      "rawMarkdown": "Thank you for your comment. Indeed, it seems that wave data may be not relevant if the score of 0.8 was obtained.",
      "votes": null
    },
    {
      "id": "1394751",
      "postDate": "07/20/2021 15:16:26",
      "content": "<p>it is better to use mixed precision, for instance torch.cuda.amp if you use pytorch.  With this, model weights are FP32 but gradient is computed using FP16.  Moreover amp rescales gradient so that max FP16 precision is kept.</p>\n<p>I train all my models with amp now.</p>",
      "rawMarkdown": "it is better to use mixed precision, for instance torch.cuda.amp if you use pytorch.  With this, model weights are FP32 but gradient is computed using FP16.  Moreover amp rescales gradient so that max FP16 precision is kept.\n\nI train all my models with amp now.",
      "votes": null
    },
    {
      "id": "1394752",
      "postDate": "07/20/2021 15:20:41",
      "content": "<p>I disagree, it is dangerous to disable autocast, as you're more likely to get NaN or wrong gradients.</p>",
      "rawMarkdown": "I disagree, it is dangerous to disable autocast, as you're more likely to get NaN or wrong gradients.",
      "votes": null
    },
    {
      "id": "1395138",
      "postDate": "07/21/2021 00:22:02",
      "content": "<p>I assumed <a href=\"https://www.kaggle.com/wabinab\" target=\"_blank\">@wabinab</a> uses precision instead of FP16 only, thus partially disabling autocast (=forcing FP32 operation) is a good option for debugging. I suppose this behaviour is due to autocast's failure to determine which cast to use for some specific operations used in nnAudio package.<br>\nI trained all my models with amp (autocast partially disabled).</p>",
      "rawMarkdown": "I assumed @wabinab uses precision instead of FP16 only, thus partially disabling autocast (=forcing FP32 operation) is a good option for debugging. I suppose this behaviour is due to autocast's failure to determine which cast to use for some specific operations used in nnAudio package.\nI trained all my models with amp (autocast partially disabled).",
      "votes": null
    },
    {
      "id": "1395331",
      "postDate": "07/21/2021 06:27:38",
      "content": "<p>Thanks for your comment <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> </p>\n<p>One uses fastai's \"fit_one_cycle\" to do the training so underlyingly one haven't check the source code of which part it uses autocast. One have been experimenting it with pretrained efficientnet one could still get 0.8+ AUC sometimes though one have difficulty to push it higher than 0.85 yet. It might be perhaps tuning learning rate (and maybe other hyperparameters but one haven't tried it out) is quite important. </p>\n<p>As for <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> yes in this case all the data are represented as png images so they're scaled to either 0-255 or 0-1 depending on convention. </p>",
      "rawMarkdown": "Thanks for your comment @analokamus \n\nOne uses fastai's \"fit_one_cycle\" to do the training so underlyingly one haven't check the source code of which part it uses autocast. One have been experimenting it with pretrained efficientnet one could still get 0.8+ AUC sometimes though one have difficulty to push it higher than 0.85 yet. It might be perhaps tuning learning rate (and maybe other hyperparameters but one haven't tried it out) is quite important. \n\nAs for @naoism yes in this case all the data are represented as png images so they're scaled to either 0-255 or 0-1 depending on convention.",
      "votes": null
    },
    {
      "id": "1395526",
      "postDate": "07/21/2021 09:35:35",
      "content": "<p><a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> Fair enough, I thought you were discussing Scaler, sorry.  </p>",
      "rawMarkdown": "analokamus Fair enough, I thought you were discussing Scaler, sorry.",
      "votes": null
    },
    {
      "id": "1395667",
      "postDate": "07/21/2021 12:43:13",
      "content": "<p>Good day<br>\nI read somewhere that the neural networks needed is one input layer two convolutions ReLu and the output… <br>\nI believe it was the MatLab model(?)…<br>\nThen add Propagation from two convolution networks and that’s what gives you your probability is using Bayesian theory<br>\nAlso, tensor flow can be used (?)</p>",
      "rawMarkdown": "Good day\nI read somewhere that the neural networks needed is one input layer two convolutions ReLu and the output… \nI believe it was the MatLab model(?)…\nThen add Propagation from two convolution networks and that’s what gives you your probability is using Bayesian theory\nAlso, tensor flow can be used (?)",
      "votes": null
    },
    {
      "id": "1480516",
      "postDate": "08/19/2021 04:02:53",
      "content": "<p>Good day, the FP16 I read could not be used in this as it causes issues of overfitting   So using FP64 is the one to use I read this on the MATLAB documents for this very first detection of a GWbiH<br>\n<a href=\"https://www.mathworks.com/company/newsletters/articles/confirming-the-first-ever-detection-of-gravitational-waves-by-analyzing-laser-interferometer-data.html\" target=\"_blank\">https://www.mathworks.com/company/newsletters/articles/confirming-the-first-ever-detection-of-gravitational-waves-by-analyzing-laser-interferometer-data.html</a></p>",
      "rawMarkdown": "Good day, the FP16 I read could not be used in this as it causes issues of overfitting   So using FP64 is the one to use I read this on the MATLAB documents for this very first detection of a GWbiH\nhttps://www.mathworks.com/company/newsletters/articles/confirming-the-first-ever-detection-of-gravitational-waves-by-analyzing-laser-interferometer-data.html",
      "votes": null
    },
    {
      "id": "1561219",
      "postDate": "10/27/2021 12:20:55",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1394493,
      "author_name": "naoism",
      "author_url": "",
      "post_date": "07/20/2021 12:10:13",
      "content": "<p>I think it is due to the effect of floating point. The value of the data in this competition is quite small (order ~ e-20), so when we change the data's dtype to FP16, they will not retain the original value. <br>\nThat can be confirmed by following simple code.</p>\n<pre><code>file_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float64)\nprint(waves)\n</code></pre>\n<pre><code>[[-5.94830548e-21 -5.84995448e-21 -5.42415169e-21 ... -6.06698987e-21\n  -5.96345722e-21 -5.75778438e-21]\n [ 9.75407048e-22  4.52586118e-22  4.58643893e-23 ... -1.09608208e-20\n  -1.09766636e-20 -1.10858129e-20]\n [-1.74871983e-21 -1.18286791e-21 -1.93223777e-21 ...  1.46502268e-21\n   2.18644864e-21  1.54085934e-21]]\n</code></pre>\n<pre><code>file_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float32)\nprint(waves)\n</code></pre>\n<pre><code>[[-5.9483055e-21 -5.8499546e-21 -5.4241516e-21 ... -6.0669897e-21\n  -5.9634572e-21 -5.7577845e-21]\n [ 9.7540710e-22  4.5258614e-22  4.5864389e-23 ... -1.0960821e-20\n  -1.0976663e-20 -1.1085813e-20]\n [-1.7487198e-21 -1.1828679e-21 -1.9322378e-21 ...  1.4650227e-21\n   2.1864486e-21  1.5408594e-21]]\n</code></pre>\n<pre><code>file_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float16)\nprint(waves)\n</code></pre>\n<pre><code>[[-0. -0. -0. ... -0. -0. -0.]\n [ 0.  0.  0. ... -0. -0. -0.]\n [-0. -0. -0. ...  0.  0.  0.]]\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1394707,
      "author_name": "analokamus",
      "author_url": "",
      "post_date": "07/20/2021 14:36:50",
      "content": "<p>As <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> pointed out, FP16 cannot store extremely small values in the order of e-20 such as the data from this competition (this is because FP16 has only 5 bits for exponent!). <br>\nBut I don't think your case is related to the wave data because at least you were able to reach score of 0,8 or above. <br>\nI recommend you partially disable AMP autocast by adding a <code>with torch.cuda.amp.autocast(enabled=False):</code> subregion to your model, and see what happens. I noticed nnAudio package sometimes return NaN with AMP enabled. <br>\nHope this helps you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1394746,
          "author_name": "naoism",
          "author_url": "",
          "post_date": "07/20/2021 15:10:43",
          "content": "<p>Thank you for your comment. Indeed, it seems that wave data may be not relevant if the score of 0.8 was obtained.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1394752,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/20/2021 15:20:41",
          "content": "<p>I disagree, it is dangerous to disable autocast, as you're more likely to get NaN or wrong gradients.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1395138,
          "author_name": "analokamus",
          "author_url": "",
          "post_date": "07/21/2021 00:22:02",
          "content": "<p>I assumed <a href=\"https://www.kaggle.com/wabinab\" target=\"_blank\">@wabinab</a> uses precision instead of FP16 only, thus partially disabling autocast (=forcing FP32 operation) is a good option for debugging. I suppose this behaviour is due to autocast's failure to determine which cast to use for some specific operations used in nnAudio package.<br>\nI trained all my models with amp (autocast partially disabled).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1395331,
          "author_name": "wabinab",
          "author_url": "",
          "post_date": "07/21/2021 06:27:38",
          "content": "<p>Thanks for your comment <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> </p>\n<p>One uses fastai's \"fit_one_cycle\" to do the training so underlyingly one haven't check the source code of which part it uses autocast. One have been experimenting it with pretrained efficientnet one could still get 0.8+ AUC sometimes though one have difficulty to push it higher than 0.85 yet. It might be perhaps tuning learning rate (and maybe other hyperparameters but one haven't tried it out) is quite important. </p>\n<p>As for <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a> yes in this case all the data are represented as png images so they're scaled to either 0-255 or 0-1 depending on convention. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1395526,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/21/2021 09:35:35",
          "content": "<p><a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> Fair enough, I thought you were discussing Scaler, sorry.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1394751,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/20/2021 15:16:26",
      "content": "<p>it is better to use mixed precision, for instance torch.cuda.amp if you use pytorch.  With this, model weights are FP32 but gradient is computed using FP16.  Moreover amp rescales gradient so that max FP16 precision is kept.</p>\n<p>I train all my models with amp now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1395667,
      "author_name": "captbullett",
      "author_url": "",
      "post_date": "07/21/2021 12:43:13",
      "content": "<p>Good day<br>\nI read somewhere that the neural networks needed is one input layer two convolutions ReLu and the output… <br>\nI believe it was the MatLab model(?)…<br>\nThen add Propagation from two convolution networks and that’s what gives you your probability is using Bayesian theory<br>\nAlso, tensor flow can be used (?)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1480516,
      "author_name": "captbullett",
      "author_url": "",
      "post_date": "08/19/2021 04:02:53",
      "content": "<p>Good day, the FP16 I read could not be used in this as it causes issues of overfitting   So using FP64 is the one to use I read this on the MATLAB documents for this very first detection of a GWbiH<br>\n<a href=\"https://www.mathworks.com/company/newsletters/articles/confirming-the-first-ever-detection-of-gravitational-waves-by-analyzing-laser-interferometer-data.html\" target=\"_blank\">https://www.mathworks.com/company/newsletters/articles/confirming-the-first-ever-detection-of-gravitational-waves-by-analyzing-laser-interferometer-data.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1561219,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 12:20:55",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1393965": "Hi guys, I was experimenting with training on FP16 on a V100 GPU comparing the same model (4 layers of convolutions + max pool, followed by flattening and 3 layers of fully-connected layers including the final output layer) with full precision on [this](https://www.kaggle.com/wabinab/g2net-as-image) dataset with batch size of 96, and somehow FP16 have (either slightly or far more) worse performance than full precision. It seems like, with a learning rate finder, FP16 shows much more flat \"loss against learning rate\" curve than bumpy FP32. Training on FP16, two of the three times being tested shows 0.6-0.7 RocAuc training, and tweaking the learning rate (since it has a flat region over a large range of learning rate) can gives me about same accuracy as TF32, however, training for 10 epochs the RocAuc seems to fluctuate around that value of 0.82-0.84 without clear direction of improvement. \n\nAnd yes, since the above has little experimentations one cannot be sure whether it is an illusion or solid. Perhaps there exists difference in improvements between different platforms (i.e. VM vs on Kaggle).",
    "1394493": "I think it is due to the effect of floating point. The value of the data in this competition is quite small (order ~ e-20), so when we change the data's dtype to FP16, they will not retain the original value. \nThat can be confirmed by following simple code.\n\n```\nfile_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float64)\nprint(waves)\n```\n```\n[[-5.94830548e-21 -5.84995448e-21 -5.42415169e-21 ... -6.06698987e-21\n  -5.96345722e-21 -5.75778438e-21]\n [ 9.75407048e-22  4.52586118e-22  4.58643893e-23 ... -1.09608208e-20\n  -1.09766636e-20 -1.10858129e-20]\n [-1.74871983e-21 -1.18286791e-21 -1.93223777e-21 ...  1.46502268e-21\n   2.18644864e-21  1.54085934e-21]]\n```\n\n```\nfile_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float32)\nprint(waves)\n```\n```\n[[-5.9483055e-21 -5.8499546e-21 -5.4241516e-21 ... -6.0669897e-21\n  -5.9634572e-21 -5.7577845e-21]\n [ 9.7540710e-22  4.5258614e-22  4.5864389e-23 ... -1.0960821e-20\n  -1.0976663e-20 -1.1085813e-20]\n [-1.7487198e-21 -1.1828679e-21 -1.9322378e-21 ...  1.4650227e-21\n   2.1864486e-21  1.5408594e-21]]\n```\n\n```\nfile_path = train[\"file_path\"][0]\nwaves = np.load(file_path).astype(np.float16)\nprint(waves)\n```\n```\n[[-0. -0. -0. ... -0. -0. -0.]\n [ 0.  0.  0. ... -0. -0. -0.]\n [-0. -0. -0. ...  0.  0.  0.]]\n```",
    "1394707": "As @naoism pointed out, FP16 cannot store extremely small values in the order of e-20 such as the data from this competition (this is because FP16 has only 5 bits for exponent!). \nBut I don't think your case is related to the wave data because at least you were able to reach score of 0,8 or above. \nI recommend you partially disable AMP autocast by adding a `with torch.cuda.amp.autocast(enabled=False):` subregion to your model, and see what happens. I noticed nnAudio package sometimes return NaN with AMP enabled. \nHope this helps you!",
    "1394746": "Thank you for your comment. Indeed, it seems that wave data may be not relevant if the score of 0.8 was obtained.",
    "1394751": "it is better to use mixed precision, for instance torch.cuda.amp if you use pytorch.  With this, model weights are FP32 but gradient is computed using FP16.  Moreover amp rescales gradient so that max FP16 precision is kept.\n\nI train all my models with amp now.",
    "1394752": "I disagree, it is dangerous to disable autocast, as you're more likely to get NaN or wrong gradients.",
    "1395138": "I assumed @wabinab uses precision instead of FP16 only, thus partially disabling autocast (=forcing FP32 operation) is a good option for debugging. I suppose this behaviour is due to autocast's failure to determine which cast to use for some specific operations used in nnAudio package.\nI trained all my models with amp (autocast partially disabled).",
    "1395331": "Thanks for your comment @analokamus \n\nOne uses fastai's \"fit_one_cycle\" to do the training so underlyingly one haven't check the source code of which part it uses autocast. One have been experimenting it with pretrained efficientnet one could still get 0.8+ AUC sometimes though one have difficulty to push it higher than 0.85 yet. It might be perhaps tuning learning rate (and maybe other hyperparameters but one haven't tried it out) is quite important. \n\nAs for @naoism yes in this case all the data are represented as png images so they're scaled to either 0-255 or 0-1 depending on convention.",
    "1395526": "analokamus Fair enough, I thought you were discussing Scaler, sorry.",
    "1395667": "Good day\nI read somewhere that the neural networks needed is one input layer two convolutions ReLu and the output… \nI believe it was the MatLab model(?)…\nThen add Propagation from two convolution networks and that’s what gives you your probability is using Bayesian theory\nAlso, tensor flow can be used (?)",
    "1480516": "Good day, the FP16 I read could not be used in this as it causes issues of overfitting   So using FP64 is the one to use I read this on the MATLAB documents for this very first detection of a GWbiH\nhttps://www.mathworks.com/company/newsletters/articles/confirming-the-first-ever-detection-of-gravitational-waves-by-analyzing-laser-interferometer-data.html",
    "1561219": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}