{
  "id": 314372,
  "title": "Embeddings contains Nan, how to deal with it?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/314372",
  "author_name": "Time Master",
  "post_date": "2022-03-22T11:28:41.622000",
  "votes": 7,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi, When I ran this in tf notebook:</p>\n<pre><code>from sklearn.neighbors import NearestNeighbors\nneigh = NearestNeighbors(n_neighbors=config.KNN,metric='cosine')\nneigh.fit(train_embeddings)\n</code></pre>\n<p>I got an error,that is, nan value is in train_embeddings.</p>\n<p>How can I deal with it ?</p>",
  "messages": [
    {
      "id": 1731446,
      "postDate": "2022-03-22T11:28:41.623Z",
      "content": "<p>Hi, When I ran this in tf notebook:</p>\n<pre><code>from sklearn.neighbors import NearestNeighbors\nneigh = NearestNeighbors(n_neighbors=config.KNN,metric='cosine')\nneigh.fit(train_embeddings)\n</code></pre>\n<p>I got an error,that is, nan value is in train_embeddings.</p>\n<p>How can I deal with it ?</p>",
      "rawMarkdown": "Hi, When I ran this in tf notebook:\n```\nfrom sklearn.neighbors import NearestNeighbors\nneigh = NearestNeighbors(n_neighbors=config.KNN,metric='cosine')\nneigh.fit(train_embeddings)\n```\nI got an error,that is, nan value is in train_embeddings.\n\nHow can I deal with it ?",
      "votes": 7
    },
    {
      "id": 1735038,
      "postDate": "2022-03-25T19:52:38.333Z",
      "content": "<p>The best solution is to investigate why NAN exist and correct the source of the problem. However, if you want a quick fix, you can use NumPy's NAN to num function:</p>\n<pre><code> import numpy as np\n train_embeddings = np.nan_to_num( train_embeddings )\n</code></pre>",
      "rawMarkdown": "The best solution is to investigate why NAN exist and correct the source of the problem. However, if you want a quick fix, you can use NumPy's NAN to num function:\n\n     import numpy as np\n     train_embeddings = np.nan_to_num( train_embeddings )",
      "votes": 3,
      "replies": [
        {
          "id": 1735265,
          "postDate": "2022-03-26T03:55:10.263Z",
          "content": "<p>In my case, NAN is in the model predictions. When I run </p>\n<pre><code>emb = model.predict(ds)\n</code></pre>\n<p>Some value in the emb is NAN. Is this the problem of input (images) or the model?</p>",
          "rawMarkdown": "In my case, NAN is in the model predictions. When I run \n\n```\nemb = model.predict(ds)\n```\n\nSome value in the emb is NAN. Is this the problem of input (images) or the model?"
        },
        {
          "id": 1737164,
          "postDate": "2022-03-28T06:58:51.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/ptran1203\" target=\"_blank\">@ptran1203</a> I'm having the same issue. How did you solve this? I use now the suggestion by Chris Deotte:<br>\n<code>embedding = np.nan_to_num( embedding )</code></p>",
          "rawMarkdown": "@ptran1203 I'm having the same issue. How did you solve this? I use now the suggestion by Chris Deotte:\n`embedding = np.nan_to_num( embedding )`"
        }
      ]
    },
    {
      "id": 1736779,
      "postDate": "2022-03-27T18:18:16.910Z",
      "content": "<p>Hi there! I also encountered this problem. I solved it by simply installing another version of tensorflow.</p>",
      "rawMarkdown": "Hi there! I also encountered this problem. I solved it by simply installing another version of tensorflow.",
      "votes": 1
    },
    {
      "id": 1735090,
      "postDate": "2022-03-25T20:59:23.467Z",
      "content": "<p>I think I also had a NaN error at that point in one particular fold. I just reduced the batch size from 4 to 2 to fix it.  </p>",
      "rawMarkdown": "I think I also had a NaN error at that point in one particular fold. I just reduced the batch size from 4 to 2 to fix it.  ",
      "votes": 1
    },
    {
      "id": 1735708,
      "postDate": "2022-03-26T14:39:04.763Z",
      "content": "<p>nan_to_num should work but I've found SimpleImputer is a bit better probably</p>\n<pre><code>from sklearn.impute import SimpleImputer\nmy_imputer = SimpleImputer()\ntrain_embeddings = my_imputer.fit_transform(train_embeddings)\n</code></pre>\n<p>you can check <a href=\"https://www.kaggle.com/code/dansbecker/handling-missing-values/notebook\" target=\"_blank\">this</a> for more details</p>",
      "rawMarkdown": "nan_to_num should work but I've found SimpleImputer is a bit better probably\n\n```\nfrom sklearn.impute import SimpleImputer\nmy_imputer = SimpleImputer()\ntrain_embeddings = my_imputer.fit_transform(train_embeddings)\n```\nyou can check [this](https://www.kaggle.com/code/dansbecker/handling-missing-values/notebook) for more details"
    },
    {
      "id": 1735614,
      "postDate": "2022-03-26T12:48:52.230Z",
      "content": "<p>my solution is add this line</p>\n<p>train_embeddings[np.isnan(train_embeddings)] = 1e-9</p>",
      "rawMarkdown": "my solution is add this line\n\ntrain_embeddings[np.isnan(train_embeddings)] = 1e-9"
    },
    {
      "id": 1731488,
      "postDate": "2022-03-22T12:16:57.753Z",
      "content": "<p>If I run on Colab it's ok but when I copy the notebook to Kaggle, It produces nan value in test embeddings</p>",
      "rawMarkdown": "If I run on Colab it's ok but when I copy the notebook to Kaggle, It produces nan value in test embeddings",
      "replies": [
        {
          "id": 1731515,
          "postDate": "2022-03-22T12:42:33.067Z",
          "content": "<p>Me,too. How do  you deal with this?</p>",
          "rawMarkdown": "Me,too. How do  you deal with this?"
        },
        {
          "id": 1732594,
          "postDate": "2022-03-23T14:41:23.920Z",
          "content": "<p>I did not solve it yet, just manually copy the weight and run on colab :v</p>",
          "rawMarkdown": "I did not solve it yet, just manually copy the weight and run on colab :v",
          "votes": 1
        },
        {
          "id": 1734522,
          "postDate": "2022-03-25T12:16:26.437Z",
          "content": "<p><a href=\"https://www.kaggle.com/rainfalllove\" target=\"_blank\">@rainfalllove</a> ,对这些出现nan的数值取均值，或者赋值一个极小的数：1e-9，顺便问一句，你还需要队友吗，我们可以一起组队往前冲吗？我目前single fold 0.773💪💪💪</p>",
          "rawMarkdown": "@rainfalllove ,对这些出现nan的数值取均值，或者赋值一个极小的数：1e-9，顺便问一句，你还需要队友吗，我们可以一起组队往前冲吗？我目前single fold 0.773💪💪💪"
        }
      ]
    },
    {
      "id": 1731462,
      "postDate": "2022-03-22T11:44:43.423Z",
      "content": "<p>I'd recommend to check the <code>train_embeddings</code> source (where you got it), there are a lot of potential reasons - from bad model prediction to bugs with numpy/torch processing </p>",
      "rawMarkdown": "I'd recommend to check the `train_embeddings` source (where you got it), there are a lot of potential reasons - from bad model prediction to bugs with numpy/torch processing "
    },
    {
      "id": 1739573,
      "postDate": "2022-03-30T06:58:47.923Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1739571,
      "postDate": "2022-03-30T06:57:33.497Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1734932,
      "postDate": "2022-03-25T18:08:50.180Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1737240,
      "postDate": "2022-03-28T09:06:50.740Z",
      "content": "<p>problem solved, thank everyone!</p>",
      "rawMarkdown": "problem solved, thank everyone!"
    }
  ],
  "comments": [
    {
      "id": 1735038,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-03-25T19:52:38.333000",
      "content": "<p>The best solution is to investigate why NAN exist and correct the source of the problem. However, if you want a quick fix, you can use NumPy's NAN to num function:</p>\n<pre><code> import numpy as np\n train_embeddings = np.nan_to_num( train_embeddings )\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 1735265,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2022-03-26T03:55:10.263000",
          "content": "<p>In my case, NAN is in the model predictions. When I run </p>\n<pre><code>emb = model.predict(ds)\n</code></pre>\n<p>Some value in the emb is NAN. Is this the problem of input (images) or the model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1737164,
          "author_name": "Meesz9",
          "author_url": "",
          "post_date": "2022-03-28T06:58:51.357000",
          "content": "<p><a href=\"https://www.kaggle.com/ptran1203\" target=\"_blank\">@ptran1203</a> I'm having the same issue. How did you solve this? I use now the suggestion by Chris Deotte:<br>\n<code>embedding = np.nan_to_num( embedding )</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1736779,
      "author_name": "Andrij",
      "author_url": "",
      "post_date": "2022-03-27T18:18:16.910000",
      "content": "<p>Hi there! I also encountered this problem. I solved it by simply installing another version of tensorflow.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1735090,
      "author_name": "Bruce Young",
      "author_url": "",
      "post_date": "2022-03-25T20:59:23.467000",
      "content": "<p>I think I also had a NaN error at that point in one particular fold. I just reduced the batch size from 4 to 2 to fix it.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1735708,
      "author_name": "Hannah B",
      "author_url": "",
      "post_date": "2022-03-26T14:39:04.763000",
      "content": "<p>nan_to_num should work but I've found SimpleImputer is a bit better probably</p>\n<pre><code>from sklearn.impute import SimpleImputer\nmy_imputer = SimpleImputer()\ntrain_embeddings = my_imputer.fit_transform(train_embeddings)\n</code></pre>\n<p>you can check <a href=\"https://www.kaggle.com/code/dansbecker/handling-missing-values/notebook\" target=\"_blank\">this</a> for more details</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1735614,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2022-03-26T12:48:52.230000",
      "content": "<p>my solution is add this line</p>\n<p>train_embeddings[np.isnan(train_embeddings)] = 1e-9</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1731488,
      "author_name": "Phat Tran",
      "author_url": "",
      "post_date": "2022-03-22T12:16:57.753000",
      "content": "<p>If I run on Colab it's ok but when I copy the notebook to Kaggle, It produces nan value in test embeddings</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1731515,
          "author_name": "Time Master",
          "author_url": "",
          "post_date": "2022-03-22T12:42:33.067000",
          "content": "<p>Me,too. How do  you deal with this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1732594,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2022-03-23T14:41:23.920000",
          "content": "<p>I did not solve it yet, just manually copy the weight and run on colab :v</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1734522,
          "author_name": "huyahuya",
          "author_url": "",
          "post_date": "2022-03-25T12:16:26.437000",
          "content": "<p><a href=\"https://www.kaggle.com/rainfalllove\" target=\"_blank\">@rainfalllove</a> ,对这些出现nan的数值取均值，或者赋值一个极小的数：1e-9，顺便问一句，你还需要队友吗，我们可以一起组队往前冲吗？我目前single fold 0.773💪💪💪</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1731462,
      "author_name": "Aleksey Alekseev",
      "author_url": "",
      "post_date": "2022-03-22T11:44:43.423000",
      "content": "<p>I'd recommend to check the <code>train_embeddings</code> source (where you got it), there are a lot of potential reasons - from bad model prediction to bugs with numpy/torch processing </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1739573,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-30T06:58:47.923000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1739571,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-30T06:57:33.497000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1734932,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-25T18:08:50.180000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1737240,
      "author_name": "Time Master",
      "author_url": "",
      "post_date": "2022-03-28T09:06:50.740000",
      "content": "<p>problem solved, thank everyone!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1731446": "Hi, When I ran this in tf notebook:\n```\nfrom sklearn.neighbors import NearestNeighbors\nneigh = NearestNeighbors(n_neighbors=config.KNN,metric='cosine')\nneigh.fit(train_embeddings)\n```\nI got an error,that is, nan value is in train_embeddings.\n\nHow can I deal with it ?",
    "1735038": "The best solution is to investigate why NAN exist and correct the source of the problem. However, if you want a quick fix, you can use NumPy's NAN to num function:\n\n     import numpy as np\n     train_embeddings = np.nan_to_num( train_embeddings )",
    "1736779": "Hi there! I also encountered this problem. I solved it by simply installing another version of tensorflow.",
    "1735090": "I think I also had a NaN error at that point in one particular fold. I just reduced the batch size from 4 to 2 to fix it.  ",
    "1735708": "nan_to_num should work but I've found SimpleImputer is a bit better probably\n\n```\nfrom sklearn.impute import SimpleImputer\nmy_imputer = SimpleImputer()\ntrain_embeddings = my_imputer.fit_transform(train_embeddings)\n```\nyou can check [this](https://www.kaggle.com/code/dansbecker/handling-missing-values/notebook) for more details",
    "1735614": "my solution is add this line\n\ntrain_embeddings[np.isnan(train_embeddings)] = 1e-9",
    "1731488": "If I run on Colab it's ok but when I copy the notebook to Kaggle, It produces nan value in test embeddings",
    "1731462": "I'd recommend to check the `train_embeddings` source (where you got it), there are a lot of potential reasons - from bad model prediction to bugs with numpy/torch processing ",
    "1739573": "",
    "1739571": "",
    "1734932": "",
    "1737240": "problem solved, thank everyone!"
  }
}