{
  "id": 58501,
  "title": "Issues with self-embedding",
  "url": "/competitions/avito-demand-prediction/discussion/58501",
  "author_name": "",
  "post_date": "2018-06-09T05:36:33.684939600Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Need one help : \nWhile trying to read the self embedding W2V file ( the code follows as of Dieter) , I am getting below error: \"'utf-8' codec can't decode byte 0x80 in position 0: invalid start byte\".</p>\n\n<p>The code :</p>\n\n<p>print('getting embeddings') \ndef get_coefs(word, *arr): \n    return word, np.asarray(arr, dtype='float32') embeddings_index = dict(get_coefs(*o.rstrip().rsplit(' ')) for o in tqdm(open(EMBEDDING_FILE,encoding='utf-8')))</p>",
  "messages": [
    {
      "id": "340411",
      "postDate": "06/09/2018 05:36:33",
      "content": "<p>Need one help : \nWhile trying to read the self embedding W2V file ( the code follows as of Dieter) , I am getting below error: \"'utf-8' codec can't decode byte 0x80 in position 0: invalid start byte\".</p>\n\n<p>The code :</p>\n\n<p>print('getting embeddings') \ndef get_coefs(word, *arr): \n    return word, np.asarray(arr, dtype='float32') embeddings_index = dict(get_coefs(*o.rstrip().rsplit(' ')) for o in tqdm(open(EMBEDDING_FILE,encoding='utf-8')))</p>",
      "rawMarkdown": "Need one help : \nWhile trying to read the self embedding W2V file ( the code follows as of Dieter) , I am getting below error: \"'utf-8' codec can't decode byte 0x80 in position 0: invalid start byte\".\n\nThe code :\n\nprint('getting embeddings') \ndef get_coefs(word, *arr): \n    return word, np.asarray(arr, dtype='float32') embeddings_index = dict(get_coefs(*o.rstrip().rsplit(' ')) for o in tqdm(open(EMBEDDING_FILE,encoding='utf-8')))",
      "votes": null
    },
    {
      "id": "340676",
      "postDate": "06/10/2018 01:22:55",
      "content": "<p>Maybe you could try <code>open(EMBEDDING_FILE,encoding='utf-8', errors='ignore')</code></p>\n\n<p>What's more, <code>gensim.models.keyedvectors</code> is a good helper</p>",
      "rawMarkdown": "Maybe you could try `open(EMBEDDING_FILE,encoding='utf-8', errors='ignore')`\n\nWhat's more, `gensim.models.keyedvectors` is a good helper",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 340676,
      "author_name": "liujilong",
      "author_url": "",
      "post_date": "06/10/2018 01:22:55",
      "content": "<p>Maybe you could try <code>open(EMBEDDING_FILE,encoding='utf-8', errors='ignore')</code></p>\n\n<p>What's more, <code>gensim.models.keyedvectors</code> is a good helper</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "340411": "Need one help : \nWhile trying to read the self embedding W2V file ( the code follows as of Dieter) , I am getting below error: \"'utf-8' codec can't decode byte 0x80 in position 0: invalid start byte\".\n\nThe code :\n\nprint('getting embeddings') \ndef get_coefs(word, *arr): \n    return word, np.asarray(arr, dtype='float32') embeddings_index = dict(get_coefs(*o.rstrip().rsplit(' ')) for o in tqdm(open(EMBEDDING_FILE,encoding='utf-8')))",
    "340676": "Maybe you could try `open(EMBEDDING_FILE,encoding='utf-8', errors='ignore')`\n\nWhat's more, `gensim.models.keyedvectors` is a good helper"
  },
  "source": "meta"
}