{
  "id": 74810,
  "title": "A Strange Problem",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74810",
  "author_name": "",
  "post_date": "2018-12-16T02:47:58.803547900Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>For reading/processing the Glove embedding data, in the kaggle's docker environment, I encountered a strange problem. Here are the code I ran and printed output.</p>\n\n<p>===========\nCode:</p>\n\n<p>EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'</p>\n\n<p>names = ['word'] + [f'f{i+1}' for i in range(300)]</p>\n\n<p>glove = pd.read_csv(EMBEDDING_FILE, sep = ' ', header = None, names = names)</p>\n\n<p>print('Done read_csv.', flush = True)</p>\n\n<p>w10 = glove['word'][0:10].tolist()</p>\n\n<p>i10 = glove.index[0:10].tolist()</p>\n\n<p>print('Simply printing a few words...', flush = True)</p>\n\n<p>print(w10)</p>\n\n<p>print(i10)</p>\n\n<p>=============\nOutput:</p>\n\n<p>Done read_csv.</p>\n\n<p>Simply printing a few words...</p>\n\n<p>Then it never finished (crashed?) and returned a black \"screen\" with an unhappy moji face.</p>",
  "messages": [
    {
      "id": "439659",
      "postDate": "12/16/2018 02:47:58",
      "content": "<p>For reading/processing the Glove embedding data, in the kaggle's docker environment, I encountered a strange problem. Here are the code I ran and printed output.</p>\n\n<p>===========\nCode:</p>\n\n<p>EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'</p>\n\n<p>names = ['word'] + [f'f{i+1}' for i in range(300)]</p>\n\n<p>glove = pd.read_csv(EMBEDDING_FILE, sep = ' ', header = None, names = names)</p>\n\n<p>print('Done read_csv.', flush = True)</p>\n\n<p>w10 = glove['word'][0:10].tolist()</p>\n\n<p>i10 = glove.index[0:10].tolist()</p>\n\n<p>print('Simply printing a few words...', flush = True)</p>\n\n<p>print(w10)</p>\n\n<p>print(i10)</p>\n\n<p>=============\nOutput:</p>\n\n<p>Done read_csv.</p>\n\n<p>Simply printing a few words...</p>\n\n<p>Then it never finished (crashed?) and returned a black \"screen\" with an unhappy moji face.</p>",
      "rawMarkdown": "For reading/processing the Glove embedding data, in the kaggle's docker environment, I encountered a strange problem. Here are the code I ran and printed output.\n\n===========\nCode:\n\nEMBEDDING\\_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n\nnames = ['word'] + [f'f{i+1}' for i in range(300)]\n\nglove = pd.read\\_csv(EMBEDDING\\_FILE, sep = ' ', header = None, names = names)\n\nprint('Done read\\_csv.', flush = True)\n\nw10 = glove['word'][0:10].tolist()\n\ni10 = glove.index[0:10].tolist()\n\nprint('Simply printing a few words...', flush = True)\n\nprint(w10)\n\nprint(i10)\n\n=============\nOutput:\n\nDone read_csv.\n\nSimply printing a few words...\n\nThen it never finished (crashed?) and returned a black \"screen\" with an unhappy moji face.",
      "votes": null
    },
    {
      "id": "439701",
      "postDate": "12/16/2018 05:41:42",
      "content": "<p>My guess was right, the pandas \"read_csv\" couldn't read the Glove embedding file correctly. As I posted in another thread, \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74524\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74524</a>\nit only reads up to 2036775 rows.</p>\n\n<p>The total #rows of the Glove embedding file is 2196017.</p>",
      "rawMarkdown": "My guess was right, the pandas \"read\\_csv\" couldn't read the Glove embedding file correctly. As I posted in another thread, \nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74524\nit only reads up to 2036775 rows.\n\nThe total #rows of the Glove embedding file is 2196017.",
      "votes": null
    },
    {
      "id": "440409",
      "postDate": "12/17/2018 14:12:15",
      "content": "<p>Have a look at some popular kernels. There is already a robust way to load embeddings floating around. No need to reinvent the wheel :)</p>",
      "rawMarkdown": "Have a look at some popular kernels. There is already a robust way to load embeddings floating around. No need to reinvent the wheel :)",
      "votes": null
    },
    {
      "id": "442541",
      "postDate": "12/20/2018 04:55:01",
      "content": "<p>Thanks. Yes, I did get around this problem by writing my own code (pretty simple indeed). The reason I tried \"pandas.read_csv\" because it's the least way to re-invent, unfortunately it was not working. It's yet a good way to quickly learn the \"status quo\" of the Kaggle's docker system ;).</p>",
      "rawMarkdown": "Thanks. Yes, I did get around this problem by writing my own code (pretty simple indeed). The reason I tried \"pandas.read_csv\" because it's the least way to re-invent, unfortunately it was not working. It's yet a good way to quickly learn the \"status quo\" of the Kaggle's docker system ;).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 439701,
      "author_name": "feynmann",
      "author_url": "",
      "post_date": "12/16/2018 05:41:42",
      "content": "<p>My guess was right, the pandas \"read_csv\" couldn't read the Glove embedding file correctly. As I posted in another thread, \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74524\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74524</a>\nit only reads up to 2036775 rows.</p>\n\n<p>The total #rows of the Glove embedding file is 2196017.</p>",
      "votes": null,
      "replies": [
        {
          "id": 440409,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/17/2018 14:12:15",
          "content": "<p>Have a look at some popular kernels. There is already a robust way to load embeddings floating around. No need to reinvent the wheel :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442541,
          "author_name": "feynmann",
          "author_url": "",
          "post_date": "12/20/2018 04:55:01",
          "content": "<p>Thanks. Yes, I did get around this problem by writing my own code (pretty simple indeed). The reason I tried \"pandas.read_csv\" because it's the least way to re-invent, unfortunately it was not working. It's yet a good way to quickly learn the \"status quo\" of the Kaggle's docker system ;).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "439659": "For reading/processing the Glove embedding data, in the kaggle's docker environment, I encountered a strange problem. Here are the code I ran and printed output.\n\n===========\nCode:\n\nEMBEDDING\\_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n\nnames = ['word'] + [f'f{i+1}' for i in range(300)]\n\nglove = pd.read\\_csv(EMBEDDING\\_FILE, sep = ' ', header = None, names = names)\n\nprint('Done read\\_csv.', flush = True)\n\nw10 = glove['word'][0:10].tolist()\n\ni10 = glove.index[0:10].tolist()\n\nprint('Simply printing a few words...', flush = True)\n\nprint(w10)\n\nprint(i10)\n\n=============\nOutput:\n\nDone read_csv.\n\nSimply printing a few words...\n\nThen it never finished (crashed?) and returned a black \"screen\" with an unhappy moji face.",
    "439701": "My guess was right, the pandas \"read\\_csv\" couldn't read the Glove embedding file correctly. As I posted in another thread, \nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74524\nit only reads up to 2036775 rows.\n\nThe total #rows of the Glove embedding file is 2196017.",
    "440409": "Have a look at some popular kernels. There is already a robust way to load embeddings floating around. No need to reinvent the wheel :)",
    "442541": "Thanks. Yes, I did get around this problem by writing my own code (pretty simple indeed). The reason I tried \"pandas.read_csv\" because it's the least way to re-invent, unfortunately it was not working. It's yet a good way to quickly learn the \"status quo\" of the Kaggle's docker system ;)."
  },
  "source": "meta"
}