{
  "id": 74524,
  "title": "Question about extra/external datasets",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74524",
  "author_name": "",
  "post_date": "2018-12-13T05:06:13.890653400Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I have a question concerning the \"Data\" page, where it states \"<em>External data sources are not allowed for this competition. We are, though, providing a number of word embeddings along with the dataset that can be used in the models.</em> \"  -- are we allowed to use <strong>any</strong> datasets available via the provided websites?</p>\n\n<p>For example, from the provided URLs:</p>\n\n<ul>\n<li>paragram_300_sl999 - <a href=\"https://cogcomp.org/page/resource_view/106\">https://cogcomp.org/page/resource_view/106</a></li>\n<li>wiki-news-300d-1M - <a href=\"https://fasttext.cc/docs/en/english-vectors.html\">https://fasttext.cc/docs/en/english-vectors.html</a></li>\n</ul>\n\n<p>we can find the datasets links:</p>\n\n<ul>\n<li><a href=\"https://cogcomp.org/page/data/\">https://cogcomp.org/page/data/</a></li>\n<li><a href=\"https://fasttext.cc/docs/en/dataset.html\">https://fasttext.cc/docs/en/dataset.html</a></li>\n</ul>\n\n<p>Are we allowed freely using <strong>any</strong> of these datasets? I understand that's a LOT of data.</p>\n\n<p>Thanks very much in advance,\nG.X.</p>",
  "messages": [
    {
      "id": "438096",
      "postDate": "12/13/2018 05:06:13",
      "content": "<p>Hi,</p>\n\n<p>I have a question concerning the \"Data\" page, where it states \"<em>External data sources are not allowed for this competition. We are, though, providing a number of word embeddings along with the dataset that can be used in the models.</em> \"  -- are we allowed to use <strong>any</strong> datasets available via the provided websites?</p>\n\n<p>For example, from the provided URLs:</p>\n\n<ul>\n<li>paragram_300_sl999 - <a href=\"https://cogcomp.org/page/resource_view/106\">https://cogcomp.org/page/resource_view/106</a></li>\n<li>wiki-news-300d-1M - <a href=\"https://fasttext.cc/docs/en/english-vectors.html\">https://fasttext.cc/docs/en/english-vectors.html</a></li>\n</ul>\n\n<p>we can find the datasets links:</p>\n\n<ul>\n<li><a href=\"https://cogcomp.org/page/data/\">https://cogcomp.org/page/data/</a></li>\n<li><a href=\"https://fasttext.cc/docs/en/dataset.html\">https://fasttext.cc/docs/en/dataset.html</a></li>\n</ul>\n\n<p>Are we allowed freely using <strong>any</strong> of these datasets? I understand that's a LOT of data.</p>\n\n<p>Thanks very much in advance,\nG.X.</p>",
      "rawMarkdown": "Hi,\n\nI have a question concerning the \"Data\" page, where it states \"*External data sources are not allowed for this competition. We are, though, providing a number of word embeddings along with the dataset that can be used in the models.* \"  -- are we allowed to use **any** datasets available via the provided websites?\n\nFor example, from the provided URLs:\n \n- paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n- wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html\n\nwe can find the datasets links:\n\n- https://cogcomp.org/page/data/\n- https://fasttext.cc/docs/en/dataset.html\n\nAre we allowed freely using **any** of these datasets? I understand that's a LOT of data.\n\nThanks very much in advance,\nG.X.",
      "votes": null
    },
    {
      "id": "438167",
      "postDate": "12/13/2018 07:51:37",
      "content": "<p>No, you are not, just the embeddings provided in the data and all packages within the docker image.</p>",
      "rawMarkdown": "No, you are not, just the embeddings provided in the data and all packages within the docker image.",
      "votes": null
    },
    {
      "id": "439320",
      "postDate": "12/15/2018 06:27:52",
      "content": "<p>@Philipp thanks!</p>\n\n<p>Just did a quick check on Glove 300-dim embeddings -- I noticed that, many common words are not even in the Glove model. Am I making some silly mistake? For example, the word \"learning\" is not in the first column (after being converted to lower cases).</p>",
      "rawMarkdown": "Philipp thanks!\n\nJust did a quick check on Glove 300-dim embeddings -- I noticed that, many common words are not even in the Glove model. Am I making some silly mistake? For example, the word \"learning\" is not in the first column (after being converted to lower cases).",
      "votes": null
    },
    {
      "id": "439489",
      "postDate": "12/15/2018 15:31:23",
      "content": "<p>Yes you are making a mistake. \"Learning\" and \"learning\" are in the Glove embeddings.</p>\n\n<p>EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'</p>\n\n<p>def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')</p>\n\n<p>embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding='utf-8'))</p>\n\n<p>print(embeddings_index['learning'])</p>",
      "rawMarkdown": "Yes you are making a mistake. \"Learning\" and \"learning\" are in the Glove embeddings.\n\nEMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n\nembeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding='utf-8'))\n\nprint(embeddings_index['learning'])",
      "votes": null
    },
    {
      "id": "439645",
      "postDate": "12/16/2018 01:53:55",
      "content": "<p>Thanks very much! You are right, I was mistaken somewhere. By the way, in your code, two typos need to be corrected to get the correct output (I just noticed, it's this discussion platform's editing issue).</p>\n\n<p>\"dict(getcoefs\" → \"dict(get_coefs\"\n\"embeddingsindex = \" → \"embeddings_index =\"</p>\n\n<p>However, I'm not sure what went wrong with the following code. Maybe the pandas has issues reading the data file and made it quietly?</p>\n\n<p>names = ['word'] + [f'f{i+1}' for i in range(300)]\nglove = pd.read_csv(EMBEDDING_FILE, sep = ' ', header = None, names = names)\nGloveWords = glove['word'].tolist()\nprint(glove.shape, len(GloveWords))\nprint('learning' in GloveWords)\nprint('Learning' in GloveWords)</p>\n\n<p>The outputs are:\n(2036775, 301) 2036775\nFalse\nFalse</p>",
      "rawMarkdown": "Thanks very much! You are right, I was mistaken somewhere. By the way, in your code, two typos need to be corrected to get the correct output (I just noticed, it's this discussion platform's editing issue).\n\n\"dict(getcoefs\" → \"dict(get\\_coefs\"\n\"embeddingsindex = \" → \"embeddings\\_index =\"\n\nHowever, I'm not sure what went wrong with the following code. Maybe the pandas has issues reading the data file and made it quietly?\n\nnames = ['word'] + [f'f{i+1}' for i in range(300)]\nglove = pd.read\\_csv(EMBEDDING\\_FILE, sep = ' ', header = None, names = names)\nGloveWords = glove['word'].tolist()\nprint(glove.shape, len(GloveWords))\nprint('learning' in GloveWords)\nprint('Learning' in GloveWords)\n\nThe outputs are:\n(2036775, 301) 2036775\nFalse\nFalse",
      "votes": null
    },
    {
      "id": "439661",
      "postDate": "12/16/2018 02:49:54",
      "content": "<p>To see how strange the issue is, I did another simple test, which I posted in a separate thread:\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74810\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74810</a></p>",
      "rawMarkdown": "To see how strange the issue is, I did another simple test, which I posted in a separate thread:\nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74810",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 438167,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "12/13/2018 07:51:37",
      "content": "<p>No, you are not, just the embeddings provided in the data and all packages within the docker image.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 439320,
      "author_name": "feynmann",
      "author_url": "",
      "post_date": "12/15/2018 06:27:52",
      "content": "<p>@Philipp thanks!</p>\n\n<p>Just did a quick check on Glove 300-dim embeddings -- I noticed that, many common words are not even in the Glove model. Am I making some silly mistake? For example, the word \"learning\" is not in the first column (after being converted to lower cases).</p>",
      "votes": null,
      "replies": [
        {
          "id": 439489,
          "author_name": "joeytaj",
          "author_url": "",
          "post_date": "12/15/2018 15:31:23",
          "content": "<p>Yes you are making a mistake. \"Learning\" and \"learning\" are in the Glove embeddings.</p>\n\n<p>EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'</p>\n\n<p>def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')</p>\n\n<p>embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding='utf-8'))</p>\n\n<p>print(embeddings_index['learning'])</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439645,
          "author_name": "feynmann",
          "author_url": "",
          "post_date": "12/16/2018 01:53:55",
          "content": "<p>Thanks very much! You are right, I was mistaken somewhere. By the way, in your code, two typos need to be corrected to get the correct output (I just noticed, it's this discussion platform's editing issue).</p>\n\n<p>\"dict(getcoefs\" → \"dict(get_coefs\"\n\"embeddingsindex = \" → \"embeddings_index =\"</p>\n\n<p>However, I'm not sure what went wrong with the following code. Maybe the pandas has issues reading the data file and made it quietly?</p>\n\n<p>names = ['word'] + [f'f{i+1}' for i in range(300)]\nglove = pd.read_csv(EMBEDDING_FILE, sep = ' ', header = None, names = names)\nGloveWords = glove['word'].tolist()\nprint(glove.shape, len(GloveWords))\nprint('learning' in GloveWords)\nprint('Learning' in GloveWords)</p>\n\n<p>The outputs are:\n(2036775, 301) 2036775\nFalse\nFalse</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439661,
          "author_name": "feynmann",
          "author_url": "",
          "post_date": "12/16/2018 02:49:54",
          "content": "<p>To see how strange the issue is, I did another simple test, which I posted in a separate thread:\n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74810\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74810</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "438096": "Hi,\n\nI have a question concerning the \"Data\" page, where it states \"*External data sources are not allowed for this competition. We are, though, providing a number of word embeddings along with the dataset that can be used in the models.* \"  -- are we allowed to use **any** datasets available via the provided websites?\n\nFor example, from the provided URLs:\n \n- paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n- wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html\n\nwe can find the datasets links:\n\n- https://cogcomp.org/page/data/\n- https://fasttext.cc/docs/en/dataset.html\n\nAre we allowed freely using **any** of these datasets? I understand that's a LOT of data.\n\nThanks very much in advance,\nG.X.",
    "438167": "No, you are not, just the embeddings provided in the data and all packages within the docker image.",
    "439320": "Philipp thanks!\n\nJust did a quick check on Glove 300-dim embeddings -- I noticed that, many common words are not even in the Glove model. Am I making some silly mistake? For example, the word \"learning\" is not in the first column (after being converted to lower cases).",
    "439489": "Yes you are making a mistake. \"Learning\" and \"learning\" are in the Glove embeddings.\n\nEMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n\nembeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding='utf-8'))\n\nprint(embeddings_index['learning'])",
    "439645": "Thanks very much! You are right, I was mistaken somewhere. By the way, in your code, two typos need to be corrected to get the correct output (I just noticed, it's this discussion platform's editing issue).\n\n\"dict(getcoefs\" → \"dict(get\\_coefs\"\n\"embeddingsindex = \" → \"embeddings\\_index =\"\n\nHowever, I'm not sure what went wrong with the following code. Maybe the pandas has issues reading the data file and made it quietly?\n\nnames = ['word'] + [f'f{i+1}' for i in range(300)]\nglove = pd.read\\_csv(EMBEDDING\\_FILE, sep = ' ', header = None, names = names)\nGloveWords = glove['word'].tolist()\nprint(glove.shape, len(GloveWords))\nprint('learning' in GloveWords)\nprint('Learning' in GloveWords)\n\nThe outputs are:\n(2036775, 301) 2036775\nFalse\nFalse",
    "439661": "To see how strange the issue is, I did another simple test, which I posted in a separate thread:\nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74810"
  },
  "source": "meta"
}