{
  "id": 75646,
  "title": "A small trick that may helps",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75646",
  "author_name": "",
  "post_date": "2018-12-24T13:16:34.392385100Z",
  "votes": 30,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Try word.capitalize() if you can't find a word in embeddings_index, it works for me!</p>\n\n<pre><code>    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: \n        embedding_matrix[i] = embedding_vector\n    else:\n        embedding_vector = embeddings_index.get(word.capitalize())\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n</code></pre>",
  "messages": [
    {
      "id": "444639",
      "postDate": "12/24/2018 13:16:34",
      "content": "<p>Try word.capitalize() if you can't find a word in embeddings_index, it works for me!</p>\n\n<pre><code>    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: \n        embedding_matrix[i] = embedding_vector\n    else:\n        embedding_vector = embeddings_index.get(word.capitalize())\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n</code></pre>",
      "rawMarkdown": "Try word.capitalize() if you can't find a word in embeddings_index, it works for me!\n\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n        else:\n            embedding_vector = embeddings_index.get(word.capitalize())\n            if embedding_vector is not None: \n                embedding_matrix[i] = embedding_vector",
      "votes": null
    },
    {
      "id": "444665",
      "postDate": "12/24/2018 14:17:26",
      "content": "<p>You can also do <code>.upper()</code> as a third options (that way, you get some abbrevations like <code>USA</code>.</p>\n\n<p>Unfortunately, this did not improve my score, possibly because my seed for the random distribution of the oov word embeddings is already too good 😐</p>",
      "rawMarkdown": "You can also do `.upper()` as a third options (that way, you get some abbrevations like `USA`.\n\nUnfortunately, this did not improve my score, possibly because my seed for the random distribution of the oov word embeddings is already too good 😐",
      "votes": null
    },
    {
      "id": "445051",
      "postDate": "12/25/2018 13:48:02",
      "content": "<pre>embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n        else:\n            embedding_vector = embeddings_index.get(word.lower())\n            if embedding_vector is not None and word != word.lower(): \n                embedding_matrix[i] = embedding_vector\n<code>\nI try this one, but the local cv f1 is same.</code></pre>",
      "rawMarkdown": "<pre>embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n        else:\n            embedding_vector = embeddings_index.get(word.lower())\n            if embedding_vector is not None and word != word.lower(): \n                embedding_matrix[i] = embedding_vector\n<code>\nI try this one, but the local cv f1 is same.</code></pre>",
      "votes": null
    },
    {
      "id": "445279",
      "postDate": "12/26/2018 05:54:44",
      "content": "<p>When you said it works for you, it that both local CV and LB improved? Or just LB?</p>",
      "rawMarkdown": "When you said it works for you, it that both local CV and LB improved? Or just LB?",
      "votes": null
    },
    {
      "id": "445291",
      "postDate": "12/26/2018 06:15:23",
      "content": "<p>I got same  CV as before</p>",
      "rawMarkdown": "I got same  CV as before",
      "votes": null
    },
    {
      "id": "445295",
      "postDate": "12/26/2018 06:41:36",
      "content": "<p>0.001 on local CV, torch model.</p>",
      "rawMarkdown": "0.001 on local CV, torch model.",
      "votes": null
    },
    {
      "id": "452809",
      "postDate": "01/09/2019 07:17:32",
      "content": "<p>Decrease the f1_score of local CV and LB... Anyone has found this helpful?</p>",
      "rawMarkdown": "Decrease the f1_score of local CV and LB... Anyone has found this helpful?",
      "votes": null
    },
    {
      "id": "452924",
      "postDate": "01/09/2019 11:17:45",
      "content": "<p>Same here.</p>",
      "rawMarkdown": "Same here.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 444665,
      "author_name": "mschumacher",
      "author_url": "",
      "post_date": "12/24/2018 14:17:26",
      "content": "<p>You can also do <code>.upper()</code> as a third options (that way, you get some abbrevations like <code>USA</code>.</p>\n\n<p>Unfortunately, this did not improve my score, possibly because my seed for the random distribution of the oov word embeddings is already too good 😐</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445051,
      "author_name": "mldevl",
      "author_url": "",
      "post_date": "12/25/2018 13:48:02",
      "content": "<pre>embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n        else:\n            embedding_vector = embeddings_index.get(word.lower())\n            if embedding_vector is not None and word != word.lower(): \n                embedding_matrix[i] = embedding_vector\n<code>\nI try this one, but the local cv f1 is same.</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445279,
      "author_name": "hengzheng",
      "author_url": "",
      "post_date": "12/26/2018 05:54:44",
      "content": "<p>When you said it works for you, it that both local CV and LB improved? Or just LB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 445295,
          "author_name": "emotionevil",
          "author_url": "",
          "post_date": "12/26/2018 06:41:36",
          "content": "<p>0.001 on local CV, torch model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 445291,
      "author_name": "kalyankkr",
      "author_url": "",
      "post_date": "12/26/2018 06:15:23",
      "content": "<p>I got same  CV as before</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 452809,
      "author_name": "sunnymarkliu",
      "author_url": "",
      "post_date": "01/09/2019 07:17:32",
      "content": "<p>Decrease the f1_score of local CV and LB... Anyone has found this helpful?</p>",
      "votes": null,
      "replies": [
        {
          "id": 452924,
          "author_name": "manuelsh",
          "author_url": "",
          "post_date": "01/09/2019 11:17:45",
          "content": "<p>Same here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "444639": "Try word.capitalize() if you can't find a word in embeddings_index, it works for me!\n\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n        else:\n            embedding_vector = embeddings_index.get(word.capitalize())\n            if embedding_vector is not None: \n                embedding_matrix[i] = embedding_vector",
    "444665": "You can also do `.upper()` as a third options (that way, you get some abbrevations like `USA`.\n\nUnfortunately, this did not improve my score, possibly because my seed for the random distribution of the oov word embeddings is already too good 😐",
    "445051": "<pre>embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n        else:\n            embedding_vector = embeddings_index.get(word.lower())\n            if embedding_vector is not None and word != word.lower(): \n                embedding_matrix[i] = embedding_vector\n<code>\nI try this one, but the local cv f1 is same.</code></pre>",
    "445279": "When you said it works for you, it that both local CV and LB improved? Or just LB?",
    "445291": "I got same  CV as before",
    "445295": "0.001 on local CV, torch model.",
    "452809": "Decrease the f1_score of local CV and LB... Anyone has found this helpful?",
    "452924": "Same here."
  },
  "source": "meta"
}