{
  "id": 76883,
  "title": "Can anyone explain Spatial Dropout in text data",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76883",
  "author_name": "Soura_env",
  "post_date": "2019-01-07T15:20:42.994000",
  "votes": 13,
  "comment_count": 10,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": 454981,
      "postDate": "2019-01-12T17:12:28.687Z",
      "content": "<p>Let me try...</p>\n\n<p>Imagine you have word embeddings with 3 dimensions, and an input sentence of 4 words. So this will be the feature matrix with each column being a word, and each row being an embedding feature. Each <code>*</code> represents a feature value:</p>\n\n<pre><code>\n  I   like  green apples\n  *     *     *     *\n  *     *     *     *\n  *     *     *     *\n</code></pre>\n\n<p>Now if you would add normal dropout of 0.5 to this, some random <em>cells</em> in the matrix would get blacked out. Here is an example with the removed values being replaced by <code>x</code></p>\n\n<pre><code>\n  I   like  green apples\n  *     x     *     x\n  x     x     *     x\n  *     *     x     *\n</code></pre>\n\n<p>With 1D spatial dropout, instead of selecting random cells, we would select random <em>rows</em> to black out. So e.g. it could look like this:</p>\n\n<pre><code>\n  I   like  green apples\n  *     *     *     *\n  x     x     x     x\n  *     *     *     *\n</code></pre>\n\n<p>This would mean the next layer in the network could not rely on the second feature at all and had to learn how to work without it completely. Which is exactly the intended effect of original dropout.</p>\n\n<p>You can see this as a slight modification to dropout so it works as expected in sequential data.</p>\n\n<p>There is even a mathematical proof somewhere that traditional dropout in sequential data just results in an effective reduction of learning rate, nothing else.</p>",
      "rawMarkdown": "Let me try...\n\nImagine you have word embeddings with 3 dimensions, and an input sentence of 4 words. So this will be the feature matrix with each column being a word, and each row being an embedding feature. Each `*` represents a feature value:\n\n<pre><code>\n  I   like  green apples\n  *     *     *     *\n  *     *     *     *\n  *     *     *     *\n</code></pre>\n\nNow if you would add normal dropout of 0.5 to this, some random *cells* in the matrix would get blacked out. Here is an example with the removed values being replaced by `x`\n\n<pre><code>\n  I   like  green apples\n  *     x     *     x\n  x     x     *     x\n  *     *     x     *\n</code></pre>\n\nWith 1D spatial dropout, instead of selecting random cells, we would select random *rows* to black out. So e.g. it could look like this:\n\n<pre><code>\n  I   like  green apples\n  *     *     *     *\n  x     x     x     x\n  *     *     *     *\n</code></pre>\n\nThis would mean the next layer in the network could not rely on the second feature at all and had to learn how to work without it completely. Which is exactly the intended effect of original dropout.\n\nYou can see this as a slight modification to dropout so it works as expected in sequential data.\n\nThere is even a mathematical proof somewhere that traditional dropout in sequential data just results in an effective reduction of learning rate, nothing else.",
      "votes": 29
    },
    {
      "id": 451729,
      "postDate": "2019-01-07T15:20:42.993Z",
      "rawMarkdown": "",
      "votes": 13
    },
    {
      "id": 451800,
      "postDate": "2019-01-07T17:17:13.973Z",
      "content": "<p>Drops dimensions from embeddings.</p>",
      "rawMarkdown": "Drops dimensions from embeddings.",
      "votes": 2,
      "replies": [
        {
          "id": 452054,
          "postDate": "2019-01-08T05:43:47.297Z",
          "content": "<p>Thanks for the reply. Can you illustrate a little bit, on how. Since the embeddings will be for words, will it remove certain ones where the vectors are similar and use one for both the words.\nIf you can share some link that will be fine as well.</p>",
          "rawMarkdown": "Thanks for the reply. Can you illustrate a little bit, on how. Since the embeddings will be for words, will it remove certain ones where the vectors are similar and use one for both the words.\nIf you can share some link that will be fine as well.",
          "votes": 1
        },
        {
          "id": 452066,
          "postDate": "2019-01-08T06:22:58.887Z",
          "content": "<p>@psi - or drop few values from dimensions of the embedding right? correct me if I am wrong.</p>",
          "rawMarkdown": "@psi - or drop few values from dimensions of the embedding right? correct me if I am wrong."
        },
        {
          "id": 452091,
          "postDate": "2019-01-08T07:16:24.463Z",
          "content": "<p>Yeah, it drops a certain percentage of dimensions from each word vector in the training sample.</p>",
          "rawMarkdown": "Yeah, it drops a certain percentage of dimensions from each word vector in the training sample.",
          "votes": 6
        },
        {
          "id": 452502,
          "postDate": "2019-01-08T20:32:46.973Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 452515,
          "postDate": "2019-01-08T20:53:58.193Z",
          "content": "<p>Dropouts are never shown in the model summary afaik.</p>",
          "rawMarkdown": "Dropouts are never shown in the model summary afaik.",
          "votes": 1
        },
        {
          "id": 452654,
          "postDate": "2019-01-09T02:04:06.310Z",
          "content": "<p>Yeah correct.Ok Thanks</p>",
          "rawMarkdown": "Yeah correct.Ok Thanks"
        },
        {
          "id": 452958,
          "postDate": "2019-01-09T12:05:01.567Z",
          "content": "<p>@Psi on what basis, certain percentage of dimensions are dropped from the embeddings?</p>",
          "rawMarkdown": "@Psi on what basis, certain percentage of dimensions are dropped from the embeddings?"
        },
        {
          "id": 452959,
          "postDate": "2019-01-09T12:06:10.860Z",
          "content": "<p>If I am not completely wrong, then yes.</p>",
          "rawMarkdown": "If I am not completely wrong, then yes."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 454981,
      "author_name": "Max Schumacher",
      "author_url": "",
      "post_date": "2019-01-12T17:12:28.687000",
      "content": "<p>Let me try...</p>\n\n<p>Imagine you have word embeddings with 3 dimensions, and an input sentence of 4 words. So this will be the feature matrix with each column being a word, and each row being an embedding feature. Each <code>*</code> represents a feature value:</p>\n\n<pre><code>\n  I   like  green apples\n  *     *     *     *\n  *     *     *     *\n  *     *     *     *\n</code></pre>\n\n<p>Now if you would add normal dropout of 0.5 to this, some random <em>cells</em> in the matrix would get blacked out. Here is an example with the removed values being replaced by <code>x</code></p>\n\n<pre><code>\n  I   like  green apples\n  *     x     *     x\n  x     x     *     x\n  *     *     x     *\n</code></pre>\n\n<p>With 1D spatial dropout, instead of selecting random cells, we would select random <em>rows</em> to black out. So e.g. it could look like this:</p>\n\n<pre><code>\n  I   like  green apples\n  *     *     *     *\n  x     x     x     x\n  *     *     *     *\n</code></pre>\n\n<p>This would mean the next layer in the network could not rely on the second feature at all and had to learn how to work without it completely. Which is exactly the intended effect of original dropout.</p>\n\n<p>You can see this as a slight modification to dropout so it works as expected in sequential data.</p>\n\n<p>There is even a mathematical proof somewhere that traditional dropout in sequential data just results in an effective reduction of learning rate, nothing else.</p>",
      "votes": 29,
      "replies": []
    },
    {
      "id": 451800,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-01-07T17:17:13.973000",
      "content": "<p>Drops dimensions from embeddings.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 452054,
          "author_name": "Soura_env",
          "author_url": "",
          "post_date": "2019-01-08T05:43:47.297000",
          "content": "<p>Thanks for the reply. Can you illustrate a little bit, on how. Since the embeddings will be for words, will it remove certain ones where the vectors are similar and use one for both the words.\nIf you can share some link that will be fine as well.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 452066,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-01-08T06:22:58.887000",
          "content": "<p>@psi - or drop few values from dimensions of the embedding right? correct me if I am wrong.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452091,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-01-08T07:16:24.463000",
          "content": "<p>Yeah, it drops a certain percentage of dimensions from each word vector in the training sample.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 452502,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-08T20:32:46.973000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452515,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-01-08T20:53:58.193000",
          "content": "<p>Dropouts are never shown in the model summary afaik.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 452654,
          "author_name": "Soura_env",
          "author_url": "",
          "post_date": "2019-01-09T02:04:06.310000",
          "content": "<p>Yeah correct.Ok Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452958,
          "author_name": "Rajesh Shreedhar",
          "author_url": "",
          "post_date": "2019-01-09T12:05:01.567000",
          "content": "<p>@Psi on what basis, certain percentage of dimensions are dropped from the embeddings?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452959,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-01-09T12:06:10.860000",
          "content": "<p>If I am not completely wrong, then yes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454981": "Let me try...\n\nImagine you have word embeddings with 3 dimensions, and an input sentence of 4 words. So this will be the feature matrix with each column being a word, and each row being an embedding feature. Each `*` represents a feature value:\n\n<pre><code>\n  I   like  green apples\n  *     *     *     *\n  *     *     *     *\n  *     *     *     *\n</code></pre>\n\nNow if you would add normal dropout of 0.5 to this, some random *cells* in the matrix would get blacked out. Here is an example with the removed values being replaced by `x`\n\n<pre><code>\n  I   like  green apples\n  *     x     *     x\n  x     x     *     x\n  *     *     x     *\n</code></pre>\n\nWith 1D spatial dropout, instead of selecting random cells, we would select random *rows* to black out. So e.g. it could look like this:\n\n<pre><code>\n  I   like  green apples\n  *     *     *     *\n  x     x     x     x\n  *     *     *     *\n</code></pre>\n\nThis would mean the next layer in the network could not rely on the second feature at all and had to learn how to work without it completely. Which is exactly the intended effect of original dropout.\n\nYou can see this as a slight modification to dropout so it works as expected in sequential data.\n\nThere is even a mathematical proof somewhere that traditional dropout in sequential data just results in an effective reduction of learning rate, nothing else.",
    "451729": "",
    "451800": "Drops dimensions from embeddings."
  }
}