{
  "id": 77976,
  "title": "Target depends on index 😲",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77976",
  "author_name": "",
  "post_date": "2019-01-18T09:36:36.512863Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey, </p>\n\n<p>Well, actually it doesn't. </p>\n\n<p>But, I was a little bit surprised that many authors in kernels shuffle the data before any testing an index dependence of a target. We might guess that index could be proportional to the time and the time-series nature of data could explain a difference between local CV and public LB and a low correlation between folds. </p>\n\n<p>So, I decided to plot target = f(index). </p>\n\n<pre><code>df['time_target'] = df['target'].rolling(window=100000).mean()\n</code></pre>\n\n<p>Also, I shuffled data several times and calculated 'time_target' on shuffled data. The results are shown below. Could you guess which is not a shuffled data? </p>\n\n<p><img src=\"https://habrastorage.org/webt/vr/qr/ta/vrqrtaik4zyk5epvvb9pwyxo8dc.png\" alt=\"enter image description here\"></p>\n\n<p>From these plots we could say that the target doesn't depend on the index (by the way, nonshuffled  data is upper right). </p>",
  "messages": [
    {
      "id": "457894",
      "postDate": "01/18/2019 09:36:36",
      "content": "<p>Hey, </p>\n\n<p>Well, actually it doesn't. </p>\n\n<p>But, I was a little bit surprised that many authors in kernels shuffle the data before any testing an index dependence of a target. We might guess that index could be proportional to the time and the time-series nature of data could explain a difference between local CV and public LB and a low correlation between folds. </p>\n\n<p>So, I decided to plot target = f(index). </p>\n\n<pre><code>df['time_target'] = df['target'].rolling(window=100000).mean()\n</code></pre>\n\n<p>Also, I shuffled data several times and calculated 'time_target' on shuffled data. The results are shown below. Could you guess which is not a shuffled data? </p>\n\n<p><img src=\"https://habrastorage.org/webt/vr/qr/ta/vrqrtaik4zyk5epvvb9pwyxo8dc.png\" alt=\"enter image description here\"></p>\n\n<p>From these plots we could say that the target doesn't depend on the index (by the way, nonshuffled  data is upper right). </p>",
      "rawMarkdown": "Hey, \n\nWell, actually it doesn't. \n\nBut, I was a little bit surprised that many authors in kernels shuffle the data before any testing an index dependence of a target. We might guess that index could be proportional to the time and the time-series nature of data could explain a difference between local CV and public LB and a low correlation between folds. \n\nSo, I decided to plot target = f(index). \n\n    df['time_target'] = df['target'].rolling(window=100000).mean()\n\nAlso, I shuffled data several times and calculated 'time_target' on shuffled data. The results are shown below. Could you guess which is not a shuffled data? \n\n![enter image description here][1]\n\nFrom these plots we could say that the target doesn't depend on the index (by the way, nonshuffled  data is upper right). \n\n\n  [1]: https://habrastorage.org/webt/vr/qr/ta/vrqrtaik4zyk5epvvb9pwyxo8dc.png",
      "votes": null
    },
    {
      "id": "457995",
      "postDate": "01/18/2019 13:59:45",
      "content": "<p>At first sight at your graphs it might seem that there is some dependancy, but when one see that Y axis range is 0.060-0.064 then everything is clear that there is none.</p>",
      "rawMarkdown": "At first sight at your graphs it might seem that there is some dependancy, but when one see that Y axis range is 0.060-0.064 then everything is clear that there is none.",
      "votes": null
    },
    {
      "id": "459497",
      "postDate": "01/21/2019 21:53:17",
      "content": "<p>Thanks for clearing that up. I saw the shuffling too and had no idea what that could be for.</p>",
      "rawMarkdown": "Thanks for clearing that up. I saw the shuffling too and had no idea what that could be for.",
      "votes": null
    },
    {
      "id": "459730",
      "postDate": "01/22/2019 08:58:01",
      "content": "<p>Could it be that shuffling between epochs in batch stochastic training may speed up the convergence? \nIn that case in next epoch the batches are totally different than in previous - so maybe model trained in that manner generalizes better?</p>",
      "rawMarkdown": "Could it be that shuffling between epochs in batch stochastic training may speed up the convergence? \nIn that case in next epoch the batches are totally different than in previous - so maybe model trained in that manner generalizes better?",
      "votes": null
    },
    {
      "id": "461809",
      "postDate": "01/27/2019 03:32:22",
      "content": "<p>Shuffling is a standard practice. It has some significance when a new batch is seen by the model, when it comes to calculate batch gradients.</p>",
      "rawMarkdown": "Shuffling is a standard practice. It has some significance when a new batch is seen by the model, when it comes to calculate batch gradients.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 457995,
      "author_name": "akuropatwinski",
      "author_url": "",
      "post_date": "01/18/2019 13:59:45",
      "content": "<p>At first sight at your graphs it might seem that there is some dependancy, but when one see that Y axis range is 0.060-0.064 then everything is clear that there is none.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459497,
      "author_name": "arigion",
      "author_url": "",
      "post_date": "01/21/2019 21:53:17",
      "content": "<p>Thanks for clearing that up. I saw the shuffling too and had no idea what that could be for.</p>",
      "votes": null,
      "replies": [
        {
          "id": 459730,
          "author_name": "akuropatwinski",
          "author_url": "",
          "post_date": "01/22/2019 08:58:01",
          "content": "<p>Could it be that shuffling between epochs in batch stochastic training may speed up the convergence? \nIn that case in next epoch the batches are totally different than in previous - so maybe model trained in that manner generalizes better?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461809,
          "author_name": "s4sarath",
          "author_url": "",
          "post_date": "01/27/2019 03:32:22",
          "content": "<p>Shuffling is a standard practice. It has some significance when a new batch is seen by the model, when it comes to calculate batch gradients.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "457894": "Hey, \n\nWell, actually it doesn't. \n\nBut, I was a little bit surprised that many authors in kernels shuffle the data before any testing an index dependence of a target. We might guess that index could be proportional to the time and the time-series nature of data could explain a difference between local CV and public LB and a low correlation between folds. \n\nSo, I decided to plot target = f(index). \n\n    df['time_target'] = df['target'].rolling(window=100000).mean()\n\nAlso, I shuffled data several times and calculated 'time_target' on shuffled data. The results are shown below. Could you guess which is not a shuffled data? \n\n![enter image description here][1]\n\nFrom these plots we could say that the target doesn't depend on the index (by the way, nonshuffled  data is upper right). \n\n\n  [1]: https://habrastorage.org/webt/vr/qr/ta/vrqrtaik4zyk5epvvb9pwyxo8dc.png",
    "457995": "At first sight at your graphs it might seem that there is some dependancy, but when one see that Y axis range is 0.060-0.064 then everything is clear that there is none.",
    "459497": "Thanks for clearing that up. I saw the shuffling too and had no idea what that could be for.",
    "459730": "Could it be that shuffling between epochs in batch stochastic training may speed up the convergence? \nIn that case in next epoch the batches are totally different than in previous - so maybe model trained in that manner generalizes better?",
    "461809": "Shuffling is a standard practice. It has some significance when a new batch is seen by the model, when it comes to calculate batch gradients."
  },
  "source": "meta"
}