{
  "id": 94435,
  "title": "Identifying batch start points",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94435",
  "author_name": "",
  "post_date": "2019-06-04T14:01:43.584605Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Interesting small detail about the data is that it is possible to identify batch starts (4096 sample points in each batch) for test segments. It is very visible on high-amplitude signals, see examples below. On the pictures the outlier point in the middle, indexed 21, is a first sample point of its batch. </p>\n\n<p>Besides the discontinuity, batch starts are special because the standard deviation of their distribution is twice smaller than that of other points. Moreover, mean of all batch starts in the train is 2.57, compared with around 4.5 for all other points. </p>\n\n<p>Based on those features, I built a model that identify batch starts index almost perfectly for test segments (99%+ on CV). Note that in 150k-point segment there are 36-37 batch starts.</p>\n\n<p>I was very excited about that model, but unfortunately it didn't help us, despite the best effort:\n1. I tried to learn test segments ordering by observing that if two test segments go exactly one after another than batchStartIdx2 = (batchStartIdx1 - 150000) % 4096, but apparently test segments have random gaps between them, about 7500 sample points on average by my estimation, so it didn't work\n2. Features build around batch start values were not significant\n3. We run FFT on batches of 4096 and averaging, motivated by the fact that FFT doesn't like \"missing data\", but apparently it doesn't help much or at all, because FFT with random batch starts were just as good.</p>\n\n<p>Can be interesting to hear others experience on that.</p>\n\n<p><img src=\"https://i.imgur.com/1CvYyWI.jpg\" alt=\"batch start example 1\">\n<img src=\"https://i.imgur.com/UgvfWjH.jpg\" alt=\"batch start example 2\"></p>",
  "messages": [
    {
      "id": "543379",
      "postDate": "06/04/2019 14:01:43",
      "content": "<p>Interesting small detail about the data is that it is possible to identify batch starts (4096 sample points in each batch) for test segments. It is very visible on high-amplitude signals, see examples below. On the pictures the outlier point in the middle, indexed 21, is a first sample point of its batch. </p>\n\n<p>Besides the discontinuity, batch starts are special because the standard deviation of their distribution is twice smaller than that of other points. Moreover, mean of all batch starts in the train is 2.57, compared with around 4.5 for all other points. </p>\n\n<p>Based on those features, I built a model that identify batch starts index almost perfectly for test segments (99%+ on CV). Note that in 150k-point segment there are 36-37 batch starts.</p>\n\n<p>I was very excited about that model, but unfortunately it didn't help us, despite the best effort:\n1. I tried to learn test segments ordering by observing that if two test segments go exactly one after another than batchStartIdx2 = (batchStartIdx1 - 150000) % 4096, but apparently test segments have random gaps between them, about 7500 sample points on average by my estimation, so it didn't work\n2. Features build around batch start values were not significant\n3. We run FFT on batches of 4096 and averaging, motivated by the fact that FFT doesn't like \"missing data\", but apparently it doesn't help much or at all, because FFT with random batch starts were just as good.</p>\n\n<p>Can be interesting to hear others experience on that.</p>\n\n<p><img src=\"https://i.imgur.com/1CvYyWI.jpg\" alt=\"batch start example 1\">\n<img src=\"https://i.imgur.com/UgvfWjH.jpg\" alt=\"batch start example 2\"></p>",
      "rawMarkdown": "Interesting small detail about the data is that it is possible to identify batch starts (4096 sample points in each batch) for test segments. It is very visible on high-amplitude signals, see examples below. On the pictures the outlier point in the middle, indexed 21, is a first sample point of its batch. \n\nBesides the discontinuity, batch starts are special because the standard deviation of their distribution is twice smaller than that of other points. Moreover, mean of all batch starts in the train is 2.57, compared with around 4.5 for all other points. \n\nBased on those features, I built a model that identify batch starts index almost perfectly for test segments (99%+ on CV). Note that in 150k-point segment there are 36-37 batch starts.\n\nI was very excited about that model, but unfortunately it didn't help us, despite the best effort:\n1. I tried to learn test segments ordering by observing that if two test segments go exactly one after another than batchStartIdx2 = (batchStartIdx1 - 150000) % 4096, but apparently test segments have random gaps between them, about 7500 sample points on average by my estimation, so it didn't work\n2. Features build around batch start values were not significant\n3. We run FFT on batches of 4096 and averaging, motivated by the fact that FFT doesn't like \"missing data\", but apparently it doesn't help much or at all, because FFT with random batch starts were just as good.\n\nCan be interesting to hear others experience on that.\n\n![batch start example 1](https://i.imgur.com/1CvYyWI.jpg)\n![batch start example 2](https://i.imgur.com/UgvfWjH.jpg)",
      "votes": null
    },
    {
      "id": "543395",
      "postDate": "06/04/2019 14:07:40",
      "content": "<p>Thanks for sharing, I thought of looking at it because, as you write, it can be used to sequence test segments if they were adjacent to each others.  </p>",
      "rawMarkdown": "Thanks for sharing, I thought of looking at it because, as you write, it can be used to sequence test segments if they were adjacent to each others.",
      "votes": null
    },
    {
      "id": "543574",
      "postDate": "06/04/2019 15:46:42",
      "content": "<p>Very interesting. I tried to run some DL models to detect sequence of segments with ZERO success. Maybe mixing with your approach could reverse engineer some of the testset ordering of the segments. It could be an enormous advantage. </p>",
      "rawMarkdown": "Very interesting. I tried to run some DL models to detect sequence of segments with ZERO success. Maybe mixing with your approach could reverse engineer some of the testset ordering of the segments. It could be an enormous advantage.",
      "votes": null
    },
    {
      "id": "546690",
      "postDate": "06/06/2019 20:11:51",
      "content": "<p>You were not alone! I went down the same rabbit hole and hoped to reverse engineer the test chunk ordering using a combination of:\n- The TTF predictions\n- A model that predicts if a sub-chunk (1,500 observations) follows  'closely' after another sub-chunk where closely was defined with a random uniform gap of 0-4,096\n- A model that predicts if a chunk (150,000 observations) follows directly after another chunk\n- A batch start point model - modeled on the raw data with an RNN</p>\n\n<p>Sadly, despite a near-perfect batch start model, it was impossible to leverage it given the random gaps between test-chunks. I did, however, notice a specific batch start pattern which probably coincides with initial test set chunks for an earthquake cycle in:\n1db8e8, 35a2d7, 35dd45, 395e0e, 62a403, 996c37, a35e7c and d1eee8\nThat nicely aligned with the 8 cycle starts of the test data according to the P4677 paper.</p>",
      "rawMarkdown": "You were not alone! I went down the same rabbit hole and hoped to reverse engineer the test chunk ordering using a combination of:\n- The TTF predictions\n- A model that predicts if a sub-chunk (1,500 observations) follows  'closely' after another sub-chunk where closely was defined with a random uniform gap of 0-4,096\n- A model that predicts if a chunk (150,000 observations) follows directly after another chunk\n- A batch start point model - modeled on the raw data with an RNN\n\nSadly, despite a near-perfect batch start model, it was impossible to leverage it given the random gaps between test-chunks. I did, however, notice a specific batch start pattern which probably coincides with initial test set chunks for an earthquake cycle in:\n1db8e8, 35a2d7, 35dd45, 395e0e, 62a403, 996c37, a35e7c and d1eee8\nThat nicely aligned with the 8 cycle starts of the test data according to the P4677 paper.",
      "votes": null
    },
    {
      "id": "546780",
      "postDate": "06/06/2019 21:46:00",
      "content": "<p>Cool, you went deeper than me on this, interesting ideas. Can you please give more detail on that last insight. Do you say that those segments are the first of their cycles? What is the pattern, and how exactly do you know that?</p>",
      "rawMarkdown": "Cool, you went deeper than me on this, interesting ideas. Can you please give more detail on that last insight. Do you say that those segments are the first of their cycles? What is the pattern, and how exactly do you know that?",
      "votes": null
    },
    {
      "id": "546792",
      "postDate": "06/06/2019 22:13:43",
      "content": "<p>Like you, I had a model that predicted if either point was the start of a new batch. When ignoring the rare 4095 batch lengths, there can only be 4096 patterns. I computed the average likelihood of each pattern for each test segment but sadly, the most likely patterns contained no structure of which the original order could be reconstructed - the count of the most likely pattern resembled a uniform draw of the 4096 patterns which can be explained by the random gaps.\nThere was no deviation from the uniform draw for all but one pattern - the one with batch starts at ids 2544+k*4096. Of those, there were 8 which is highly unlikely assuming independent draws. The likelihood of those patterns also exceeds the ones of the second most likely patterns by at least 7 orders of magnitude. Since those test chunks are all associated with very high TTF predictions I assumed they flag the start of a cycle. </p>",
      "rawMarkdown": "Like you, I had a model that predicted if either point was the start of a new batch. When ignoring the rare 4095 batch lengths, there can only be 4096 patterns. I computed the average likelihood of each pattern for each test segment but sadly, the most likely patterns contained no structure of which the original order could be reconstructed - the count of the most likely pattern resembled a uniform draw of the 4096 patterns which can be explained by the random gaps.\nThere was no deviation from the uniform draw for all but one pattern - the one with batch starts at ids 2544+k*4096. Of those, there were 8 which is highly unlikely assuming independent draws. The likelihood of those patterns also exceeds the ones of the second most likely patterns by at least 7 orders of magnitude. Since those test chunks are all associated with very high TTF predictions I assumed they flag the start of a cycle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543395,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 14:07:40",
      "content": "<p>Thanks for sharing, I thought of looking at it because, as you write, it can be used to sequence test segments if they were adjacent to each others.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543574,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 15:46:42",
      "content": "<p>Very interesting. I tried to run some DL models to detect sequence of segments with ZERO success. Maybe mixing with your approach could reverse engineer some of the testset ordering of the segments. It could be an enormous advantage. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546690,
      "author_name": "tvdwiele",
      "author_url": "",
      "post_date": "06/06/2019 20:11:51",
      "content": "<p>You were not alone! I went down the same rabbit hole and hoped to reverse engineer the test chunk ordering using a combination of:\n- The TTF predictions\n- A model that predicts if a sub-chunk (1,500 observations) follows  'closely' after another sub-chunk where closely was defined with a random uniform gap of 0-4,096\n- A model that predicts if a chunk (150,000 observations) follows directly after another chunk\n- A batch start point model - modeled on the raw data with an RNN</p>\n\n<p>Sadly, despite a near-perfect batch start model, it was impossible to leverage it given the random gaps between test-chunks. I did, however, notice a specific batch start pattern which probably coincides with initial test set chunks for an earthquake cycle in:\n1db8e8, 35a2d7, 35dd45, 395e0e, 62a403, 996c37, a35e7c and d1eee8\nThat nicely aligned with the 8 cycle starts of the test data according to the P4677 paper.</p>",
      "votes": null,
      "replies": [
        {
          "id": 546780,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "06/06/2019 21:46:00",
          "content": "<p>Cool, you went deeper than me on this, interesting ideas. Can you please give more detail on that last insight. Do you say that those segments are the first of their cycles? What is the pattern, and how exactly do you know that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546792,
          "author_name": "tvdwiele",
          "author_url": "",
          "post_date": "06/06/2019 22:13:43",
          "content": "<p>Like you, I had a model that predicted if either point was the start of a new batch. When ignoring the rare 4095 batch lengths, there can only be 4096 patterns. I computed the average likelihood of each pattern for each test segment but sadly, the most likely patterns contained no structure of which the original order could be reconstructed - the count of the most likely pattern resembled a uniform draw of the 4096 patterns which can be explained by the random gaps.\nThere was no deviation from the uniform draw for all but one pattern - the one with batch starts at ids 2544+k*4096. Of those, there were 8 which is highly unlikely assuming independent draws. The likelihood of those patterns also exceeds the ones of the second most likely patterns by at least 7 orders of magnitude. Since those test chunks are all associated with very high TTF predictions I assumed they flag the start of a cycle. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "543379": "Interesting small detail about the data is that it is possible to identify batch starts (4096 sample points in each batch) for test segments. It is very visible on high-amplitude signals, see examples below. On the pictures the outlier point in the middle, indexed 21, is a first sample point of its batch. \n\nBesides the discontinuity, batch starts are special because the standard deviation of their distribution is twice smaller than that of other points. Moreover, mean of all batch starts in the train is 2.57, compared with around 4.5 for all other points. \n\nBased on those features, I built a model that identify batch starts index almost perfectly for test segments (99%+ on CV). Note that in 150k-point segment there are 36-37 batch starts.\n\nI was very excited about that model, but unfortunately it didn't help us, despite the best effort:\n1. I tried to learn test segments ordering by observing that if two test segments go exactly one after another than batchStartIdx2 = (batchStartIdx1 - 150000) % 4096, but apparently test segments have random gaps between them, about 7500 sample points on average by my estimation, so it didn't work\n2. Features build around batch start values were not significant\n3. We run FFT on batches of 4096 and averaging, motivated by the fact that FFT doesn't like \"missing data\", but apparently it doesn't help much or at all, because FFT with random batch starts were just as good.\n\nCan be interesting to hear others experience on that.\n\n![batch start example 1](https://i.imgur.com/1CvYyWI.jpg)\n![batch start example 2](https://i.imgur.com/UgvfWjH.jpg)",
    "543395": "Thanks for sharing, I thought of looking at it because, as you write, it can be used to sequence test segments if they were adjacent to each others.",
    "543574": "Very interesting. I tried to run some DL models to detect sequence of segments with ZERO success. Maybe mixing with your approach could reverse engineer some of the testset ordering of the segments. It could be an enormous advantage.",
    "546690": "You were not alone! I went down the same rabbit hole and hoped to reverse engineer the test chunk ordering using a combination of:\n- The TTF predictions\n- A model that predicts if a sub-chunk (1,500 observations) follows  'closely' after another sub-chunk where closely was defined with a random uniform gap of 0-4,096\n- A model that predicts if a chunk (150,000 observations) follows directly after another chunk\n- A batch start point model - modeled on the raw data with an RNN\n\nSadly, despite a near-perfect batch start model, it was impossible to leverage it given the random gaps between test-chunks. I did, however, notice a specific batch start pattern which probably coincides with initial test set chunks for an earthquake cycle in:\n1db8e8, 35a2d7, 35dd45, 395e0e, 62a403, 996c37, a35e7c and d1eee8\nThat nicely aligned with the 8 cycle starts of the test data according to the P4677 paper.",
    "546780": "Cool, you went deeper than me on this, interesting ideas. Can you please give more detail on that last insight. Do you say that those segments are the first of their cycles? What is the pattern, and how exactly do you know that?",
    "546792": "Like you, I had a model that predicted if either point was the start of a new batch. When ignoring the rare 4095 batch lengths, there can only be 4096 patterns. I computed the average likelihood of each pattern for each test segment but sadly, the most likely patterns contained no structure of which the original order could be reconstructed - the count of the most likely pattern resembled a uniform draw of the 4096 patterns which can be explained by the random gaps.\nThere was no deviation from the uniform draw for all but one pattern - the one with batch starts at ids 2544+k*4096. Of those, there were 8 which is highly unlikely assuming independent draws. The likelihood of those patterns also exceeds the ones of the second most likely patterns by at least 7 orders of magnitude. Since those test chunks are all associated with very high TTF predictions I assumed they flag the start of a cycle."
  },
  "source": "meta"
}