{
  "id": 92168,
  "title": "Looking for teammates to improve Conv1D model",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/92168",
  "author_name": "",
  "post_date": "2019-05-14T02:03:37.576952900Z",
  "votes": 2,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I have build a very simple Conv1D and tested with few training epoch. \nIt gave 1.71 LB. \nIf we work as a team to overcome the challenges arise due to the size of data, we may be able to improve a model for better result. </p>",
  "messages": [
    {
      "id": "530931",
      "postDate": "05/14/2019 02:03:37",
      "content": "<p>I have build a very simple Conv1D and tested with few training epoch. \nIt gave 1.71 LB. \nIf we work as a team to overcome the challenges arise due to the size of data, we may be able to improve a model for better result. </p>",
      "rawMarkdown": "I have build a very simple Conv1D and tested with few training epoch. \nIt gave 1.71 LB. \nIf we work as a team to overcome the challenges arise due to the size of data, we may be able to improve a model for better result.",
      "votes": null
    },
    {
      "id": "531404",
      "postDate": "05/14/2019 20:31:31",
      "content": "<p>what is your CV? <a href=\"/isaranja\">@isaranja</a> </p>",
      "rawMarkdown": "what is your CV? @isaranja",
      "votes": null
    },
    {
      "id": "531479",
      "postDate": "05/15/2019 01:47:38",
      "content": "<p>LB : 1.71, CV : 1.81</p>",
      "rawMarkdown": "LB : 1.71, CV : 1.81",
      "votes": null
    },
    {
      "id": "532119",
      "postDate": "05/16/2019 08:31:55",
      "content": "<p>Let's work together <a href=\"/isaranja\">@isaranja</a> </p>",
      "rawMarkdown": "Let's work together @isaranja",
      "votes": null
    },
    {
      "id": "532177",
      "postDate": "05/16/2019 10:48:09",
      "content": "<p>I got a simple Conv1D model that got an LB of 1.54. Feel free to take a look:\n<a href=\"https://www.kaggle.com/pedrormarques/signal-convolution-v4\">https://www.kaggle.com/pedrormarques/signal-convolution-v4</a></p>",
      "rawMarkdown": "I got a simple Conv1D model that got an LB of 1.54. Feel free to take a look:\nhttps://www.kaggle.com/pedrormarques/signal-convolution-v4",
      "votes": null
    },
    {
      "id": "532288",
      "postDate": "05/16/2019 15:19:19",
      "content": "<p>Thank you. It's great pleasure to work with you. Do you have a team that I can join ?</p>",
      "rawMarkdown": "Thank you. It's great pleasure to work with you. Do you have a team that I can join ?",
      "votes": null
    },
    {
      "id": "532290",
      "postDate": "05/16/2019 15:19:44",
      "content": "<p>Thank you. I am going through it now.</p>",
      "rawMarkdown": "Thank you. I am going through it now.",
      "votes": null
    },
    {
      "id": "532378",
      "postDate": "05/16/2019 19:12:26",
      "content": "<p>I have sent you merge request <a href=\"/isaranja\">@isaranja</a> .</p>",
      "rawMarkdown": "I have sent you merge request @isaranja .",
      "votes": null
    },
    {
      "id": "533415",
      "postDate": "05/19/2019 07:18:56",
      "content": "<p><a href=\"/pedrormarques\">@pedrormarques</a> \nThanks for posting this - it is quite interesting. Can I ask a question about it?</p>\n\n<p>Near the output, you permute the indices to enable a dense layer with output size 8. The permutation means that the dense acts on the \"time series dimension\" rather than on the \"channel dimension\" (I think - I am a pytorch user, not Keras). This is followed by an additional Dense after flattening that brings you to a single float output - the ttf. </p>\n\n<p>The question is: why not just combine the two dense layers by first flattening and then a single Dense from (32,16) to 1. In other words, what is the purpose of the first Dense, which specifically operates on a chosen dimension?</p>",
      "rawMarkdown": "pedrormarques \nThanks for posting this - it is quite interesting. Can I ask a question about it?\n\nNear the output, you permute the indices to enable a dense layer with output size 8. The permutation means that the dense acts on the \"time series dimension\" rather than on the \"channel dimension\" (I think - I am a pytorch user, not Keras). This is followed by an additional Dense after flattening that brings you to a single float output - the ttf. \n\nThe question is: why not just combine the two dense layers by first flattening and then a single Dense from (32,16) to 1. In other words, what is the purpose of the first Dense, which specifically operates on a chosen dimension?",
      "votes": null
    },
    {
      "id": "533496",
      "postDate": "05/19/2019 11:29:43",
      "content": "<p><a href=\"/petewills\">@petewills</a> your interpretation is correct, the permutation makes the Dense layer operate on the time domain rather than channels / features. So it ends up a layer that is creating a linear function across the time series for each of its input features. In that way it works in a way similar to a convolution.</p>\n\n<p>The last dense layer has a softmax activation. 'ttf' ranges from [0, 16]; I wanted to create approx. 16 to 32 classes to steer the algorithm to categorise the input shapes in buckets that would approximate the TTF. ConvNets are often used for categorisation and the hypothesis is that this problem is more of a categorisation problem that one where individual signal features contribute to the final 'ttf' value.</p>\n\n<p>The layer before last is simply a way to reduce the dimensionality of the last layer.</p>\n\n<p>I've to say, I'm not sure that what I'm doing is correct or that has a solid justification.</p>\n\n<p>I've a separate notebook that visualises the outputs of the conv layers (<a href=\"https://www.kaggle.com/pedrormarques/signal-model-explain\">https://www.kaggle.com/pedrormarques/signal-model-explain</a>). Some of the shapes of the intermediate layers do seem to follow patterns that one would consider logical... i.e. they follow peaks or the shape of some parts of the wave.</p>\n\n<p>Of course that this is far from being perfect since my CV score is around 2.0; which is not great. And you can clearly see that the network is seriously overfitting.</p>\n\n<p>What I've tried to do:\n-  The data generator allows one to generate more data by specifying strides over the input data. I.e. one can get training examples not just each 150_000 interval but the generator can create an example by advancing a 150_000 window by  over the train / test data. This helps... but it doesn't by itself fix overfitting.</p>\n\n<ul>\n<li><p>I tend to try to initially overfit a single batch of training samples when I do a modification on the network. That lets you eliminate experiments that don't really capture the input data.</p></li>\n<li><p>Visualize the shap explanations and the conv layers outputs.</p></li>\n</ul>",
      "rawMarkdown": "petewills your interpretation is correct, the permutation makes the Dense layer operate on the time domain rather than channels / features. So it ends up a layer that is creating a linear function across the time series for each of its input features. In that way it works in a way similar to a convolution.\n\nThe last dense layer has a softmax activation. 'ttf' ranges from [0, 16]; I wanted to create approx. 16 to 32 classes to steer the algorithm to categorise the input shapes in buckets that would approximate the TTF. ConvNets are often used for categorisation and the hypothesis is that this problem is more of a categorisation problem that one where individual signal features contribute to the final 'ttf' value.\n\nThe layer before last is simply a way to reduce the dimensionality of the last layer.\n\nI've to say, I'm not sure that what I'm doing is correct or that has a solid justification.\n\nI've a separate notebook that visualises the outputs of the conv layers (https://www.kaggle.com/pedrormarques/signal-model-explain). Some of the shapes of the intermediate layers do seem to follow patterns that one would consider logical... i.e. they follow peaks or the shape of some parts of the wave.\n\nOf course that this is far from being perfect since my CV score is around 2.0; which is not great. And you can clearly see that the network is seriously overfitting.\n\nWhat I've tried to do:\n-  The data generator allows one to generate more data by specifying strides over the input data. I.e. one can get training examples not just each 150_000 interval but the generator can create an example by advancing a 150_000 window by",
      "votes": null
    },
    {
      "id": "533891",
      "postDate": "05/20/2019 07:43:34",
      "content": "<p>Sounds like a reasonable design to me. As you say, overfitting seems a major issue - due to the small data set. That needs a magic bullet.  </p>\n\n<p>Looks like the evaluation runs are spiky. Would you also attribute that to small test data sets?</p>",
      "rawMarkdown": "Sounds like a reasonable design to me. As you say, overfitting seems a major issue - due to the small data set. That needs a magic bullet.  \n\nLooks like the evaluation runs are spiky. Would you also attribute that to small test data sets?",
      "votes": null
    },
    {
      "id": "536725",
      "postDate": "05/25/2019 05:03:27",
      "content": "<p>One further question: As you say the cv is around 2 but the LB is ~ 1.5. </p>\n\n<p>Can I ask why the difference, especially as the cv is overfitting? I am used to getting LB being equal or worse than cv.</p>",
      "rawMarkdown": "One further question: As you say the cv is around 2 but the LB is ~ 1.5. \n\nCan I ask why the difference, especially as the cv is overfitting? I am used to getting LB being equal or worse than cv.",
      "votes": null
    },
    {
      "id": "537562",
      "postDate": "05/27/2019 08:42:45",
      "content": "<p>I'm getting an LB score that is similar to my CV score if I only consider segments with a time_to_failure of [0.5, 8.0). If you take a look at <a href=\"https://www.kaggle.com/pedrormarques/signal-convolution-v6\">https://www.kaggle.com/pedrormarques/signal-convolution-v6</a>, you can see at the end a graph of MAE vs ttf on the validation set.</p>\n\n<p>I'm using the following script <a href=\"https://www.kaggle.com/pedrormarques/lanl-generator-py\">https://www.kaggle.com/pedrormarques/lanl-generator-py</a> in order to generate folds that have the same percentage of samples across 32 bins of TTF values (each 0.5 secs). This gives me validation sets with the same distribution as train datasets.</p>\n\n<p>This gives me that the impression that the LB scores are currently calculated only with the \"easier\" subset of samples... i.e. the interval [0.5, 8.0) for which there are more training examples and for which the problem is simpler to formulate. The examples very close to failure are harder to classify correctly and so are the examples \"far from failure\". i.e. after a failure it is hard to say whether the next failure will come in 8 secs or 16 ( the longest example in the training data).</p>\n\n<p>This is just speculation, of course. But my suggestion would be to plot MAE vs ttf on the validation set. That may be helpful to understand the LB vs CV score.</p>",
      "rawMarkdown": "I'm getting an LB score that is similar to my CV score if I only consider segments with a time_to_failure of [0.5, 8.0). If you take a look at https://www.kaggle.com/pedrormarques/signal-convolution-v6, you can see at the end a graph of MAE vs ttf on the validation set.\n\nI'm using the following script https://www.kaggle.com/pedrormarques/lanl-generator-py in order to generate folds that have the same percentage of samples across 32 bins of TTF values (each 0.5 secs). This gives me validation sets with the same distribution as train datasets.\n\nThis gives me that the impression that the LB scores are currently calculated only with the \"easier\" subset of samples... i.e. the interval [0.5, 8.0) for which there are more training examples and for which the problem is simpler to formulate. The examples very close to failure are harder to classify correctly and so are the examples \"far from failure\". i.e. after a failure it is hard to say whether the next failure will come in 8 secs or 16 ( the longest example in the training data).\n\nThis is just speculation, of course. But my suggestion would be to plot MAE vs ttf on the validation set. That may be helpful to understand the LB vs CV score.",
      "votes": null
    },
    {
      "id": "537615",
      "postDate": "05/27/2019 10:37:08",
      "content": "<p>Thanks for sharing. This should be obvious for numerous reasons. Let me explain. pub LB is 13%. there are 8-9 full EQ's in test. That means public test is 1 full low TTF earthquake, not \"easier\" subset. It has been said to treat LB as just another \"fold\". Sure, you can optimize your CV and LB to move in the same direction, but doing so has been the challenge for most people. The only thing I think will be helpful is being able to predict high ttf and overcoming mini-quakes. It is very clear that the entire train is scattered with unpredictable non-EQ spikes so just be careful. LB and CV should be taken seriously if and only if you can manage to reliably overcome the 2 challenges mentioned above.</p>\n\n<p><img src=\"https://i.imgur.com/KlnPi5D.png\" alt=\"Full test\"></p>",
      "rawMarkdown": "Thanks for sharing. This should be obvious for numerous reasons. Let me explain. pub LB is 13%. there are 8-9 full EQ's in test. That means public test is 1 full low TTF earthquake, not \"easier\" subset. It has been said to treat LB as just another \"fold\". Sure, you can optimize your CV and LB to move in the same direction, but doing so has been the challenge for most people. The only thing I think will be helpful is being able to predict high ttf and overcoming mini-quakes. It is very clear that the entire train is scattered with unpredictable non-EQ spikes so just be careful. LB and CV should be taken seriously if and only if you can manage to reliably overcome the 2 challenges mentioned above.\n\n![Full test](https://i.imgur.com/KlnPi5D.png)",
      "votes": null
    },
    {
      "id": "538687",
      "postDate": "05/29/2019 01:30:09",
      "content": "<p>\"public test is 1 full low TTF earthquake\"</p>\n\n<p>Actually, we don't know that the public is 1 full low TTF earthquake as we don't know how they chose the samples to make public. They could have sampled randomly from the full data set as far as we know. </p>",
      "rawMarkdown": "\"public test is 1 full low TTF earthquake\"\n\nActually, we don't know that the public is 1 full low TTF earthquake as we don't know how they chose the samples to make public. They could have sampled randomly from the full data set as far as we know.",
      "votes": null
    },
    {
      "id": "538959",
      "postDate": "05/29/2019 10:48:48",
      "content": "<p>True. You're right.</p>",
      "rawMarkdown": "True. You're right.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 531404,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "05/14/2019 20:31:31",
      "content": "<p>what is your CV? <a href=\"/isaranja\">@isaranja</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 531479,
          "author_name": "isaranja",
          "author_url": "",
          "post_date": "05/15/2019 01:47:38",
          "content": "<p>LB : 1.71, CV : 1.81</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532119,
          "author_name": "karanjakhar",
          "author_url": "",
          "post_date": "05/16/2019 08:31:55",
          "content": "<p>Let's work together <a href=\"/isaranja\">@isaranja</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532288,
          "author_name": "isaranja",
          "author_url": "",
          "post_date": "05/16/2019 15:19:19",
          "content": "<p>Thank you. It's great pleasure to work with you. Do you have a team that I can join ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532378,
          "author_name": "karanjakhar",
          "author_url": "",
          "post_date": "05/16/2019 19:12:26",
          "content": "<p>I have sent you merge request <a href=\"/isaranja\">@isaranja</a> .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 532177,
      "author_name": "pedrormarques",
      "author_url": "",
      "post_date": "05/16/2019 10:48:09",
      "content": "<p>I got a simple Conv1D model that got an LB of 1.54. Feel free to take a look:\n<a href=\"https://www.kaggle.com/pedrormarques/signal-convolution-v4\">https://www.kaggle.com/pedrormarques/signal-convolution-v4</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 532290,
          "author_name": "isaranja",
          "author_url": "",
          "post_date": "05/16/2019 15:19:44",
          "content": "<p>Thank you. I am going through it now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 533415,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "05/19/2019 07:18:56",
          "content": "<p><a href=\"/pedrormarques\">@pedrormarques</a> \nThanks for posting this - it is quite interesting. Can I ask a question about it?</p>\n\n<p>Near the output, you permute the indices to enable a dense layer with output size 8. The permutation means that the dense acts on the \"time series dimension\" rather than on the \"channel dimension\" (I think - I am a pytorch user, not Keras). This is followed by an additional Dense after flattening that brings you to a single float output - the ttf. </p>\n\n<p>The question is: why not just combine the two dense layers by first flattening and then a single Dense from (32,16) to 1. In other words, what is the purpose of the first Dense, which specifically operates on a chosen dimension?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 533496,
          "author_name": "pedrormarques",
          "author_url": "",
          "post_date": "05/19/2019 11:29:43",
          "content": "<p><a href=\"/petewills\">@petewills</a> your interpretation is correct, the permutation makes the Dense layer operate on the time domain rather than channels / features. So it ends up a layer that is creating a linear function across the time series for each of its input features. In that way it works in a way similar to a convolution.</p>\n\n<p>The last dense layer has a softmax activation. 'ttf' ranges from [0, 16]; I wanted to create approx. 16 to 32 classes to steer the algorithm to categorise the input shapes in buckets that would approximate the TTF. ConvNets are often used for categorisation and the hypothesis is that this problem is more of a categorisation problem that one where individual signal features contribute to the final 'ttf' value.</p>\n\n<p>The layer before last is simply a way to reduce the dimensionality of the last layer.</p>\n\n<p>I've to say, I'm not sure that what I'm doing is correct or that has a solid justification.</p>\n\n<p>I've a separate notebook that visualises the outputs of the conv layers (<a href=\"https://www.kaggle.com/pedrormarques/signal-model-explain\">https://www.kaggle.com/pedrormarques/signal-model-explain</a>). Some of the shapes of the intermediate layers do seem to follow patterns that one would consider logical... i.e. they follow peaks or the shape of some parts of the wave.</p>\n\n<p>Of course that this is far from being perfect since my CV score is around 2.0; which is not great. And you can clearly see that the network is seriously overfitting.</p>\n\n<p>What I've tried to do:\n-  The data generator allows one to generate more data by specifying strides over the input data. I.e. one can get training examples not just each 150_000 interval but the generator can create an example by advancing a 150_000 window by  over the train / test data. This helps... but it doesn't by itself fix overfitting.</p>\n\n<ul>\n<li><p>I tend to try to initially overfit a single batch of training samples when I do a modification on the network. That lets you eliminate experiments that don't really capture the input data.</p></li>\n<li><p>Visualize the shap explanations and the conv layers outputs.</p></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 533891,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "05/20/2019 07:43:34",
          "content": "<p>Sounds like a reasonable design to me. As you say, overfitting seems a major issue - due to the small data set. That needs a magic bullet.  </p>\n\n<p>Looks like the evaluation runs are spiky. Would you also attribute that to small test data sets?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536725,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "05/25/2019 05:03:27",
          "content": "<p>One further question: As you say the cv is around 2 but the LB is ~ 1.5. </p>\n\n<p>Can I ask why the difference, especially as the cv is overfitting? I am used to getting LB being equal or worse than cv.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 537562,
          "author_name": "pedrormarques",
          "author_url": "",
          "post_date": "05/27/2019 08:42:45",
          "content": "<p>I'm getting an LB score that is similar to my CV score if I only consider segments with a time_to_failure of [0.5, 8.0). If you take a look at <a href=\"https://www.kaggle.com/pedrormarques/signal-convolution-v6\">https://www.kaggle.com/pedrormarques/signal-convolution-v6</a>, you can see at the end a graph of MAE vs ttf on the validation set.</p>\n\n<p>I'm using the following script <a href=\"https://www.kaggle.com/pedrormarques/lanl-generator-py\">https://www.kaggle.com/pedrormarques/lanl-generator-py</a> in order to generate folds that have the same percentage of samples across 32 bins of TTF values (each 0.5 secs). This gives me validation sets with the same distribution as train datasets.</p>\n\n<p>This gives me that the impression that the LB scores are currently calculated only with the \"easier\" subset of samples... i.e. the interval [0.5, 8.0) for which there are more training examples and for which the problem is simpler to formulate. The examples very close to failure are harder to classify correctly and so are the examples \"far from failure\". i.e. after a failure it is hard to say whether the next failure will come in 8 secs or 16 ( the longest example in the training data).</p>\n\n<p>This is just speculation, of course. But my suggestion would be to plot MAE vs ttf on the validation set. That may be helpful to understand the LB vs CV score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 537615,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "05/27/2019 10:37:08",
          "content": "<p>Thanks for sharing. This should be obvious for numerous reasons. Let me explain. pub LB is 13%. there are 8-9 full EQ's in test. That means public test is 1 full low TTF earthquake, not \"easier\" subset. It has been said to treat LB as just another \"fold\". Sure, you can optimize your CV and LB to move in the same direction, but doing so has been the challenge for most people. The only thing I think will be helpful is being able to predict high ttf and overcoming mini-quakes. It is very clear that the entire train is scattered with unpredictable non-EQ spikes so just be careful. LB and CV should be taken seriously if and only if you can manage to reliably overcome the 2 challenges mentioned above.</p>\n\n<p><img src=\"https://i.imgur.com/KlnPi5D.png\" alt=\"Full test\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538687,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "05/29/2019 01:30:09",
          "content": "<p>\"public test is 1 full low TTF earthquake\"</p>\n\n<p>Actually, we don't know that the public is 1 full low TTF earthquake as we don't know how they chose the samples to make public. They could have sampled randomly from the full data set as far as we know. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538959,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "05/29/2019 10:48:48",
          "content": "<p>True. You're right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "530931": "I have build a very simple Conv1D and tested with few training epoch. \nIt gave 1.71 LB. \nIf we work as a team to overcome the challenges arise due to the size of data, we may be able to improve a model for better result.",
    "531404": "what is your CV? @isaranja",
    "531479": "LB : 1.71, CV : 1.81",
    "532119": "Let's work together @isaranja",
    "532177": "I got a simple Conv1D model that got an LB of 1.54. Feel free to take a look:\nhttps://www.kaggle.com/pedrormarques/signal-convolution-v4",
    "532288": "Thank you. It's great pleasure to work with you. Do you have a team that I can join ?",
    "532290": "Thank you. I am going through it now.",
    "532378": "I have sent you merge request @isaranja .",
    "533415": "pedrormarques \nThanks for posting this - it is quite interesting. Can I ask a question about it?\n\nNear the output, you permute the indices to enable a dense layer with output size 8. The permutation means that the dense acts on the \"time series dimension\" rather than on the \"channel dimension\" (I think - I am a pytorch user, not Keras). This is followed by an additional Dense after flattening that brings you to a single float output - the ttf. \n\nThe question is: why not just combine the two dense layers by first flattening and then a single Dense from (32,16) to 1. In other words, what is the purpose of the first Dense, which specifically operates on a chosen dimension?",
    "533496": "petewills your interpretation is correct, the permutation makes the Dense layer operate on the time domain rather than channels / features. So it ends up a layer that is creating a linear function across the time series for each of its input features. In that way it works in a way similar to a convolution.\n\nThe last dense layer has a softmax activation. 'ttf' ranges from [0, 16]; I wanted to create approx. 16 to 32 classes to steer the algorithm to categorise the input shapes in buckets that would approximate the TTF. ConvNets are often used for categorisation and the hypothesis is that this problem is more of a categorisation problem that one where individual signal features contribute to the final 'ttf' value.\n\nThe layer before last is simply a way to reduce the dimensionality of the last layer.\n\nI've to say, I'm not sure that what I'm doing is correct or that has a solid justification.\n\nI've a separate notebook that visualises the outputs of the conv layers (https://www.kaggle.com/pedrormarques/signal-model-explain). Some of the shapes of the intermediate layers do seem to follow patterns that one would consider logical... i.e. they follow peaks or the shape of some parts of the wave.\n\nOf course that this is far from being perfect since my CV score is around 2.0; which is not great. And you can clearly see that the network is seriously overfitting.\n\nWhat I've tried to do:\n-  The data generator allows one to generate more data by specifying strides over the input data. I.e. one can get training examples not just each 150_000 interval but the generator can create an example by advancing a 150_000 window by",
    "533891": "Sounds like a reasonable design to me. As you say, overfitting seems a major issue - due to the small data set. That needs a magic bullet.  \n\nLooks like the evaluation runs are spiky. Would you also attribute that to small test data sets?",
    "536725": "One further question: As you say the cv is around 2 but the LB is ~ 1.5. \n\nCan I ask why the difference, especially as the cv is overfitting? I am used to getting LB being equal or worse than cv.",
    "537562": "I'm getting an LB score that is similar to my CV score if I only consider segments with a time_to_failure of [0.5, 8.0). If you take a look at https://www.kaggle.com/pedrormarques/signal-convolution-v6, you can see at the end a graph of MAE vs ttf on the validation set.\n\nI'm using the following script https://www.kaggle.com/pedrormarques/lanl-generator-py in order to generate folds that have the same percentage of samples across 32 bins of TTF values (each 0.5 secs). This gives me validation sets with the same distribution as train datasets.\n\nThis gives me that the impression that the LB scores are currently calculated only with the \"easier\" subset of samples... i.e. the interval [0.5, 8.0) for which there are more training examples and for which the problem is simpler to formulate. The examples very close to failure are harder to classify correctly and so are the examples \"far from failure\". i.e. after a failure it is hard to say whether the next failure will come in 8 secs or 16 ( the longest example in the training data).\n\nThis is just speculation, of course. But my suggestion would be to plot MAE vs ttf on the validation set. That may be helpful to understand the LB vs CV score.",
    "537615": "Thanks for sharing. This should be obvious for numerous reasons. Let me explain. pub LB is 13%. there are 8-9 full EQ's in test. That means public test is 1 full low TTF earthquake, not \"easier\" subset. It has been said to treat LB as just another \"fold\". Sure, you can optimize your CV and LB to move in the same direction, but doing so has been the challenge for most people. The only thing I think will be helpful is being able to predict high ttf and overcoming mini-quakes. It is very clear that the entire train is scattered with unpredictable non-EQ spikes so just be careful. LB and CV should be taken seriously if and only if you can manage to reliably overcome the 2 challenges mentioned above.\n\n![Full test](https://i.imgur.com/KlnPi5D.png)",
    "538687": "\"public test is 1 full low TTF earthquake\"\n\nActually, we don't know that the public is 1 full low TTF earthquake as we don't know how they chose the samples to make public. They could have sampled randomly from the full data set as far as we know.",
    "538959": "True. You're right."
  },
  "source": "meta"
}