{
  "id": 91333,
  "title": "Are CNN or RNN good enough?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91333",
  "author_name": "GuillemDelgado",
  "post_date": "2019-05-03T10:19:46.266000",
  "votes": 6,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I have been working mainly with RNN and CNN for Earthquake Prediction and this is the field I am more comfortable with. The best result I was able to reach is LS - 1.538 using LSTMs with different types of architectures and features.</p>\n\n<p>I know that RNNs might not be the best option for this challenge as it can easily suffer from the lack of samples and also vanishing gradient but I strongly believe we can still improve the results in different ways.</p>\n\n<p>As I have not seen many discussions/kernels about RNN or CNN I wanted to start one so we can share what we tried and did not work or what we tried and improved the results.</p>\n\n<p>I'll start by saying that, I saw that using multiple-stacked LSTM improved my results by a bit using few hand-crafted features. As stacked LSTM increase the complexity, this allows to tackle the problem at different time scales. Still, Id like to use CNN to extract the features but I have not heard if this could improve a RNN approach.</p>",
  "messages": [
    {
      "id": 532333,
      "postDate": "2019-05-16T16:50:25.863Z",
      "content": "<p>My best LSTM scores CV: 2.01 and LB: 1.50.</p>",
      "rawMarkdown": "My best LSTM scores CV: 2.01 and LB: 1.50.",
      "votes": 5,
      "replies": [
        {
          "id": 532429,
          "postDate": "2019-05-16T23:57:48.837Z",
          "content": "<p>It seems like that is pretty much the floor. Hard to beat hand crafted features. Have you tried utilizing external data from other experiments?</p>",
          "rawMarkdown": "It seems like that is pretty much the floor. Hard to beat hand crafted features. Have you tried utilizing external data from other experiments?"
        },
        {
          "id": 532654,
          "postDate": "2019-05-17T13:27:00.713Z",
          "content": "<p>No I didn't try yet.</p>",
          "rawMarkdown": "No I didn't try yet."
        }
      ]
    },
    {
      "id": 526584,
      "postDate": "2019-05-03T10:19:46.267Z",
      "content": "<p>I have been working mainly with RNN and CNN for Earthquake Prediction and this is the field I am more comfortable with. The best result I was able to reach is LS - 1.538 using LSTMs with different types of architectures and features.</p>\n\n<p>I know that RNNs might not be the best option for this challenge as it can easily suffer from the lack of samples and also vanishing gradient but I strongly believe we can still improve the results in different ways.</p>\n\n<p>As I have not seen many discussions/kernels about RNN or CNN I wanted to start one so we can share what we tried and did not work or what we tried and improved the results.</p>\n\n<p>I'll start by saying that, I saw that using multiple-stacked LSTM improved my results by a bit using few hand-crafted features. As stacked LSTM increase the complexity, this allows to tackle the problem at different time scales. Still, Id like to use CNN to extract the features but I have not heard if this could improve a RNN approach.</p>",
      "rawMarkdown": "I have been working mainly with RNN and CNN for Earthquake Prediction and this is the field I am more comfortable with. The best result I was able to reach is LS - 1.538 using LSTMs with different types of architectures and features.\n\nI know that RNNs might not be the best option for this challenge as it can easily suffer from the lack of samples and also vanishing gradient but I strongly believe we can still improve the results in different ways.\n\nAs I have not seen many discussions/kernels about RNN or CNN I wanted to start one so we can share what we tried and did not work or what we tried and improved the results.\n\nI'll start by saying that, I saw that using multiple-stacked LSTM improved my results by a bit using few hand-crafted features. As stacked LSTM increase the complexity, this allows to tackle the problem at different time scales. Still, Id like to use CNN to extract the features but I have not heard if this could improve a RNN approach.",
      "votes": 5
    },
    {
      "id": 526842,
      "postDate": "2019-05-03T23:03:51.733Z",
      "content": "<p>I think the key is not to give up too early when CNNs and RNN's don't appear to be working. I've tried both, GPU RAM is really limiting what kind of architectures you can try though. That is the major downside. Kind of have to be creative while working within the limitations of RAM though. I would consider the architectures posted on the kernels as well as the ones I've experimented with are fairly simplistic. Can't go very deep or do crazy feature engineering. The other thing is probably not to focus too much on LB fitting. I have a feeling CNN/RNN's will translate well into unseen data. I'm still trying to figure out what a good CV strategy is. I'm getting a mix of overfitting and underfitting depending on which datapoints the model are trained and validated on. And from what I gather, CNN/RNN's are severely underfitting the LB, which indicates that there is a lot of room for improvement if GPU RAM wasn't a limiting factor in building deeper and more complex architectures. To be honest, there's still a lot of stuff I haven't tried considering I've only started trying about 2 weeks ago. List of things to be tried: spectrograms, attention, splitting up signal into smaller chunks, etc... </p>\n\n<p>Based on my experience with this dataset, CNNs appear to be performing better than RNNs. That's my personal experience. Maybe others here will disagree.</p>",
      "rawMarkdown": "I think the key is not to give up too early when CNNs and RNN's don't appear to be working. I've tried both, GPU RAM is really limiting what kind of architectures you can try though. That is the major downside. Kind of have to be creative while working within the limitations of RAM though. I would consider the architectures posted on the kernels as well as the ones I've experimented with are fairly simplistic. Can't go very deep or do crazy feature engineering. The other thing is probably not to focus too much on LB fitting. I have a feeling CNN/RNN's will translate well into unseen data. I'm still trying to figure out what a good CV strategy is. I'm getting a mix of overfitting and underfitting depending on which datapoints the model are trained and validated on. And from what I gather, CNN/RNN's are severely underfitting the LB, which indicates that there is a lot of room for improvement if GPU RAM wasn't a limiting factor in building deeper and more complex architectures. To be honest, there's still a lot of stuff I haven't tried considering I've only started trying about 2 weeks ago. List of things to be tried: spectrograms, attention, splitting up signal into smaller chunks, etc... \n\nBased on my experience with this dataset, CNNs appear to be performing better than RNNs. That's my personal experience. Maybe others here will disagree.",
      "votes": 2,
      "replies": [
        {
          "id": 527008,
          "postDate": "2019-05-04T11:15:01.037Z",
          "content": "<p>There is something really interesting of what you said. CNN and RNN don't appear to be working but they might work better unseen samples. Generalizing correctly might not reflect on the LB score but on the whole test set. As people now is concentrating in improving this 13% of test set which might lead to some overfitting at later stages.</p>\n\n<p>I agree, CNN are easier to be trained and perform better than RNN but I believe that the combination of CNN and RNN should be the way to go. Maybe after finding the best architecture, CNN feature visualization could be interesting to study the features and adding an attention layer to focus on the important ones.</p>\n\n<p>When you are referring to GPU RAM issues, what are exactly those? Cant you fit the networks on the GPU?</p>",
          "rawMarkdown": "There is something really interesting of what you said. CNN and RNN don't appear to be working but they might work better unseen samples. Generalizing correctly might not reflect on the LB score but on the whole test set. As people now is concentrating in improving this 13% of test set which might lead to some overfitting at later stages.\n\nI agree, CNN are easier to be trained and perform better than RNN but I believe that the combination of CNN and RNN should be the way to go. Maybe after finding the best architecture, CNN feature visualization could be interesting to study the features and adding an attention layer to focus on the important ones.\n\nWhen you are referring to GPU RAM issues, what are exactly those? Cant you fit the networks on the GPU?",
          "votes": 2
        },
        {
          "id": 527424,
          "postDate": "2019-05-05T13:11:20.517Z",
          "content": "<p>Well, try playing around with deeper or wider CNN/RNN, add feature engineering, increase batch size, and feed it more data and you will see out of memory errors. Ultimately, I wouldn't pay too much attention to LB. Just when I thought I learned to not chase the LB, I ended up doing just that on another competition. The reason I've opted for CNN/RNN for this particular comp is because I just came from a similar comp where my public LB looked very good. And I deceived myself into thinking there would not be a LB shakeup. I won't be fooled this time. I also looked at high performing solutions after the competition and they used NN architectures. I played around with them to try to figure out how they performed well. And if I'm being honest, I was one of the people that gave up on NN architectures then for a much higher performing non-NN model on LB. Not to mention, a lot of the people who post public kernels that achieve somewhat high (LB) scores do so mainly for the up-votes. That is something I've noticed. </p>",
          "rawMarkdown": "Well, try playing around with deeper or wider CNN/RNN, add feature engineering, increase batch size, and feed it more data and you will see out of memory errors. Ultimately, I wouldn't pay too much attention to LB. Just when I thought I learned to not chase the LB, I ended up doing just that on another competition. The reason I've opted for CNN/RNN for this particular comp is because I just came from a similar comp where my public LB looked very good. And I deceived myself into thinking there would not be a LB shakeup. I won't be fooled this time. I also looked at high performing solutions after the competition and they used NN architectures. I played around with them to try to figure out how they performed well. And if I'm being honest, I was one of the people that gave up on NN architectures then for a much higher performing non-NN model on LB. Not to mention, a lot of the people who post public kernels that achieve somewhat high (LB) scores do so mainly for the up-votes. That is something I've noticed. ",
          "votes": 1
        },
        {
          "id": 527537,
          "postDate": "2019-05-05T18:57:26.780Z",
          "content": "<p>Indeed, depends on how deep you are building you architecture and so on you might get out of memory. However, I still have not used the whole GPU memory in Kaggle so far with my CNN+RNN, so I am kinda fine.</p>\n\n<p>That's a good point. I'll keep researching the proper architecture, let's see if I find any interesting papers.</p>",
          "rawMarkdown": "Indeed, depends on how deep you are building you architecture and so on you might get out of memory. However, I still have not used the whole GPU memory in Kaggle so far with my CNN+RNN, so I am kinda fine.\n\nThat's a good point. I'll keep researching the proper architecture, let's see if I find any interesting papers.",
          "votes": 1
        }
      ]
    },
    {
      "id": 526710,
      "postDate": "2019-05-03T15:14:52.300Z",
      "content": "<p><a href=\"/guillemdelgado\">@guillemdelgado</a> I think one of the problems here -which I think is a problema for the reasearch purposes- is that the test data was divided into small shunks, so we can just provide an instant as input to the models.\nI wonder how much the predictions could be improved if instead of those \"instants\" we could give more information about what has happend in the experiment, and how would perform LSTM and CNN in that case. My bet is that the prediction would improve and that more patterns could be understood.</p>",
      "rawMarkdown": "@guillemdelgado I think one of the problems here -which I think is a problema for the reasearch purposes- is that the test data was divided into small shunks, so we can just provide an instant as input to the models.\nI wonder how much the predictions could be improved if instead of those \"instants\" we could give more information about what has happend in the experiment, and how would perform LSTM and CNN in that case. My bet is that the prediction would improve and that more patterns could be understood.",
      "votes": 2,
      "replies": [
        {
          "id": 526757,
          "postDate": "2019-05-03T17:17:18.437Z",
          "content": "<p>I agree, at the end you would have more test samples with more context which definitely would help. However, that would mean that we might have longer sequences to train which will produce more parameters to train or even a different approach should be used.</p>",
          "rawMarkdown": "I agree, at the end you would have more test samples with more context which definitely would help. However, that would mean that we might have longer sequences to train which will produce more parameters to train or even a different approach should be used.",
          "votes": 1
        },
        {
          "id": 526876,
          "postDate": "2019-05-04T02:05:36.220Z",
          "content": "<p>Not necessarily it means more data. You can sample a longer period of time but not so densely sampled</p>",
          "rawMarkdown": "Not necessarily it means more data. You can sample a longer period of time but not so densely sampled",
          "votes": 1
        },
        {
          "id": 527010,
          "postDate": "2019-05-04T11:20:28.193Z",
          "content": "<p>True! I thought about down-sampling the input signal for RNN but as the important part of the signal (the peak) usually is only on some samples, I was afraid of not providing enough information to the network so I discarded the idea.</p>",
          "rawMarkdown": "True! I thought about down-sampling the input signal for RNN but as the important part of the signal (the peak) usually is only on some samples, I was afraid of not providing enough information to the network so I discarded the idea.",
          "votes": 1
        }
      ]
    },
    {
      "id": 531483,
      "postDate": "2019-05-15T02:03:26.837Z",
      "content": "<p>I recently start this competition and I'm using CNN. Mi idea is to train the model using a sequence of 150.000 consecutive point of the data as one sample. I feel this approach could avoid overfitting and is in accordance with the test data so I won't care of the LB, specially as It only reflects 13% of the data.</p>",
      "rawMarkdown": "I recently start this competition and I'm using CNN. Mi idea is to train the model using a sequence of 150.000 consecutive point of the data as one sample. I feel this approach could avoid overfitting and is in accordance with the test data so I won't care of the LB, specially as It only reflects 13% of the data.",
      "replies": [
        {
          "id": 531892,
          "postDate": "2019-05-15T17:57:48.320Z",
          "content": "<p>I feel the same way. I believe they only give 13% because otherwise competition sponsor would end up getting garbage because of over-fit and leaky models. Someone put up the re-created RF models used in the experiments. I think that is a good benchmark to aim to hit with a NN model. It seems many people have shared that they are getting 1.5 ball park area with NN models. I'm a bit skeptical if there are any audio feature engineering out there. Although, if there is, it is usually just a handful of guys and they are probably domain experts.</p>",
          "rawMarkdown": "I feel the same way. I believe they only give 13% because otherwise competition sponsor would end up getting garbage because of over-fit and leaky models. Someone put up the re-created RF models used in the experiments. I think that is a good benchmark to aim to hit with a NN model. It seems many people have shared that they are getting 1.5 ball park area with NN models. I'm a bit skeptical if there are any audio feature engineering out there. Although, if there is, it is usually just a handful of guys and they are probably domain experts.",
          "votes": 1
        }
      ]
    },
    {
      "id": 526596,
      "postDate": "2019-05-03T11:03:35.983Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 526752,
          "postDate": "2019-05-03T17:03:47.230Z",
          "content": "<p>Exactly, it is not feasible to use the raw data directly to the RNN but with a combination of CNN + RNN can extract more useful features rather than the hand-crafted ones that I was using. As selecting the proper features could take lots of time, using CNN could solve this issue. </p>\n\n<p>I think the worst part is finding the correct hyper-parameters and the correct topology of the net (how many CNN layers, how many LSTM layers and all of it's parameters). However, as there is a month to go I guess it could still be feasible to fine-tune everything</p>",
          "rawMarkdown": "Exactly, it is not feasible to use the raw data directly to the RNN but with a combination of CNN + RNN can extract more useful features rather than the hand-crafted ones that I was using. As selecting the proper features could take lots of time, using CNN could solve this issue. \n\nI think the worst part is finding the correct hyper-parameters and the correct topology of the net (how many CNN layers, how many LSTM layers and all of it's parameters). However, as there is a month to go I guess it could still be feasible to fine-tune everything",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 532333,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-05-16T16:50:25.863000",
      "content": "<p>My best LSTM scores CV: 2.01 and LB: 1.50.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 532429,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-05-16T23:57:48.837000",
          "content": "<p>It seems like that is pretty much the floor. Hard to beat hand crafted features. Have you tried utilizing external data from other experiments?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532654,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-05-17T13:27:00.713000",
          "content": "<p>No I didn't try yet.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 526842,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-05-03T23:03:51.733000",
      "content": "<p>I think the key is not to give up too early when CNNs and RNN's don't appear to be working. I've tried both, GPU RAM is really limiting what kind of architectures you can try though. That is the major downside. Kind of have to be creative while working within the limitations of RAM though. I would consider the architectures posted on the kernels as well as the ones I've experimented with are fairly simplistic. Can't go very deep or do crazy feature engineering. The other thing is probably not to focus too much on LB fitting. I have a feeling CNN/RNN's will translate well into unseen data. I'm still trying to figure out what a good CV strategy is. I'm getting a mix of overfitting and underfitting depending on which datapoints the model are trained and validated on. And from what I gather, CNN/RNN's are severely underfitting the LB, which indicates that there is a lot of room for improvement if GPU RAM wasn't a limiting factor in building deeper and more complex architectures. To be honest, there's still a lot of stuff I haven't tried considering I've only started trying about 2 weeks ago. List of things to be tried: spectrograms, attention, splitting up signal into smaller chunks, etc... </p>\n\n<p>Based on my experience with this dataset, CNNs appear to be performing better than RNNs. That's my personal experience. Maybe others here will disagree.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 527008,
          "author_name": "GuillemDelgado",
          "author_url": "",
          "post_date": "2019-05-04T11:15:01.037000",
          "content": "<p>There is something really interesting of what you said. CNN and RNN don't appear to be working but they might work better unseen samples. Generalizing correctly might not reflect on the LB score but on the whole test set. As people now is concentrating in improving this 13% of test set which might lead to some overfitting at later stages.</p>\n\n<p>I agree, CNN are easier to be trained and perform better than RNN but I believe that the combination of CNN and RNN should be the way to go. Maybe after finding the best architecture, CNN feature visualization could be interesting to study the features and adding an attention layer to focus on the important ones.</p>\n\n<p>When you are referring to GPU RAM issues, what are exactly those? Cant you fit the networks on the GPU?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 527424,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-05-05T13:11:20.517000",
          "content": "<p>Well, try playing around with deeper or wider CNN/RNN, add feature engineering, increase batch size, and feed it more data and you will see out of memory errors. Ultimately, I wouldn't pay too much attention to LB. Just when I thought I learned to not chase the LB, I ended up doing just that on another competition. The reason I've opted for CNN/RNN for this particular comp is because I just came from a similar comp where my public LB looked very good. And I deceived myself into thinking there would not be a LB shakeup. I won't be fooled this time. I also looked at high performing solutions after the competition and they used NN architectures. I played around with them to try to figure out how they performed well. And if I'm being honest, I was one of the people that gave up on NN architectures then for a much higher performing non-NN model on LB. Not to mention, a lot of the people who post public kernels that achieve somewhat high (LB) scores do so mainly for the up-votes. That is something I've noticed. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527537,
          "author_name": "GuillemDelgado",
          "author_url": "",
          "post_date": "2019-05-05T18:57:26.780000",
          "content": "<p>Indeed, depends on how deep you are building you architecture and so on you might get out of memory. However, I still have not used the whole GPU memory in Kaggle so far with my CNN+RNN, so I am kinda fine.</p>\n\n<p>That's a good point. I'll keep researching the proper architecture, let's see if I find any interesting papers.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 526710,
      "author_name": "Carlos Prades K.",
      "author_url": "",
      "post_date": "2019-05-03T15:14:52.300000",
      "content": "<p><a href=\"/guillemdelgado\">@guillemdelgado</a> I think one of the problems here -which I think is a problema for the reasearch purposes- is that the test data was divided into small shunks, so we can just provide an instant as input to the models.\nI wonder how much the predictions could be improved if instead of those \"instants\" we could give more information about what has happend in the experiment, and how would perform LSTM and CNN in that case. My bet is that the prediction would improve and that more patterns could be understood.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 526757,
          "author_name": "GuillemDelgado",
          "author_url": "",
          "post_date": "2019-05-03T17:17:18.437000",
          "content": "<p>I agree, at the end you would have more test samples with more context which definitely would help. However, that would mean that we might have longer sequences to train which will produce more parameters to train or even a different approach should be used.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526876,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-05-04T02:05:36.220000",
          "content": "<p>Not necessarily it means more data. You can sample a longer period of time but not so densely sampled</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527010,
          "author_name": "GuillemDelgado",
          "author_url": "",
          "post_date": "2019-05-04T11:20:28.193000",
          "content": "<p>True! I thought about down-sampling the input signal for RNN but as the important part of the signal (the peak) usually is only on some samples, I was afraid of not providing enough information to the network so I discarded the idea.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 531483,
      "author_name": "AndresGonzalez",
      "author_url": "",
      "post_date": "2019-05-15T02:03:26.837000",
      "content": "<p>I recently start this competition and I'm using CNN. Mi idea is to train the model using a sequence of 150.000 consecutive point of the data as one sample. I feel this approach could avoid overfitting and is in accordance with the test data so I won't care of the LB, specially as It only reflects 13% of the data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 531892,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-05-15T17:57:48.320000",
          "content": "<p>I feel the same way. I believe they only give 13% because otherwise competition sponsor would end up getting garbage because of over-fit and leaky models. Someone put up the re-created RF models used in the experiments. I think that is a good benchmark to aim to hit with a NN model. It seems many people have shared that they are getting 1.5 ball park area with NN models. I'm a bit skeptical if there are any audio feature engineering out there. Although, if there is, it is usually just a handful of guys and they are probably domain experts.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 526596,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-03T11:03:35.983000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 526752,
          "author_name": "GuillemDelgado",
          "author_url": "",
          "post_date": "2019-05-03T17:03:47.230000",
          "content": "<p>Exactly, it is not feasible to use the raw data directly to the RNN but with a combination of CNN + RNN can extract more useful features rather than the hand-crafted ones that I was using. As selecting the proper features could take lots of time, using CNN could solve this issue. </p>\n\n<p>I think the worst part is finding the correct hyper-parameters and the correct topology of the net (how many CNN layers, how many LSTM layers and all of it's parameters). However, as there is a month to go I guess it could still be feasible to fine-tune everything</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "532333": "My best LSTM scores CV: 2.01 and LB: 1.50.",
    "526584": "I have been working mainly with RNN and CNN for Earthquake Prediction and this is the field I am more comfortable with. The best result I was able to reach is LS - 1.538 using LSTMs with different types of architectures and features.\n\nI know that RNNs might not be the best option for this challenge as it can easily suffer from the lack of samples and also vanishing gradient but I strongly believe we can still improve the results in different ways.\n\nAs I have not seen many discussions/kernels about RNN or CNN I wanted to start one so we can share what we tried and did not work or what we tried and improved the results.\n\nI'll start by saying that, I saw that using multiple-stacked LSTM improved my results by a bit using few hand-crafted features. As stacked LSTM increase the complexity, this allows to tackle the problem at different time scales. Still, Id like to use CNN to extract the features but I have not heard if this could improve a RNN approach.",
    "526842": "I think the key is not to give up too early when CNNs and RNN's don't appear to be working. I've tried both, GPU RAM is really limiting what kind of architectures you can try though. That is the major downside. Kind of have to be creative while working within the limitations of RAM though. I would consider the architectures posted on the kernels as well as the ones I've experimented with are fairly simplistic. Can't go very deep or do crazy feature engineering. The other thing is probably not to focus too much on LB fitting. I have a feeling CNN/RNN's will translate well into unseen data. I'm still trying to figure out what a good CV strategy is. I'm getting a mix of overfitting and underfitting depending on which datapoints the model are trained and validated on. And from what I gather, CNN/RNN's are severely underfitting the LB, which indicates that there is a lot of room for improvement if GPU RAM wasn't a limiting factor in building deeper and more complex architectures. To be honest, there's still a lot of stuff I haven't tried considering I've only started trying about 2 weeks ago. List of things to be tried: spectrograms, attention, splitting up signal into smaller chunks, etc... \n\nBased on my experience with this dataset, CNNs appear to be performing better than RNNs. That's my personal experience. Maybe others here will disagree.",
    "526710": "@guillemdelgado I think one of the problems here -which I think is a problema for the reasearch purposes- is that the test data was divided into small shunks, so we can just provide an instant as input to the models.\nI wonder how much the predictions could be improved if instead of those \"instants\" we could give more information about what has happend in the experiment, and how would perform LSTM and CNN in that case. My bet is that the prediction would improve and that more patterns could be understood.",
    "531483": "I recently start this competition and I'm using CNN. Mi idea is to train the model using a sequence of 150.000 consecutive point of the data as one sample. I feel this approach could avoid overfitting and is in accordance with the test data so I won't care of the LB, specially as It only reflects 13% of the data.",
    "526596": ""
  }
}