{
  "id": 78799,
  "title": "Sample size too small",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/78799",
  "author_name": "",
  "post_date": "2019-01-28T01:56:20.072422400Z",
  "votes": 5,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I tried to implement the method mentioned in this paper : &lt; <a href=\"https://arxiv.org/pdf/1803.03211.pdf\">PhaseNet: A Deep-Neural-Network-Based Seismic Arrival Time Picking Method</a> &gt; , a U-net like neural network. In the paper, it was mentioned \"We use stratified sampling based on stations to divide this dataset into training, validation and test datasets, with 623,054, 77,866 and 78,592 samples respectively\". While in this competition only 4194 samples are given. Any suggestions to get a network well-trained with limited sample size ?</p>",
  "messages": [
    {
      "id": "462264",
      "postDate": "01/28/2019 01:56:20",
      "content": "<p>I tried to implement the method mentioned in this paper : &lt; <a href=\"https://arxiv.org/pdf/1803.03211.pdf\">PhaseNet: A Deep-Neural-Network-Based Seismic Arrival Time Picking Method</a> &gt; , a U-net like neural network. In the paper, it was mentioned \"We use stratified sampling based on stations to divide this dataset into training, validation and test datasets, with 623,054, 77,866 and 78,592 samples respectively\". While in this competition only 4194 samples are given. Any suggestions to get a network well-trained with limited sample size ?</p>",
      "rawMarkdown": "I tried to implement the method mentioned in this paper : &lt; [PhaseNet: A Deep-Neural-Network-Based Seismic Arrival Time Picking Method][1] &gt; , a U-net like neural network. In the paper, it was mentioned \"We use stratified sampling based on stations to divide this dataset into training, validation and test datasets, with 623,054, 77,866 and 78,592 samples respectively\". While in this competition only 4194 samples are given. Any suggestions to get a network well-trained with limited sample size ?\n\n  [1]: https://arxiv.org/pdf/1803.03211.pdf",
      "votes": null
    },
    {
      "id": "462292",
      "postDate": "01/28/2019 04:04:14",
      "content": "<p>I wonder if you tried requesting data from them?</p>",
      "rawMarkdown": "I wonder if you tried requesting data from them?",
      "votes": null
    },
    {
      "id": "462315",
      "postDate": "01/28/2019 05:06:23",
      "content": "<p>I don't think there's any problem with having partially overlapping samples for this dataset. The first <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525\">paper</a> posted by the competition organizers does this. Their samples are 1.8s long but are only separated by 0.18s.</p>",
      "rawMarkdown": "I don't think there's any problem with having partially overlapping samples for this dataset. The first [paper](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525) posted by the competition organizers does this. Their samples are 1.8s long but are only separated by 0.18s.",
      "votes": null
    },
    {
      "id": "462322",
      "postDate": "01/28/2019 05:27:07",
      "content": "<p>Try that, select that submission with extremely great CV as the final submission, thr private LB is revealed, and you’ll regret to death, I’m sure. </p>",
      "rawMarkdown": "Try that, select that submission with extremely great CV as the final submission, thr private LB is revealed, and you’ll regret to death, I’m sure.",
      "votes": null
    },
    {
      "id": "462325",
      "postDate": "01/28/2019 05:42:29",
      "content": "<p>with this sample size, and you want to have some overlapping for augmentation? well, be really careful though</p>",
      "rawMarkdown": "with this sample size, and you want to have some overlapping for augmentation? well, be really careful though",
      "votes": null
    },
    {
      "id": "462326",
      "postDate": "01/28/2019 05:47:21",
      "content": "<p>The model for this paper not only builds on the 623k samples, it is also, by its nature, a mapping to 3x3001. \nThis huge dimension prediction vectors, itself, serve as a huge regularization matrix.</p>\n\n<p>In our case, even at the max sample segment length of 150k, it's span in \"time to failure\" is extremely small.\nEven if we build a model with output vector of length 150k, it is still effectively 1.</p>\n\n<p>The model would suffer.</p>\n\n<p>This model, though, will serve as a great tool for the competition host.\nOnce they collect our models, they can use our results, together with this neural net, and with longer segment size, \nthey would get a badass model with surprising performance</p>",
      "rawMarkdown": "The model for this paper not only builds on the 623k samples, it is also, by its nature, a mapping to 3x3001. \nThis huge dimension prediction vectors, itself, serve as a huge regularization matrix.\n\nIn our case, even at the max sample segment length of 150k, it's span in \"time to failure\" is extremely small.\nEven if we build a model with output vector of length 150k, it is still effectively 1.\n\nThe model would suffer.\n\nThis model, though, will serve as a great tool for the competition host.\nOnce they collect our models, they can use our results, together with this neural net, and with longer segment size, \nthey would get a badass model with surprising performance",
      "votes": null
    },
    {
      "id": "462331",
      "postDate": "01/28/2019 06:07:19",
      "content": "<p>Sorry... I'm pretty new here. What does LB mean?</p>",
      "rawMarkdown": "Sorry... I'm pretty new here. What does LB mean?",
      "votes": null
    },
    {
      "id": "462342",
      "postDate": "01/28/2019 06:35:40",
      "content": "<p>LB: learderboard</p>",
      "rawMarkdown": "LB: learderboard",
      "votes": null
    },
    {
      "id": "462468",
      "postDate": "01/28/2019 10:42:57",
      "content": "<p>Overlapping is ok if works, but keep the overlapping samples in the same train/valid split </p>",
      "rawMarkdown": "Overlapping is ok if works, but keep the overlapping samples in the same train/valid split",
      "votes": null
    },
    {
      "id": "462798",
      "postDate": "01/28/2019 23:11:18",
      "content": "<p>&gt; <strong>DavidS wrote</strong>\n&gt; \n&gt; &gt; Overlapping is ok if works, but keep the overlapping samples in the same train/valid split </p>\n\n<p>Yeah... I should have clarified that. </p>\n\n<p>&gt; <strong>Elliot wrote</strong>\n&gt; \n&gt; &gt; with this sample size, and you want to have some overlapping for augmentation? well, be really careful though</p>\n\n<p>I haven't really taken a through look into it, but I naively would think as long the separation between two samples is significantly larger than the mean auto-correlation time they would act as independent samples. Or am i missing something?</p>",
      "rawMarkdown": "&gt; **DavidS wrote**\n&gt; \n&gt; &gt; Overlapping is ok if works, but keep the overlapping samples in the same train/valid split \n\nYeah... I should have clarified that. \n\n&gt; **Elliot wrote**\n&gt; \n&gt; &gt; with this sample size, and you want to have some overlapping for augmentation? well, be really careful though\n\nI haven't really taken a through look into it, but I naively would think as long the separation between two samples is significantly larger than the mean auto-correlation time they would act as independent samples. Or am i missing something?",
      "votes": null
    },
    {
      "id": "463295",
      "postDate": "01/29/2019 18:56:25",
      "content": "<p>Why did you think it's only 4194 samples? Isn't there 600 millions plus rows, even after subtracting the 1st 150k rows? How do you tell which row is a beginning (or end) of a sample? </p>",
      "rawMarkdown": "Why did you think it's only 4194 samples? Isn't there 600 millions plus rows, even after subtracting the 1st 150k rows? How do you tell which row is a beginning (or end) of a sample?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 462292,
      "author_name": "syedzs",
      "author_url": "",
      "post_date": "01/28/2019 04:04:14",
      "content": "<p>I wonder if you tried requesting data from them?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 462315,
      "author_name": "varunvai",
      "author_url": "",
      "post_date": "01/28/2019 05:06:23",
      "content": "<p>I don't think there's any problem with having partially overlapping samples for this dataset. The first <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525\">paper</a> posted by the competition organizers does this. Their samples are 1.8s long but are only separated by 0.18s.</p>",
      "votes": null,
      "replies": [
        {
          "id": 462322,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "01/28/2019 05:27:07",
          "content": "<p>Try that, select that submission with extremely great CV as the final submission, thr private LB is revealed, and you’ll regret to death, I’m sure. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462325,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/28/2019 05:42:29",
          "content": "<p>with this sample size, and you want to have some overlapping for augmentation? well, be really careful though</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462331,
          "author_name": "varunvai",
          "author_url": "",
          "post_date": "01/28/2019 06:07:19",
          "content": "<p>Sorry... I'm pretty new here. What does LB mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462342,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/28/2019 06:35:40",
          "content": "<p>LB: learderboard</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462468,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "01/28/2019 10:42:57",
          "content": "<p>Overlapping is ok if works, but keep the overlapping samples in the same train/valid split </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462798,
          "author_name": "varunvai",
          "author_url": "",
          "post_date": "01/28/2019 23:11:18",
          "content": "<p>&gt; <strong>DavidS wrote</strong>\n&gt; \n&gt; &gt; Overlapping is ok if works, but keep the overlapping samples in the same train/valid split </p>\n\n<p>Yeah... I should have clarified that. </p>\n\n<p>&gt; <strong>Elliot wrote</strong>\n&gt; \n&gt; &gt; with this sample size, and you want to have some overlapping for augmentation? well, be really careful though</p>\n\n<p>I haven't really taken a through look into it, but I naively would think as long the separation between two samples is significantly larger than the mean auto-correlation time they would act as independent samples. Or am i missing something?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 462326,
      "author_name": "tclf90",
      "author_url": "",
      "post_date": "01/28/2019 05:47:21",
      "content": "<p>The model for this paper not only builds on the 623k samples, it is also, by its nature, a mapping to 3x3001. \nThis huge dimension prediction vectors, itself, serve as a huge regularization matrix.</p>\n\n<p>In our case, even at the max sample segment length of 150k, it's span in \"time to failure\" is extremely small.\nEven if we build a model with output vector of length 150k, it is still effectively 1.</p>\n\n<p>The model would suffer.</p>\n\n<p>This model, though, will serve as a great tool for the competition host.\nOnce they collect our models, they can use our results, together with this neural net, and with longer segment size, \nthey would get a badass model with surprising performance</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 463295,
      "author_name": "azureblue83",
      "author_url": "",
      "post_date": "01/29/2019 18:56:25",
      "content": "<p>Why did you think it's only 4194 samples? Isn't there 600 millions plus rows, even after subtracting the 1st 150k rows? How do you tell which row is a beginning (or end) of a sample? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "462264": "I tried to implement the method mentioned in this paper : &lt; [PhaseNet: A Deep-Neural-Network-Based Seismic Arrival Time Picking Method][1] &gt; , a U-net like neural network. In the paper, it was mentioned \"We use stratified sampling based on stations to divide this dataset into training, validation and test datasets, with 623,054, 77,866 and 78,592 samples respectively\". While in this competition only 4194 samples are given. Any suggestions to get a network well-trained with limited sample size ?\n\n  [1]: https://arxiv.org/pdf/1803.03211.pdf",
    "462292": "I wonder if you tried requesting data from them?",
    "462315": "I don't think there's any problem with having partially overlapping samples for this dataset. The first [paper](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525) posted by the competition organizers does this. Their samples are 1.8s long but are only separated by 0.18s.",
    "462322": "Try that, select that submission with extremely great CV as the final submission, thr private LB is revealed, and you’ll regret to death, I’m sure.",
    "462325": "with this sample size, and you want to have some overlapping for augmentation? well, be really careful though",
    "462326": "The model for this paper not only builds on the 623k samples, it is also, by its nature, a mapping to 3x3001. \nThis huge dimension prediction vectors, itself, serve as a huge regularization matrix.\n\nIn our case, even at the max sample segment length of 150k, it's span in \"time to failure\" is extremely small.\nEven if we build a model with output vector of length 150k, it is still effectively 1.\n\nThe model would suffer.\n\nThis model, though, will serve as a great tool for the competition host.\nOnce they collect our models, they can use our results, together with this neural net, and with longer segment size, \nthey would get a badass model with surprising performance",
    "462331": "Sorry... I'm pretty new here. What does LB mean?",
    "462342": "LB: learderboard",
    "462468": "Overlapping is ok if works, but keep the overlapping samples in the same train/valid split",
    "462798": "&gt; **DavidS wrote**\n&gt; \n&gt; &gt; Overlapping is ok if works, but keep the overlapping samples in the same train/valid split \n\nYeah... I should have clarified that. \n\n&gt; **Elliot wrote**\n&gt; \n&gt; &gt; with this sample size, and you want to have some overlapping for augmentation? well, be really careful though\n\nI haven't really taken a through look into it, but I naively would think as long the separation between two samples is significantly larger than the mean auto-correlation time they would act as independent samples. Or am i missing something?",
    "463295": "Why did you think it's only 4194 samples? Isn't there 600 millions plus rows, even after subtracting the 1st 150k rows? How do you tell which row is a beginning (or end) of a sample?"
  },
  "source": "meta"
}