{
  "id": 93574,
  "title": "Stupid question",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/93574",
  "author_name": "",
  "post_date": "2019-05-28T11:03:42.068526300Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Stupid question, better safe than sorry:</p>\n\n<p>Kernel wont be re-runned on the private 87% test set(plus 13% what we already saw) but 87% of ttf that we predicted will be swaped with some new ones. But what new ones? If I had a sample that indicated ttf 2 and i predicted 2.5 and it is swaped by some ttf of 15 what are we checking here, generalisation?</p>\n\n<p>If that is true than new ttf-s that are going to be swaped will be random and not in the ball park of for example 2?</p>\n\n<p><strong>I guess my final question (than I have to be missing something-sorry!) how can a good lb(with this percentage of private/public) be a good thing?</strong></p>",
  "messages": [
    {
      "id": "538276",
      "postDate": "05/28/2019 11:03:42",
      "content": "<p>Stupid question, better safe than sorry:</p>\n\n<p>Kernel wont be re-runned on the private 87% test set(plus 13% what we already saw) but 87% of ttf that we predicted will be swaped with some new ones. But what new ones? If I had a sample that indicated ttf 2 and i predicted 2.5 and it is swaped by some ttf of 15 what are we checking here, generalisation?</p>\n\n<p>If that is true than new ttf-s that are going to be swaped will be random and not in the ball park of for example 2?</p>\n\n<p><strong>I guess my final question (than I have to be missing something-sorry!) how can a good lb(with this percentage of private/public) be a good thing?</strong></p>",
      "rawMarkdown": "Stupid question, better safe than sorry:\n\nKernel wont be re-runned on the private 87% test set(plus 13% what we already saw) but 87% of ttf that we predicted will be swaped with some new ones. But what new ones? If I had a sample that indicated ttf 2 and i predicted 2.5 and it is swaped by some ttf of 15 what are we checking here, generalisation?\n\nIf that is true than new ttf-s that are going to be swaped will be random and not in the ball park of for example 2?\n\n\n\n**I guess my final question (than I have to be missing something-sorry!) how can a good lb(with this percentage of private/public) be a good thing?**",
      "votes": null
    },
    {
      "id": "538281",
      "postDate": "05/28/2019 11:06:00",
      "content": "<p>You see the private test data and you made predictions for it.  What you don't see is which part of test is public and which part is private.  There will be no swap of data at all.</p>",
      "rawMarkdown": "You see the private test data and you made predictions for it.  What you don't see is which part of test is public and which part is private.  There will be no swap of data at all.",
      "votes": null
    },
    {
      "id": "538283",
      "postDate": "05/28/2019 11:06:23",
      "content": "<p>This is not a kernel competition, there is no re-running of kernels. What you submit is what counts and contains both public as well as private test set data.</p>",
      "rawMarkdown": "This is not a kernel competition, there is no re-running of kernels. What you submit is what counts and contains both public as well as private test set data.",
      "votes": null
    },
    {
      "id": "538300",
      "postDate": "05/28/2019 11:22:39",
      "content": "<p>Thank you!</p>\n\n<p>Does not that than leave space to look a the most prominent samples in the test data, find similiar in the train. Augment/Up-sample these in the train and let the algorithm focus on that?</p>",
      "rawMarkdown": "Thank you!\n\nDoes not that than leave space to look a the most prominent samples in the test data, find similiar in the train. Augment/Up-sample these in the train and let the algorithm focus on that?",
      "votes": null
    },
    {
      "id": "539176",
      "postDate": "05/29/2019 16:33:18",
      "content": "<p><strong>Are you suggesting to take more similar samples from the train data which aligns with the given test data and run an algorithm on it. This means that we are over-fitting right?</strong></p>",
      "rawMarkdown": "**Are you suggesting to take more similar samples from the train data which aligns with the given test data and run an algorithm on it. This means that we are over-fitting right?**",
      "votes": null
    },
    {
      "id": "539281",
      "postDate": "05/29/2019 20:16:58",
      "content": "<p>How do you define align?\nThe way I see it what I suggested is exactly finding a good CV strategy...</p>",
      "rawMarkdown": "How do you define align?\nThe way I see it what I suggested is exactly finding a good CV strategy...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 538281,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/28/2019 11:06:00",
      "content": "<p>You see the private test data and you made predictions for it.  What you don't see is which part of test is public and which part is private.  There will be no swap of data at all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 538283,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "05/28/2019 11:06:23",
      "content": "<p>This is not a kernel competition, there is no re-running of kernels. What you submit is what counts and contains both public as well as private test set data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 538300,
      "author_name": "zikazika",
      "author_url": "",
      "post_date": "05/28/2019 11:22:39",
      "content": "<p>Thank you!</p>\n\n<p>Does not that than leave space to look a the most prominent samples in the test data, find similiar in the train. Augment/Up-sample these in the train and let the algorithm focus on that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 539176,
          "author_name": "akashravichandran",
          "author_url": "",
          "post_date": "05/29/2019 16:33:18",
          "content": "<p><strong>Are you suggesting to take more similar samples from the train data which aligns with the given test data and run an algorithm on it. This means that we are over-fitting right?</strong></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539281,
          "author_name": "zikazika",
          "author_url": "",
          "post_date": "05/29/2019 20:16:58",
          "content": "<p>How do you define align?\nThe way I see it what I suggested is exactly finding a good CV strategy...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "538276": "Stupid question, better safe than sorry:\n\nKernel wont be re-runned on the private 87% test set(plus 13% what we already saw) but 87% of ttf that we predicted will be swaped with some new ones. But what new ones? If I had a sample that indicated ttf 2 and i predicted 2.5 and it is swaped by some ttf of 15 what are we checking here, generalisation?\n\nIf that is true than new ttf-s that are going to be swaped will be random and not in the ball park of for example 2?\n\n\n\n**I guess my final question (than I have to be missing something-sorry!) how can a good lb(with this percentage of private/public) be a good thing?**",
    "538281": "You see the private test data and you made predictions for it.  What you don't see is which part of test is public and which part is private.  There will be no swap of data at all.",
    "538283": "This is not a kernel competition, there is no re-running of kernels. What you submit is what counts and contains both public as well as private test set data.",
    "538300": "Thank you!\n\nDoes not that than leave space to look a the most prominent samples in the test data, find similiar in the train. Augment/Up-sample these in the train and let the algorithm focus on that?",
    "539176": "**Are you suggesting to take more similar samples from the train data which aligns with the given test data and run an algorithm on it. This means that we are over-fitting right?**",
    "539281": "How do you define align?\nThe way I see it what I suggested is exactly finding a good CV strategy..."
  },
  "source": "meta"
}