{
  "id": 94569,
  "title": "Good models without test data",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94569",
  "author_name": "",
  "post_date": "2019-06-05T09:11:43.969050300Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Yes, you had to use test estimates to finish in top 10.  But you could get good models that do not use any test data.  I have several of them that would get a gold here:</p>\n\n<p>A knn that scores 2.382 on private\n4 gam (general additive model) that score between 2.34 and 2.39</p>\n\n<p>I think that gradient boosted machines (lgb, xgb, catboost, etc) are less interesting because they cannot extrapolate linear relationship to range no seen in training data.</p>\n\n<p>In our team we only used test data to derive sample weights, and neither sklearn KNN regressor nor pygam support sample weights.  This is why I am sure the above models do not depend on test data, and would generalize well to other test data than the one we had.</p>\n\n<p>I am pretty sure these would be useful for EQ research community.</p>",
  "messages": [
    {
      "id": "544211",
      "postDate": "06/05/2019 09:11:43",
      "content": "<p>Yes, you had to use test estimates to finish in top 10.  But you could get good models that do not use any test data.  I have several of them that would get a gold here:</p>\n\n<p>A knn that scores 2.382 on private\n4 gam (general additive model) that score between 2.34 and 2.39</p>\n\n<p>I think that gradient boosted machines (lgb, xgb, catboost, etc) are less interesting because they cannot extrapolate linear relationship to range no seen in training data.</p>\n\n<p>In our team we only used test data to derive sample weights, and neither sklearn KNN regressor nor pygam support sample weights.  This is why I am sure the above models do not depend on test data, and would generalize well to other test data than the one we had.</p>\n\n<p>I am pretty sure these would be useful for EQ research community.</p>",
      "rawMarkdown": "Yes, you had to use test estimates to finish in top 10.  But you could get good models that do not use any test data.  I have several of them that would get a gold here:\n\nA knn that scores 2.382 on private\n4 gam (general additive model) that score between 2.34 and 2.39\n\nI think that gradient boosted machines (lgb, xgb, catboost, etc) are less interesting because they cannot extrapolate linear relationship to range no seen in training data.\n\nIn our team we only used test data to derive sample weights, and neither sklearn KNN regressor nor pygam support sample weights.  This is why I am sure the above models do not depend on test data, and would generalize well to other test data than the one we had.\n\nI am pretty sure these would be useful for EQ research community.",
      "votes": null
    },
    {
      "id": "544233",
      "postDate": "06/05/2019 09:45:19",
      "content": "<p>I can just confirm. I didn't use any test data and finished 15th.</p>",
      "rawMarkdown": "I can just confirm. I didn't use any test data and finished 15th.",
      "votes": null
    },
    {
      "id": "544272",
      "postDate": "06/05/2019 10:46:35",
      "content": "<p>I guess using the whole supplied train data without any modifications?</p>",
      "rawMarkdown": "I guess using the whole supplied train data without any modifications?",
      "votes": null
    },
    {
      "id": "544279",
      "postDate": "06/05/2019 10:52:17",
      "content": "<p>Using the same features as for other models with one modification: standard scaler fit on training data.</p>",
      "rawMarkdown": "Using the same features as for other models with one modification: standard scaler fit on training data.",
      "votes": null
    },
    {
      "id": "544285",
      "postDate": "06/05/2019 11:03:15",
      "content": "<p>Okay, but using the training data supplied as is without mergin EQs, or removing some etc. Then this is pretty good indeed. </p>",
      "rawMarkdown": "Okay, but using the training data supplied as is without mergin EQs, or removing some etc. Then this is pretty good indeed.",
      "votes": null
    },
    {
      "id": "544287",
      "postDate": "06/05/2019 11:04:07",
      "content": "<p>No, using the all segments, 12k or so.</p>",
      "rawMarkdown": "No, using the all segments, 12k or so.",
      "votes": null
    },
    {
      "id": "544637",
      "postDate": "06/05/2019 18:13:32",
      "content": "<p>Amazing achievement,  what's your method ? </p>",
      "rawMarkdown": "Amazing achievement,  what's your method ?",
      "votes": null
    },
    {
      "id": "544665",
      "postDate": "06/05/2019 18:40:11",
      "content": "<p>I've described it here: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94556#latest-544294\">15th Place Memo and CV Scheme</a></p>",
      "rawMarkdown": "I've described it here: [15th Place Memo and CV Scheme](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94556#latest-544294)",
      "votes": null
    },
    {
      "id": "544758",
      "postDate": "06/05/2019 21:55:27",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "552940",
      "postDate": "06/14/2019 20:40:59",
      "content": "<p>no test data, 57th, silver</p>",
      "rawMarkdown": "no test data, 57th, silver",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 544233,
      "author_name": "bernir",
      "author_url": "",
      "post_date": "06/05/2019 09:45:19",
      "content": "<p>I can just confirm. I didn't use any test data and finished 15th.</p>",
      "votes": null,
      "replies": [
        {
          "id": 544637,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "06/05/2019 18:13:32",
          "content": "<p>Amazing achievement,  what's your method ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544665,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "06/05/2019 18:40:11",
          "content": "<p>I've described it here: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94556#latest-544294\">15th Place Memo and CV Scheme</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544758,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "06/05/2019 21:55:27",
          "content": "<p>Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544272,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "06/05/2019 10:46:35",
      "content": "<p>I guess using the whole supplied train data without any modifications?</p>",
      "votes": null,
      "replies": [
        {
          "id": 544279,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 10:52:17",
          "content": "<p>Using the same features as for other models with one modification: standard scaler fit on training data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544285,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/05/2019 11:03:15",
          "content": "<p>Okay, but using the training data supplied as is without mergin EQs, or removing some etc. Then this is pretty good indeed. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544287,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 11:04:07",
          "content": "<p>No, using the all segments, 12k or so.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 552940,
      "author_name": "frtgnn",
      "author_url": "",
      "post_date": "06/14/2019 20:40:59",
      "content": "<p>no test data, 57th, silver</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "544211": "Yes, you had to use test estimates to finish in top 10.  But you could get good models that do not use any test data.  I have several of them that would get a gold here:\n\nA knn that scores 2.382 on private\n4 gam (general additive model) that score between 2.34 and 2.39\n\nI think that gradient boosted machines (lgb, xgb, catboost, etc) are less interesting because they cannot extrapolate linear relationship to range no seen in training data.\n\nIn our team we only used test data to derive sample weights, and neither sklearn KNN regressor nor pygam support sample weights.  This is why I am sure the above models do not depend on test data, and would generalize well to other test data than the one we had.\n\nI am pretty sure these would be useful for EQ research community.",
    "544233": "I can just confirm. I didn't use any test data and finished 15th.",
    "544272": "I guess using the whole supplied train data without any modifications?",
    "544279": "Using the same features as for other models with one modification: standard scaler fit on training data.",
    "544285": "Okay, but using the training data supplied as is without mergin EQs, or removing some etc. Then this is pretty good indeed.",
    "544287": "No, using the all segments, 12k or so.",
    "544637": "Amazing achievement,  what's your method ?",
    "544665": "I've described it here: [15th Place Memo and CV Scheme](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94556#latest-544294)",
    "544758": "Thanks",
    "552940": "no test data, 57th, silver"
  },
  "source": "meta"
}