{
  "id": 596578,
  "title": "LinearRegression",
  "url": "/competitions/drw-crypto-market-prediction/writeups/linearregression",
  "author_name": "",
  "post_date": "2025-08-04T12:28:50.153Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As noted by some fellow kagglers, LinearRegression can achieve good ranking on the Leaderboard.</p>\n<p><strong>Solution:</strong></p>\n<ol>\n<li>Remove single value and duplicate features. </li>\n<li>Using 29 features from X752 to X780, train 12 LinearRegression models in sklearn 1.7.x, each by reserving one month of data as validation. Public: 0.10630, Private: 0.10664.</li>\n<li>Feature selection is done by removing one feature at a time from the given features (according to order from the train.parquet file) and fit with multiple LinearRegression, each using data of one month and validation is the data of the next month except for the final month where validation becomes the data of the first month. Validation score (PearsonR) from all these models are averaged, set of features that gave the best score is selected. </li>\n</ol>\n<p><strong>Remarks:</strong></p>\n<ol>\n<li>There is a discrepancy between the prediction output for LinearRegression model of scikit-learn 1.2.2 (default on kaggle notebook) and 1.7.x. Version 1.2.2 performs worse.</li>\n<li>Using only the last 50 features in the given order of train.parquet file and training one LinearRegression model on the whole train dataset, gives a better score on private leaderboard. Public score: 0.09877, Private score: 0.11127.</li>\n</ol>\n<p>Thanks everyone!</p>",
  "messages": [
    {
      "id": "3262892",
      "postDate": "08/04/2025 12:28:33",
      "content": "<p>As noted by some fellow kagglers, LinearRegression can achieve good ranking on the Leaderboard.</p>\n<p><strong>Solution:</strong></p>\n<ol>\n<li>Remove single value and duplicate features. </li>\n<li>Using 29 features from X752 to X780, train 12 LinearRegression models in sklearn 1.7.x, each by reserving one month of data as validation. Public: 0.10630, Private: 0.10664.</li>\n<li>Feature selection is done by removing one feature at a time from the given features (according to order from the train.parquet file) and fit with multiple LinearRegression, each using data of one month and validation is the data of the next month except for the final month where validation becomes the data of the first month. Validation score (PearsonR) from all these models are averaged, set of features that gave the best score is selected. </li>\n</ol>\n<p><strong>Remarks:</strong></p>\n<ol>\n<li>There is a discrepancy between the prediction output for LinearRegression model of scikit-learn 1.2.2 (default on kaggle notebook) and 1.7.x. Version 1.2.2 performs worse.</li>\n<li>Using only the last 50 features in the given order of train.parquet file and training one LinearRegression model on the whole train dataset, gives a better score on private leaderboard. Public score: 0.09877, Private score: 0.11127.</li>\n</ol>\n<p>Thanks everyone!</p>",
      "rawMarkdown": "As noted by some fellow kagglers, LinearRegression can achieve good ranking on the Leaderboard.\n\n**Solution:**\n1. Remove single value and duplicate features. \n2. Using 29 features from X752 to X780, train 12 LinearRegression models in sklearn 1.7.x, each by reserving one month of data as validation. Public: 0.10630, Private: 0.10664.\n3. Feature selection is done by removing one feature at a time from the given features (according to order from the train.parquet file) and fit with multiple LinearRegression, each using data of one month and validation is the data of the next month except for the final month where validation becomes the data of the first month. Validation score (PearsonR) from all these models are averaged, set of features that gave the best score is selected. \n\n**Remarks:**\n1. There is a discrepancy between the prediction output for LinearRegression model of scikit-learn 1.2.2 (default on kaggle notebook) and 1.7.x. Version 1.2.2 performs worse.\n2. Using only the last 50 features in the given order of train.parquet file and training one LinearRegression model on the whole train dataset, gives a better score on private leaderboard. Public score: 0.09877, Private score: 0.11127.\n \nThanks everyone!",
      "votes": null
    },
    {
      "id": "3265413",
      "postDate": "08/07/2025 12:40:16",
      "content": "<p>Can you elaborate on the feature selection method? Till what point did you remove features.</p>",
      "rawMarkdown": "Can you elaborate on the feature selection method? Till what point did you remove features.",
      "votes": null
    },
    {
      "id": "3268118",
      "postDate": "08/12/2025 12:38:06",
      "content": "<ol>\n<li>Remove feature one by one, example after n loops, 29 feature left</li>\n</ol>\n<pre><code>[\n     ', ', ', ', ', \n     ', ', ', ', ', \n     ', ', ', ', ', \n     ', ', ', ', ', \n     ', ', ', ', ', \n     ', ', ', '\n ]\n</code></pre>\n<ol>\n<li>For the next two loops, remove feature 'X752' (28 features left), remove 'X753' (27 features left). Continue until only one feature left, i.e. 'X780'.</li>\n<li>For each set of features left after removal, fit LinearRegression for each month, use the data of next month as validation, for the final month, use the data of the first month. Take the average validation score for each month as the score for that feature set.</li>\n<li>Select the set of features with the best score.</li>\n</ol>",
      "rawMarkdown": "1. Remove feature one by one, example after n loops, 29 feature left\n```\n[\n     'X752', 'X753', 'X754', 'X755', 'X756', \n     'X757', 'X758', 'X759', 'X760', 'X761', \n     'X762', 'X763', 'X764', 'X765', 'X766', \n     'X767', 'X768', 'X769', 'X770', 'X771', \n     'X772', 'X773', 'X774', 'X775', 'X776', \n     'X777', 'X778', 'X779', 'X780'\n ]\n```\n2. For the next two loops, remove feature 'X752' (28 features left), remove 'X753' (27 features left). Continue until only one feature left, i.e. 'X780'.\n3. For each set of features left after removal, fit LinearRegression for each month, use the data of next month as validation, for the final month, use the data of the first month. Take the average validation score for each month as the score for that feature set.\n4. Select the set of features with the best score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3265413,
      "author_name": "aravindproml",
      "author_url": "",
      "post_date": "08/07/2025 12:40:16",
      "content": "<p>Can you elaborate on the feature selection method? Till what point did you remove features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3268118,
          "author_name": "cheowthianliang",
          "author_url": "",
          "post_date": "08/12/2025 12:38:06",
          "content": "<ol>\n<li>Remove feature one by one, example after n loops, 29 feature left</li>\n</ol>\n<pre><code>[\n     ', ', ', ', ', \n     ', ', ', ', ', \n     ', ', ', ', ', \n     ', ', ', ', ', \n     ', ', ', ', ', \n     ', ', ', '\n ]\n</code></pre>\n<ol>\n<li>For the next two loops, remove feature 'X752' (28 features left), remove 'X753' (27 features left). Continue until only one feature left, i.e. 'X780'.</li>\n<li>For each set of features left after removal, fit LinearRegression for each month, use the data of next month as validation, for the final month, use the data of the first month. Take the average validation score for each month as the score for that feature set.</li>\n<li>Select the set of features with the best score.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3262892": "As noted by some fellow kagglers, LinearRegression can achieve good ranking on the Leaderboard.\n\n**Solution:**\n1. Remove single value and duplicate features. \n2. Using 29 features from X752 to X780, train 12 LinearRegression models in sklearn 1.7.x, each by reserving one month of data as validation. Public: 0.10630, Private: 0.10664.\n3. Feature selection is done by removing one feature at a time from the given features (according to order from the train.parquet file) and fit with multiple LinearRegression, each using data of one month and validation is the data of the next month except for the final month where validation becomes the data of the first month. Validation score (PearsonR) from all these models are averaged, set of features that gave the best score is selected. \n\n**Remarks:**\n1. There is a discrepancy between the prediction output for LinearRegression model of scikit-learn 1.2.2 (default on kaggle notebook) and 1.7.x. Version 1.2.2 performs worse.\n2. Using only the last 50 features in the given order of train.parquet file and training one LinearRegression model on the whole train dataset, gives a better score on private leaderboard. Public score: 0.09877, Private score: 0.11127.\n \nThanks everyone!",
    "3265413": "Can you elaborate on the feature selection method? Till what point did you remove features.",
    "3268118": "1. Remove feature one by one, example after n loops, 29 feature left\n```\n[\n     'X752', 'X753', 'X754', 'X755', 'X756', \n     'X757', 'X758', 'X759', 'X760', 'X761', \n     'X762', 'X763', 'X764', 'X765', 'X766', \n     'X767', 'X768', 'X769', 'X770', 'X771', \n     'X772', 'X773', 'X774', 'X775', 'X776', \n     'X777', 'X778', 'X779', 'X780'\n ]\n```\n2. For the next two loops, remove feature 'X752' (28 features left), remove 'X753' (27 features left). Continue until only one feature left, i.e. 'X780'.\n3. For each set of features left after removal, fit LinearRegression for each month, use the data of next month as validation, for the final month, use the data of the first month. Take the average validation score for each month as the score for that feature set.\n4. Select the set of features with the best score."
  },
  "source": "meta"
}