{
  "id": 377697,
  "title": "Question for the top 100 of the competition",
  "url": "/competitions/otto-recommender-system/discussion/377697",
  "author_name": "",
  "post_date": "2023-01-12T13:10:35.627615700Z",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello community, I was wondering how many features the top 1% of the LB are using ? This question is quite important regarding the memory consumption and all the chunk processing we've to deal with.</p>\n<p>Did someone try dimensionality reduction of the features ? </p>\n<p>I personnally did  on some features using incremental PCA to process it with mini-batch, ( all hypothesis were respected ), but the recall decreased   even with 95% of explained inertia of the PCA.</p>\n<p>Once again, I investigated and respected all the hypothesis of the PCA before applied it :)</p>",
  "messages": [
    {
      "id": "2097088",
      "postDate": "01/12/2023 13:10:35",
      "content": "<p>Hello community, I was wondering how many features the top 1% of the LB are using ? This question is quite important regarding the memory consumption and all the chunk processing we've to deal with.</p>\n<p>Did someone try dimensionality reduction of the features ? </p>\n<p>I personnally did  on some features using incremental PCA to process it with mini-batch, ( all hypothesis were respected ), but the recall decreased   even with 95% of explained inertia of the PCA.</p>\n<p>Once again, I investigated and respected all the hypothesis of the PCA before applied it :)</p>",
      "rawMarkdown": "Hello community, I was wondering how many features the top 1% of the LB are using ? This question is quite important regarding the memory consumption and all the chunk processing we've to deal with.\n\nDid someone try dimensionality reduction of the features ? \n\n\nI personnally did  on some features using incremental PCA to process it with mini-batch, ( all hypothesis were respected ), but the recall decreased   even with 95% of explained inertia of the PCA.\n\nOnce again, I investigated and respected all the hypothesis of the PCA before applied it :)",
      "votes": null
    },
    {
      "id": "2097132",
      "postDate": "01/12/2023 13:50:59",
      "content": "<p>I'm only using 10 features (no PCA), feeding into XGB - I know others have said they are using several hundred?!</p>",
      "rawMarkdown": "I'm only using 10 features (no PCA), feeding into XGB - I know others have said they are using several hundred?!",
      "votes": null
    },
    {
      "id": "2097208",
      "postDate": "01/12/2023 14:16:55",
      "content": "<p><a href=\"https://www.kaggle.com/johnwakefield\" target=\"_blank\">@johnwakefield</a> 10 features is quite low for such amazing LB, must be really good features. Are we talking user-item interaction features?</p>",
      "rawMarkdown": "johnwakefield 10 features is quite low for such amazing LB, must be really good features. Are we talking user-item interaction features?",
      "votes": null
    },
    {
      "id": "2097289",
      "postDate": "01/12/2023 15:02:41",
      "content": "<p>98 features for now, will increase for sure. And only use 0.8 week data for training(0.2 for validation), so the memory consumption is quite low, around 5GB for orders. But of course feature generation consumes more RAM.</p>",
      "rawMarkdown": "98 features for now, will increase for sure. And only use 0.8 week data for training(0.2 for validation), so the memory consumption is quite low, around 5GB for orders. But of course feature generation consumes more RAM.",
      "votes": null
    },
    {
      "id": "2097320",
      "postDate": "01/12/2023 15:23:09",
      "content": "<p>Nice one! If you don't mind: what were your local validation scores? </p>",
      "rawMarkdown": "Nice one! If you don't mind: what were your local validation scores?",
      "votes": null
    },
    {
      "id": "2097328",
      "postDate": "01/12/2023 15:34:13",
      "content": "<p>0.581, but mine might be lower than it should be, because I am using the old version host script to generate validation set. So there is unseen items in valid set. Since it still correlates quite well with LB, I just do not change it as I am tired of doing it all over again.</p>",
      "rawMarkdown": "0.581, but mine might be lower than it should be, because I am using the old version host script to generate validation set. So there is unseen items in valid set. Since it still correlates quite well with LB, I just do not change it as I am tired of doing it all over again.",
      "votes": null
    },
    {
      "id": "2097368",
      "postDate": "01/12/2023 16:08:23",
      "content": "<p>Thanks a ton for sharing! </p>",
      "rawMarkdown": "Thanks a ton for sharing!",
      "votes": null
    },
    {
      "id": "2097508",
      "postDate": "01/12/2023 18:15:56",
      "content": "<p>5 user (session) features and 5 item features - I would like more, but anything I add just upsets my score 🙃</p>",
      "rawMarkdown": "5 user (session) features and 5 item features - I would like more, but anything I add just upsets my score 🙃",
      "votes": null
    },
    {
      "id": "2098392",
      "postDate": "01/13/2023 13:56:27",
      "content": "<p>Great !  <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a>  if you don't mind : </p>\n<ul>\n<li>when you talk about the data used for training you mean that you train your model on the  last week of train data ? ( of course the features you create will use all the data ).</li>\n<li>Did you extract candidates for the validation and the test (inference ) from the same co-visitatation matrix ?</li>\n</ul>\n<p>Thanks for your clarification !</p>",
      "rawMarkdown": "Great !  @buumoo  if you don't mind : \n- when you talk about the data used for training you mean that you train your model on the  last week of train data ? ( of course the features you create will use all the data ).\n- Did you extract candidates for the validation and the test (inference ) from the same co-visitatation matrix ?\n\nThanks for your clarification !",
      "votes": null
    },
    {
      "id": "2098393",
      "postDate": "01/13/2023 14:00:24",
      "content": "<p><a href=\"https://www.kaggle.com/johnwakefield\" target=\"_blank\">@johnwakefield</a>  Amazing ! did you use some kind of regularizations for your model ? I'm personally overfitting on my train data, despite trying some regularizations. I also noticed that the number of boosting iterations didn't help increasing the score.</p>\n<p>Plus, from this discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/radek\" target=\"_blank\">@radek</a> explained that using a small number of estimators helped. </p>",
      "rawMarkdown": "johnwakefield  Amazing ! did you use some kind of regularizations for your model ? I'm personally overfitting on my train data, despite trying some regularizations. I also noticed that the number of boosting iterations didn't help increasing the score.\n\nPlus, from this discussion [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278), @radek explained that using a small number of estimators helped.",
      "votes": null
    },
    {
      "id": "2098425",
      "postDate": "01/13/2023 14:50:12",
      "content": "<p>I try to build the most constrained and simple models I can - this is a philosophy I use everyday at work for our commercial ml models - that naturally combats overfitting.</p>",
      "rawMarkdown": "I try to build the most constrained and simple models I can - this is a philosophy I use everyday at work for our commercial ml models - that naturally combats overfitting.",
      "votes": null
    },
    {
      "id": "2098512",
      "postDate": "01/13/2023 16:18:37",
      "content": "<ol>\n<li>for ranker model, yes.</li>\n<li>No. Losing the information contained in test set is a big loss, since the host confirmed it is okay to use the test data for training.</li>\n</ol>",
      "rawMarkdown": "1. for ranker model, yes.\n2. No. Losing the information contained in test set is a big loss, since the host confirmed it is okay to use the test data for training.",
      "votes": null
    },
    {
      "id": "2110149",
      "postDate": "01/22/2023 02:26:52",
      "content": "<p>Using around 230 features, haven't try feature selection for now, why are you trying to apply PCA on features, if you want to lower the dimension,  feature selections should be better</p>",
      "rawMarkdown": "Using around 230 features, haven't try feature selection for now, why are you trying to apply PCA on features, if you want to lower the dimension,  feature selections should be better",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2097132,
      "author_name": "johnwakefield",
      "author_url": "",
      "post_date": "01/12/2023 13:50:59",
      "content": "<p>I'm only using 10 features (no PCA), feeding into XGB - I know others have said they are using several hundred?!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2097208,
          "author_name": "parthpankajtiwary",
          "author_url": "",
          "post_date": "01/12/2023 14:16:55",
          "content": "<p><a href=\"https://www.kaggle.com/johnwakefield\" target=\"_blank\">@johnwakefield</a> 10 features is quite low for such amazing LB, must be really good features. Are we talking user-item interaction features?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2097508,
              "author_name": "johnwakefield",
              "author_url": "",
              "post_date": "01/12/2023 18:15:56",
              "content": "<p>5 user (session) features and 5 item features - I would like more, but anything I add just upsets my score 🙃</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2098393,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/13/2023 14:00:24",
          "content": "<p><a href=\"https://www.kaggle.com/johnwakefield\" target=\"_blank\">@johnwakefield</a>  Amazing ! did you use some kind of regularizations for your model ? I'm personally overfitting on my train data, despite trying some regularizations. I also noticed that the number of boosting iterations didn't help increasing the score.</p>\n<p>Plus, from this discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/radek\" target=\"_blank\">@radek</a> explained that using a small number of estimators helped. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2098425,
              "author_name": "johnwakefield",
              "author_url": "",
              "post_date": "01/13/2023 14:50:12",
              "content": "<p>I try to build the most constrained and simple models I can - this is a philosophy I use everyday at work for our commercial ml models - that naturally combats overfitting.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2097289,
      "author_name": "buumoo",
      "author_url": "",
      "post_date": "01/12/2023 15:02:41",
      "content": "<p>98 features for now, will increase for sure. And only use 0.8 week data for training(0.2 for validation), so the memory consumption is quite low, around 5GB for orders. But of course feature generation consumes more RAM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2097320,
          "author_name": "parthpankajtiwary",
          "author_url": "",
          "post_date": "01/12/2023 15:23:09",
          "content": "<p>Nice one! If you don't mind: what were your local validation scores? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2097328,
              "author_name": "buumoo",
              "author_url": "",
              "post_date": "01/12/2023 15:34:13",
              "content": "<p>0.581, but mine might be lower than it should be, because I am using the old version host script to generate validation set. So there is unseen items in valid set. Since it still correlates quite well with LB, I just do not change it as I am tired of doing it all over again.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2097368,
                  "author_name": "parthpankajtiwary",
                  "author_url": "",
                  "post_date": "01/12/2023 16:08:23",
                  "content": "<p>Thanks a ton for sharing! </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2098392,
                      "author_name": "rayanaay",
                      "author_url": "",
                      "post_date": "01/13/2023 13:56:27",
                      "content": "<p>Great !  <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a>  if you don't mind : </p>\n<ul>\n<li>when you talk about the data used for training you mean that you train your model on the  last week of train data ? ( of course the features you create will use all the data ).</li>\n<li>Did you extract candidates for the validation and the test (inference ) from the same co-visitatation matrix ?</li>\n</ul>\n<p>Thanks for your clarification !</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2098512,
                          "author_name": "buumoo",
                          "author_url": "",
                          "post_date": "01/13/2023 16:18:37",
                          "content": "<ol>\n<li>for ranker model, yes.</li>\n<li>No. Losing the information contained in test set is a big loss, since the host confirmed it is okay to use the test data for training.</li>\n</ol>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2110149,
      "author_name": "kellenxu",
      "author_url": "",
      "post_date": "01/22/2023 02:26:52",
      "content": "<p>Using around 230 features, haven't try feature selection for now, why are you trying to apply PCA on features, if you want to lower the dimension,  feature selections should be better</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2097088": "Hello community, I was wondering how many features the top 1% of the LB are using ? This question is quite important regarding the memory consumption and all the chunk processing we've to deal with.\n\nDid someone try dimensionality reduction of the features ? \n\n\nI personnally did  on some features using incremental PCA to process it with mini-batch, ( all hypothesis were respected ), but the recall decreased   even with 95% of explained inertia of the PCA.\n\nOnce again, I investigated and respected all the hypothesis of the PCA before applied it :)",
    "2097132": "I'm only using 10 features (no PCA), feeding into XGB - I know others have said they are using several hundred?!",
    "2097208": "johnwakefield 10 features is quite low for such amazing LB, must be really good features. Are we talking user-item interaction features?",
    "2097289": "98 features for now, will increase for sure. And only use 0.8 week data for training(0.2 for validation), so the memory consumption is quite low, around 5GB for orders. But of course feature generation consumes more RAM.",
    "2097320": "Nice one! If you don't mind: what were your local validation scores?",
    "2097328": "0.581, but mine might be lower than it should be, because I am using the old version host script to generate validation set. So there is unseen items in valid set. Since it still correlates quite well with LB, I just do not change it as I am tired of doing it all over again.",
    "2097368": "Thanks a ton for sharing!",
    "2097508": "5 user (session) features and 5 item features - I would like more, but anything I add just upsets my score 🙃",
    "2098392": "Great !  @buumoo  if you don't mind : \n- when you talk about the data used for training you mean that you train your model on the  last week of train data ? ( of course the features you create will use all the data ).\n- Did you extract candidates for the validation and the test (inference ) from the same co-visitatation matrix ?\n\nThanks for your clarification !",
    "2098393": "johnwakefield  Amazing ! did you use some kind of regularizations for your model ? I'm personally overfitting on my train data, despite trying some regularizations. I also noticed that the number of boosting iterations didn't help increasing the score.\n\nPlus, from this discussion [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278), @radek explained that using a small number of estimators helped.",
    "2098425": "I try to build the most constrained and simple models I can - this is a philosophy I use everyday at work for our commercial ml models - that naturally combats overfitting.",
    "2098512": "1. for ranker model, yes.\n2. No. Losing the information contained in test set is a big loss, since the host confirmed it is okay to use the test data for training.",
    "2110149": "Using around 230 features, haven't try feature selection for now, why are you trying to apply PCA on features, if you want to lower the dimension,  feature selections should be better"
  },
  "source": "meta"
}