{
  "id": 497512,
  "title": "Almost all the public notebooks are just copies?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/497512",
  "author_name": "",
  "post_date": "2024-04-24T22:22:00.011990600Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Looking at many public notebooks from this competition I didn't met a single one that process the data in a different way than in the host's example. Everywhere the same copy/paste blocks of codes or entire processing cells with classes based on Polars  with just small differences between them. Or it's because it's still rare who knows Polars and for big data like this nobody bothers to use Pandas so people are just using what's working by copying the code?<br>\nJust from curiosity I was able to preprocess all the files in Pandas without exceeding the memory and got ~900 features. After this I re-wrote this in Polars just from curiosity to learn some Polars. It seams really strange for me and I don't think I'm the only crazy that wrote from scratch all the processing in an absolutely other way than the 'public template'. Maybe you saw other interesting notebooks with alternative ways and not just a variation of the 'template'? </p>",
  "messages": [
    {
      "id": "2773782",
      "postDate": "04/24/2024 22:22:00",
      "content": "<p>Looking at many public notebooks from this competition I didn't met a single one that process the data in a different way than in the host's example. Everywhere the same copy/paste blocks of codes or entire processing cells with classes based on Polars  with just small differences between them. Or it's because it's still rare who knows Polars and for big data like this nobody bothers to use Pandas so people are just using what's working by copying the code?<br>\nJust from curiosity I was able to preprocess all the files in Pandas without exceeding the memory and got ~900 features. After this I re-wrote this in Polars just from curiosity to learn some Polars. It seams really strange for me and I don't think I'm the only crazy that wrote from scratch all the processing in an absolutely other way than the 'public template'. Maybe you saw other interesting notebooks with alternative ways and not just a variation of the 'template'? </p>",
      "rawMarkdown": "Looking at many public notebooks from this competition I didn't met a single one that process the data in a different way than in the host's example. Everywhere the same copy/paste blocks of codes or entire processing cells with classes based on Polars  with just small differences between them. Or it's because it's still rare who knows Polars and for big data like this nobody bothers to use Pandas so people are just using what's working by copying the code?\nJust from curiosity I was able to preprocess all the files in Pandas without exceeding the memory and got ~900 features. After this I re-wrote this in Polars just from curiosity to learn some Polars. It seams really strange for me and I don't think I'm the only crazy that wrote from scratch all the processing in an absolutely other way than the 'public template'. Maybe you saw other interesting notebooks with alternative ways and not just a variation of the 'template'?",
      "votes": null
    },
    {
      "id": "2773815",
      "postDate": "04/24/2024 23:24:57",
      "content": "<p>I did it myself, although I followed the idea of open source code and after integrating open source catboost, I was able to achieve 0.582.</p>\n<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-inference\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-inference</a></p>",
      "rawMarkdown": "I did it myself, although I followed the idea of open source code and after integrating open source catboost, I was able to achieve 0.582.\n\nhttps://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-inference",
      "votes": null
    },
    {
      "id": "2773817",
      "postDate": "04/24/2024 23:31:09",
      "content": "<p>Try to read from parquet - should be quicker and it reads the columns' dtypes as parquet holds this info</p>",
      "rawMarkdown": "Try to read from parquet - should be quicker and it reads the columns' dtypes as parquet holds this info",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2773815,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "04/24/2024 23:24:57",
      "content": "<p>I did it myself, although I followed the idea of open source code and after integrating open source catboost, I was able to achieve 0.582.</p>\n<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-inference\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-inference</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2773817,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "04/24/2024 23:31:09",
          "content": "<p>Try to read from parquet - should be quicker and it reads the columns' dtypes as parquet holds this info</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2773782": "Looking at many public notebooks from this competition I didn't met a single one that process the data in a different way than in the host's example. Everywhere the same copy/paste blocks of codes or entire processing cells with classes based on Polars  with just small differences between them. Or it's because it's still rare who knows Polars and for big data like this nobody bothers to use Pandas so people are just using what's working by copying the code?\nJust from curiosity I was able to preprocess all the files in Pandas without exceeding the memory and got ~900 features. After this I re-wrote this in Polars just from curiosity to learn some Polars. It seams really strange for me and I don't think I'm the only crazy that wrote from scratch all the processing in an absolutely other way than the 'public template'. Maybe you saw other interesting notebooks with alternative ways and not just a variation of the 'template'?",
    "2773815": "I did it myself, although I followed the idea of open source code and after integrating open source catboost, I was able to achieve 0.582.\n\nhttps://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-inference",
    "2773817": "Try to read from parquet - should be quicker and it reads the columns' dtypes as parquet holds this info"
  },
  "source": "meta"
}