{
  "id": 420281,
  "title": "Top 0.5% Efficiency Leaderboard with datatable",
  "url": "/competitions/predict-student-performance-from-game-play/writeups/oleksiy-kononenko-top-0-5-efficiency-leaderboard-w",
  "author_name": "",
  "post_date": "2023-07-26T11:25:54.837Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Before jumping into details, I would like to say thank you to the organizers, Kaggle team and community.</p>\n<p>This was actually my first competition, though I had some limited Kaggle experience. Several years&nbsp;ago I benchmarked my newly developed models here, but at that time I didn't really look into the data and compete.&nbsp;</p>\n<p>Surprisingly, with the first attempt my solution was ranked 7th out of 2051 on <a href=\"https://www.kaggle.com/code/philculliton/student-performance-efficiency-leaderboard/notebook\" target=\"_blank\">the Efficiency Leaderboard</a>.</p>\n<h2>Background</h2>\n<p>When I joined about a month ago I was impressed by the fact, that simply <a href=\"https://www.kaggle.com/code/cpmpml/random-submission/\" target=\"_blank\">submitting mean values</a> could bring your score to&nbsp;<code>0.659</code> LB, while the most advanced models were at the level of ~<code>0.7</code>.&nbsp;</p>\n<p>So I decided to compete on the efficiency LB only and at the same&nbsp;time give a try to <a href=\"https://datatable.readthedocs.io/en/latest/api/models/linear_model.html\" target=\"_blank\">LinearModel</a>&nbsp;GLM I've recently developed. </p>\n<p>For data munging and feature engineering I've been using Python <a href=\"https://datatable.readthedocs.io/en/latest/index.html\" target=\"_blank\">datatable</a>,&nbsp;a package similar to pandas, but with a specific emphasis on speed and big data support.</p>\n<p>First, I have designed and shared <a href=\"https://www.kaggle.com/code/kononenko/datatable-linearmodel-0-676lb-in-6-seconds\" target=\"_blank\">a simple baseline</a>, that scored <code>0.676</code> and was pretty high on the efficiency LB.&nbsp;My next goal was to improve feature engineering and the overall code performance.</p>\n<h2>Feature engineering</h2>\n<p>Moving&nbsp;forward I've ended up with the following numeric features</p>\n<ul>\n<li>number of events per session, i.e. <code>sessions_id.count()</code>;</li>\n<li>session duration, i.e. <code>elapsed_time.max()</code>;</li>\n<li>mean level, i.e. <code>level.mean()</code>;</li>\n<li>screen x/y range, i.e. <code>screen_coor_x/y.max() - screen_coor_x/y.min()</code>.</li>\n</ul>\n<p>For categorical columns, I started with a number of unique values per a&nbsp;column. Then, created additional&nbsp;features for each of the values. Even though it worked locally on my CV, it didn't work on the public LB, so I had to employ feature selection based on the importances and also picked different features for different level groups.</p>\n<h2>Performance tuning</h2>\n<p>I have also&nbsp;employed some overall code tuning</p>\n<ul>\n<li>since <code>LinearModel</code> is fully&nbsp;parallel, both <code>.fit()</code> and <code>.predict()</code> methods, I have adjusted the number of threads to <code>2</code> to match the number of CPUs. By default, 4 threads were detected that could lead to over-parallelization;</li>\n<li>avoided <code>!pip install</code> in the inference code by pre-installing packages in a separate&nbsp;notebook. This saved me at least 30 seconds per submission;</li>\n<li>disabled \"Persistence\", so that no additional time is spent when the inference notebook is starting.</li>\n</ul>\n<h2>Final model</h2>\n<p><strong>Pros</strong></p>\n<ul>\n<li>a small number of features;</li>\n<li>less than one minute to be trained;</li>\n<li>highly interpretable;</li>\n<li>pretty robust: no errors due to the API changes, no submission errors, public LB scores are exactly the same as the private scores.</li>\n</ul>\n<p><strong>Cons</strong></p>\n<ul>\n<li>since at the end it is just a logistic regression, I don't think there is a huge room to further improve the model's score.</li>\n</ul>\n<h2>Conclusions</h2>\n<p>The best submission, I have selected, scores <code>0.687</code> LB with the scoring time of around 2-3 minutes. Notebook, that includes the training and the inference parts, is available <a href=\"https://www.kaggle.com/code/kononenko/top-1-public-efficiency-lb-with-linearmodel\" target=\"_blank\">here</a>.&nbsp;</p>",
  "messages": [
    {
      "id": "2323644",
      "postDate": "06/30/2023 05:15:34",
      "content": "<p>Before jumping into details, I would like to say thank you to the organizers, Kaggle team and community.</p>\n<p>This was actually my first competition, though I had some limited Kaggle experience. Several years&nbsp;ago I benchmarked my newly developed models here, but at that time I didn't really look into the data and compete.&nbsp;</p>\n<p>Surprisingly, with the first attempt my solution was ranked 7th out of 2051 on <a href=\"https://www.kaggle.com/code/philculliton/student-performance-efficiency-leaderboard/notebook\" target=\"_blank\">the Efficiency Leaderboard</a>.</p>\n<h2>Background</h2>\n<p>When I joined about a month ago I was impressed by the fact, that simply <a href=\"https://www.kaggle.com/code/cpmpml/random-submission/\" target=\"_blank\">submitting mean values</a> could bring your score to&nbsp;<code>0.659</code> LB, while the most advanced models were at the level of ~<code>0.7</code>.&nbsp;</p>\n<p>So I decided to compete on the efficiency LB only and at the same&nbsp;time give a try to <a href=\"https://datatable.readthedocs.io/en/latest/api/models/linear_model.html\" target=\"_blank\">LinearModel</a>&nbsp;GLM I've recently developed. </p>\n<p>For data munging and feature engineering I've been using Python <a href=\"https://datatable.readthedocs.io/en/latest/index.html\" target=\"_blank\">datatable</a>,&nbsp;a package similar to pandas, but with a specific emphasis on speed and big data support.</p>\n<p>First, I have designed and shared <a href=\"https://www.kaggle.com/code/kononenko/datatable-linearmodel-0-676lb-in-6-seconds\" target=\"_blank\">a simple baseline</a>, that scored <code>0.676</code> and was pretty high on the efficiency LB.&nbsp;My next goal was to improve feature engineering and the overall code performance.</p>\n<h2>Feature engineering</h2>\n<p>Moving&nbsp;forward I've ended up with the following numeric features</p>\n<ul>\n<li>number of events per session, i.e. <code>sessions_id.count()</code>;</li>\n<li>session duration, i.e. <code>elapsed_time.max()</code>;</li>\n<li>mean level, i.e. <code>level.mean()</code>;</li>\n<li>screen x/y range, i.e. <code>screen_coor_x/y.max() - screen_coor_x/y.min()</code>.</li>\n</ul>\n<p>For categorical columns, I started with a number of unique values per a&nbsp;column. Then, created additional&nbsp;features for each of the values. Even though it worked locally on my CV, it didn't work on the public LB, so I had to employ feature selection based on the importances and also picked different features for different level groups.</p>\n<h2>Performance tuning</h2>\n<p>I have also&nbsp;employed some overall code tuning</p>\n<ul>\n<li>since <code>LinearModel</code> is fully&nbsp;parallel, both <code>.fit()</code> and <code>.predict()</code> methods, I have adjusted the number of threads to <code>2</code> to match the number of CPUs. By default, 4 threads were detected that could lead to over-parallelization;</li>\n<li>avoided <code>!pip install</code> in the inference code by pre-installing packages in a separate&nbsp;notebook. This saved me at least 30 seconds per submission;</li>\n<li>disabled \"Persistence\", so that no additional time is spent when the inference notebook is starting.</li>\n</ul>\n<h2>Final model</h2>\n<p><strong>Pros</strong></p>\n<ul>\n<li>a small number of features;</li>\n<li>less than one minute to be trained;</li>\n<li>highly interpretable;</li>\n<li>pretty robust: no errors due to the API changes, no submission errors, public LB scores are exactly the same as the private scores.</li>\n</ul>\n<p><strong>Cons</strong></p>\n<ul>\n<li>since at the end it is just a logistic regression, I don't think there is a huge room to further improve the model's score.</li>\n</ul>\n<h2>Conclusions</h2>\n<p>The best submission, I have selected, scores <code>0.687</code> LB with the scoring time of around 2-3 minutes. Notebook, that includes the training and the inference parts, is available <a href=\"https://www.kaggle.com/code/kononenko/top-1-public-efficiency-lb-with-linearmodel\" target=\"_blank\">here</a>.&nbsp;</p>",
      "rawMarkdown": "Before jumping into details, I would like to say thank you to the organizers, Kaggle team and community.\n\nThis was actually my first competition, though I had some limited Kaggle experience. Several years ago I benchmarked my newly developed models here, but at that time I didn't really look into the data and compete. \n\nSurprisingly, with the first attempt my solution was ranked 7th out of 2051 on [the Efficiency Leaderboard](https://www.kaggle.com/code/philculliton/student-performance-efficiency-leaderboard/notebook).\n\n## Background\nWhen I joined about a month ago I was impressed by the fact, that simply [submitting mean values](https://www.kaggle.com/code/cpmpml/random-submission/) could bring your score to `0.659` LB, while the most advanced models were at the level of ~`0.7`. \n\nSo I decided to compete on the efficiency LB only and at the same time give a try to [LinearModel](https://datatable.readthedocs.io/en/latest/api/models/linear_model.html) GLM I've recently developed. \n\nFor data munging and feature engineering I've been using Python [datatable](https://datatable.readthedocs.io/en/latest/index.html), a package similar to pandas, but with a specific emphasis on speed and big data support.\n\nFirst, I have designed and shared [a simple baseline](https://www.kaggle.com/code/kononenko/datatable-linearmodel-0-676lb-in-6-seconds), that scored `0.676` and was pretty high on the efficiency LB. My next goal was to improve feature engineering and the overall code performance.\n\n## Feature engineering\n\nMoving forward I've ended up with the following numeric features\n- number of events per session, i.e. `sessions_id.count()`;\n- session duration, i.e. `elapsed_time.max()`;\n- mean level, i.e. `level.mean()`;\n- screen x/y range, i.e. `screen_coor_x/y.max() - screen_coor_x/y.min()`.\n\nFor categorical columns, I started with a number of unique values per a column. Then, created additional features for each of the values. Even though it worked locally on my CV, it didn't work on the public LB, so I had to employ feature selection based on the importances and also picked different features for different level groups.\n\n## Performance tuning\n\nI have also employed some overall code tuning\n- since `LinearModel` is fully parallel, both `.fit()` and `.predict()` methods, I have adjusted the number of threads to `2` to match the number of CPUs. By default, 4 threads were detected that could lead to over-parallelization;\n- avoided `!pip install` in the inference code by pre-installing packages in a separate notebook. This saved me at least 30 seconds per submission;\n- disabled \"Persistence\", so that no additional time is spent when the inference notebook is starting.\n\n## Final model\n\n**Pros**\n- a small number of features;\n- less than one minute to be trained;\n- highly interpretable;\n- pretty robust: no errors due to the API changes, no submission errors, public LB scores are exactly the same as the private scores.\n\n**Cons**\n- since at the end it is just a logistic regression, I don't think there is a huge room to further improve the model's score.\n\n## Conclusions\n\nThe best submission, I have selected, scores `0.687` LB with the scoring time of around 2-3 minutes. Notebook, that includes the training and the inference parts, is available [here](https://www.kaggle.com/code/kononenko/top-1-public-efficiency-lb-with-linearmodel).",
      "votes": null
    },
    {
      "id": "2323671",
      "postDate": "06/30/2023 05:31:22",
      "content": "<p>Thank you for sharing! 😃 </p>",
      "rawMarkdown": "Thank you for sharing! 😃",
      "votes": null
    },
    {
      "id": "2323696",
      "postDate": "06/30/2023 06:01:12",
      "content": "<p>An elegant solution and a summary that is easy to read, thank you for sharing!</p>",
      "rawMarkdown": "An elegant solution and a summary that is easy to read, thank you for sharing!",
      "votes": null
    },
    {
      "id": "2323842",
      "postDate": "06/30/2023 07:35:14",
      "content": "<p>Thanks for sharing this beautifully written pointwise solution and approach <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> </p>",
      "rawMarkdown": "Thanks for sharing this beautifully written pointwise solution and approach @kononenko",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2323671,
      "author_name": "belalemadhussein",
      "author_url": "",
      "post_date": "06/30/2023 05:31:22",
      "content": "<p>Thank you for sharing! 😃 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323696,
      "author_name": "ikonon",
      "author_url": "",
      "post_date": "06/30/2023 06:01:12",
      "content": "<p>An elegant solution and a summary that is easy to read, thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323842,
      "author_name": "swapnilchowdhury",
      "author_url": "",
      "post_date": "06/30/2023 07:35:14",
      "content": "<p>Thanks for sharing this beautifully written pointwise solution and approach <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2323644": "Before jumping into details, I would like to say thank you to the organizers, Kaggle team and community.\n\nThis was actually my first competition, though I had some limited Kaggle experience. Several years ago I benchmarked my newly developed models here, but at that time I didn't really look into the data and compete. \n\nSurprisingly, with the first attempt my solution was ranked 7th out of 2051 on [the Efficiency Leaderboard](https://www.kaggle.com/code/philculliton/student-performance-efficiency-leaderboard/notebook).\n\n## Background\nWhen I joined about a month ago I was impressed by the fact, that simply [submitting mean values](https://www.kaggle.com/code/cpmpml/random-submission/) could bring your score to `0.659` LB, while the most advanced models were at the level of ~`0.7`. \n\nSo I decided to compete on the efficiency LB only and at the same time give a try to [LinearModel](https://datatable.readthedocs.io/en/latest/api/models/linear_model.html) GLM I've recently developed. \n\nFor data munging and feature engineering I've been using Python [datatable](https://datatable.readthedocs.io/en/latest/index.html), a package similar to pandas, but with a specific emphasis on speed and big data support.\n\nFirst, I have designed and shared [a simple baseline](https://www.kaggle.com/code/kononenko/datatable-linearmodel-0-676lb-in-6-seconds), that scored `0.676` and was pretty high on the efficiency LB. My next goal was to improve feature engineering and the overall code performance.\n\n## Feature engineering\n\nMoving forward I've ended up with the following numeric features\n- number of events per session, i.e. `sessions_id.count()`;\n- session duration, i.e. `elapsed_time.max()`;\n- mean level, i.e. `level.mean()`;\n- screen x/y range, i.e. `screen_coor_x/y.max() - screen_coor_x/y.min()`.\n\nFor categorical columns, I started with a number of unique values per a column. Then, created additional features for each of the values. Even though it worked locally on my CV, it didn't work on the public LB, so I had to employ feature selection based on the importances and also picked different features for different level groups.\n\n## Performance tuning\n\nI have also employed some overall code tuning\n- since `LinearModel` is fully parallel, both `.fit()` and `.predict()` methods, I have adjusted the number of threads to `2` to match the number of CPUs. By default, 4 threads were detected that could lead to over-parallelization;\n- avoided `!pip install` in the inference code by pre-installing packages in a separate notebook. This saved me at least 30 seconds per submission;\n- disabled \"Persistence\", so that no additional time is spent when the inference notebook is starting.\n\n## Final model\n\n**Pros**\n- a small number of features;\n- less than one minute to be trained;\n- highly interpretable;\n- pretty robust: no errors due to the API changes, no submission errors, public LB scores are exactly the same as the private scores.\n\n**Cons**\n- since at the end it is just a logistic regression, I don't think there is a huge room to further improve the model's score.\n\n## Conclusions\n\nThe best submission, I have selected, scores `0.687` LB with the scoring time of around 2-3 minutes. Notebook, that includes the training and the inference parts, is available [here](https://www.kaggle.com/code/kononenko/top-1-public-efficiency-lb-with-linearmodel).",
    "2323671": "Thank you for sharing! 😃",
    "2323696": "An elegant solution and a summary that is easy to read, thank you for sharing!",
    "2323842": "Thanks for sharing this beautifully written pointwise solution and approach @kononenko"
  },
  "source": "meta"
}