{
  "id": 392157,
  "title": "Do this tip when loading the dataset to save 50% of the RAM",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/392157",
  "author_name": "Nael Aqel",
  "post_date": "2023-03-03T21:12:49.294000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Dears,</p>\n<p>I developed a module for cleaning data frames, I tested it on the data set in this competition without any <strong>accelerator</strong>, you can view the <a href=\"https://www.kaggle.com/code/naelaqel/saving-50-of-memory-tip/notebook\" target=\"_blank\">notebook here</a>.</p>\n<p><strong>As a summery</strong><br>\nWhen loading the data frame specify this two parameters with values mentioned below:<br>\n1- <code>usecols=['session_id', 'index', 'elapsed_time', 'event_name', 'name', 'level', 'page', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'hover_duration', 'text', 'fqid', 'room_fqid', 'text_fqid', 'level_group']</code><br>\n2- <code>dtype={'session_id': np.uint64, 'index': np.uint16, 'elapsed_time': np.uint32, 'event_name': 'category',\n'name': 'category', 'level': np.uint8, 'room_fqid': 'category', 'level_group': 'category'}</code></p>\n<p><strong>Note:</strong> these columns have the potential to change its datatype after <strong>filling the missing values</strong>, <br>\n<code>{'page': numpy.uint8, 'room_coor_x': numpy.float16, 'room_coor_y': numpy.float16, 'screen_coor_x': numpy.float16, 'screen_coor_y': numpy.float16, 'hover_duration': numpy.uint32}</code> </p>\n<p>Good luck in the competition …. 🙂</p>",
  "messages": [
    {
      "id": 2168078,
      "postDate": "2023-03-03T21:12:49.293Z",
      "content": "<p>Dears,</p>\n<p>I developed a module for cleaning data frames, I tested it on the data set in this competition without any <strong>accelerator</strong>, you can view the <a href=\"https://www.kaggle.com/code/naelaqel/saving-50-of-memory-tip/notebook\" target=\"_blank\">notebook here</a>.</p>\n<p><strong>As a summery</strong><br>\nWhen loading the data frame specify this two parameters with values mentioned below:<br>\n1- <code>usecols=['session_id', 'index', 'elapsed_time', 'event_name', 'name', 'level', 'page', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'hover_duration', 'text', 'fqid', 'room_fqid', 'text_fqid', 'level_group']</code><br>\n2- <code>dtype={'session_id': np.uint64, 'index': np.uint16, 'elapsed_time': np.uint32, 'event_name': 'category',\n'name': 'category', 'level': np.uint8, 'room_fqid': 'category', 'level_group': 'category'}</code></p>\n<p><strong>Note:</strong> these columns have the potential to change its datatype after <strong>filling the missing values</strong>, <br>\n<code>{'page': numpy.uint8, 'room_coor_x': numpy.float16, 'room_coor_y': numpy.float16, 'screen_coor_x': numpy.float16, 'screen_coor_y': numpy.float16, 'hover_duration': numpy.uint32}</code> </p>\n<p>Good luck in the competition …. 🙂</p>",
      "rawMarkdown": "Dears,\n\nI developed a module for cleaning data frames, I tested it on the data set in this competition without any **accelerator**, you can view the [notebook here](https://www.kaggle.com/code/naelaqel/saving-50-of-memory-tip/notebook).\n\n**As a summery**\nWhen loading the data frame specify this two parameters with values mentioned below:\n1- `usecols=['session_id', 'index', 'elapsed_time', 'event_name', 'name', 'level', 'page', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'hover_duration', 'text', 'fqid', 'room_fqid', 'text_fqid', 'level_group']`\n2- `dtype={'session_id': np.uint64, 'index': np.uint16, 'elapsed_time': np.uint32, 'event_name': 'category',\n'name': 'category', 'level': np.uint8, 'room_fqid': 'category', 'level_group': 'category'}`\n\n**Note:** these columns have the potential to change its datatype after **filling the missing values**, \n`{'page': numpy.uint8, 'room_coor_x': numpy.float16, 'room_coor_y': numpy.float16, 'screen_coor_x': numpy.float16, 'screen_coor_y': numpy.float16, 'hover_duration': numpy.uint32}` \n\nGood luck in the competition .... 🙂",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2168078": "Dears,\n\nI developed a module for cleaning data frames, I tested it on the data set in this competition without any **accelerator**, you can view the [notebook here](https://www.kaggle.com/code/naelaqel/saving-50-of-memory-tip/notebook).\n\n**As a summery**\nWhen loading the data frame specify this two parameters with values mentioned below:\n1- `usecols=['session_id', 'index', 'elapsed_time', 'event_name', 'name', 'level', 'page', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'hover_duration', 'text', 'fqid', 'room_fqid', 'text_fqid', 'level_group']`\n2- `dtype={'session_id': np.uint64, 'index': np.uint16, 'elapsed_time': np.uint32, 'event_name': 'category',\n'name': 'category', 'level': np.uint8, 'room_fqid': 'category', 'level_group': 'category'}`\n\n**Note:** these columns have the potential to change its datatype after **filling the missing values**, \n`{'page': numpy.uint8, 'room_coor_x': numpy.float16, 'room_coor_y': numpy.float16, 'screen_coor_x': numpy.float16, 'screen_coor_y': numpy.float16, 'hover_duration': numpy.uint32}` \n\nGood luck in the competition .... 🙂"
  }
}