{
  "id": 425798,
  "title": "An essay for a PSPFGP comepetition.",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/425798",
  "author_name": "DongYK",
  "post_date": "2023-07-20T12:14:40.879000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In this competition, I faced my limitation as usual. So far, I joined competitions which are using much less features than this. The solution of the last competition was to use only one feature out of about 1000 features. Except for 15 competitors out of 1800, the most of solutions were overfitted. My soultion was too. I got 20th in a leaderboard, but 800th after closing the competition, which makes me have some fear to use many features. I couldn't use features until they were proven strictly. I did feature engineering much more carefully until the last week of this competition and I didn't get satisfactory result.<br>\nThat's why I felt difficulty in this competition.</p>\n<p>I have a habit to analyze my failure thoroughly and accept several novel ideas from others after a competition. So, this post is for my personal study. This post is written based on:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2323109\" target=\"_blank\">1st Place Solution</a></li>\n<li><a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining\" target=\"_blank\">PSPFGP 1st Place - NN Pretraining</a></li>\n<li><a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training\" target=\"_blank\">PSPFGP 1st Place - NN Training</a></li>\n<li><a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference\" target=\"_blank\">PSPFGP 1st Place - Inference</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565#2203292\" target=\"_blank\">Achieve [CV 0.6914 | LB 0.694] with Only Four Features</a></li>\n<li><a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">[LB 0.694] Event-Aware TConv with Only 4 Features</a></li>\n</ol>\n<p>I learned:</p>\n<ol>\n<li>pipeline (make a dataset, pretrain, train, and deploy over different notebooks)</li>\n<li>computation f1 fast</li>\n<li>various embedding techniques</li>\n<li>implementation of freezing</li>\n<li>combining different models into one model for early stopping with a custom metric.</li>\n<li>edge ML (implementation of tflite)</li>\n<li>Polars library</li>\n</ol>\n<p><strong>Embedding</strong> <br>\nI tried two versions. The first approach is from <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2323109\" target=\"_blank\">1st Place Solution</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F4af79cf4b37d55be7a375f7ef50912cd%2Fembedding1.png?generation=1689856648430770&amp;alt=media\" alt=\"\"><br>\nAnd this is from <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">[LB 0.694] Event-Aware TConv with Only 4 Features</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F6b8f1f9d1c72beed4d114eb776b354fe%2Fembedding2.png?generation=1689856658744403&amp;alt=media\" alt=\"\"><br>\nThe second one is more straightforward to me and I got a little improvement. However, the training time was increased by 1.3 times. It outweighed the improvement. So, I chose the first one. It was great for reducing training time but getting simliar results.</p>\n<p><strong>Optimizer</strong></p>\n<ol>\n<li>SGD with Nesterov and Cyclical Learning Rate</li>\n<li>Adam</li>\n<li>Nadam</li>\n<li>Adam with Cyclical Learning Rate</li>\n</ol>\n<p>I liked the first approach, because I believed that it's a flawless approach for escaping bad local minima. However, it couldn't find even any good local minima. the model sometimes predicted all 0 or 1. As the way I see it, It's very bad way for imbalanced targets.  The second and third one were good.<br>\nIn <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining\" target=\"_blank\">PSPFGP 1st Place - NN Pretraining</a> and <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training\" target=\"_blank\">PSPFGP 1st Place - NN Training</a>, there is a LearningRateSchedulerCallback which increase the learning rate at some specific point, I thought that it's for escaping bad local minima. So, I used the fourth one for it. It works well and improves CV +0.002. Additionally, <code>kernel_initializer='he_uniform'</code> improves CV +0.001.</p>\n<p>The whole figure of the NN architecture is from <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166\" target=\"_blank\">1st Place Solution</a></p>\n<p><strong>Edge AI (tflite)</strong><br>\nIn <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference\" target=\"_blank\">PSPFGP 1st Place - Inference</a>, I found that they used tflite for eficiency. These two posts was helpful to understand why do I use tflite:</p>\n<ol>\n<li><a href=\"https://viso.ai/edge-ai/tensorflow-lite/#:~:text=TensorFlow%20Lite%20(TFLite)%20is%20a,more%20than%204%20billion%20devices\" target=\"_blank\">TensorFlow Lite – Real-Time Computer Vision on Edge Devices (2022)</a></li>\n<li><a href=\"https://viso.ai/edge-ai/edge-ai-applications-and-trends/\" target=\"_blank\">Edge AI – Driving Next-Gen AI Applications in 2023</a></li>\n</ol>\n<p>After I used tflite, the inference time became <strong>6 minutes</strong>. Wow.</p>\n<p><strong>Notebooks</strong><br>\nYou can read my notebooks:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/dongyk/pspfgp-nn-dataset\" target=\"_blank\">PSPFGP_NN_dataset</a></li>\n<li><a href=\"https://www.kaggle.com/code/dongyk/pspfgp-nn-pretrain\" target=\"_blank\">PSPFGP_NN_Pretrain</a></li>\n<li><a href=\"https://www.kaggle.com/code/dongyk/pspfgp-nn-train\" target=\"_blank\">PSPFGP_NN_Train</a></li>\n<li><a href=\"https://www.kaggle.com/code/dongyk/pspfgp-nn-inference\" target=\"_blank\">PSPFGP_NN_Inference</a></li>\n</ol>\n<table>\n<thead>\n<tr>\n<th>Public</th>\n<th>Private</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.698</td>\n<td>0.697</td>\n<td>0.698</td>\n</tr>\n</tbody>\n</table>\n<p>I didn't use supplemental data.</p>\n<p>Again, thank you everybody for providing wonderful solutions and notebooks.</p>\n<p>Thank you for your patience.</p>",
  "messages": [
    {
      "id": 2351836,
      "postDate": "2023-07-20T12:14:40.880Z",
      "content": "<p>In this competition, I faced my limitation as usual. So far, I joined competitions which are using much less features than this. The solution of the last competition was to use only one feature out of about 1000 features. Except for 15 competitors out of 1800, the most of solutions were overfitted. My soultion was too. I got 20th in a leaderboard, but 800th after closing the competition, which makes me have some fear to use many features. I couldn't use features until they were proven strictly. I did feature engineering much more carefully until the last week of this competition and I didn't get satisfactory result.<br>\nThat's why I felt difficulty in this competition.</p>\n<p>I have a habit to analyze my failure thoroughly and accept several novel ideas from others after a competition. So, this post is for my personal study. This post is written based on:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2323109\" target=\"_blank\">1st Place Solution</a></li>\n<li><a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining\" target=\"_blank\">PSPFGP 1st Place - NN Pretraining</a></li>\n<li><a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training\" target=\"_blank\">PSPFGP 1st Place - NN Training</a></li>\n<li><a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference\" target=\"_blank\">PSPFGP 1st Place - Inference</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565#2203292\" target=\"_blank\">Achieve [CV 0.6914 | LB 0.694] with Only Four Features</a></li>\n<li><a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">[LB 0.694] Event-Aware TConv with Only 4 Features</a></li>\n</ol>\n<p>I learned:</p>\n<ol>\n<li>pipeline (make a dataset, pretrain, train, and deploy over different notebooks)</li>\n<li>computation f1 fast</li>\n<li>various embedding techniques</li>\n<li>implementation of freezing</li>\n<li>combining different models into one model for early stopping with a custom metric.</li>\n<li>edge ML (implementation of tflite)</li>\n<li>Polars library</li>\n</ol>\n<p><strong>Embedding</strong> <br>\nI tried two versions. The first approach is from <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2323109\" target=\"_blank\">1st Place Solution</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F4af79cf4b37d55be7a375f7ef50912cd%2Fembedding1.png?generation=1689856648430770&amp;alt=media\" alt=\"\"><br>\nAnd this is from <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">[LB 0.694] Event-Aware TConv with Only 4 Features</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F6b8f1f9d1c72beed4d114eb776b354fe%2Fembedding2.png?generation=1689856658744403&amp;alt=media\" alt=\"\"><br>\nThe second one is more straightforward to me and I got a little improvement. However, the training time was increased by 1.3 times. It outweighed the improvement. So, I chose the first one. It was great for reducing training time but getting simliar results.</p>\n<p><strong>Optimizer</strong></p>\n<ol>\n<li>SGD with Nesterov and Cyclical Learning Rate</li>\n<li>Adam</li>\n<li>Nadam</li>\n<li>Adam with Cyclical Learning Rate</li>\n</ol>\n<p>I liked the first approach, because I believed that it's a flawless approach for escaping bad local minima. However, it couldn't find even any good local minima. the model sometimes predicted all 0 or 1. As the way I see it, It's very bad way for imbalanced targets.  The second and third one were good.<br>\nIn <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining\" target=\"_blank\">PSPFGP 1st Place - NN Pretraining</a> and <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training\" target=\"_blank\">PSPFGP 1st Place - NN Training</a>, there is a LearningRateSchedulerCallback which increase the learning rate at some specific point, I thought that it's for escaping bad local minima. So, I used the fourth one for it. It works well and improves CV +0.002. Additionally, <code>kernel_initializer='he_uniform'</code> improves CV +0.001.</p>\n<p>The whole figure of the NN architecture is from <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166\" target=\"_blank\">1st Place Solution</a></p>\n<p><strong>Edge AI (tflite)</strong><br>\nIn <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference\" target=\"_blank\">PSPFGP 1st Place - Inference</a>, I found that they used tflite for eficiency. These two posts was helpful to understand why do I use tflite:</p>\n<ol>\n<li><a href=\"https://viso.ai/edge-ai/tensorflow-lite/#:~:text=TensorFlow%20Lite%20(TFLite)%20is%20a,more%20than%204%20billion%20devices\" target=\"_blank\">TensorFlow Lite – Real-Time Computer Vision on Edge Devices (2022)</a></li>\n<li><a href=\"https://viso.ai/edge-ai/edge-ai-applications-and-trends/\" target=\"_blank\">Edge AI – Driving Next-Gen AI Applications in 2023</a></li>\n</ol>\n<p>After I used tflite, the inference time became <strong>6 minutes</strong>. Wow.</p>\n<p><strong>Notebooks</strong><br>\nYou can read my notebooks:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/dongyk/pspfgp-nn-dataset\" target=\"_blank\">PSPFGP_NN_dataset</a></li>\n<li><a href=\"https://www.kaggle.com/code/dongyk/pspfgp-nn-pretrain\" target=\"_blank\">PSPFGP_NN_Pretrain</a></li>\n<li><a href=\"https://www.kaggle.com/code/dongyk/pspfgp-nn-train\" target=\"_blank\">PSPFGP_NN_Train</a></li>\n<li><a href=\"https://www.kaggle.com/code/dongyk/pspfgp-nn-inference\" target=\"_blank\">PSPFGP_NN_Inference</a></li>\n</ol>\n<table>\n<thead>\n<tr>\n<th>Public</th>\n<th>Private</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.698</td>\n<td>0.697</td>\n<td>0.698</td>\n</tr>\n</tbody>\n</table>\n<p>I didn't use supplemental data.</p>\n<p>Again, thank you everybody for providing wonderful solutions and notebooks.</p>\n<p>Thank you for your patience.</p>",
      "rawMarkdown": "In this competition, I faced my limitation as usual. So far, I joined competitions which are using much less features than this. The solution of the last competition was to use only one feature out of about 1000 features. Except for 15 competitors out of 1800, the most of solutions were overfitted. My soultion was too. I got 20th in a leaderboard, but 800th after closing the competition, which makes me have some fear to use many features. I couldn't use features until they were proven strictly. I did feature engineering much more carefully until the last week of this competition and I didn't get satisfactory result.\nThat's why I felt difficulty in this competition.\n\nI have a habit to analyze my failure thoroughly and accept several novel ideas from others after a competition. So, this post is for my personal study. This post is written based on:\n1. [1st Place Solution](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2323109)\n2. [PSPFGP 1st Place - NN Pretraining](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining)\n3. [PSPFGP 1st Place - NN Training](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training)\n4. [PSPFGP 1st Place - Inference](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference)\n5. [Achieve [CV 0.6914 | LB 0.694] with Only Four Features](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565#2203292)\n6. [[LB 0.694] Event-Aware TConv with Only 4 Features](https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features)\n\nI learned:\n1. pipeline (make a dataset, pretrain, train, and deploy over different notebooks)\n2. computation f1 fast\n3. various embedding techniques\n4. implementation of freezing\n5. combining different models into one model for early stopping with a custom metric.\n6. edge ML (implementation of tflite)\n7. Polars library\n\n**Embedding** \nI tried two versions. The first approach is from [1st Place Solution](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2323109)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F4af79cf4b37d55be7a375f7ef50912cd%2Fembedding1.png?generation=1689856648430770&alt=media)\nAnd this is from [[LB 0.694] Event-Aware TConv with Only 4 Features](https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F6b8f1f9d1c72beed4d114eb776b354fe%2Fembedding2.png?generation=1689856658744403&alt=media)\nThe second one is more straightforward to me and I got a little improvement. However, the training time was increased by 1.3 times. It outweighed the improvement. So, I chose the first one. It was great for reducing training time but getting simliar results.\n\n**Optimizer**\n\n1. SGD with Nesterov and Cyclical Learning Rate\n2. Adam\n3. Nadam\n4. Adam with Cyclical Learning Rate\n\nI liked the first approach, because I believed that it's a flawless approach for escaping bad local minima. However, it couldn't find even any good local minima. the model sometimes predicted all 0 or 1. As the way I see it, It's very bad way for imbalanced targets.  The second and third one were good.\nIn [PSPFGP 1st Place - NN Pretraining](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining) and [PSPFGP 1st Place - NN Training](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training), there is a LearningRateSchedulerCallback which increase the learning rate at some specific point, I thought that it's for escaping bad local minima. So, I used the fourth one for it. It works well and improves CV +0.002. Additionally, `kernel_initializer='he_uniform'` improves CV +0.001.\n\nThe whole figure of the NN architecture is from [1st Place Solution](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166)\n\n**Edge AI (tflite)**\nIn [PSPFGP 1st Place - Inference](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference), I found that they used tflite for eficiency. These two posts was helpful to understand why do I use tflite:\n1. [TensorFlow Lite – Real-Time Computer Vision on Edge Devices (2022)](https://viso.ai/edge-ai/tensorflow-lite/#:~:text=TensorFlow%20Lite%20(TFLite)%20is%20a,more%20than%204%20billion%20devices)\n2. [Edge AI – Driving Next-Gen AI Applications in 2023](https://viso.ai/edge-ai/edge-ai-applications-and-trends/)\n\nAfter I used tflite, the inference time became **6 minutes**. Wow.\n\n**Notebooks**\nYou can read my notebooks:\n1. [PSPFGP_NN_dataset](https://www.kaggle.com/code/dongyk/pspfgp-nn-dataset)\n2. [PSPFGP_NN_Pretrain](https://www.kaggle.com/code/dongyk/pspfgp-nn-pretrain)\n3. [PSPFGP_NN_Train](https://www.kaggle.com/code/dongyk/pspfgp-nn-train)\n4. [PSPFGP_NN_Inference](https://www.kaggle.com/code/dongyk/pspfgp-nn-inference)\n\n| Public | Private | CV |\n| --- | --- | --- |\n| 0.698 | 0.697 | 0.698 |\n\nI didn't use supplemental data.\n\nAgain, thank you everybody for providing wonderful solutions and notebooks.\n\nThank you for your patience.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2351836": "In this competition, I faced my limitation as usual. So far, I joined competitions which are using much less features than this. The solution of the last competition was to use only one feature out of about 1000 features. Except for 15 competitors out of 1800, the most of solutions were overfitted. My soultion was too. I got 20th in a leaderboard, but 800th after closing the competition, which makes me have some fear to use many features. I couldn't use features until they were proven strictly. I did feature engineering much more carefully until the last week of this competition and I didn't get satisfactory result.\nThat's why I felt difficulty in this competition.\n\nI have a habit to analyze my failure thoroughly and accept several novel ideas from others after a competition. So, this post is for my personal study. This post is written based on:\n1. [1st Place Solution](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2323109)\n2. [PSPFGP 1st Place - NN Pretraining](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining)\n3. [PSPFGP 1st Place - NN Training](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training)\n4. [PSPFGP 1st Place - Inference](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference)\n5. [Achieve [CV 0.6914 | LB 0.694] with Only Four Features](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565#2203292)\n6. [[LB 0.694] Event-Aware TConv with Only 4 Features](https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features)\n\nI learned:\n1. pipeline (make a dataset, pretrain, train, and deploy over different notebooks)\n2. computation f1 fast\n3. various embedding techniques\n4. implementation of freezing\n5. combining different models into one model for early stopping with a custom metric.\n6. edge ML (implementation of tflite)\n7. Polars library\n\n**Embedding** \nI tried two versions. The first approach is from [1st Place Solution](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2323109)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F4af79cf4b37d55be7a375f7ef50912cd%2Fembedding1.png?generation=1689856648430770&alt=media)\nAnd this is from [[LB 0.694] Event-Aware TConv with Only 4 Features](https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F6b8f1f9d1c72beed4d114eb776b354fe%2Fembedding2.png?generation=1689856658744403&alt=media)\nThe second one is more straightforward to me and I got a little improvement. However, the training time was increased by 1.3 times. It outweighed the improvement. So, I chose the first one. It was great for reducing training time but getting simliar results.\n\n**Optimizer**\n\n1. SGD with Nesterov and Cyclical Learning Rate\n2. Adam\n3. Nadam\n4. Adam with Cyclical Learning Rate\n\nI liked the first approach, because I believed that it's a flawless approach for escaping bad local minima. However, it couldn't find even any good local minima. the model sometimes predicted all 0 or 1. As the way I see it, It's very bad way for imbalanced targets.  The second and third one were good.\nIn [PSPFGP 1st Place - NN Pretraining](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining) and [PSPFGP 1st Place - NN Training](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training), there is a LearningRateSchedulerCallback which increase the learning rate at some specific point, I thought that it's for escaping bad local minima. So, I used the fourth one for it. It works well and improves CV +0.002. Additionally, `kernel_initializer='he_uniform'` improves CV +0.001.\n\nThe whole figure of the NN architecture is from [1st Place Solution](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166)\n\n**Edge AI (tflite)**\nIn [PSPFGP 1st Place - Inference](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference), I found that they used tflite for eficiency. These two posts was helpful to understand why do I use tflite:\n1. [TensorFlow Lite – Real-Time Computer Vision on Edge Devices (2022)](https://viso.ai/edge-ai/tensorflow-lite/#:~:text=TensorFlow%20Lite%20(TFLite)%20is%20a,more%20than%204%20billion%20devices)\n2. [Edge AI – Driving Next-Gen AI Applications in 2023](https://viso.ai/edge-ai/edge-ai-applications-and-trends/)\n\nAfter I used tflite, the inference time became **6 minutes**. Wow.\n\n**Notebooks**\nYou can read my notebooks:\n1. [PSPFGP_NN_dataset](https://www.kaggle.com/code/dongyk/pspfgp-nn-dataset)\n2. [PSPFGP_NN_Pretrain](https://www.kaggle.com/code/dongyk/pspfgp-nn-pretrain)\n3. [PSPFGP_NN_Train](https://www.kaggle.com/code/dongyk/pspfgp-nn-train)\n4. [PSPFGP_NN_Inference](https://www.kaggle.com/code/dongyk/pspfgp-nn-inference)\n\n| Public | Private | CV |\n| --- | --- | --- |\n| 0.698 | 0.697 | 0.698 |\n\nI didn't use supplemental data.\n\nAgain, thank you everybody for providing wonderful solutions and notebooks.\n\nThank you for your patience."
  }
}