{
  "id": 457873,
  "title": "53th place solution: Fast or Slow with Data Prunning and Pretraining",
  "url": "/competitions/predict-ai-model-runtime/discussion/457873",
  "author_name": "belgraviton",
  "post_date": "2023-11-27T09:20:41.742000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Dataset preprocessing</h1>\n<p>We have merged parts of the layout dataset for pretraining purposes:</p>\n<ul>\n<li>NLP = NLP.RANDOM + NLP.DEFAULT   -&gt;    Pretrain data for NLP models</li>\n<li>XLA = XLA.RANDOM + XLA.DEFAULT   -&gt;    Pretrain data for XLA models</li>\n</ul>\n<p>Pretrained and finetuned models use fixed mean/std values which are calculated as average values of  ‘random’ and ‘default’ datasets.</p>\n<p>Unuseful (zero) data was deleted from the dataset for training speed up:</p>\n<ul>\n<li>Truncated features in LAYOUT.NLP dataset: nodes 140 -&gt; 40, edges 18 -&gt; 8 </li>\n<li>Truncated features in LAYOUT.XLA dataset: nodes 140 -&gt; 112, edges 18 -&gt; 14</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554956%2F7bc2b9470906e2c8d343fdd71b26f9de%2Fdata_stat.png?generation=1701075640215878&amp;alt=media\" alt=\"\"></p>\n<h1>Training</h1>\n<h2>LAYOUT dataset (<a href=\"https://github.com/belgraviton/gst_tpu/tree/main/scripts\" target=\"_blank\">github</a>):</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554956%2Fea563e44aec6ecbd67c573a0a352f8e6%2Fglob_pool.png?generation=1701075656696317&amp;alt=media\" alt=\"\"></p>\n<p>Solution:</p>\n<ul>\n<li>pretraining on full NLP and XLA datasets and finetune on specific ones</li>\n<li>number of configs 64</li>\n<li>loss margin strategy: 0.01 -&gt; 0.1 -&gt; 0.5</li>\n<li>clip gradients for XLA:RANDOM</li>\n<li>dropout 0.0</li>\n<li>global pooling improvement with concatenate operation</li>\n<li>hidden dim = 256</li>\n<li>1 k epochs</li>\n<li>CosineLRScheduler</li>\n<li>Adam, lr = 1e-4</li>\n</ul>\n<p>Results: Kendall tau: NLP.DEFAULT 48.3%, NLP.RANDOM 85.2%, XLA.DEFAULT 29.6%, XLA.RANDOM 37.3%</p>\n<h2>TILE dataset (<a href=\"https://github.com/belgraviton/tpupredict/blob/main/scripts/t03_99_MSEtl_10k.sh\" target=\"_blank\">github</a>):</h2>\n<p>Solution:</p>\n<ul>\n<li>batchsize 16</li>\n<li>number of configs 512</li>\n<li>2 k epochs with early stop 200 epochs</li>\n<li>Adam, lr = 1e-3</li>\n</ul>\n<p>Results: validation OPA: 90.82%</p>\n<h2>Other experiments that did NOT work:</h2>\n<p>GENConv+Transformer hybrid model from GraphGPS repo</p>\n<p>Architecture design:</p>\n<ul>\n<li>number of layers</li>\n<li>hidden dimension</li>\n</ul>",
  "messages": [
    {
      "id": 2539705,
      "postDate": "2023-11-27T09:20:41.743Z",
      "content": "<h1>Dataset preprocessing</h1>\n<p>We have merged parts of the layout dataset for pretraining purposes:</p>\n<ul>\n<li>NLP = NLP.RANDOM + NLP.DEFAULT   -&gt;    Pretrain data for NLP models</li>\n<li>XLA = XLA.RANDOM + XLA.DEFAULT   -&gt;    Pretrain data for XLA models</li>\n</ul>\n<p>Pretrained and finetuned models use fixed mean/std values which are calculated as average values of  ‘random’ and ‘default’ datasets.</p>\n<p>Unuseful (zero) data was deleted from the dataset for training speed up:</p>\n<ul>\n<li>Truncated features in LAYOUT.NLP dataset: nodes 140 -&gt; 40, edges 18 -&gt; 8 </li>\n<li>Truncated features in LAYOUT.XLA dataset: nodes 140 -&gt; 112, edges 18 -&gt; 14</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554956%2F7bc2b9470906e2c8d343fdd71b26f9de%2Fdata_stat.png?generation=1701075640215878&amp;alt=media\" alt=\"\"></p>\n<h1>Training</h1>\n<h2>LAYOUT dataset (<a href=\"https://github.com/belgraviton/gst_tpu/tree/main/scripts\" target=\"_blank\">github</a>):</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554956%2Fea563e44aec6ecbd67c573a0a352f8e6%2Fglob_pool.png?generation=1701075656696317&amp;alt=media\" alt=\"\"></p>\n<p>Solution:</p>\n<ul>\n<li>pretraining on full NLP and XLA datasets and finetune on specific ones</li>\n<li>number of configs 64</li>\n<li>loss margin strategy: 0.01 -&gt; 0.1 -&gt; 0.5</li>\n<li>clip gradients for XLA:RANDOM</li>\n<li>dropout 0.0</li>\n<li>global pooling improvement with concatenate operation</li>\n<li>hidden dim = 256</li>\n<li>1 k epochs</li>\n<li>CosineLRScheduler</li>\n<li>Adam, lr = 1e-4</li>\n</ul>\n<p>Results: Kendall tau: NLP.DEFAULT 48.3%, NLP.RANDOM 85.2%, XLA.DEFAULT 29.6%, XLA.RANDOM 37.3%</p>\n<h2>TILE dataset (<a href=\"https://github.com/belgraviton/tpupredict/blob/main/scripts/t03_99_MSEtl_10k.sh\" target=\"_blank\">github</a>):</h2>\n<p>Solution:</p>\n<ul>\n<li>batchsize 16</li>\n<li>number of configs 512</li>\n<li>2 k epochs with early stop 200 epochs</li>\n<li>Adam, lr = 1e-3</li>\n</ul>\n<p>Results: validation OPA: 90.82%</p>\n<h2>Other experiments that did NOT work:</h2>\n<p>GENConv+Transformer hybrid model from GraphGPS repo</p>\n<p>Architecture design:</p>\n<ul>\n<li>number of layers</li>\n<li>hidden dimension</li>\n</ul>",
      "rawMarkdown": "# Dataset preprocessing\nWe have merged parts of the layout dataset for pretraining purposes:\n-  NLP = NLP.RANDOM + NLP.DEFAULT   ->    Pretrain data for NLP models\n-  XLA = XLA.RANDOM + XLA.DEFAULT   ->    Pretrain data for XLA models\n\nPretrained and finetuned models use fixed mean/std values which are calculated as average values of  ‘random’ and ‘default’ datasets.\n\nUnuseful (zero) data was deleted from the dataset for training speed up:\n- Truncated features in LAYOUT.NLP dataset: nodes 140 -> 40, edges 18 -> 8 \n- Truncated features in LAYOUT.XLA dataset: nodes 140 -> 112, edges 18 -> 14\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554956%2F7bc2b9470906e2c8d343fdd71b26f9de%2Fdata_stat.png?generation=1701075640215878&alt=media)\n\n# Training\n\n## LAYOUT dataset ([github](https://github.com/belgraviton/gst_tpu/tree/main/scripts)):\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554956%2Fea563e44aec6ecbd67c573a0a352f8e6%2Fglob_pool.png?generation=1701075656696317&alt=media)\n\nSolution:\n- pretraining on full NLP and XLA datasets and finetune on specific ones\n- number of configs 64\n- loss margin strategy: 0.01 -> 0.1 -> 0.5\n- clip gradients for XLA:RANDOM\n- dropout 0.0\n- global pooling improvement with concatenate operation\n- hidden dim = 256\n- 1 k epochs\n- CosineLRScheduler\n- Adam, lr = 1e-4\n\nResults: Kendall tau: NLP.DEFAULT 48.3%, NLP.RANDOM 85.2%, XLA.DEFAULT 29.6%, XLA.RANDOM 37.3%\n\n## TILE dataset ([github](https://github.com/belgraviton/tpupredict/blob/main/scripts/t03_99_MSEtl_10k.sh)):\n\nSolution:\n- batchsize 16\n- number of configs 512\n- 2 k epochs with early stop 200 epochs\n- Adam, lr = 1e-3\n\nResults: validation OPA: 90.82%\n\n\n## Other experiments that did NOT work:\n\nGENConv+Transformer hybrid model from GraphGPS repo\n\nArchitecture design:\n- number of layers\n- hidden dimension\n\n\n\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2539705": "# Dataset preprocessing\nWe have merged parts of the layout dataset for pretraining purposes:\n-  NLP = NLP.RANDOM + NLP.DEFAULT   ->    Pretrain data for NLP models\n-  XLA = XLA.RANDOM + XLA.DEFAULT   ->    Pretrain data for XLA models\n\nPretrained and finetuned models use fixed mean/std values which are calculated as average values of  ‘random’ and ‘default’ datasets.\n\nUnuseful (zero) data was deleted from the dataset for training speed up:\n- Truncated features in LAYOUT.NLP dataset: nodes 140 -> 40, edges 18 -> 8 \n- Truncated features in LAYOUT.XLA dataset: nodes 140 -> 112, edges 18 -> 14\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554956%2F7bc2b9470906e2c8d343fdd71b26f9de%2Fdata_stat.png?generation=1701075640215878&alt=media)\n\n# Training\n\n## LAYOUT dataset ([github](https://github.com/belgraviton/gst_tpu/tree/main/scripts)):\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F554956%2Fea563e44aec6ecbd67c573a0a352f8e6%2Fglob_pool.png?generation=1701075656696317&alt=media)\n\nSolution:\n- pretraining on full NLP and XLA datasets and finetune on specific ones\n- number of configs 64\n- loss margin strategy: 0.01 -> 0.1 -> 0.5\n- clip gradients for XLA:RANDOM\n- dropout 0.0\n- global pooling improvement with concatenate operation\n- hidden dim = 256\n- 1 k epochs\n- CosineLRScheduler\n- Adam, lr = 1e-4\n\nResults: Kendall tau: NLP.DEFAULT 48.3%, NLP.RANDOM 85.2%, XLA.DEFAULT 29.6%, XLA.RANDOM 37.3%\n\n## TILE dataset ([github](https://github.com/belgraviton/tpupredict/blob/main/scripts/t03_99_MSEtl_10k.sh)):\n\nSolution:\n- batchsize 16\n- number of configs 512\n- 2 k epochs with early stop 200 epochs\n- Adam, lr = 1e-3\n\nResults: validation OPA: 90.82%\n\n\n## Other experiments that did NOT work:\n\nGENConv+Transformer hybrid model from GraphGPS repo\n\nArchitecture design:\n- number of layers\n- hidden dimension\n\n\n\n"
  }
}