{
  "id": 456645,
  "title": "8th place solution",
  "url": "/competitions/predict-ai-model-runtime/writeups/30crmnsia-senkin13-8th-place-solution",
  "author_name": "",
  "post_date": "2023-11-21T03:46:21.183Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to competition organizer giving us such a interseting competition.Our team joined very late we don't have time to deepdive but we found a simple solution very useful.</p>\n<h1>Data Processing &amp; Feature Engineering</h1>\n<p>Tile: count of config&amp;node&amp;edge, mean,max,std,last of node_feat, mean,max,std of config_feat<br>\nLayout: flatten node_config_feat -&gt; remove unique value columns -&gt; remove duplicate columns</p>\n<h1>Train Data</h1>\n<p>The keypoint of our layout solution is finding the most similar train data for each test data. We can observe some data have almost same edge&amp;node number and can guess they are same model type except size,the test data should be the same model with different batch size.<br>\ne.g.</p>\n<pre><code>train: small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_16_test\nvalid: small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_32_test\ntest(same edge&amp;node number) should be the small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_64_test\n</code></pre>\n<p>We cannot find all the similar train data for each test data, but enough to get a good result.And it's very fast for iteration.</p>\n<h1>Model</h1>\n<p>transform the target to minmaxscaler , use xentropy as the loss function, lightgbm as model, we don't use validation just give a fix round for each model.</p>",
  "messages": [
    {
      "id": "2532412",
      "postDate": "11/21/2023 03:44:11",
      "content": "<p>Thanks to competition organizer giving us such a interseting competition.Our team joined very late we don't have time to deepdive but we found a simple solution very useful.</p>\n<h1>Data Processing &amp; Feature Engineering</h1>\n<p>Tile: count of config&amp;node&amp;edge, mean,max,std,last of node_feat, mean,max,std of config_feat<br>\nLayout: flatten node_config_feat -&gt; remove unique value columns -&gt; remove duplicate columns</p>\n<h1>Train Data</h1>\n<p>The keypoint of our layout solution is finding the most similar train data for each test data. We can observe some data have almost same edge&amp;node number and can guess they are same model type except size,the test data should be the same model with different batch size.<br>\ne.g.</p>\n<pre><code>train: small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_16_test\nvalid: small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_32_test\ntest(same edge&amp;node number) should be the small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_64_test\n</code></pre>\n<p>We cannot find all the similar train data for each test data, but enough to get a good result.And it's very fast for iteration.</p>\n<h1>Model</h1>\n<p>transform the target to minmaxscaler , use xentropy as the loss function, lightgbm as model, we don't use validation just give a fix round for each model.</p>",
      "rawMarkdown": "Thanks to competition organizer giving us such a interseting competition.Our team joined very late we don't have time to deepdive but we found a simple solution very useful.\n\n# Data Processing & Feature Engineering\nTile: count of config&node&edge, mean,max,std,last of node_feat, mean,max,std of config_feat\nLayout: flatten node_config_feat -> remove unique value columns -> remove duplicate columns\n\n# Train Data\nThe keypoint of our layout solution is finding the most similar train data for each test data. We can observe some data have almost same edge&node number and can guess they are same model type except size,the test data should be the same model with different batch size.\ne.g.\n```python\ntrain: small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_16_test\nvalid: small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_32_test\ntest(same edge&node number) should be the small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_64_test\n```\nWe cannot find all the similar train data for each test data, but enough to get a good result.And it's very fast for iteration.\n\n# Model\ntransform the target to minmaxscaler , use xentropy as the loss function, lightgbm as model, we don't use validation just give a fix round for each model.",
      "votes": null
    },
    {
      "id": "2533172",
      "postDate": "11/21/2023 16:34:50",
      "content": "<p>Thank you for sharing your solution and congratulations for your position Senkin13-san.</p>",
      "rawMarkdown": "Thank you for sharing your solution and congratulations for your position Senkin13-san.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2533172,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "11/21/2023 16:34:50",
      "content": "<p>Thank you for sharing your solution and congratulations for your position Senkin13-san.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2532412": "Thanks to competition organizer giving us such a interseting competition.Our team joined very late we don't have time to deepdive but we found a simple solution very useful.\n\n# Data Processing & Feature Engineering\nTile: count of config&node&edge, mean,max,std,last of node_feat, mean,max,std of config_feat\nLayout: flatten node_config_feat -> remove unique value columns -> remove duplicate columns\n\n# Train Data\nThe keypoint of our layout solution is finding the most similar train data for each test data. We can observe some data have almost same edge&node number and can guess they are same model type except size,the test data should be the same model with different batch size.\ne.g.\n```python\ntrain: small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_16_test\nvalid: small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_32_test\ntest(same edge&node number) should be the small_bert_bert_en_uncased_L-12_H-768_A-12_batch_size_64_test\n```\nWe cannot find all the similar train data for each test data, but enough to get a good result.And it's very fast for iteration.\n\n# Model\ntransform the target to minmaxscaler , use xentropy as the loss function, lightgbm as model, we don't use validation just give a fix round for each model.",
    "2533172": "Thank you for sharing your solution and congratulations for your position Senkin13-san."
  },
  "source": "meta"
}