{
  "id": 456092,
  "title": "11th place solution: LightGBM",
  "url": "/competitions/predict-ai-model-runtime/writeups/shun-pi-11th-place-solution-lightgbm",
  "author_name": "",
  "post_date": "2023-11-19T22:47:52.917Z",
  "votes": 27,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thanks to the host and congratulations to the winners.<br>\nI'm glad to win my 4th gold medal on this competition.<br>\nI did not use GNN, but used LightGBM with feature engineering.<br>\nThe reason why I chose LightGBM is that I thought that the runtime of the model in a single TPU is equal to the sum of the runtime of each node, and learning the aggregated statistics from the graph would be sufficient to achieve good score without learning the graph structures.</p>\n<h2>Layout</h2>\n<h3>Features</h3>\n<p>I extracted the following features from the graph.</p>\n<ul>\n<li>Number of nodes of each type for conv/dot/reshape when using the following classification method<ul>\n<li>Conv: If any one of the input/output number of dimensions or config values, or the 93~106th values of node_feat is different, it is of a different type</li>\n<li>Dot: If any one of the input/output number of dimensions or config values, or the index features of dot operation (extracted from the .pb file on its own) is different, then it is of a different type</li>\n<li>Reshape: If any one of the input/output number of dimensions or config values, or the inconsistency on the element products in a set of a certain dimension, then it is of a different type</li></ul></li>\n<li>Number of times the element is copied<ul>\n<li>Determine if a copy is needed by comparing the layout of each configurable node and the configurable nodes connected to it.</li></ul></li>\n<li>Counting binning the size of the 1st/2nd minor dimension</li>\n<li>Sum of padding generated at each configurable node</li>\n<li>(only default) Occurrence rate of config for each node relative to the total data for that model<ul>\n<li>Because genetic algorithms are used in the search, the more frequently a config pattern appears, the faster the config runtime will tend to be.</li></ul></li>\n</ul>\n<h3>Models</h3>\n<p>I trained the pointwise and pairwise models.</p>\n<ul>\n<li>Pointwise LightGBM<ul>\n<li>Target: Normalized rankings(0~1)</li>\n<li>Loss: MAE Loss</li>\n<li>Public LB: 0.715 (Private LB: 0.680)</li></ul></li>\n<li>Pairwise LightGBM<ul>\n<li>Loss: Binary</li>\n<li>Randomly select the same number of pairs as the number of configurations and generate train/valid.</li>\n<li>Inference on test data predicts for all pairs of 1000^2 and then sort configs using the sum of the predictions.</li>\n<li>Input features are as follows: <ul>\n<li>The features for one config of the pair</li>\n<li>The difference between the features of the two configs</li></ul></li>\n<li>Public LB: 0.728 (Private LB:0.701)</li></ul></li>\n</ul>\n<p>I used the same train/valid given by the host. (Failure to devise a better CV may have been the cause of the shake down)</p>\n<h2>Tile</h2>\n<h3>Features</h3>\n<ul>\n<li>Config features</li>\n<li>Node features averaged over all nodes</li>\n</ul>\n<h3>Models</h3>\n<p>I trained the pointwise LightGBM model.</p>\n<ul>\n<li>Target: Normalized rankings(0~1)</li>\n<li>Loss: MAE Loss</li>\n<li>Public LB(only tile): 0.198 (Private LB(only tile): 0.195)</li>\n</ul>",
  "messages": [
    {
      "id": "2529177",
      "postDate": "11/18/2023 02:36:13",
      "content": "<p>Thanks to the host and congratulations to the winners.<br>\nI'm glad to win my 4th gold medal on this competition.<br>\nI did not use GNN, but used LightGBM with feature engineering.<br>\nThe reason why I chose LightGBM is that I thought that the runtime of the model in a single TPU is equal to the sum of the runtime of each node, and learning the aggregated statistics from the graph would be sufficient to achieve good score without learning the graph structures.</p>\n<h2>Layout</h2>\n<h3>Features</h3>\n<p>I extracted the following features from the graph.</p>\n<ul>\n<li>Number of nodes of each type for conv/dot/reshape when using the following classification method<ul>\n<li>Conv: If any one of the input/output number of dimensions or config values, or the 93~106th values of node_feat is different, it is of a different type</li>\n<li>Dot: If any one of the input/output number of dimensions or config values, or the index features of dot operation (extracted from the .pb file on its own) is different, then it is of a different type</li>\n<li>Reshape: If any one of the input/output number of dimensions or config values, or the inconsistency on the element products in a set of a certain dimension, then it is of a different type</li></ul></li>\n<li>Number of times the element is copied<ul>\n<li>Determine if a copy is needed by comparing the layout of each configurable node and the configurable nodes connected to it.</li></ul></li>\n<li>Counting binning the size of the 1st/2nd minor dimension</li>\n<li>Sum of padding generated at each configurable node</li>\n<li>(only default) Occurrence rate of config for each node relative to the total data for that model<ul>\n<li>Because genetic algorithms are used in the search, the more frequently a config pattern appears, the faster the config runtime will tend to be.</li></ul></li>\n</ul>\n<h3>Models</h3>\n<p>I trained the pointwise and pairwise models.</p>\n<ul>\n<li>Pointwise LightGBM<ul>\n<li>Target: Normalized rankings(0~1)</li>\n<li>Loss: MAE Loss</li>\n<li>Public LB: 0.715 (Private LB: 0.680)</li></ul></li>\n<li>Pairwise LightGBM<ul>\n<li>Loss: Binary</li>\n<li>Randomly select the same number of pairs as the number of configurations and generate train/valid.</li>\n<li>Inference on test data predicts for all pairs of 1000^2 and then sort configs using the sum of the predictions.</li>\n<li>Input features are as follows: <ul>\n<li>The features for one config of the pair</li>\n<li>The difference between the features of the two configs</li></ul></li>\n<li>Public LB: 0.728 (Private LB:0.701)</li></ul></li>\n</ul>\n<p>I used the same train/valid given by the host. (Failure to devise a better CV may have been the cause of the shake down)</p>\n<h2>Tile</h2>\n<h3>Features</h3>\n<ul>\n<li>Config features</li>\n<li>Node features averaged over all nodes</li>\n</ul>\n<h3>Models</h3>\n<p>I trained the pointwise LightGBM model.</p>\n<ul>\n<li>Target: Normalized rankings(0~1)</li>\n<li>Loss: MAE Loss</li>\n<li>Public LB(only tile): 0.198 (Private LB(only tile): 0.195)</li>\n</ul>",
      "rawMarkdown": "Thanks to the host and congratulations to the winners.\nI'm glad to win my 4th gold medal on this competition.\nI did not use GNN, but used LightGBM with feature engineering.\nThe reason why I chose LightGBM is that I thought that the runtime of the model in a single TPU is equal to the sum of the runtime of each node, and learning the aggregated statistics from the graph would be sufficient to achieve good score without learning the graph structures.\n\n## Layout\n\n### Features\n\nI extracted the following features from the graph.\n- Number of nodes of each type for conv/dot/reshape when using the following classification method\n\t- Conv: If any one of the input/output number of dimensions or config values, or the 93~106th values of node_feat is different, it is of a different type\n\t- Dot: If any one of the input/output number of dimensions or config values, or the index features of dot operation (extracted from the .pb file on its own) is different, then it is of a different type\n\t- Reshape: If any one of the input/output number of dimensions or config values, or the inconsistency on the element products in a set of a certain dimension, then it is of a different type\n- Number of times the element is copied\n\t- Determine if a copy is needed by comparing the layout of each configurable node and the configurable nodes connected to it.\n- Counting binning the size of the 1st/2nd minor dimension\n- Sum of padding generated at each configurable node\n- (only default) Occurrence rate of config for each node relative to the total data for that model\n\t- Because genetic algorithms are used in the search, the more frequently a config pattern appears, the faster the config runtime will tend to be.\n\n### Models\n\nI trained the pointwise and pairwise models.\n\n- Pointwise LightGBM\n\t- Target: Normalized rankings(0~1)\n\t- Loss: MAE Loss\n\t- Public LB: 0.715 (Private LB: 0.680)\n- Pairwise LightGBM\n\t- Loss: Binary\n\t- Randomly select the same number of pairs as the number of configurations and generate train/valid.\n\t- Inference on test data predicts for all pairs of 1000^2 and then sort configs using the sum of the predictions.\n\t- Input features are as follows: \n\t\t- The features for one config of the pair\n\t\t- The difference between the features of the two configs\n\t- Public LB: 0.728 (Private LB:0.701)\n\nI used the same train/valid given by the host. (Failure to devise a better CV may have been the cause of the shake down)\n\n## Tile\n\n### Features\n\n- Config features\n- Node features averaged over all nodes\n### Models\n\nI trained the pointwise LightGBM model.\n- Target: Normalized rankings(0~1)\n- Loss: MAE Loss\n- Public LB(only tile): 0.198 (Private LB(only tile): 0.195)",
      "votes": null
    },
    {
      "id": "2529198",
      "postDate": "11/18/2023 03:20:24",
      "content": "<p>Quite different methodology, thanks for sharing.</p>",
      "rawMarkdown": "Quite different methodology, thanks for sharing.",
      "votes": null
    },
    {
      "id": "2529267",
      "postDate": "11/18/2023 04:53:33",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2529893",
      "postDate": "11/18/2023 16:21:41",
      "content": "<p>Congratulations on achieving 11th position. Interesting to observe that you have used LightGBM.</p>",
      "rawMarkdown": "Congratulations on achieving 11th position. Interesting to observe that you have used LightGBM.",
      "votes": null
    },
    {
      "id": "2529935",
      "postDate": "11/18/2023 16:59:07",
      "content": "<p>Good work!</p>\n<p>what surprises me is that MAE Loss works!</p>",
      "rawMarkdown": "Good work!\n\nwhat surprises me is that MAE Loss works!",
      "votes": null
    },
    {
      "id": "2530918",
      "postDate": "11/19/2023 16:34:50",
      "content": "<p>Congratulations on the solution, really interesting approach! Also, very smart observation:</p>\n<blockquote>\n  <p>Because genetic algorithms are used in the search, the more frequently a config pattern appears, the faster the config runtime will tend to be.</p>\n</blockquote>",
      "rawMarkdown": "Congratulations on the solution, really interesting approach! Also, very smart observation:\n>Because genetic algorithms are used in the search, the more frequently a config pattern appears, the faster the config runtime will tend to be.",
      "votes": null
    },
    {
      "id": "2531139",
      "postDate": "11/19/2023 23:14:42",
      "content": "<p>this is so cool</p>",
      "rawMarkdown": "this is so cool",
      "votes": null
    },
    {
      "id": "2531157",
      "postDate": "11/20/2023 00:46:08",
      "content": "<p>Great work. Congrats</p>",
      "rawMarkdown": "Great work. Congrats",
      "votes": null
    },
    {
      "id": "2537326",
      "postDate": "11/25/2023 04:13:58",
      "content": "<p>Congratulations on achieving 11th position, would you mind making the notebook public? you have such a unique solution to the problem it would be of help to the community and have you tried XGB ranker too ?</p>",
      "rawMarkdown": "Congratulations on achieving 11th position, would you mind making the notebook public? you have such a unique solution to the problem it would be of help to the community and have you tried XGB ranker too ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2529198,
      "author_name": "goelyash",
      "author_url": "",
      "post_date": "11/18/2023 03:20:24",
      "content": "<p>Quite different methodology, thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2529267,
      "author_name": "mylychee",
      "author_url": "",
      "post_date": "11/18/2023 04:53:33",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2529893,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "11/18/2023 16:21:41",
      "content": "<p>Congratulations on achieving 11th position. Interesting to observe that you have used LightGBM.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2529935,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/18/2023 16:59:07",
      "content": "<p>Good work!</p>\n<p>what surprises me is that MAE Loss works!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2530918,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "11/19/2023 16:34:50",
      "content": "<p>Congratulations on the solution, really interesting approach! Also, very smart observation:</p>\n<blockquote>\n  <p>Because genetic algorithms are used in the search, the more frequently a config pattern appears, the faster the config runtime will tend to be.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2531139,
      "author_name": "lewington",
      "author_url": "",
      "post_date": "11/19/2023 23:14:42",
      "content": "<p>this is so cool</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2531157,
      "author_name": "indrasn0wal",
      "author_url": "",
      "post_date": "11/20/2023 00:46:08",
      "content": "<p>Great work. Congrats</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2537326,
      "author_name": "arunsensei",
      "author_url": "",
      "post_date": "11/25/2023 04:13:58",
      "content": "<p>Congratulations on achieving 11th position, would you mind making the notebook public? you have such a unique solution to the problem it would be of help to the community and have you tried XGB ranker too ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2529177": "Thanks to the host and congratulations to the winners.\nI'm glad to win my 4th gold medal on this competition.\nI did not use GNN, but used LightGBM with feature engineering.\nThe reason why I chose LightGBM is that I thought that the runtime of the model in a single TPU is equal to the sum of the runtime of each node, and learning the aggregated statistics from the graph would be sufficient to achieve good score without learning the graph structures.\n\n## Layout\n\n### Features\n\nI extracted the following features from the graph.\n- Number of nodes of each type for conv/dot/reshape when using the following classification method\n\t- Conv: If any one of the input/output number of dimensions or config values, or the 93~106th values of node_feat is different, it is of a different type\n\t- Dot: If any one of the input/output number of dimensions or config values, or the index features of dot operation (extracted from the .pb file on its own) is different, then it is of a different type\n\t- Reshape: If any one of the input/output number of dimensions or config values, or the inconsistency on the element products in a set of a certain dimension, then it is of a different type\n- Number of times the element is copied\n\t- Determine if a copy is needed by comparing the layout of each configurable node and the configurable nodes connected to it.\n- Counting binning the size of the 1st/2nd minor dimension\n- Sum of padding generated at each configurable node\n- (only default) Occurrence rate of config for each node relative to the total data for that model\n\t- Because genetic algorithms are used in the search, the more frequently a config pattern appears, the faster the config runtime will tend to be.\n\n### Models\n\nI trained the pointwise and pairwise models.\n\n- Pointwise LightGBM\n\t- Target: Normalized rankings(0~1)\n\t- Loss: MAE Loss\n\t- Public LB: 0.715 (Private LB: 0.680)\n- Pairwise LightGBM\n\t- Loss: Binary\n\t- Randomly select the same number of pairs as the number of configurations and generate train/valid.\n\t- Inference on test data predicts for all pairs of 1000^2 and then sort configs using the sum of the predictions.\n\t- Input features are as follows: \n\t\t- The features for one config of the pair\n\t\t- The difference between the features of the two configs\n\t- Public LB: 0.728 (Private LB:0.701)\n\nI used the same train/valid given by the host. (Failure to devise a better CV may have been the cause of the shake down)\n\n## Tile\n\n### Features\n\n- Config features\n- Node features averaged over all nodes\n### Models\n\nI trained the pointwise LightGBM model.\n- Target: Normalized rankings(0~1)\n- Loss: MAE Loss\n- Public LB(only tile): 0.198 (Private LB(only tile): 0.195)",
    "2529198": "Quite different methodology, thanks for sharing.",
    "2529267": "Thanks for sharing.",
    "2529893": "Congratulations on achieving 11th position. Interesting to observe that you have used LightGBM.",
    "2529935": "Good work!\n\nwhat surprises me is that MAE Loss works!",
    "2530918": "Congratulations on the solution, really interesting approach! Also, very smart observation:\n>Because genetic algorithms are used in the search, the more frequently a config pattern appears, the faster the config runtime will tend to be.",
    "2531139": "this is so cool",
    "2531157": "Great work. Congrats",
    "2537326": "Congratulations on achieving 11th position, would you mind making the notebook public? you have such a unique solution to the problem it would be of help to the community and have you tried XGB ranker too ?"
  },
  "source": "meta"
}