{
  "id": 436741,
  "title": "[place holder] my experimental results",
  "url": "/competitions/predict-ai-model-runtime/discussion/436741",
  "author_name": "hengck23",
  "post_date": "2023-09-04T00:31:38.926000",
  "votes": 20,
  "comment_count": 14,
  "views": 0,
  "content": "<p>to be updated as i progress along.<br>\nthe first steps are to repeat papers results.<br>\ne.g. GraphSAGE, segmented graph training …etc</p>",
  "messages": [
    {
      "id": 2422354,
      "postDate": "2023-09-04T00:31:38.927Z",
      "content": "<p>to be updated as i progress along.<br>\nthe first steps are to repeat papers results.<br>\ne.g. GraphSAGE, segmented graph training …etc</p>",
      "rawMarkdown": "to be updated as i progress along.\nthe first steps are to repeat papers results.\ne.g. GraphSAGE, segmented graph training ...etc",
      "votes": 17
    },
    {
      "id": 2422377,
      "postDate": "2023-09-04T01:10:03.977Z",
      "content": "<p>Suggestion for <em>pure</em> beginner</p>\n<p>[1] start with GraphSAGE (there are many resources)</p>\n<ul>\n<li>google for blogs, youtube, code from scratch, pytorch Graph NN lib (e.g. torch geometric), kaggle kernels/notebook</li>\n<li>understand how it work, create some mini project (finished in a days or two) and train model on the dataset in the tutorial you found<br>\n(e.g. <a href=\"https://towardsdatascience.com/a-comprehensive-case-study-of-graphsage-algorithm-with-hands-on-experience-using-pytorchgeometric-6fc631ab1067\" target=\"_blank\">https://towardsdatascience.com/a-comprehensive-case-study-of-graphsage-algorithm-with-hands-on-experience-using-pytorchgeometric-6fc631ab1067</a>)</li>\n</ul>\n<p>[2] learned generally about computational graph compilation<br>\n(faster way is to read blogs and youtube)</p>\n<p>[3] go to tensorflow XLA and HLO website and read generally</p>\n<p>[4]  Read competition papers and repeat results (use the github baseline to cross check your paper understanding and implementation)</p>\n<p>if the workload is too heavy, find some buddy … study in a group …<br>\n[2],[3] are domain knowledge, so just read briefly.<br>\nfocus on [1],[4]</p>",
      "rawMarkdown": "Suggestion for *pure* beginner\n\n[1] start with GraphSAGE (there are many resources)\n- google for blogs, youtube, code from scratch, pytorch Graph NN lib (e.g. torch geometric), kaggle kernels/notebook\n- understand how it work, create some mini project (finished in a days or two) and train model on the dataset in the tutorial you found\n(e.g. https://towardsdatascience.com/a-comprehensive-case-study-of-graphsage-algorithm-with-hands-on-experience-using-pytorchgeometric-6fc631ab1067)\n\n[2] learned generally about computational graph compilation\n(faster way is to read blogs and youtube)\n\n[3] go to tensorflow XLA and HLO website and read generally\n\n[4]  Read competition papers and repeat results (use the github baseline to cross check your paper understanding and implementation)\n\nif the workload is too heavy, find some buddy ... study in a group ...\n[2],[3] are domain knowledge, so just read briefly.\nfocus on [1],[4]\n",
      "votes": 7,
      "replies": [
        {
          "id": 2425903,
          "postDate": "2023-09-06T09:09:31.270Z",
          "content": "<p>To start with GNNs, I just started with that blog and it's pretty nice:</p>\n<p><a href=\"https://mlabonne.github.io/blog/posts/2022-03-09-Graph_Attention_Network.html\" target=\"_blank\">https://mlabonne.github.io/blog/posts/2022-03-09-Graph_Attention_Network.html</a></p>",
          "rawMarkdown": "To start with GNNs, I just started with that blog and it's pretty nice:\n\nhttps://mlabonne.github.io/blog/posts/2022-03-09-Graph_Attention_Network.html\n",
          "votes": 2,
          "replies": [
            {
              "id": 2466253,
              "postDate": "2023-10-03T17:06:11.730Z",
              "content": "<p>have you did any changes or as it is you wrote</p>",
              "rawMarkdown": "have you did any changes or as it is you wrote",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2515323,
      "postDate": "2023-11-06T21:10:03.217Z",
      "content": "<p>xla-default  cross validation results:</p>\n<hr>\n<ol>\n<li>tf2_bert_pretrain_dynamic_batch_size.npz SignificanceResult(statistic=0.41437652242798656, pvalue=0.0)</li>\n<li>bert_pretraining.4x4.fp16.npz SignificanceResult(statistic=0.5085575478666131, pvalue=0.0)</li>\n<li>mlperf_bert_batch_24_2x2.npz SignificanceResult(statistic=0.5105886065939758, pvalue=0.0)</li>\n<li>inception_v3_batch_128_train.npz SignificanceResult(statistic=0.6576461322736008, pvalue=0.0)</li>\n<li>resnet_v1_50_official_batch_128_bf16.npz SignificanceResult(statistic=0.503077787092724, pvalue=0.0)</li>\n<li>resnet50.4x4.fp16.npz SignificanceResult(statistic=0.566724142955186, pvalue=0.0)</li>\n<li>unet_3d.4x4.bf16.npz SignificanceResult(statistic=0.12675177953744027, pvalue=3.864550340557695e-17)</li>\n</ol>\n<p>i haven't made submission yet, not sure if the results are correct.<br>\nmy models are modification from [1], the reference 3xSAGEConv with my own gnn normalisation</p>\n<p>[1]Learning Large Graph Property Prediction via Graph Segment Training<br>\n<a href=\"https://github.com/kaidic/GST/tree/main\" target=\"_blank\">https://github.com/kaidic/GST/tree/main</a></p>",
      "rawMarkdown": "xla-default  cross validation results:\n\n-----\n\n0. tf2_bert_pretrain_dynamic_batch_size.npz SignificanceResult(statistic=0.41437652242798656, pvalue=0.0)\n1. bert_pretraining.4x4.fp16.npz SignificanceResult(statistic=0.5085575478666131, pvalue=0.0)\n2. mlperf_bert_batch_24_2x2.npz SignificanceResult(statistic=0.5105886065939758, pvalue=0.0)\n3. inception_v3_batch_128_train.npz SignificanceResult(statistic=0.6576461322736008, pvalue=0.0)\n4. resnet_v1_50_official_batch_128_bf16.npz SignificanceResult(statistic=0.503077787092724, pvalue=0.0)\n5. resnet50.4x4.fp16.npz SignificanceResult(statistic=0.566724142955186, pvalue=0.0)\n6. unet_3d.4x4.bf16.npz SignificanceResult(statistic=0.12675177953744027, pvalue=3.864550340557695e-17)\n\ni haven't made submission yet, not sure if the results are correct.\nmy models are modification from [1], the reference 3xSAGEConv with my own gnn normalisation\n\n[1]Learning Large Graph Property Prediction via Graph Segment Training\nhttps://github.com/kaidic/GST/tree/main",
      "votes": 2,
      "replies": [
        {
          "id": 2516876,
          "postDate": "2023-11-08T03:05:29.693Z",
          "content": "<p>Are these kendall Tau score calculated using entire config runtime?? Amazing scores. </p>",
          "rawMarkdown": "Are these kendall Tau score calculated using entire config runtime?? Amazing scores. ",
          "replies": [
            {
              "id": 2516880,
              "postDate": "2023-11-08T03:14:04.093Z",
              "content": "<p>I assume these are Pearson, which would mean the Kendalls would be even higher</p>",
              "rawMarkdown": "I assume these are Pearson, which would mean the Kendalls would be even higher",
              "votes": 1
            },
            {
              "id": 2517172,
              "postDate": "2023-11-08T09:06:48.237Z",
              "content": "<p>\"Are these kendall Tau score calculated using entire config runtime?? Amazing scores.\"<br>\nyes.</p>\n<p>i think these are typical scores.<br>\noverall, i got current lb 0.675 with this</p>\n<pre><code>    truth = result\n    predict = result\n    corr, _ = (predict, truth)\n\n    (, f,f,)\n</code></pre>",
              "rawMarkdown": "\"Are these kendall Tau score calculated using entire config runtime?? Amazing scores.\"\nyes.\n\ni think these are typical scores.\noverall, i got current lb 0.675 with this\n\n```\n\n\ttruth = result[i].truth\n\tpredict = result[i].predict\n\tcorr, _ = kendalltau(predict, truth)\n\t\n\tprint(i, f'{corr:0.6f}',f'{result[i][\"file\"]:128s}',)\n\n\n\n```",
              "votes": 2
            },
            {
              "id": 2517176,
              "postDate": "2023-11-08T09:09:30.363Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. Impressive Cv and nice jump on leaderboard. </p>",
              "rawMarkdown": "Thanks @hengck23. Impressive Cv and nice jump on leaderboard. ",
              "votes": 2
            },
            {
              "id": 2520895,
              "postDate": "2023-11-11T09:10:52.990Z",
              "content": "<p>normalisation  normalisation normalisation normalisation …. 😀</p>",
              "rawMarkdown": "normalisation  normalisation normalisation normalisation .... 😀"
            }
          ]
        }
      ]
    },
    {
      "id": 2424524,
      "postDate": "2023-09-05T10:54:13.193Z",
      "content": "<p>Thanks for sharing, very useful information</p>",
      "rawMarkdown": "Thanks for sharing, very useful information"
    },
    {
      "id": 2422358,
      "postDate": "2023-09-04T00:39:00.763Z",
      "content": "<p>i am also working on a tool to visualize the results, e.g.  (like dot for graphviz, or html) to draw the best configuration + tensor graph.</p>",
      "rawMarkdown": "i am also working on a tool to visualize the results, e.g.  (like dot for graphviz, or html) to draw the best configuration + tensor graph."
    },
    {
      "id": 2422357,
      "postDate": "2023-09-04T00:36:13.233Z",
      "content": "<p>here are some good papers related to model runtime estimation for HLO:</p>\n<p>1)  OPTIMIZING DNN COMPILATION FOR DISTRIBUTED TRAINING WITH JOINT OP AND TENSOR FUSION<br>\n<a href=\"https://arxiv.org/pdf/2209.12769.pdf\" target=\"_blank\">https://arxiv.org/pdf/2209.12769.pdf</a></p>\n<ul>\n<li>not quite the same as our competition, but a good source for related references<br>\n\"A Fused Op Estimator is built based on a graph neural network (GNN) model to predict the execution time of fused ops. An efficient simulator is created to estimate the end-to-end execution time of a distributed DNN training graph using the Fused Op Estimator, and serves as a cost model to our search algorithm\"</li>\n</ul>",
      "rawMarkdown": "here are some good papers related to model runtime estimation for HLO:\n\n1)  OPTIMIZING DNN COMPILATION FOR DISTRIBUTED TRAINING WITH JOINT OP AND TENSOR FUSION\nhttps://arxiv.org/pdf/2209.12769.pdf\n\n- not quite the same as our competition, but a good source for related references\n\"A Fused Op Estimator is built based on a graph neural network (GNN) model to predict the execution time of fused ops. An efficient simulator is created to estimate the end-to-end execution time of a distributed DNN training graph using the Fused Op Estimator, and serves as a cost model to our search algorithm\"",
      "replies": [
        {
          "id": 2422369,
          "postDate": "2023-09-04T00:57:00.037Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7db639da289e00a67fac54c8e855ff75%2FSelection_999(3038).png?generation=1693789004947926&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"http://mlforsystems.org/assets/papers/neurips2019/learned_tpu_kaufman_2019.pdf\" target=\"_blank\">http://mlforsystems.org/assets/papers/neurips2019/learned_tpu_kaufman_2019.pdf</a><br>\nLearned TPU Cost Model for XLA Tensor Programs</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7db639da289e00a67fac54c8e855ff75%2FSelection_999(3038).png?generation=1693789004947926&alt=media)\n\nhttp://mlforsystems.org/assets/papers/neurips2019/learned_tpu_kaufman_2019.pdf\nLearned TPU Cost Model for XLA Tensor Programs",
          "votes": 2,
          "replies": [
            {
              "id": 2433235,
              "postDate": "2023-09-11T12:26:32.367Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2422377,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-09-04T01:10:03.977000",
      "content": "<p>Suggestion for <em>pure</em> beginner</p>\n<p>[1] start with GraphSAGE (there are many resources)</p>\n<ul>\n<li>google for blogs, youtube, code from scratch, pytorch Graph NN lib (e.g. torch geometric), kaggle kernels/notebook</li>\n<li>understand how it work, create some mini project (finished in a days or two) and train model on the dataset in the tutorial you found<br>\n(e.g. <a href=\"https://towardsdatascience.com/a-comprehensive-case-study-of-graphsage-algorithm-with-hands-on-experience-using-pytorchgeometric-6fc631ab1067\" target=\"_blank\">https://towardsdatascience.com/a-comprehensive-case-study-of-graphsage-algorithm-with-hands-on-experience-using-pytorchgeometric-6fc631ab1067</a>)</li>\n</ul>\n<p>[2] learned generally about computational graph compilation<br>\n(faster way is to read blogs and youtube)</p>\n<p>[3] go to tensorflow XLA and HLO website and read generally</p>\n<p>[4]  Read competition papers and repeat results (use the github baseline to cross check your paper understanding and implementation)</p>\n<p>if the workload is too heavy, find some buddy … study in a group …<br>\n[2],[3] are domain knowledge, so just read briefly.<br>\nfocus on [1],[4]</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2425903,
          "author_name": "FabienDaniel",
          "author_url": "",
          "post_date": "2023-09-06T09:09:31.270000",
          "content": "<p>To start with GNNs, I just started with that blog and it's pretty nice:</p>\n<p><a href=\"https://mlabonne.github.io/blog/posts/2022-03-09-Graph_Attention_Network.html\" target=\"_blank\">https://mlabonne.github.io/blog/posts/2022-03-09-Graph_Attention_Network.html</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2466253,
              "author_name": "Satheesh Bhukya",
              "author_url": "",
              "post_date": "2023-10-03T17:06:11.730000",
              "content": "<p>have you did any changes or as it is you wrote</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2515323,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-11-06T21:10:03.217000",
      "content": "<p>xla-default  cross validation results:</p>\n<hr>\n<ol>\n<li>tf2_bert_pretrain_dynamic_batch_size.npz SignificanceResult(statistic=0.41437652242798656, pvalue=0.0)</li>\n<li>bert_pretraining.4x4.fp16.npz SignificanceResult(statistic=0.5085575478666131, pvalue=0.0)</li>\n<li>mlperf_bert_batch_24_2x2.npz SignificanceResult(statistic=0.5105886065939758, pvalue=0.0)</li>\n<li>inception_v3_batch_128_train.npz SignificanceResult(statistic=0.6576461322736008, pvalue=0.0)</li>\n<li>resnet_v1_50_official_batch_128_bf16.npz SignificanceResult(statistic=0.503077787092724, pvalue=0.0)</li>\n<li>resnet50.4x4.fp16.npz SignificanceResult(statistic=0.566724142955186, pvalue=0.0)</li>\n<li>unet_3d.4x4.bf16.npz SignificanceResult(statistic=0.12675177953744027, pvalue=3.864550340557695e-17)</li>\n</ol>\n<p>i haven't made submission yet, not sure if the results are correct.<br>\nmy models are modification from [1], the reference 3xSAGEConv with my own gnn normalisation</p>\n<p>[1]Learning Large Graph Property Prediction via Graph Segment Training<br>\n<a href=\"https://github.com/kaidic/GST/tree/main\" target=\"_blank\">https://github.com/kaidic/GST/tree/main</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2516876,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-11-08T03:05:29.693000",
          "content": "<p>Are these kendall Tau score calculated using entire config runtime?? Amazing scores. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2516880,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2023-11-08T03:14:04.093000",
              "content": "<p>I assume these are Pearson, which would mean the Kendalls would be even higher</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2517172,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-11-08T09:06:48.237000",
              "content": "<p>\"Are these kendall Tau score calculated using entire config runtime?? Amazing scores.\"<br>\nyes.</p>\n<p>i think these are typical scores.<br>\noverall, i got current lb 0.675 with this</p>\n<pre><code>    truth = result\n    predict = result\n    corr, _ = (predict, truth)\n\n    (, f,f,)\n</code></pre>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2517176,
              "author_name": "Nischay Dhankhar",
              "author_url": "",
              "post_date": "2023-11-08T09:09:30.363000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. Impressive Cv and nice jump on leaderboard. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2520895,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-11-11T09:10:52.990000",
              "content": "<p>normalisation  normalisation normalisation normalisation …. 😀</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2424524,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-09-05T10:54:13.193000",
      "content": "<p>Thanks for sharing, very useful information</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2422358,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-09-04T00:39:00.763000",
      "content": "<p>i am also working on a tool to visualize the results, e.g.  (like dot for graphviz, or html) to draw the best configuration + tensor graph.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2422357,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-09-04T00:36:13.233000",
      "content": "<p>here are some good papers related to model runtime estimation for HLO:</p>\n<p>1)  OPTIMIZING DNN COMPILATION FOR DISTRIBUTED TRAINING WITH JOINT OP AND TENSOR FUSION<br>\n<a href=\"https://arxiv.org/pdf/2209.12769.pdf\" target=\"_blank\">https://arxiv.org/pdf/2209.12769.pdf</a></p>\n<ul>\n<li>not quite the same as our competition, but a good source for related references<br>\n\"A Fused Op Estimator is built based on a graph neural network (GNN) model to predict the execution time of fused ops. An efficient simulator is created to estimate the end-to-end execution time of a distributed DNN training graph using the Fused Op Estimator, and serves as a cost model to our search algorithm\"</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 2422369,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-09-04T00:57:00.037000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7db639da289e00a67fac54c8e855ff75%2FSelection_999(3038).png?generation=1693789004947926&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"http://mlforsystems.org/assets/papers/neurips2019/learned_tpu_kaufman_2019.pdf\" target=\"_blank\">http://mlforsystems.org/assets/papers/neurips2019/learned_tpu_kaufman_2019.pdf</a><br>\nLearned TPU Cost Model for XLA Tensor Programs</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2433235,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-11T12:26:32.367000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2422354": "to be updated as i progress along.\nthe first steps are to repeat papers results.\ne.g. GraphSAGE, segmented graph training ...etc",
    "2422377": "Suggestion for *pure* beginner\n\n[1] start with GraphSAGE (there are many resources)\n- google for blogs, youtube, code from scratch, pytorch Graph NN lib (e.g. torch geometric), kaggle kernels/notebook\n- understand how it work, create some mini project (finished in a days or two) and train model on the dataset in the tutorial you found\n(e.g. https://towardsdatascience.com/a-comprehensive-case-study-of-graphsage-algorithm-with-hands-on-experience-using-pytorchgeometric-6fc631ab1067)\n\n[2] learned generally about computational graph compilation\n(faster way is to read blogs and youtube)\n\n[3] go to tensorflow XLA and HLO website and read generally\n\n[4]  Read competition papers and repeat results (use the github baseline to cross check your paper understanding and implementation)\n\nif the workload is too heavy, find some buddy ... study in a group ...\n[2],[3] are domain knowledge, so just read briefly.\nfocus on [1],[4]\n",
    "2515323": "xla-default  cross validation results:\n\n-----\n\n0. tf2_bert_pretrain_dynamic_batch_size.npz SignificanceResult(statistic=0.41437652242798656, pvalue=0.0)\n1. bert_pretraining.4x4.fp16.npz SignificanceResult(statistic=0.5085575478666131, pvalue=0.0)\n2. mlperf_bert_batch_24_2x2.npz SignificanceResult(statistic=0.5105886065939758, pvalue=0.0)\n3. inception_v3_batch_128_train.npz SignificanceResult(statistic=0.6576461322736008, pvalue=0.0)\n4. resnet_v1_50_official_batch_128_bf16.npz SignificanceResult(statistic=0.503077787092724, pvalue=0.0)\n5. resnet50.4x4.fp16.npz SignificanceResult(statistic=0.566724142955186, pvalue=0.0)\n6. unet_3d.4x4.bf16.npz SignificanceResult(statistic=0.12675177953744027, pvalue=3.864550340557695e-17)\n\ni haven't made submission yet, not sure if the results are correct.\nmy models are modification from [1], the reference 3xSAGEConv with my own gnn normalisation\n\n[1]Learning Large Graph Property Prediction via Graph Segment Training\nhttps://github.com/kaidic/GST/tree/main",
    "2424524": "Thanks for sharing, very useful information",
    "2422358": "i am also working on a tool to visualize the results, e.g.  (like dot for graphviz, or html) to draw the best configuration + tensor graph.",
    "2422357": "here are some good papers related to model runtime estimation for HLO:\n\n1)  OPTIMIZING DNN COMPILATION FOR DISTRIBUTED TRAINING WITH JOINT OP AND TENSOR FUSION\nhttps://arxiv.org/pdf/2209.12769.pdf\n\n- not quite the same as our competition, but a good source for related references\n\"A Fused Op Estimator is built based on a graph neural network (GNN) model to predict the execution time of fused ops. An efficient simulator is created to estimate the end-to-end execution time of a distributed DNN training graph using the Fused Op Estimator, and serves as a cost model to our search algorithm\""
  }
}