{
  "id": 436224,
  "title": "Power BI EDA - [01.09.2023, Topology Analysis & Runtime analysis]",
  "url": "/competitions/predict-ai-model-runtime/discussion/436224",
  "author_name": "",
  "post_date": "2023-09-01T12:45:12.259773400Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<h1>POWER BI EDA [01.09.2023 STATE]</h1>\n<p>Hey everyone!<br>\nI've decided to support our competition activities by performing Exploratory Data Analysis on our dataset. It's not fully comprehensive yet, as I've mainly focused on NLP in Layout right now, but I'll try to update this post regularly with more parts of the analysis! :D</p>\n<p>I did the analysis, along with the visualisations in Power Bi, and for the Topology Analysis I used python to do bulk calculations on the files, concerning the nodes.</p>\n<p>Additionally, all the output that came out of my analysis will be posted in a dataset on kaggle (including where I converted the files to parquet, so as not to clutter the forum unnecessarily with more datasets ).</p>\n<p>Below is a brief description of the initial EDA:</p>\n<p><strong>Metrics Included:</strong></p>\n<p><strong>Edge Density</strong><br>\nThis metric provides insights into how densely the nodes in each model are connected. Lower values often suggest sparser connections, possibly indicating less complexity.</p>\n<p><strong>Node Count</strong><br>\nThe Node Count metric reveals the number of nodes present in each model. Higher numbers often indicate more complex models.</p>\n<p><strong>Edge Count</strong><br>\nEdge Count shows the number of edges or connections between nodes in the model. A higher count may suggest a greater degree of interconnectedness.</p>\n<p><strong>Average Degree</strong><br>\nThis measure shows the average number of connections for nodes in a model. A higher average degree might suggest a more intricate internal structure.</p>\n<p><strong>Common Nodes</strong><br>\nThe number of nodes that are common between the source and target nodes in a model. Higher numbers could indicate redundancy or shared structures.</p>\n<p><strong>Possible Edges</strong><br>\nThis is a theoretical metric showing the maximum number of edges that could exist in a model. A higher number of possible edges with a lower actual edge count could indicate room for optimization.</p>\n<p>For a deeper dive into each metric and model, please refer to the full <strong>GoogleSpeedEDA 1.0 report</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F74c114db2861bab299ab750bb83eb05b%2FGoogleSpeedEDA_1.0-1.png?generation=1693572051005617&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F22db196cb192efe108a4e59ee17351d8%2FGoogleSpeedEDA_1.0-2.png?generation=1693572064088152&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fd9835dfc7c0534a690029b4512b73c47%2FGoogleSpeedEDA_1.0-3.png?generation=1693572073034337&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F0c72a19ca952df2bc77154c374f4717d%2FGoogleSpeedEDA_1.0-4.png?generation=1693572081812752&amp;alt=media\" alt=\"\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F263f634214bf5e739577c96ab08fca73%2FGoogleSpeedEDA_1.0-5.png?generation=1693572091738217&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fde194e8bf14b7ca2d6431f5c849116b6%2FGoogleSpeedEDA_1.0-6.png?generation=1693572102830262&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fa1eaf9f1b16eb31eda2f38111cdcf7a9%2FGoogleSpeedEDA_1.0-7.png?generation=1693572113431852&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fa5933c9ee582cff0d8076d72eacf3c59%2FGoogleSpeedEDA_1.0-8.png?generation=1693572122382197&amp;alt=media\" alt=\"\"></p>\n<p>Below is the link to the Kaggle dataset, where I will be storing the output of my analysis:<br>\n<a href=\"https://www.kaggle.com/datasets/sebastianbarry55/google-competition-dataset-in-parquet-and-csv\" target=\"_blank\">https://www.kaggle.com/datasets/sebastianbarry55/google-competition-dataset-in-parquet-and-csv</a></p>\n<p>If you have any questions, feel free to ask!<br>\nHope you enjoy! </p>",
  "messages": [
    {
      "id": "2418639",
      "postDate": "09/01/2023 12:45:12",
      "content": "<h1>POWER BI EDA [01.09.2023 STATE]</h1>\n<p>Hey everyone!<br>\nI've decided to support our competition activities by performing Exploratory Data Analysis on our dataset. It's not fully comprehensive yet, as I've mainly focused on NLP in Layout right now, but I'll try to update this post regularly with more parts of the analysis! :D</p>\n<p>I did the analysis, along with the visualisations in Power Bi, and for the Topology Analysis I used python to do bulk calculations on the files, concerning the nodes.</p>\n<p>Additionally, all the output that came out of my analysis will be posted in a dataset on kaggle (including where I converted the files to parquet, so as not to clutter the forum unnecessarily with more datasets ).</p>\n<p>Below is a brief description of the initial EDA:</p>\n<p><strong>Metrics Included:</strong></p>\n<p><strong>Edge Density</strong><br>\nThis metric provides insights into how densely the nodes in each model are connected. Lower values often suggest sparser connections, possibly indicating less complexity.</p>\n<p><strong>Node Count</strong><br>\nThe Node Count metric reveals the number of nodes present in each model. Higher numbers often indicate more complex models.</p>\n<p><strong>Edge Count</strong><br>\nEdge Count shows the number of edges or connections between nodes in the model. A higher count may suggest a greater degree of interconnectedness.</p>\n<p><strong>Average Degree</strong><br>\nThis measure shows the average number of connections for nodes in a model. A higher average degree might suggest a more intricate internal structure.</p>\n<p><strong>Common Nodes</strong><br>\nThe number of nodes that are common between the source and target nodes in a model. Higher numbers could indicate redundancy or shared structures.</p>\n<p><strong>Possible Edges</strong><br>\nThis is a theoretical metric showing the maximum number of edges that could exist in a model. A higher number of possible edges with a lower actual edge count could indicate room for optimization.</p>\n<p>For a deeper dive into each metric and model, please refer to the full <strong>GoogleSpeedEDA 1.0 report</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F74c114db2861bab299ab750bb83eb05b%2FGoogleSpeedEDA_1.0-1.png?generation=1693572051005617&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F22db196cb192efe108a4e59ee17351d8%2FGoogleSpeedEDA_1.0-2.png?generation=1693572064088152&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fd9835dfc7c0534a690029b4512b73c47%2FGoogleSpeedEDA_1.0-3.png?generation=1693572073034337&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F0c72a19ca952df2bc77154c374f4717d%2FGoogleSpeedEDA_1.0-4.png?generation=1693572081812752&amp;alt=media\" alt=\"\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F263f634214bf5e739577c96ab08fca73%2FGoogleSpeedEDA_1.0-5.png?generation=1693572091738217&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fde194e8bf14b7ca2d6431f5c849116b6%2FGoogleSpeedEDA_1.0-6.png?generation=1693572102830262&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fa1eaf9f1b16eb31eda2f38111cdcf7a9%2FGoogleSpeedEDA_1.0-7.png?generation=1693572113431852&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fa5933c9ee582cff0d8076d72eacf3c59%2FGoogleSpeedEDA_1.0-8.png?generation=1693572122382197&amp;alt=media\" alt=\"\"></p>\n<p>Below is the link to the Kaggle dataset, where I will be storing the output of my analysis:<br>\n<a href=\"https://www.kaggle.com/datasets/sebastianbarry55/google-competition-dataset-in-parquet-and-csv\" target=\"_blank\">https://www.kaggle.com/datasets/sebastianbarry55/google-competition-dataset-in-parquet-and-csv</a></p>\n<p>If you have any questions, feel free to ask!<br>\nHope you enjoy! </p>",
      "rawMarkdown": "# POWER BI EDA [01.09.2023 STATE]\n\nHey everyone!\nI've decided to support our competition activities by performing Exploratory Data Analysis on our dataset. It's not fully comprehensive yet, as I've mainly focused on NLP in Layout right now, but I'll try to update this post regularly with more parts of the analysis! :D\n\nI did the analysis, along with the visualisations in Power Bi, and for the Topology Analysis I used python to do bulk calculations on the files, concerning the nodes.\n\nAdditionally, all the output that came out of my analysis will be posted in a dataset on kaggle (including where I converted the files to parquet, so as not to clutter the forum unnecessarily with more datasets ).\n\nBelow is a brief description of the initial EDA:\n\n**Metrics Included:**\n\n**Edge Density**\nThis metric provides insights into how densely the nodes in each model are connected. Lower values often suggest sparser connections, possibly indicating less complexity.\n\n**Node Count**\nThe Node Count metric reveals the number of nodes present in each model. Higher numbers often indicate more complex models.\n\n**Edge Count**\nEdge Count shows the number of edges or connections between nodes in the model. A higher count may suggest a greater degree of interconnectedness.\n\n**Average Degree**\nThis measure shows the average number of connections for nodes in a model. A higher average degree might suggest a more intricate internal structure.\n\n**Common Nodes**\nThe number of nodes that are common between the source and target nodes in a model. Higher numbers could indicate redundancy or shared structures.\n\n**Possible Edges**\nThis is a theoretical metric showing the maximum number of edges that could exist in a model. A higher number of possible edges with a lower actual edge count could indicate room for optimization.\n\n\nFor a deeper dive into each metric and model, please refer to the full **GoogleSpeedEDA 1.0 report**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F74c114db2861bab299ab750bb83eb05b%2FGoogleSpeedEDA_1.0-1.png?generation=1693572051005617&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F22db196cb192efe108a4e59ee17351d8%2FGoogleSpeedEDA_1.0-2.png?generation=1693572064088152&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fd9835dfc7c0534a690029b4512b73c47%2FGoogleSpeedEDA_1.0-3.png?generation=1693572073034337&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F0c72a19ca952df2bc77154c374f4717d%2FGoogleSpeedEDA_1.0-4.png?generation=1693572081812752&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F263f634214bf5e739577c96ab08fca73%2FGoogleSpeedEDA_1.0-5.png?generation=1693572091738217&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fde194e8bf14b7ca2d6431f5c849116b6%2FGoogleSpeedEDA_1.0-6.png?generation=1693572102830262&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fa1eaf9f1b16eb31eda2f38111cdcf7a9%2FGoogleSpeedEDA_1.0-7.png?generation=1693572113431852&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fa5933c9ee582cff0d8076d72eacf3c59%2FGoogleSpeedEDA_1.0-8.png?generation=1693572122382197&alt=media)\n\n\n\nBelow is the link to the Kaggle dataset, where I will be storing the output of my analysis:\nhttps://www.kaggle.com/datasets/sebastianbarry55/google-competition-dataset-in-parquet-and-csv\n\nIf you have any questions, feel free to ask!\nHope you enjoy!",
      "votes": null
    },
    {
      "id": "2419124",
      "postDate": "09/01/2023 17:35:21",
      "content": "<p>Thank you for sharing this!</p>",
      "rawMarkdown": "Thank you for sharing this!",
      "votes": null
    },
    {
      "id": "2426997",
      "postDate": "09/07/2023 03:16:48",
      "content": "<p>Thank you for sharing this! 😘</p>",
      "rawMarkdown": "Thank you for sharing this! 😘",
      "votes": null
    },
    {
      "id": "2427961",
      "postDate": "09/07/2023 15:03:29",
      "content": "<p>Good work!</p>",
      "rawMarkdown": "Good work!",
      "votes": null
    },
    {
      "id": "2428528",
      "postDate": "09/08/2023 01:12:21",
      "content": "<p>The NLP graphs are truly huge. I wonder if we can even make models to decently converge in the first place……</p>",
      "rawMarkdown": "The NLP graphs are truly huge. I wonder if we can even make models to decently converge in the first place......",
      "votes": null
    },
    {
      "id": "2428625",
      "postDate": "09/08/2023 04:00:53",
      "content": "<p>This is our attempt on dealing with such large graphs: <a href=\"https://arxiv.org/pdf/2305.12322.pdf\" target=\"_blank\">https://arxiv.org/pdf/2305.12322.pdf</a></p>",
      "rawMarkdown": "This is our attempt on dealing with such large graphs: https://arxiv.org/pdf/2305.12322.pdf",
      "votes": null
    },
    {
      "id": "2429907",
      "postDate": "09/08/2023 22:38:29",
      "content": "<p>Thanks! Why are the hyper-parameters different between the GST used in the tpu_graphs paper and the original paper? Also, the GST paper reports up to 90% test OPA without specifying which collection, which seems to be much higher than what's reported in the tpu_graphs work (granted only Top-k error is reported, but some rough ballpark conversion is assumed here).</p>",
      "rawMarkdown": "Thanks! Why are the hyper-parameters different between the GST used in the tpu_graphs paper and the original paper? Also, the GST paper reports up to 90% test OPA without specifying which collection, which seems to be much higher than what's reported in the tpu_graphs work (granted only Top-k error is reported, but some rough ballpark conversion is assumed here).",
      "votes": null
    },
    {
      "id": "2429951",
      "postDate": "09/08/2023 23:49:03",
      "content": "<p>The XLA dataset used in the GST paper (Google internal) is different from the opensource dataset we use in this competition, so the numbers reported in the GST paper won't be an apple to apple comparison to what we get in this competition.</p>",
      "rawMarkdown": "The XLA dataset used in the GST paper (Google internal) is different from the opensource dataset we use in this competition, so the numbers reported in the GST paper won't be an apple to apple comparison to what we get in this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2419124,
      "author_name": "mangpophothilimthana",
      "author_url": "",
      "post_date": "09/01/2023 17:35:21",
      "content": "<p>Thank you for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2426997,
      "author_name": "lillteicebear",
      "author_url": "",
      "post_date": "09/07/2023 03:16:48",
      "content": "<p>Thank you for sharing this! 😘</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2427961,
      "author_name": "aparunov",
      "author_url": "",
      "post_date": "09/07/2023 15:03:29",
      "content": "<p>Good work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2428528,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "09/08/2023 01:12:21",
      "content": "<p>The NLP graphs are truly huge. I wonder if we can even make models to decently converge in the first place……</p>",
      "votes": null,
      "replies": [
        {
          "id": 2428625,
          "author_name": "mangpophothilimthana",
          "author_url": "",
          "post_date": "09/08/2023 04:00:53",
          "content": "<p>This is our attempt on dealing with such large graphs: <a href=\"https://arxiv.org/pdf/2305.12322.pdf\" target=\"_blank\">https://arxiv.org/pdf/2305.12322.pdf</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2429907,
              "author_name": "alexanderliao",
              "author_url": "",
              "post_date": "09/08/2023 22:38:29",
              "content": "<p>Thanks! Why are the hyper-parameters different between the GST used in the tpu_graphs paper and the original paper? Also, the GST paper reports up to 90% test OPA without specifying which collection, which seems to be much higher than what's reported in the tpu_graphs work (granted only Top-k error is reported, but some rough ballpark conversion is assumed here).</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2429951,
                  "author_name": "mangpophothilimthana",
                  "author_url": "",
                  "post_date": "09/08/2023 23:49:03",
                  "content": "<p>The XLA dataset used in the GST paper (Google internal) is different from the opensource dataset we use in this competition, so the numbers reported in the GST paper won't be an apple to apple comparison to what we get in this competition.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2418639": "# POWER BI EDA [01.09.2023 STATE]\n\nHey everyone!\nI've decided to support our competition activities by performing Exploratory Data Analysis on our dataset. It's not fully comprehensive yet, as I've mainly focused on NLP in Layout right now, but I'll try to update this post regularly with more parts of the analysis! :D\n\nI did the analysis, along with the visualisations in Power Bi, and for the Topology Analysis I used python to do bulk calculations on the files, concerning the nodes.\n\nAdditionally, all the output that came out of my analysis will be posted in a dataset on kaggle (including where I converted the files to parquet, so as not to clutter the forum unnecessarily with more datasets ).\n\nBelow is a brief description of the initial EDA:\n\n**Metrics Included:**\n\n**Edge Density**\nThis metric provides insights into how densely the nodes in each model are connected. Lower values often suggest sparser connections, possibly indicating less complexity.\n\n**Node Count**\nThe Node Count metric reveals the number of nodes present in each model. Higher numbers often indicate more complex models.\n\n**Edge Count**\nEdge Count shows the number of edges or connections between nodes in the model. A higher count may suggest a greater degree of interconnectedness.\n\n**Average Degree**\nThis measure shows the average number of connections for nodes in a model. A higher average degree might suggest a more intricate internal structure.\n\n**Common Nodes**\nThe number of nodes that are common between the source and target nodes in a model. Higher numbers could indicate redundancy or shared structures.\n\n**Possible Edges**\nThis is a theoretical metric showing the maximum number of edges that could exist in a model. A higher number of possible edges with a lower actual edge count could indicate room for optimization.\n\n\nFor a deeper dive into each metric and model, please refer to the full **GoogleSpeedEDA 1.0 report**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F74c114db2861bab299ab750bb83eb05b%2FGoogleSpeedEDA_1.0-1.png?generation=1693572051005617&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F22db196cb192efe108a4e59ee17351d8%2FGoogleSpeedEDA_1.0-2.png?generation=1693572064088152&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fd9835dfc7c0534a690029b4512b73c47%2FGoogleSpeedEDA_1.0-3.png?generation=1693572073034337&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F0c72a19ca952df2bc77154c374f4717d%2FGoogleSpeedEDA_1.0-4.png?generation=1693572081812752&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2F263f634214bf5e739577c96ab08fca73%2FGoogleSpeedEDA_1.0-5.png?generation=1693572091738217&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fde194e8bf14b7ca2d6431f5c849116b6%2FGoogleSpeedEDA_1.0-6.png?generation=1693572102830262&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fa1eaf9f1b16eb31eda2f38111cdcf7a9%2FGoogleSpeedEDA_1.0-7.png?generation=1693572113431852&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14649780%2Fa5933c9ee582cff0d8076d72eacf3c59%2FGoogleSpeedEDA_1.0-8.png?generation=1693572122382197&alt=media)\n\n\n\nBelow is the link to the Kaggle dataset, where I will be storing the output of my analysis:\nhttps://www.kaggle.com/datasets/sebastianbarry55/google-competition-dataset-in-parquet-and-csv\n\nIf you have any questions, feel free to ask!\nHope you enjoy!",
    "2419124": "Thank you for sharing this!",
    "2426997": "Thank you for sharing this! 😘",
    "2427961": "Good work!",
    "2428528": "The NLP graphs are truly huge. I wonder if we can even make models to decently converge in the first place......",
    "2428625": "This is our attempt on dealing with such large graphs: https://arxiv.org/pdf/2305.12322.pdf",
    "2429907": "Thanks! Why are the hyper-parameters different between the GST used in the tpu_graphs paper and the original paper? Also, the GST paper reports up to 90% test OPA without specifying which collection, which seems to be much higher than what's reported in the tpu_graphs work (granted only Top-k error is reported, but some rough ballpark conversion is assumed here).",
    "2429951": "The XLA dataset used in the GST paper (Google internal) is different from the opensource dataset we use in this competition, so the numbers reported in the GST paper won't be an apple to apple comparison to what we get in this competition."
  },
  "source": "meta"
}