{
  "id": 435640,
  "title": "Can we extract additional features from the graph?",
  "url": "/competitions/predict-ai-model-runtime/discussion/435640",
  "author_name": "Peiyuan Liao",
  "post_date": "2023-08-30T09:01:01.295000",
  "votes": 14,
  "comment_count": 16,
  "views": 0,
  "content": "<p>It seems like the training data is generated by the featurizer here: <a href=\"https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/featurizers.h#L542\" target=\"_blank\">https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/featurizers.h#L542</a> . Is it possible &amp; are we allowed to modify this to generate additional features?</p>",
  "messages": [
    {
      "id": 2415356,
      "postDate": "2023-08-30T09:01:01.297Z",
      "content": "<p>It seems like the training data is generated by the featurizer here: <a href=\"https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/featurizers.h#L542\" target=\"_blank\">https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/featurizers.h#L542</a> . Is it possible &amp; are we allowed to modify this to generate additional features?</p>",
      "rawMarkdown": "It seems like the training data is generated by the featurizer here: https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/featurizers.h#L542 . Is it possible & are we allowed to modify this to generate additional features?",
      "votes": 11
    },
    {
      "id": 2427023,
      "postDate": "2023-09-07T03:59:09.477Z",
      "content": "<p>We recently uploaded an additional <code>pb</code> directory containing raw graph data! You can find more information about how to extract additional graph features or extract them differently in Section \"Custom Feature Extraction\" at the bottom of the data tab.</p>",
      "rawMarkdown": "We recently uploaded an additional `pb` directory containing raw graph data! You can find more information about how to extract additional graph features or extract them differently in Section \"Custom Feature Extraction\" at the bottom of the data tab.",
      "votes": 2,
      "replies": [
        {
          "id": 2428525,
          "postDate": "2023-09-08T01:06:04.597Z",
          "content": "<p>This is wonderful! Thank you so much!</p>",
          "rawMarkdown": "This is wonderful! Thank you so much!"
        },
        {
          "id": 2431821,
          "postDate": "2023-09-10T12:39:31.637Z",
          "content": "<p><a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a>, is it correct to say that the information contained in a pb file is the same as the corresponding npz file?</p>\n<p>i.e., the pb file does not contain something which cannot be inferred from the npz file and supporting guidance on how to interpret the npz data?</p>",
          "rawMarkdown": "@mangpophothilimthana, is it correct to say that the information contained in a pb file is the same as the corresponding npz file?\n\ni.e., the pb file does not contain something which cannot be inferred from the npz file and supporting guidance on how to interpret the npz data?",
          "replies": [
            {
              "id": 2443334,
              "postDate": "2023-09-17T16:51:20.697Z",
              "content": "<p>Sorry for the late reply! I was on vacation for a week. </p>\n<p>Actually, a pb file does contain more information than npz file. A pb file stores a graph using XLA HLO representation in protobuf format. A node (instruction) in HLO representation has some attributes that npz data doesn't include. In the past, we analyzed the data and concluded that those attributes do not matter so much, but it wasn't a very extensive experiment. </p>",
              "rawMarkdown": "Sorry for the late reply! I was on vacation for a week. \n\nActually, a pb file does contain more information than npz file. A pb file stores a graph using XLA HLO representation in protobuf format. A node (instruction) in HLO representation has some attributes that npz data doesn't include. In the past, we analyzed the data and concluded that those attributes do not matter so much, but it wasn't a very extensive experiment. "
            },
            {
              "id": 2443700,
              "postDate": "2023-09-17T22:00:22.437Z",
              "content": "<p>It seems like the hlo.proto definition referred in the repository expired. Can we refer to the r2.14 one instead?</p>\n<p><a href=\"https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto\" target=\"_blank\">https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto</a></p>\n<p>Additionally, there might be a typo in the layout data explanation. Should it be \"input\" instead of \"intput\"?</p>",
              "rawMarkdown": "It seems like the hlo.proto definition referred in the repository expired. Can we refer to the r2.14 one instead?\n\nhttps://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto\n\nAdditionally, there might be a typo in the layout data explanation. Should it be \"input\" instead of \"intput\"?"
            },
            {
              "id": 2447383,
              "postDate": "2023-09-20T03:37:26.017Z",
              "content": "<p>Thank you for pointing these out! Yes, you can use <a href=\"https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto\" target=\"_blank\">https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto</a>. I'll update the description.</p>",
              "rawMarkdown": "Thank you for pointing these out! Yes, you can use https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto. I'll update the description."
            }
          ]
        },
        {
          "id": 2458281,
          "postDate": "2023-09-27T13:33:58.330Z",
          "content": "<p>when i load the pb file via tf1.15, return the  following bug </p>\n<p>String field 'tensorflow.NodeDef.input' contains invalid UTF-8 data when parsing a protocol buffer. Use the 'bytes' type if you intend to send raw bytes.</p>\n<p>could you please give some insturction how to load the pb file using tf1 or tf2 </p>",
          "rawMarkdown": "when i load the pb file via tf1.15, return the  following bug \n\nString field 'tensorflow.NodeDef.input' contains invalid UTF-8 data when parsing a protocol buffer. Use the 'bytes' type if you intend to send raw bytes.\n\ncould you please give some insturction how to load the pb file using tf1 or tf2 ",
          "replies": [
            {
              "id": 2460765,
              "postDate": "2023-09-29T05:14:59.333Z",
              "content": "<p>Can you try this: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/435640#2458281\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/435640#2458281</a> ?</p>",
              "rawMarkdown": "Can you try this: https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/435640#2458281 ?"
            }
          ]
        }
      ]
    },
    {
      "id": 2417889,
      "postDate": "2023-09-01T00:54:20.893Z",
      "content": "<p>Very good question! We want participants to be able to extract graph features their own ways if preferred. However, it is more complicated. We're working on the code that would allow participants to do so, but it will take us some time to get everything ready.</p>",
      "rawMarkdown": "Very good question! We want participants to be able to extract graph features their own ways if preferred. However, it is more complicated. We're working on the code that would allow participants to do so, but it will take us some time to get everything ready.",
      "votes": 2,
      "replies": [
        {
          "id": 2418156,
          "postDate": "2023-09-01T07:09:29.560Z",
          "content": "<p>Thank you for your comments! I also realized that the test set have anonymized names, does this imply that the workloads are anonymized as well?</p>",
          "rawMarkdown": "Thank you for your comments! I also realized that the test set have anonymized names, does this imply that the workloads are anonymized as well?"
        }
      ]
    },
    {
      "id": 2440299,
      "postDate": "2023-09-15T12:23:40.857Z",
      "content": "<p>Hi! I'm new to this field of AI/ML/DL I'm still studying it and gaining experiences through working with projects and some datasets. I'm wondering about feature extraction, from my prev knowledge I thought that most of the time ( maybe this is an exception case for this competition? ) we only do feature extraction for ML pipeline since ML models are good for structured data? I assumed that when we extract features we put them in columns so in the end we kinda get structured data format. however, for DL models I understand that it's kinda converted the feature extraction problem into feature 'learning' problem where the models are trained end-to-end and you have some layers to learn the features like how it works in CNN, the convolutions act as a feature extraction layer but the parameters of filters are being learned by gradients-based optimization like typical gradients descent of whatever. so I'm wodering since we're using GNNs, why do we still need to extract more features? can't we just try to make the architecture stronger/ more complexity and be able to learn the features for us like how we do in images domain? what's the inspiration and 'take away knowledge' here? and how do I know which project I need to extract some more features even though I might using DL models already which are trained end-to-end for unstructured data and features are assumed to be learned through this process already.</p>",
      "rawMarkdown": "Hi! I'm new to this field of AI/ML/DL I'm still studying it and gaining experiences through working with projects and some datasets. I'm wondering about feature extraction, from my prev knowledge I thought that most of the time ( maybe this is an exception case for this competition? ) we only do feature extraction for ML pipeline since ML models are good for structured data? I assumed that when we extract features we put them in columns so in the end we kinda get structured data format. however, for DL models I understand that it's kinda converted the feature extraction problem into feature 'learning' problem where the models are trained end-to-end and you have some layers to learn the features like how it works in CNN, the convolutions act as a feature extraction layer but the parameters of filters are being learned by gradients-based optimization like typical gradients descent of whatever. so I'm wodering since we're using GNNs, why do we still need to extract more features? can't we just try to make the architecture stronger/ more complexity and be able to learn the features for us like how we do in images domain? what's the inspiration and 'take away knowledge' here? and how do I know which project I need to extract some more features even though I might using DL models already which are trained end-to-end for unstructured data and features are assumed to be learned through this process already.",
      "replies": [
        {
          "id": 2443345,
          "postDate": "2023-09-17T16:58:59.453Z",
          "content": "<p>Your understanding is right. In DL domain, we rely on DL models to automatically learn to extract features from the raw input data, so humans don't have to do \"feature engineering\" ourselves. We use this approach for our baseline models. That's being said, some features maybe hard to learn especially if the model is not powerful enough (i.e. not big enough, too few layers, etc). Therefore, we're curious if some feature engineering may make the learning easier and make the model achieve better accuracy.</p>",
          "rawMarkdown": "Your understanding is right. In DL domain, we rely on DL models to automatically learn to extract features from the raw input data, so humans don't have to do \"feature engineering\" ourselves. We use this approach for our baseline models. That's being said, some features maybe hard to learn especially if the model is not powerful enough (i.e. not big enough, too few layers, etc). Therefore, we're curious if some feature engineering may make the learning easier and make the model achieve better accuracy.",
          "replies": [
            {
              "id": 2447435,
              "postDate": "2023-09-20T04:58:41.970Z",
              "content": "<p>Thank you for your response krub. So, can we say that even tho we're using NNs sometimes manual feature engineering is still going to help too? and also I'm curious that once we extracted the features we're still feeding them into GNNs right? and commonly for graph domain we can only extract nodes/edges features ? and the structure is representing by some matrix like adjacency, laplacian, etc. or we can extract some structural features as well and feed to GNNs like maybe naively using spectral clustering to get some sort of community structure? <br>\nyour replies will be much appreciated! I'm a college student at Thammasat at the moment, and this research topic is really cool! never knew that we could apply ML for this domain as well, specially like ML for ML's compiler , that's like some Inception stuff 🙀. Anyways, it's really cool and I really like this topic of research, it's such a great/super cool way to start learning about GNNs! </p>",
              "rawMarkdown": "Thank you for your response krub. So, can we say that even tho we're using NNs sometimes manual feature engineering is still going to help too? and also I'm curious that once we extracted the features we're still feeding them into GNNs right? and commonly for graph domain we can only extract nodes/edges features ? and the structure is representing by some matrix like adjacency, laplacian, etc. or we can extract some structural features as well and feed to GNNs like maybe naively using spectral clustering to get some sort of community structure? \nyour replies will be much appreciated! I'm a college student at Thammasat at the moment, and this research topic is really cool! never knew that we could apply ML for this domain as well, specially like ML for ML's compiler , that's like some Inception stuff 🙀. Anyways, it's really cool and I really like this topic of research, it's such a great/super cool way to start learning about GNNs! "
            },
            {
              "id": 2455887,
              "postDate": "2023-09-25T18:45:28.790Z",
              "content": "<p>Yes, even though we're using NNs, sometimes manual feature engineering could help. One can feed the extracted features to any kind of models, not limited to GNNs. For example, people use Transformer models in other coding applications.</p>\n<p>Extracting structural features could be useful if you don't use GNNs. However, for this problem specifically, I think the high level structure itself (e.g. average degree) is not as important as knowing which nodes depends on which nodes.</p>\n<p>I'm glad you found this problem very interesting!</p>",
              "rawMarkdown": "Yes, even though we're using NNs, sometimes manual feature engineering could help. One can feed the extracted features to any kind of models, not limited to GNNs. For example, people use Transformer models in other coding applications.\n\nExtracting structural features could be useful if you don't use GNNs. However, for this problem specifically, I think the high level structure itself (e.g. average degree) is not as important as knowing which nodes depends on which nodes.\n\nI'm glad you found this problem very interesting!"
            }
          ]
        }
      ]
    },
    {
      "id": 2419651,
      "postDate": "2023-09-02T06:51:20.817Z",
      "content": "<p>good finding man</p>",
      "rawMarkdown": "good finding man"
    },
    {
      "id": 2419006,
      "postDate": "2023-09-01T16:31:46.987Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2427023,
      "author_name": "Mangpo Phothilimthana",
      "author_url": "",
      "post_date": "2023-09-07T03:59:09.477000",
      "content": "<p>We recently uploaded an additional <code>pb</code> directory containing raw graph data! You can find more information about how to extract additional graph features or extract them differently in Section \"Custom Feature Extraction\" at the bottom of the data tab.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2428525,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2023-09-08T01:06:04.597000",
          "content": "<p>This is wonderful! Thank you so much!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2431821,
          "author_name": "Piyush3dB",
          "author_url": "",
          "post_date": "2023-09-10T12:39:31.637000",
          "content": "<p><a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a>, is it correct to say that the information contained in a pb file is the same as the corresponding npz file?</p>\n<p>i.e., the pb file does not contain something which cannot be inferred from the npz file and supporting guidance on how to interpret the npz data?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2443334,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-09-17T16:51:20.697000",
              "content": "<p>Sorry for the late reply! I was on vacation for a week. </p>\n<p>Actually, a pb file does contain more information than npz file. A pb file stores a graph using XLA HLO representation in protobuf format. A node (instruction) in HLO representation has some attributes that npz data doesn't include. In the past, we analyzed the data and concluded that those attributes do not matter so much, but it wasn't a very extensive experiment. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2443700,
              "author_name": "Peiyuan Liao",
              "author_url": "",
              "post_date": "2023-09-17T22:00:22.437000",
              "content": "<p>It seems like the hlo.proto definition referred in the repository expired. Can we refer to the r2.14 one instead?</p>\n<p><a href=\"https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto\" target=\"_blank\">https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto</a></p>\n<p>Additionally, there might be a typo in the layout data explanation. Should it be \"input\" instead of \"intput\"?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2447383,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-09-20T03:37:26.017000",
              "content": "<p>Thank you for pointing these out! Yes, you can use <a href=\"https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto\" target=\"_blank\">https://github.com/tensorflow/tensorflow/blob/r2.14/tensorflow/compiler/xla/service/hlo.proto</a>. I'll update the description.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2458281,
          "author_name": "HZM",
          "author_url": "",
          "post_date": "2023-09-27T13:33:58.330000",
          "content": "<p>when i load the pb file via tf1.15, return the  following bug </p>\n<p>String field 'tensorflow.NodeDef.input' contains invalid UTF-8 data when parsing a protocol buffer. Use the 'bytes' type if you intend to send raw bytes.</p>\n<p>could you please give some insturction how to load the pb file using tf1 or tf2 </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2460765,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-09-29T05:14:59.333000",
              "content": "<p>Can you try this: <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/435640#2458281\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/435640#2458281</a> ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2417889,
      "author_name": "Mangpo Phothilimthana",
      "author_url": "",
      "post_date": "2023-09-01T00:54:20.893000",
      "content": "<p>Very good question! We want participants to be able to extract graph features their own ways if preferred. However, it is more complicated. We're working on the code that would allow participants to do so, but it will take us some time to get everything ready.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2418156,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2023-09-01T07:09:29.560000",
          "content": "<p>Thank you for your comments! I also realized that the test set have anonymized names, does this imply that the workloads are anonymized as well?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2440299,
      "author_name": "Choko",
      "author_url": "",
      "post_date": "2023-09-15T12:23:40.857000",
      "content": "<p>Hi! I'm new to this field of AI/ML/DL I'm still studying it and gaining experiences through working with projects and some datasets. I'm wondering about feature extraction, from my prev knowledge I thought that most of the time ( maybe this is an exception case for this competition? ) we only do feature extraction for ML pipeline since ML models are good for structured data? I assumed that when we extract features we put them in columns so in the end we kinda get structured data format. however, for DL models I understand that it's kinda converted the feature extraction problem into feature 'learning' problem where the models are trained end-to-end and you have some layers to learn the features like how it works in CNN, the convolutions act as a feature extraction layer but the parameters of filters are being learned by gradients-based optimization like typical gradients descent of whatever. so I'm wodering since we're using GNNs, why do we still need to extract more features? can't we just try to make the architecture stronger/ more complexity and be able to learn the features for us like how we do in images domain? what's the inspiration and 'take away knowledge' here? and how do I know which project I need to extract some more features even though I might using DL models already which are trained end-to-end for unstructured data and features are assumed to be learned through this process already.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2443345,
          "author_name": "Mangpo Phothilimthana",
          "author_url": "",
          "post_date": "2023-09-17T16:58:59.453000",
          "content": "<p>Your understanding is right. In DL domain, we rely on DL models to automatically learn to extract features from the raw input data, so humans don't have to do \"feature engineering\" ourselves. We use this approach for our baseline models. That's being said, some features maybe hard to learn especially if the model is not powerful enough (i.e. not big enough, too few layers, etc). Therefore, we're curious if some feature engineering may make the learning easier and make the model achieve better accuracy.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2447435,
              "author_name": "Choko",
              "author_url": "",
              "post_date": "2023-09-20T04:58:41.970000",
              "content": "<p>Thank you for your response krub. So, can we say that even tho we're using NNs sometimes manual feature engineering is still going to help too? and also I'm curious that once we extracted the features we're still feeding them into GNNs right? and commonly for graph domain we can only extract nodes/edges features ? and the structure is representing by some matrix like adjacency, laplacian, etc. or we can extract some structural features as well and feed to GNNs like maybe naively using spectral clustering to get some sort of community structure? <br>\nyour replies will be much appreciated! I'm a college student at Thammasat at the moment, and this research topic is really cool! never knew that we could apply ML for this domain as well, specially like ML for ML's compiler , that's like some Inception stuff 🙀. Anyways, it's really cool and I really like this topic of research, it's such a great/super cool way to start learning about GNNs! </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2455887,
              "author_name": "Mangpo Phothilimthana",
              "author_url": "",
              "post_date": "2023-09-25T18:45:28.790000",
              "content": "<p>Yes, even though we're using NNs, sometimes manual feature engineering could help. One can feed the extracted features to any kind of models, not limited to GNNs. For example, people use Transformer models in other coding applications.</p>\n<p>Extracting structural features could be useful if you don't use GNNs. However, for this problem specifically, I think the high level structure itself (e.g. average degree) is not as important as knowing which nodes depends on which nodes.</p>\n<p>I'm glad you found this problem very interesting!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2419651,
      "author_name": "Bot_developer11",
      "author_url": "",
      "post_date": "2023-09-02T06:51:20.817000",
      "content": "<p>good finding man</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2419006,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-01T16:31:46.987000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2415356": "It seems like the training data is generated by the featurizer here: https://github.com/google-research-datasets/tpu_graphs/blob/main/tpu_graphs/process_data/xla/featurizers.h#L542 . Is it possible & are we allowed to modify this to generate additional features?",
    "2427023": "We recently uploaded an additional `pb` directory containing raw graph data! You can find more information about how to extract additional graph features or extract them differently in Section \"Custom Feature Extraction\" at the bottom of the data tab.",
    "2417889": "Very good question! We want participants to be able to extract graph features their own ways if preferred. However, it is more complicated. We're working on the code that would allow participants to do so, but it will take us some time to get everything ready.",
    "2440299": "Hi! I'm new to this field of AI/ML/DL I'm still studying it and gaining experiences through working with projects and some datasets. I'm wondering about feature extraction, from my prev knowledge I thought that most of the time ( maybe this is an exception case for this competition? ) we only do feature extraction for ML pipeline since ML models are good for structured data? I assumed that when we extract features we put them in columns so in the end we kinda get structured data format. however, for DL models I understand that it's kinda converted the feature extraction problem into feature 'learning' problem where the models are trained end-to-end and you have some layers to learn the features like how it works in CNN, the convolutions act as a feature extraction layer but the parameters of filters are being learned by gradients-based optimization like typical gradients descent of whatever. so I'm wodering since we're using GNNs, why do we still need to extract more features? can't we just try to make the architecture stronger/ more complexity and be able to learn the features for us like how we do in images domain? what's the inspiration and 'take away knowledge' here? and how do I know which project I need to extract some more features even though I might using DL models already which are trained end-to-end for unstructured data and features are assumed to be learned through this process already.",
    "2419651": "good finding man",
    "2419006": ""
  }
}