{
  "id": 439002,
  "title": "Building a model to learn layout data",
  "url": "/competitions/predict-ai-model-runtime/discussion/439002",
  "author_name": "",
  "post_date": "2023-09-13T10:34:58.117280200Z",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Did anyone manage to build a model to learn the layout data (I tried training on only one small file from Layout xla:default) without the TF script?<br>\nThings worked out really well in the tile:xla dataset but for some reason my model doesn't seem to learn anything in layout data (Loss not decreasing, cosine similarity between embedding vectors is 1).<br>\nIs it a gradient problem or something else? Did anyone else experience this problem?</p>\n<p>I would love to get some insight on this.<br>\nThanks a lot.</p>",
  "messages": [
    {
      "id": "2436081",
      "postDate": "09/13/2023 10:34:58",
      "content": "<p>Did anyone manage to build a model to learn the layout data (I tried training on only one small file from Layout xla:default) without the TF script?<br>\nThings worked out really well in the tile:xla dataset but for some reason my model doesn't seem to learn anything in layout data (Loss not decreasing, cosine similarity between embedding vectors is 1).<br>\nIs it a gradient problem or something else? Did anyone else experience this problem?</p>\n<p>I would love to get some insight on this.<br>\nThanks a lot.</p>",
      "rawMarkdown": "Did anyone manage to build a model to learn the layout data (I tried training on only one small file from Layout xla:default) without the TF script?\nThings worked out really well in the tile:xla dataset but for some reason my model doesn't seem to learn anything in layout data (Loss not decreasing, cosine similarity between embedding vectors is 1).\nIs it a gradient problem or something else? Did anyone else experience this problem?\n\nI would love to get some insight on this.\nThanks a lot.",
      "votes": null
    },
    {
      "id": "2436613",
      "postDate": "09/13/2023 17:01:43",
      "content": "<p>I am on the same boat, doesn't even overfit on a batch consiting of some of the models with less that 2000 nodes and only taking in to consideration the first 100 configs</p>",
      "rawMarkdown": "I am on the same boat, doesn't even overfit on a batch consiting of some of the models with less that 2000 nodes and only taking in to consideration the first 100 configs",
      "votes": null
    },
    {
      "id": "2436939",
      "postDate": "09/13/2023 21:54:53",
      "content": "<p>Yeah, the exact same thing happened to me. I tried to train it on only 100 configs but it doesn't seem to learn anything. I also tried many different models and loss functions to solve this. I thought I had a bug in my code.</p>",
      "rawMarkdown": "Yeah, the exact same thing happened to me. I tried to train it on only 100 configs but it doesn't seem to learn anything. I also tried many different models and loss functions to solve this. I thought I had a bug in my code.",
      "votes": null
    },
    {
      "id": "2438161",
      "postDate": "09/14/2023 05:04:26",
      "content": "<p>In their dataset paper, the competition hosts described the different GNNs used for layout and tile problems. I don't think it's viable to use the tile model for the layout problem since the data structure is quite different.</p>",
      "rawMarkdown": "In their dataset paper, the competition hosts described the different GNNs used for layout and tile problems. I don't think it's viable to use the tile model for the layout problem since the data structure is quite different.",
      "votes": null
    },
    {
      "id": "2439118",
      "postDate": "09/14/2023 16:40:57",
      "content": "<p>You're right. But I tried the things presented in this paper (without GST, but I tried on a file with 372 nodes. In the paper they say that they use a segment size of 1000 so I suppose it shouldn't matter much) and I didn't get any learning (not even on 100 examples out of the file).</p>",
      "rawMarkdown": "You're right. But I tried the things presented in this paper (without GST, but I tried on a file with 372 nodes. In the paper they say that they use a segment size of 1000 so I suppose it shouldn't matter much) and I didn't get any learning (not even on 100 examples out of the file).",
      "votes": null
    },
    {
      "id": "2443471",
      "postDate": "09/17/2023 18:15:13",
      "content": "<p>In our experiments, we found a couple important things:</p>\n<ul>\n<li>Using pairwise/listwise rank loss (instead of a typical MSE/MAPE loss) is important.</li>\n<li>Combine the layout configuration with node features before feeding into GNN is important.</li>\n<li>When concatenating layout config with node features, the model learns better when we scale up the layout portion, e.g. concat(node_features, SCALE * layout_config), where SCALE is something like 100.</li>\n</ul>\n<p>I hope this helps!</p>",
      "rawMarkdown": "In our experiments, we found a couple important things:\n- Using pairwise/listwise rank loss (instead of a typical MSE/MAPE loss) is important.\n- Combine the layout configuration with node features before feeding into GNN is important.\n- When concatenating layout config with node features, the model learns better when we scale up the layout portion, e.g. concat(node_features, SCALE * layout_config), where SCALE is something like 100.\n\nI hope this helps!",
      "votes": null
    },
    {
      "id": "2444485",
      "postDate": "09/18/2023 09:43:21",
      "content": "<p>Hello, thanks for your tips, but I have several questions. When I combine the layout configuration with node features, it can't work, because c is too big, and I need combine c times. May I ask if you have any solution.</p>",
      "rawMarkdown": "Hello, thanks for your tips, but I have several questions. When I combine the layout configuration with node features, it can't work, because c is too big, and I need combine c times. May I ask if you have any solution.",
      "votes": null
    },
    {
      "id": "2447388",
      "postDate": "09/20/2023 03:48:14",
      "content": "<p><code>c</code> is the number of difference configurations for a graph. When training a model, we don't need to train all <code>c</code> samples in a single training batch.</p>",
      "rawMarkdown": "`c` is the number of difference configurations for a graph. When training a model, we don't need to train all `c` samples in a single training batch.",
      "votes": null
    },
    {
      "id": "2447979",
      "postDate": "09/20/2023 10:59:48",
      "content": "<p>Thank you very much! It helps me a lot.</p>",
      "rawMarkdown": "Thank you very much! It helps me a lot.",
      "votes": null
    },
    {
      "id": "2451273",
      "postDate": "09/22/2023 12:33:48",
      "content": "<p>Thanks a lot! I have a few more questions:</p>\n<ol>\n<li><p>In the GST paper, it is written that for the dataset TPUGraphs \"the prediction head is part of F, and F' is simply a summation function.\" So how is the \"Stale Embedding Dropout\" implemented if the output of the model is a size 1 vector? Do we zero the entire output of the segment? Or is the output of the model an embedding vector (the TPUGraphs dataset paper says it outputs an embedding vector so it's quite puzzling).</p></li>\n<li><p>Also in the GST paper it is not mentioned how many segments are trained for each batch. For N different segments, approximately how many segments will require gradient, and how many will be fetched from the historical embedding table?</p></li>\n<li><p>What is the process training? In the GST paper it says that for each epoch only one batch is calculated. Is that batch taken from only one file that is sampled randomly? All the possible pairs in the batch are computed to produce a loss (for example N=8 we compute 8*7/2 = 28 pairs)? </p></li>\n</ol>",
      "rawMarkdown": "Thanks a lot! I have a few more questions:\n1. In the GST paper, it is written that for the dataset TPUGraphs \"the prediction head is part of F, and F' is simply a summation function.\" So how is the \"Stale Embedding Dropout\" implemented if the output of the model is a size 1 vector? Do we zero the entire output of the segment? Or is the output of the model an embedding vector (the TPUGraphs dataset paper says it outputs an embedding vector so it's quite puzzling).\n\n2. Also in the GST paper it is not mentioned how many segments are trained for each batch. For N different segments, approximately how many segments will require gradient, and how many will be fetched from the historical embedding table?\n\n3. What is the process training? In the GST paper it says that for each epoch only one batch is calculated. Is that batch taken from only one file that is sampled randomly? All the possible pairs in the batch are computed to produce a loss (for example N=8 we compute 8*7/2 = 28 pairs)?",
      "votes": null
    },
    {
      "id": "2455927",
      "postDate": "09/25/2023 19:18:25",
      "content": "<ol>\n<li><p>F outputs an embedding vector.</p></li>\n<li><p>The number of segments are variable, depending on the graph size. We partition a graph such that a segment has &lt; 1000 nodes.</p></li>\n<li><p>In each training step, one batch of data is used (one graph and multiple configs, selected randomly). Within each batch, we compute either <a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/PairwiseHingeLoss\" target=\"_blank\">pairwise hinge loss</a>, <a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/PairwiseLogisticLoss\" target=\"_blank\">pairwise logistic loss</a>, or <a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ListMLELoss\" target=\"_blank\">list MLE</a>.</p></li>\n</ol>",
      "rawMarkdown": "1. F outputs an embedding vector.\n\n2. The number of segments are variable, depending on the graph size. We partition a graph such that a segment has < 1000 nodes.\n\n3. In each training step, one batch of data is used (one graph and multiple configs, selected randomly). Within each batch, we compute either [pairwise hinge loss](https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/PairwiseHingeLoss), [pairwise logistic loss](https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/PairwiseLogisticLoss), or [list MLE](https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ListMLELoss).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2436613,
      "author_name": "ksmcg90",
      "author_url": "",
      "post_date": "09/13/2023 17:01:43",
      "content": "<p>I am on the same boat, doesn't even overfit on a batch consiting of some of the models with less that 2000 nodes and only taking in to consideration the first 100 configs</p>",
      "votes": null,
      "replies": [
        {
          "id": 2436939,
          "author_name": "amitaharoni",
          "author_url": "",
          "post_date": "09/13/2023 21:54:53",
          "content": "<p>Yeah, the exact same thing happened to me. I tried to train it on only 100 configs but it doesn't seem to learn anything. I also tried many different models and loss functions to solve this. I thought I had a bug in my code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2438161,
      "author_name": "passengerc07",
      "author_url": "",
      "post_date": "09/14/2023 05:04:26",
      "content": "<p>In their dataset paper, the competition hosts described the different GNNs used for layout and tile problems. I don't think it's viable to use the tile model for the layout problem since the data structure is quite different.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2439118,
          "author_name": "amitaharoni",
          "author_url": "",
          "post_date": "09/14/2023 16:40:57",
          "content": "<p>You're right. But I tried the things presented in this paper (without GST, but I tried on a file with 372 nodes. In the paper they say that they use a segment size of 1000 so I suppose it shouldn't matter much) and I didn't get any learning (not even on 100 examples out of the file).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2443471,
      "author_name": "mangpophothilimthana",
      "author_url": "",
      "post_date": "09/17/2023 18:15:13",
      "content": "<p>In our experiments, we found a couple important things:</p>\n<ul>\n<li>Using pairwise/listwise rank loss (instead of a typical MSE/MAPE loss) is important.</li>\n<li>Combine the layout configuration with node features before feeding into GNN is important.</li>\n<li>When concatenating layout config with node features, the model learns better when we scale up the layout portion, e.g. concat(node_features, SCALE * layout_config), where SCALE is something like 100.</li>\n</ul>\n<p>I hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2444485,
          "author_name": "ljjsxx",
          "author_url": "",
          "post_date": "09/18/2023 09:43:21",
          "content": "<p>Hello, thanks for your tips, but I have several questions. When I combine the layout configuration with node features, it can't work, because c is too big, and I need combine c times. May I ask if you have any solution.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2447388,
              "author_name": "mangpophothilimthana",
              "author_url": "",
              "post_date": "09/20/2023 03:48:14",
              "content": "<p><code>c</code> is the number of difference configurations for a graph. When training a model, we don't need to train all <code>c</code> samples in a single training batch.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2447979,
                  "author_name": "ljjsxx",
                  "author_url": "",
                  "post_date": "09/20/2023 10:59:48",
                  "content": "<p>Thank you very much! It helps me a lot.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2451273,
          "author_name": "amitaharoni",
          "author_url": "",
          "post_date": "09/22/2023 12:33:48",
          "content": "<p>Thanks a lot! I have a few more questions:</p>\n<ol>\n<li><p>In the GST paper, it is written that for the dataset TPUGraphs \"the prediction head is part of F, and F' is simply a summation function.\" So how is the \"Stale Embedding Dropout\" implemented if the output of the model is a size 1 vector? Do we zero the entire output of the segment? Or is the output of the model an embedding vector (the TPUGraphs dataset paper says it outputs an embedding vector so it's quite puzzling).</p></li>\n<li><p>Also in the GST paper it is not mentioned how many segments are trained for each batch. For N different segments, approximately how many segments will require gradient, and how many will be fetched from the historical embedding table?</p></li>\n<li><p>What is the process training? In the GST paper it says that for each epoch only one batch is calculated. Is that batch taken from only one file that is sampled randomly? All the possible pairs in the batch are computed to produce a loss (for example N=8 we compute 8*7/2 = 28 pairs)? </p></li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2455927,
              "author_name": "mangpophothilimthana",
              "author_url": "",
              "post_date": "09/25/2023 19:18:25",
              "content": "<ol>\n<li><p>F outputs an embedding vector.</p></li>\n<li><p>The number of segments are variable, depending on the graph size. We partition a graph such that a segment has &lt; 1000 nodes.</p></li>\n<li><p>In each training step, one batch of data is used (one graph and multiple configs, selected randomly). Within each batch, we compute either <a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/PairwiseHingeLoss\" target=\"_blank\">pairwise hinge loss</a>, <a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/PairwiseLogisticLoss\" target=\"_blank\">pairwise logistic loss</a>, or <a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ListMLELoss\" target=\"_blank\">list MLE</a>.</p></li>\n</ol>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2436081": "Did anyone manage to build a model to learn the layout data (I tried training on only one small file from Layout xla:default) without the TF script?\nThings worked out really well in the tile:xla dataset but for some reason my model doesn't seem to learn anything in layout data (Loss not decreasing, cosine similarity between embedding vectors is 1).\nIs it a gradient problem or something else? Did anyone else experience this problem?\n\nI would love to get some insight on this.\nThanks a lot.",
    "2436613": "I am on the same boat, doesn't even overfit on a batch consiting of some of the models with less that 2000 nodes and only taking in to consideration the first 100 configs",
    "2436939": "Yeah, the exact same thing happened to me. I tried to train it on only 100 configs but it doesn't seem to learn anything. I also tried many different models and loss functions to solve this. I thought I had a bug in my code.",
    "2438161": "In their dataset paper, the competition hosts described the different GNNs used for layout and tile problems. I don't think it's viable to use the tile model for the layout problem since the data structure is quite different.",
    "2439118": "You're right. But I tried the things presented in this paper (without GST, but I tried on a file with 372 nodes. In the paper they say that they use a segment size of 1000 so I suppose it shouldn't matter much) and I didn't get any learning (not even on 100 examples out of the file).",
    "2443471": "In our experiments, we found a couple important things:\n- Using pairwise/listwise rank loss (instead of a typical MSE/MAPE loss) is important.\n- Combine the layout configuration with node features before feeding into GNN is important.\n- When concatenating layout config with node features, the model learns better when we scale up the layout portion, e.g. concat(node_features, SCALE * layout_config), where SCALE is something like 100.\n\nI hope this helps!",
    "2444485": "Hello, thanks for your tips, but I have several questions. When I combine the layout configuration with node features, it can't work, because c is too big, and I need combine c times. May I ask if you have any solution.",
    "2447388": "`c` is the number of difference configurations for a graph. When training a model, we don't need to train all `c` samples in a single training batch.",
    "2447979": "Thank you very much! It helps me a lot.",
    "2451273": "Thanks a lot! I have a few more questions:\n1. In the GST paper, it is written that for the dataset TPUGraphs \"the prediction head is part of F, and F' is simply a summation function.\" So how is the \"Stale Embedding Dropout\" implemented if the output of the model is a size 1 vector? Do we zero the entire output of the segment? Or is the output of the model an embedding vector (the TPUGraphs dataset paper says it outputs an embedding vector so it's quite puzzling).\n\n2. Also in the GST paper it is not mentioned how many segments are trained for each batch. For N different segments, approximately how many segments will require gradient, and how many will be fetched from the historical embedding table?\n\n3. What is the process training? In the GST paper it says that for each epoch only one batch is calculated. Is that batch taken from only one file that is sampled randomly? All the possible pairs in the batch are computed to produce a loss (for example N=8 we compute 8*7/2 = 28 pairs)?",
    "2455927": "1. F outputs an embedding vector.\n\n2. The number of segments are variable, depending on the graph size. We partition a graph such that a segment has < 1000 nodes.\n\n3. In each training step, one batch of data is used (one graph and multiple configs, selected randomly). Within each batch, we compute either [pairwise hinge loss](https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/PairwiseHingeLoss), [pairwise logistic loss](https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/PairwiseLogisticLoss), or [list MLE](https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ListMLELoss)."
  },
  "source": "meta"
}