{
  "id": 206291,
  "title": "What is the meaning of continuous embedding in Saint+",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206291",
  "author_name": "",
  "post_date": "2020-12-24T00:45:40.001842300Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello fellow contestants,</p>\n<p>I am relatively new to this industry and I've been reading the Saint+ paper, and I am having a hard time understanding the continuous embedding it proposes for elapsed and lag time. </p>\n<blockquote>\n  <p>Similar to the elapsed time embedding, we use continuous embedding and categorical embedding<br>\n  for lag time. In continuous embedding, a latent embedding vector for a lag time lt is computed as<br>\n  <code>v_lt = lt · w_lag_time</code>, where <code>w_lag_time</code> is a trainable vector.</p>\n</blockquote>\n<p>so if we are having a batch of responses, lag fed to the decoder like the following:<br>\n<code>x_res = [1,0,0,1,1]</code><br>\n<code>x_lag = [1,0,4,80,30]</code></p>\n<p>in categorical embedding, I would simply feed it to an Embedding module (after label encoding), but in continuous, am I supposed to feed it to a linear layer without bias? is the out dim of the linear what is regarded as the embedding vector for the lag time which gets summed to the embeddings of x_res?</p>\n<p>if my understanding is incorrect, a code example would be highly appreciated.</p>",
  "messages": [
    {
      "id": "1124453",
      "postDate": "12/24/2020 00:45:40",
      "content": "<p>Hello fellow contestants,</p>\n<p>I am relatively new to this industry and I've been reading the Saint+ paper, and I am having a hard time understanding the continuous embedding it proposes for elapsed and lag time. </p>\n<blockquote>\n  <p>Similar to the elapsed time embedding, we use continuous embedding and categorical embedding<br>\n  for lag time. In continuous embedding, a latent embedding vector for a lag time lt is computed as<br>\n  <code>v_lt = lt · w_lag_time</code>, where <code>w_lag_time</code> is a trainable vector.</p>\n</blockquote>\n<p>so if we are having a batch of responses, lag fed to the decoder like the following:<br>\n<code>x_res = [1,0,0,1,1]</code><br>\n<code>x_lag = [1,0,4,80,30]</code></p>\n<p>in categorical embedding, I would simply feed it to an Embedding module (after label encoding), but in continuous, am I supposed to feed it to a linear layer without bias? is the out dim of the linear what is regarded as the embedding vector for the lag time which gets summed to the embeddings of x_res?</p>\n<p>if my understanding is incorrect, a code example would be highly appreciated.</p>",
      "rawMarkdown": "Hello fellow contestants,\n\nI am relatively new to this industry and I've been reading the Saint+ paper, and I am having a hard time understanding the continuous embedding it proposes for elapsed and lag time. \n\n> Similar to the elapsed time embedding, we use continuous embedding and categorical embedding\nfor lag time. In continuous embedding, a latent embedding vector for a lag time lt is computed as\n`v_lt = lt · w_lag_time`, where `w_lag_time` is a trainable vector.\n\nso if we are having a batch of responses, lag fed to the decoder like the following:\n`x_res = [1,0,0,1,1]`\n`x_lag = [1,0,4,80,30]`\n\nin categorical embedding, I would simply feed it to an Embedding module (after label encoding), but in continuous, am I supposed to feed it to a linear layer without bias? is the out dim of the linear what is regarded as the embedding vector for the lag time which gets summed to the embeddings of x_res?\n\nif my understanding is incorrect, a code example would be highly appreciated.",
      "votes": null
    },
    {
      "id": "1124459",
      "postDate": "12/24/2020 01:00:22",
      "content": "<p>My understanding is <code>linear layer without bias</code>, as you say. </p>",
      "rawMarkdown": "My understanding is `linear layer without bias`, as you say.",
      "votes": null
    },
    {
      "id": "1125384",
      "postDate": "12/24/2020 17:07:05",
      "content": "<p>When I used continuous embedding, I get 0.66x cv from the model. However, I used category embedding, the cv is 0.78x.<br>\nDoes anyone have the same experience?</p>",
      "rawMarkdown": "When I used continuous embedding, I get 0.66x cv from the model. However, I used category embedding, the cv is 0.78x.\nDoes anyone have the same experience?",
      "votes": null
    },
    {
      "id": "1128444",
      "postDate": "12/27/2020 12:53:56",
      "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> </p>\n<p>1)to what max time lag time can be truncated is it 1400 minutes or 1400*60 seconds..?<br>\n2) What it means by categorical embedding  ,is it each clipped integer lag time  given a time bucket ? if it falls in some range </p>",
      "rawMarkdown": "m10515009 \n \n1)to what max time lag time can be truncated is it 1400 minutes or 1400*60 seconds..?\n2) What it means by categorical embedding  ,is it each clipped integer lag time  given a time bucket ? if it falls in some range",
      "votes": null
    },
    {
      "id": "1128524",
      "postDate": "12/27/2020 14:13:44",
      "content": "<p>You can refer to the paper.<br>\n<a href=\"url\" target=\"_blank\">https://arxiv.org/abs/2010.12042</a></p>",
      "rawMarkdown": "You can refer to the paper.\n[https://arxiv.org/abs/2010.12042](url)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1124459,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/24/2020 01:00:22",
      "content": "<p>My understanding is <code>linear layer without bias</code>, as you say. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1125384,
      "author_name": "m10515009",
      "author_url": "",
      "post_date": "12/24/2020 17:07:05",
      "content": "<p>When I used continuous embedding, I get 0.66x cv from the model. However, I used category embedding, the cv is 0.78x.<br>\nDoes anyone have the same experience?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1128444,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/27/2020 12:53:56",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> </p>\n<p>1)to what max time lag time can be truncated is it 1400 minutes or 1400*60 seconds..?<br>\n2) What it means by categorical embedding  ,is it each clipped integer lag time  given a time bucket ? if it falls in some range </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128524,
          "author_name": "m10515009",
          "author_url": "",
          "post_date": "12/27/2020 14:13:44",
          "content": "<p>You can refer to the paper.<br>\n<a href=\"url\" target=\"_blank\">https://arxiv.org/abs/2010.12042</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1124453": "Hello fellow contestants,\n\nI am relatively new to this industry and I've been reading the Saint+ paper, and I am having a hard time understanding the continuous embedding it proposes for elapsed and lag time. \n\n> Similar to the elapsed time embedding, we use continuous embedding and categorical embedding\nfor lag time. In continuous embedding, a latent embedding vector for a lag time lt is computed as\n`v_lt = lt · w_lag_time`, where `w_lag_time` is a trainable vector.\n\nso if we are having a batch of responses, lag fed to the decoder like the following:\n`x_res = [1,0,0,1,1]`\n`x_lag = [1,0,4,80,30]`\n\nin categorical embedding, I would simply feed it to an Embedding module (after label encoding), but in continuous, am I supposed to feed it to a linear layer without bias? is the out dim of the linear what is regarded as the embedding vector for the lag time which gets summed to the embeddings of x_res?\n\nif my understanding is incorrect, a code example would be highly appreciated.",
    "1124459": "My understanding is `linear layer without bias`, as you say.",
    "1125384": "When I used continuous embedding, I get 0.66x cv from the model. However, I used category embedding, the cv is 0.78x.\nDoes anyone have the same experience?",
    "1128444": "m10515009 \n \n1)to what max time lag time can be truncated is it 1400 minutes or 1400*60 seconds..?\n2) What it means by categorical embedding  ,is it each clipped integer lag time  given a time bucket ? if it falls in some range",
    "1128524": "You can refer to the paper.\n[https://arxiv.org/abs/2010.12042](url)"
  },
  "source": "meta"
}