{
  "id": 212201,
  "title": "SOFT-DTW Consideration",
  "url": "/competitions/rfcx-species-audio-detection/discussion/212201",
  "author_name": "",
  "post_date": "2021-01-18T02:19:40.570184700Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>\"The celebrated dynamic time warping (DTW) defines the discrepancy between two time series, of possibly variable length, as their minimal alignment cost. Although the number of possible alignments is exponential in the length of the two time series,  showed that DTW can be computed in only quadratic time using dynamic programming.</p>\n<p>Soft-DTW proposes to replace this minimum by a soft minimum. Like the original DTW, soft-DTW can be computed in quadratic time using dynamic programming. However, the main advantage of soft-DTW stems from the fact that it is differentiable everywhere and that its gradient can also be computed in quadratic time. This enables us to use soft-DTW for time series averaging or as a loss function, between a ground-truth time series and a time series predicted by a neural network, trained end-to-end using backpropagation.\"</p>\n<p>Interesting applications at:<br>\n<a href=\"https://github.com/qdenisq/Soft-DTW-AWE\" target=\"_blank\">https://github.com/qdenisq/Soft-DTW-AWE</a><br>\n<a href=\"https://github.com/mblondel/soft-dtw\" target=\"_blank\">https://github.com/mblondel/soft-dtw</a></p>",
  "messages": [
    {
      "id": "1157550",
      "postDate": "01/18/2021 02:19:40",
      "content": "<p>\"The celebrated dynamic time warping (DTW) defines the discrepancy between two time series, of possibly variable length, as their minimal alignment cost. Although the number of possible alignments is exponential in the length of the two time series,  showed that DTW can be computed in only quadratic time using dynamic programming.</p>\n<p>Soft-DTW proposes to replace this minimum by a soft minimum. Like the original DTW, soft-DTW can be computed in quadratic time using dynamic programming. However, the main advantage of soft-DTW stems from the fact that it is differentiable everywhere and that its gradient can also be computed in quadratic time. This enables us to use soft-DTW for time series averaging or as a loss function, between a ground-truth time series and a time series predicted by a neural network, trained end-to-end using backpropagation.\"</p>\n<p>Interesting applications at:<br>\n<a href=\"https://github.com/qdenisq/Soft-DTW-AWE\" target=\"_blank\">https://github.com/qdenisq/Soft-DTW-AWE</a><br>\n<a href=\"https://github.com/mblondel/soft-dtw\" target=\"_blank\">https://github.com/mblondel/soft-dtw</a></p>",
      "rawMarkdown": "\"The celebrated dynamic time warping (DTW) defines the discrepancy between two time series, of possibly variable length, as their minimal alignment cost. Although the number of possible alignments is exponential in the length of the two time series,  showed that DTW can be computed in only quadratic time using dynamic programming.\n\nSoft-DTW proposes to replace this minimum by a soft minimum. Like the original DTW, soft-DTW can be computed in quadratic time using dynamic programming. However, the main advantage of soft-DTW stems from the fact that it is differentiable everywhere and that its gradient can also be computed in quadratic time. This enables us to use soft-DTW for time series averaging or as a loss function, between a ground-truth time series and a time series predicted by a neural network, trained end-to-end using backpropagation.\"\n\nInteresting applications at:\nhttps://github.com/qdenisq/Soft-DTW-AWE\nhttps://github.com/mblondel/soft-dtw",
      "votes": null
    },
    {
      "id": "1159061",
      "postDate": "01/19/2021 03:01:11",
      "content": "<p>seems like it would be good for optimizing SED</p>",
      "rawMarkdown": "seems like it would be good for optimizing SED",
      "votes": null
    },
    {
      "id": "1160831",
      "postDate": "01/20/2021 06:56:22",
      "content": "<p>For DTW to work, I think the template has to be consistent (eg. the each species has to say the same word each time, since it is treated as time series). I'm a bit skeptical that is the case here</p>",
      "rawMarkdown": "For DTW to work, I think the template has to be consistent (eg. the each species has to say the same word each time, since it is treated as time series). I'm a bit skeptical that is the case here",
      "votes": null
    },
    {
      "id": "1161646",
      "postDate": "01/20/2021 17:03:39",
      "content": "<p>I thought it was about calculating misalignment and distance between sequences regardless of signal size. Which I figure is good for audio segmentation because it calculates loss between the outputted sequence and the original robustly (because you can remove noisy entries or padding and calculate the loss on only the key parts of the audio/outputted sequence). What do you think?</p>",
      "rawMarkdown": "I thought it was about calculating misalignment and distance between sequences regardless of signal size. Which I figure is good for audio segmentation because it calculates loss between the outputted sequence and the original robustly (because you can remove noisy entries or padding and calculate the loss on only the key parts of the audio/outputted sequence). What do you think?",
      "votes": null
    },
    {
      "id": "1161705",
      "postDate": "01/20/2021 17:40:44",
      "content": "<p>I'm lost. Sorry, I think I need to read up more to know how to use it for audio segmentation.<br>\nTo me DTW works if you have consistent template, for example in the repo it is given for \"speech command\" so you will have many sample voice saying the same word. Then it's possible to calculate the distance between same words vs different words regardless of signal size and misalignment. <br>\nIn our train sample, the birds does not say the same word (same species doesn't always sing the same tune).  I would not know how to interpret the time warping distance in that case. </p>",
      "rawMarkdown": "I'm lost. Sorry, I think I need to read up more to know how to use it for audio segmentation.\nTo me DTW works if you have consistent template, for example in the repo it is given for \"speech command\" so you will have many sample voice saying the same word. Then it's possible to calculate the distance between same words vs different words regardless of signal size and misalignment. \nIn our train sample, the birds does not say the same word (same species doesn't always sing the same tune).  I would not know how to interpret the time warping distance in that case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1159061,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "01/19/2021 03:01:11",
      "content": "<p>seems like it would be good for optimizing SED</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1160831,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "01/20/2021 06:56:22",
      "content": "<p>For DTW to work, I think the template has to be consistent (eg. the each species has to say the same word each time, since it is treated as time series). I'm a bit skeptical that is the case here</p>",
      "votes": null,
      "replies": [
        {
          "id": 1161646,
          "author_name": "eladwar",
          "author_url": "",
          "post_date": "01/20/2021 17:03:39",
          "content": "<p>I thought it was about calculating misalignment and distance between sequences regardless of signal size. Which I figure is good for audio segmentation because it calculates loss between the outputted sequence and the original robustly (because you can remove noisy entries or padding and calculate the loss on only the key parts of the audio/outputted sequence). What do you think?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1161705,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "01/20/2021 17:40:44",
          "content": "<p>I'm lost. Sorry, I think I need to read up more to know how to use it for audio segmentation.<br>\nTo me DTW works if you have consistent template, for example in the repo it is given for \"speech command\" so you will have many sample voice saying the same word. Then it's possible to calculate the distance between same words vs different words regardless of signal size and misalignment. <br>\nIn our train sample, the birds does not say the same word (same species doesn't always sing the same tune).  I would not know how to interpret the time warping distance in that case. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1157550": "\"The celebrated dynamic time warping (DTW) defines the discrepancy between two time series, of possibly variable length, as their minimal alignment cost. Although the number of possible alignments is exponential in the length of the two time series,  showed that DTW can be computed in only quadratic time using dynamic programming.\n\nSoft-DTW proposes to replace this minimum by a soft minimum. Like the original DTW, soft-DTW can be computed in quadratic time using dynamic programming. However, the main advantage of soft-DTW stems from the fact that it is differentiable everywhere and that its gradient can also be computed in quadratic time. This enables us to use soft-DTW for time series averaging or as a loss function, between a ground-truth time series and a time series predicted by a neural network, trained end-to-end using backpropagation.\"\n\nInteresting applications at:\nhttps://github.com/qdenisq/Soft-DTW-AWE\nhttps://github.com/mblondel/soft-dtw",
    "1159061": "seems like it would be good for optimizing SED",
    "1160831": "For DTW to work, I think the template has to be consistent (eg. the each species has to say the same word each time, since it is treated as time series). I'm a bit skeptical that is the case here",
    "1161646": "I thought it was about calculating misalignment and distance between sequences regardless of signal size. Which I figure is good for audio segmentation because it calculates loss between the outputted sequence and the original robustly (because you can remove noisy entries or padding and calculate the loss on only the key parts of the audio/outputted sequence). What do you think?",
    "1161705": "I'm lost. Sorry, I think I need to read up more to know how to use it for audio segmentation.\nTo me DTW works if you have consistent template, for example in the repo it is given for \"speech command\" so you will have many sample voice saying the same word. Then it's possible to calculate the distance between same words vs different words regardless of signal size and misalignment. \nIn our train sample, the birds does not say the same word (same species doesn't always sing the same tune).  I would not know how to interpret the time warping distance in that case."
  },
  "source": "meta"
}