{
  "id": 45682,
  "title": "Unknown unknowns",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45682",
  "author_name": "",
  "post_date": "2017-12-14T15:34:00.603953Z",
  "votes": 17,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello!</p>\n\n<p>Do you have a special approach for handling the unknown unknowns? I.e. the words \"learn\", \"follow\", \"backward\" ... which are not part of the labeled public dataset? Handling the known unknowns (\"cat\", \"bird\", ...) already works well. You can just use them while training and map these classes to \"unknown\" when making the submission. However, looking at my submission I found that \"backward\" is constantly classified as \"left\", \"follow\" as \"off\" and so on.</p>\n\n<p>I think that this is also the reason for the differences between validation and leaderboard scores. I tried thresholding based on the softmax probabilities which didn't work and found this <a href=\"http://vast.uccs.edu/~abendale/papers/0348.pdf\">paper</a>. Though, the code has a few dependencies (caffe, LibMR) which I'd like to avoid.</p>",
  "messages": [
    {
      "id": "257595",
      "postDate": "12/14/2017 15:34:00",
      "content": "<p>Hello!</p>\n\n<p>Do you have a special approach for handling the unknown unknowns? I.e. the words \"learn\", \"follow\", \"backward\" ... which are not part of the labeled public dataset? Handling the known unknowns (\"cat\", \"bird\", ...) already works well. You can just use them while training and map these classes to \"unknown\" when making the submission. However, looking at my submission I found that \"backward\" is constantly classified as \"left\", \"follow\" as \"off\" and so on.</p>\n\n<p>I think that this is also the reason for the differences between validation and leaderboard scores. I tried thresholding based on the softmax probabilities which didn't work and found this <a href=\"http://vast.uccs.edu/~abendale/papers/0348.pdf\">paper</a>. Though, the code has a few dependencies (caffe, LibMR) which I'd like to avoid.</p>",
      "rawMarkdown": "Hello!\n\nDo you have a special approach for handling the unknown unknowns? I.e. the words \"learn\", \"follow\", \"backward\" ... which are not part of the labeled public dataset? Handling the known unknowns (\"cat\", \"bird\", ...) already works well. You can just use them while training and map these classes to \"unknown\" when making the submission. However, looking at my submission I found that \"backward\" is constantly classified as \"left\", \"follow\" as \"off\" and so on.\n\nI think that this is also the reason for the differences between validation and leaderboard scores. I tried thresholding based on the softmax probabilities which didn't work and found this [paper][1]. Though, the code has a few dependencies (caffe, LibMR) which I'd like to avoid.\n\n[1]: http://vast.uccs.edu/%7Eabendale/papers/0348.pdf",
      "votes": null
    },
    {
      "id": "257619",
      "postDate": "12/14/2017 16:37:58",
      "content": "<p>I also think the unknown unknowns are important to consider. For validation holding out several of the known unknowns is likely a good idea. I haven't had much time to work on this competition yet, but addressing the above issue is in my opinion one of the main interesting things to explore(along with running nets on raspberry pi). </p>\n\n<p>As a note though, a good approach to this problem may be a major differentiating factor between teams, so be careful sharing things you find;)</p>",
      "rawMarkdown": "I also think the unknown unknowns are important to consider. For validation holding out several of the known unknowns is likely a good idea. I haven't had much time to work on this competition yet, but addressing the above issue is in my opinion one of the main interesting things to explore(along with running nets on raspberry pi). \n\nAs a note though, a good approach to this problem may be a major differentiating factor between teams, so be careful sharing things you find;)",
      "votes": null
    },
    {
      "id": "257847",
      "postDate": "12/15/2017 01:41:03",
      "content": "<p>Wow , I haven't notice that ...</p>\n\n<p>But to me that sounds like a \"One-class classification\" or \"outlier detection, anomaly detection, novelty detection\"</p>",
      "rawMarkdown": "Wow , I haven't notice that ...\n\nBut to me that sounds like a \"One-class classification\" or \"outlier detection, anomaly detection, novelty detection\"",
      "votes": null
    },
    {
      "id": "259278",
      "postDate": "12/18/2017 03:47:52",
      "content": "<p>I'm facing the same problem. My plan is </p>\n\n<ol>\n<li>use 12 categories rather than 31 to get a more compatible \"unknown\" classification</li>\n<li>use model trained before to predict the test audios, and join some of the test data which are predicted with higher probability like more than 98% into the train data set and retrain the model. This seems strange for a mixture of train data and test data, but as we all know, it is better for speech-recognition to use a larger train data set and the train data is not that good. So consider the probability threshold as a parameter.</li>\n<li>label the test audios to unknown, when they are predicted without probability higher than like 50%.</li>\n</ol>\n\n<p>I'm just going to implement these plans, hoping for a progress.</p>",
      "rawMarkdown": "I'm facing the same problem. My plan is \n\n 1. use 12 categories rather than 31 to get a more compatible \"unknown\" classification\n 2. use model trained before to predict the test audios, and join some of the test data which are predicted with higher probability like more than 98% into the train data set and retrain the model. This seems strange for a mixture of train data and test data, but as we all know, it is better for speech-recognition to use a larger train data set and the train data is not that good. So consider the probability threshold as a parameter.\n 3. label the test audios to unknown, when they are predicted without probability higher than like 50%.\n\nI'm just going to implement these plans, hoping for a progress.",
      "votes": null
    },
    {
      "id": "261634",
      "postDate": "12/23/2017 09:50:49",
      "content": "<p>Hi See--,</p>\n\n<p>I tried sklearn's OneClassSVM and Isolation Forest, but so far, they aren't contributing to a higher score when I apply them :( This is such an important problem to tackle, but I'm stuck :(</p>",
      "rawMarkdown": "Hi See--,\n\nI tried sklearn's OneClassSVM and Isolation Forest, but so far, they aren't contributing to a higher score when I apply them :( This is such an important problem to tackle, but I'm stuck :(",
      "votes": null
    },
    {
      "id": "265384",
      "postDate": "01/05/2018 12:59:12",
      "content": "<p>Did anyone had success in implementing this article? Any method for unknown unknown that had impact on the LB scores?</p>",
      "rawMarkdown": "Did anyone had success in implementing this article? Any method for unknown unknown that had impact on the LB scores?",
      "votes": null
    },
    {
      "id": "265457",
      "postDate": "01/05/2018 15:16:36",
      "content": "<p>I implemented the first algorithm described in the paper, tried a couple different thresholds but all of my attempts actually lowered my LB score</p>\n\n<p>The second algo in that paper (the one they recommend) might hold more promise, but when the first algo didn't yield any improvements I abandoned the idea</p>\n\n<p>Frustrating, bc I feel like theoretically the ideas in the paper make a whole lot of sense for this problem, and thank you to @See-- for posting it</p>\n\n<p>Anybody else try and have more success?</p>",
      "rawMarkdown": "I implemented the first algorithm described in the paper, tried a couple different thresholds but all of my attempts actually lowered my LB score\n\nThe second algo in that paper (the one they recommend) might hold more promise, but when the first algo didn't yield any improvements I abandoned the idea\n\nFrustrating, bc I feel like theoretically the ideas in the paper make a whole lot of sense for this problem, and thank you to @See-- for posting it\n\nAnybody else try and have more success?",
      "votes": null
    },
    {
      "id": "265738",
      "postDate": "01/06/2018 14:23:05",
      "content": "<p>Can you tell us if you use some other approach? Does the \"unknown unknown\" are really such a big problem that require special care? (I see you reach LB score of  0.9 and I'm trying to get some tips ;) )</p>",
      "rawMarkdown": "Can you tell us if you use some other approach? Does the \"unknown unknown\" are really such a big problem that require special care? (I see you reach LB score of  0.9 and I'm trying to get some tips ;) )",
      "votes": null
    },
    {
      "id": "265786",
      "postDate": "01/06/2018 17:48:08",
      "content": "<p>So far everything I've tried to account for the unknown unknowns has failed. My best models so far use a regular 12-class final layer with a random sampling of the \"known unknowns\" for the unknown label</p>",
      "rawMarkdown": "So far everything I've tried to account for the unknown unknowns has failed. My best models so far use a regular 12-class final layer with a random sampling of the \"known unknowns\" for the unknown label",
      "votes": null
    },
    {
      "id": "266262",
      "postDate": "01/08/2018 09:11:32",
      "content": "<p>I see, that sounds great, to pick from known unknowns randomly. Thank you, Thomas :) I'll try this!</p>",
      "rawMarkdown": "I see, that sounds great, to pick from known unknowns randomly. Thank you, Thomas :) I'll try this!",
      "votes": null
    },
    {
      "id": "266726",
      "postDate": "01/09/2018 15:26:48",
      "content": "<p>Thomas, I was able to get plus 1% on LeaderBoard with the method. Thank you for sharing ^_^</p>",
      "rawMarkdown": "Thomas, I was able to get plus 1% on LeaderBoard with the method. Thank you for sharing ^_^",
      "votes": null
    },
    {
      "id": "266741",
      "postDate": "01/09/2018 16:03:56",
      "content": "<p>Indeed, Thomas, I forgot to thanks you for sharing!</p>",
      "rawMarkdown": "Indeed, Thomas, I forgot to thanks you for sharing!",
      "votes": null
    },
    {
      "id": "267362",
      "postDate": "01/11/2018 06:27:14",
      "content": "<p>Since unknown can be seen as any speech not in the 10 speech containing classes, I augmented some unknown data by slicing and re-combining the wavs of different unknowns in training set. </p>\n\n<p>Helped in making the unknown-detection in test set better. </p>",
      "rawMarkdown": "Since unknown can be seen as any speech not in the 10 speech containing classes, I augmented some unknown data by slicing and re-combining the wavs of different unknowns in training set. \n\nHelped in making the unknown-detection in test set better.",
      "votes": null
    },
    {
      "id": "3184291",
      "postDate": "04/21/2025 23:34:25",
      "content": "<p>Hi Dieter (@kagglingdieter),<br>\nI’m a med school student passionate about AI and neural network development. I’ve learned some concepts but struggle to apply them in competitions like those on Kaggle. <br>\nCongrats on topping the leaderboard for the BYU - Locating Bacterial Flagellar Motors 2025 competition! Your work is inspiring. I have two questions:  <br>\nHow do you approach building top-performing models for such complex tasks?  </p>\n<p>How did you learn these skills, and could you share guidance for someone like me looking to improve?<br>\nThank you for your time!  <br>\nWarm regards,<br>\nSanto Paul  </p>",
      "rawMarkdown": "Hi Dieter (@kagglingdieter),\nI’m a med school student passionate about AI and neural network development. I’ve learned some concepts but struggle to apply them in competitions like those on Kaggle. \nCongrats on topping the leaderboard for the BYU - Locating Bacterial Flagellar Motors 2025 competition! Your work is inspiring. I have two questions:  \nHow do you approach building top-performing models for such complex tasks?  \n\nHow did you learn these skills, and could you share guidance for someone like me looking to improve?\nThank you for your time!  \nWarm regards,\nSanto Paul",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 257619,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "12/14/2017 16:37:58",
      "content": "<p>I also think the unknown unknowns are important to consider. For validation holding out several of the known unknowns is likely a good idea. I haven't had much time to work on this competition yet, but addressing the above issue is in my opinion one of the main interesting things to explore(along with running nets on raspberry pi). </p>\n\n<p>As a note though, a good approach to this problem may be a major differentiating factor between teams, so be careful sharing things you find;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 257847,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "12/15/2017 01:41:03",
      "content": "<p>Wow , I haven't notice that ...</p>\n\n<p>But to me that sounds like a \"One-class classification\" or \"outlier detection, anomaly detection, novelty detection\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 259278,
      "author_name": "zsq2420",
      "author_url": "",
      "post_date": "12/18/2017 03:47:52",
      "content": "<p>I'm facing the same problem. My plan is </p>\n\n<ol>\n<li>use 12 categories rather than 31 to get a more compatible \"unknown\" classification</li>\n<li>use model trained before to predict the test audios, and join some of the test data which are predicted with higher probability like more than 98% into the train data set and retrain the model. This seems strange for a mixture of train data and test data, but as we all know, it is better for speech-recognition to use a larger train data set and the train data is not that good. So consider the probability threshold as a parameter.</li>\n<li>label the test audios to unknown, when they are predicted without probability higher than like 50%.</li>\n</ol>\n\n<p>I'm just going to implement these plans, hoping for a progress.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 261634,
      "author_name": "kumiko",
      "author_url": "",
      "post_date": "12/23/2017 09:50:49",
      "content": "<p>Hi See--,</p>\n\n<p>I tried sklearn's OneClassSVM and Isolation Forest, but so far, they aren't contributing to a higher score when I apply them :( This is such an important problem to tackle, but I'm stuck :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 265384,
      "author_name": "ori226",
      "author_url": "",
      "post_date": "01/05/2018 12:59:12",
      "content": "<p>Did anyone had success in implementing this article? Any method for unknown unknown that had impact on the LB scores?</p>",
      "votes": null,
      "replies": [
        {
          "id": 265457,
          "author_name": "omalleyt",
          "author_url": "",
          "post_date": "01/05/2018 15:16:36",
          "content": "<p>I implemented the first algorithm described in the paper, tried a couple different thresholds but all of my attempts actually lowered my LB score</p>\n\n<p>The second algo in that paper (the one they recommend) might hold more promise, but when the first algo didn't yield any improvements I abandoned the idea</p>\n\n<p>Frustrating, bc I feel like theoretically the ideas in the paper make a whole lot of sense for this problem, and thank you to @See-- for posting it</p>\n\n<p>Anybody else try and have more success?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265738,
          "author_name": "ori226",
          "author_url": "",
          "post_date": "01/06/2018 14:23:05",
          "content": "<p>Can you tell us if you use some other approach? Does the \"unknown unknown\" are really such a big problem that require special care? (I see you reach LB score of  0.9 and I'm trying to get some tips ;) )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265786,
          "author_name": "omalleyt",
          "author_url": "",
          "post_date": "01/06/2018 17:48:08",
          "content": "<p>So far everything I've tried to account for the unknown unknowns has failed. My best models so far use a regular 12-class final layer with a random sampling of the \"known unknowns\" for the unknown label</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266262,
          "author_name": "kumiko",
          "author_url": "",
          "post_date": "01/08/2018 09:11:32",
          "content": "<p>I see, that sounds great, to pick from known unknowns randomly. Thank you, Thomas :) I'll try this!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266726,
          "author_name": "kumiko",
          "author_url": "",
          "post_date": "01/09/2018 15:26:48",
          "content": "<p>Thomas, I was able to get plus 1% on LeaderBoard with the method. Thank you for sharing ^_^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266741,
          "author_name": "ori226",
          "author_url": "",
          "post_date": "01/09/2018 16:03:56",
          "content": "<p>Indeed, Thomas, I forgot to thanks you for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267362,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "01/11/2018 06:27:14",
      "content": "<p>Since unknown can be seen as any speech not in the 10 speech containing classes, I augmented some unknown data by slicing and re-combining the wavs of different unknowns in training set. </p>\n\n<p>Helped in making the unknown-detection in test set better. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3184291,
          "author_name": "santopaul",
          "author_url": "",
          "post_date": "04/21/2025 23:34:25",
          "content": "<p>Hi Dieter (@kagglingdieter),<br>\nI’m a med school student passionate about AI and neural network development. I’ve learned some concepts but struggle to apply them in competitions like those on Kaggle. <br>\nCongrats on topping the leaderboard for the BYU - Locating Bacterial Flagellar Motors 2025 competition! Your work is inspiring. I have two questions:  <br>\nHow do you approach building top-performing models for such complex tasks?  </p>\n<p>How did you learn these skills, and could you share guidance for someone like me looking to improve?<br>\nThank you for your time!  <br>\nWarm regards,<br>\nSanto Paul  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "257595": "Hello!\n\nDo you have a special approach for handling the unknown unknowns? I.e. the words \"learn\", \"follow\", \"backward\" ... which are not part of the labeled public dataset? Handling the known unknowns (\"cat\", \"bird\", ...) already works well. You can just use them while training and map these classes to \"unknown\" when making the submission. However, looking at my submission I found that \"backward\" is constantly classified as \"left\", \"follow\" as \"off\" and so on.\n\nI think that this is also the reason for the differences between validation and leaderboard scores. I tried thresholding based on the softmax probabilities which didn't work and found this [paper][1]. Though, the code has a few dependencies (caffe, LibMR) which I'd like to avoid.\n\n[1]: http://vast.uccs.edu/%7Eabendale/papers/0348.pdf",
    "257619": "I also think the unknown unknowns are important to consider. For validation holding out several of the known unknowns is likely a good idea. I haven't had much time to work on this competition yet, but addressing the above issue is in my opinion one of the main interesting things to explore(along with running nets on raspberry pi). \n\nAs a note though, a good approach to this problem may be a major differentiating factor between teams, so be careful sharing things you find;)",
    "257847": "Wow , I haven't notice that ...\n\nBut to me that sounds like a \"One-class classification\" or \"outlier detection, anomaly detection, novelty detection\"",
    "259278": "I'm facing the same problem. My plan is \n\n 1. use 12 categories rather than 31 to get a more compatible \"unknown\" classification\n 2. use model trained before to predict the test audios, and join some of the test data which are predicted with higher probability like more than 98% into the train data set and retrain the model. This seems strange for a mixture of train data and test data, but as we all know, it is better for speech-recognition to use a larger train data set and the train data is not that good. So consider the probability threshold as a parameter.\n 3. label the test audios to unknown, when they are predicted without probability higher than like 50%.\n\nI'm just going to implement these plans, hoping for a progress.",
    "261634": "Hi See--,\n\nI tried sklearn's OneClassSVM and Isolation Forest, but so far, they aren't contributing to a higher score when I apply them :( This is such an important problem to tackle, but I'm stuck :(",
    "265384": "Did anyone had success in implementing this article? Any method for unknown unknown that had impact on the LB scores?",
    "265457": "I implemented the first algorithm described in the paper, tried a couple different thresholds but all of my attempts actually lowered my LB score\n\nThe second algo in that paper (the one they recommend) might hold more promise, but when the first algo didn't yield any improvements I abandoned the idea\n\nFrustrating, bc I feel like theoretically the ideas in the paper make a whole lot of sense for this problem, and thank you to @See-- for posting it\n\nAnybody else try and have more success?",
    "265738": "Can you tell us if you use some other approach? Does the \"unknown unknown\" are really such a big problem that require special care? (I see you reach LB score of  0.9 and I'm trying to get some tips ;) )",
    "265786": "So far everything I've tried to account for the unknown unknowns has failed. My best models so far use a regular 12-class final layer with a random sampling of the \"known unknowns\" for the unknown label",
    "266262": "I see, that sounds great, to pick from known unknowns randomly. Thank you, Thomas :) I'll try this!",
    "266726": "Thomas, I was able to get plus 1% on LeaderBoard with the method. Thank you for sharing ^_^",
    "266741": "Indeed, Thomas, I forgot to thanks you for sharing!",
    "267362": "Since unknown can be seen as any speech not in the 10 speech containing classes, I augmented some unknown data by slicing and re-combining the wavs of different unknowns in training set. \n\nHelped in making the unknown-detection in test set better.",
    "3184291": "Hi Dieter (@kagglingdieter),\nI’m a med school student passionate about AI and neural network development. I’ve learned some concepts but struggle to apply them in competitions like those on Kaggle. \nCongrats on topping the leaderboard for the BYU - Locating Bacterial Flagellar Motors 2025 competition! Your work is inspiring. I have two questions:  \nHow do you approach building top-performing models for such complex tasks?  \n\nHow did you learn these skills, and could you share guidance for someone like me looking to improve?\nThank you for your time!  \nWarm regards,\nSanto Paul"
  },
  "source": "meta"
}