{
  "id": 72519,
  "title": "Interesting paper that measures the difficulty of text classification task",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72519",
  "author_name": "Shujian Liu",
  "post_date": "2018-11-24T04:16:18.353000",
  "votes": 35,
  "comment_count": 13,
  "views": 0,
  "content": "<p>There are already some really good EDA kernels but I would like to add some extra information. A few weeks ago, I have read a super interesting paper from CoNLL 2018. The title is: Evolutionary Data Measures: Understanding the Difficulty of Text Classification Task.</p>\n\n<p>In that paper, the authors investigated which characteristics of a dataset best determine the difficulties that dataset is for text classification task.</p>\n\n<p>The result on this competition’s data is in this kernel (only used 2% of data): <a href=\"https://www.kaggle.com/shujian/test-the-difficulty-of-this-classification-tasks\">https://www.kaggle.com/shujian/test-the-difficulty-of-this-classification-tasks</a></p>\n\n<p>“Class Diversity”, “Max. Hellinger Similarity” and “Distinct Words : Total Words” are predicted as “GOOD”.</p>\n\n<p>The overall “Difficulty” is also predicted as “GOOD”.</p>\n\n<p>“Mutual Information” is marked as “SOMEWHAT HIGH”, while “Class Imbalance” is marked as “HIGH”.</p>\n\n<p>The authors provided some suggestions for imbalance:</p>\n\n<blockquote>\n  <p>Class Imbalance can be addressed with data augmentation such as thesaurus based methods (Zhang et al., 2015) or word embedding perturbation (Zhang and Yang, 2018). Under- and oversampling can also be utilised (Chawla et al., 2002) or more data gathered. Another option is transfer learning where knowledge from high data domains can be transferred to those with little data (Jaech et al., 2016).</p>\n</blockquote>\n\n<p>If you have interest, feel free to read more details in their paper or source code.</p>\n\n<p>Reference:</p>\n\n<ul>\n<li>Paper:&nbsp;<a href=\"https://arxiv.org/abs/1811.01910\">https://arxiv.org/abs/1811.01910</a></li>\n<li>Code:&nbsp;<a href=\"https://github.com/Wluper/edm\">https://github.com/Wluper/edm</a></li>\n</ul>",
  "messages": [
    {
      "id": 426884,
      "postDate": "2018-11-24T04:16:18.353Z",
      "content": "<p>There are already some really good EDA kernels but I would like to add some extra information. A few weeks ago, I have read a super interesting paper from CoNLL 2018. The title is: Evolutionary Data Measures: Understanding the Difficulty of Text Classification Task.</p>\n\n<p>In that paper, the authors investigated which characteristics of a dataset best determine the difficulties that dataset is for text classification task.</p>\n\n<p>The result on this competition’s data is in this kernel (only used 2% of data): <a href=\"https://www.kaggle.com/shujian/test-the-difficulty-of-this-classification-tasks\">https://www.kaggle.com/shujian/test-the-difficulty-of-this-classification-tasks</a></p>\n\n<p>“Class Diversity”, “Max. Hellinger Similarity” and “Distinct Words : Total Words” are predicted as “GOOD”.</p>\n\n<p>The overall “Difficulty” is also predicted as “GOOD”.</p>\n\n<p>“Mutual Information” is marked as “SOMEWHAT HIGH”, while “Class Imbalance” is marked as “HIGH”.</p>\n\n<p>The authors provided some suggestions for imbalance:</p>\n\n<blockquote>\n  <p>Class Imbalance can be addressed with data augmentation such as thesaurus based methods (Zhang et al., 2015) or word embedding perturbation (Zhang and Yang, 2018). Under- and oversampling can also be utilised (Chawla et al., 2002) or more data gathered. Another option is transfer learning where knowledge from high data domains can be transferred to those with little data (Jaech et al., 2016).</p>\n</blockquote>\n\n<p>If you have interest, feel free to read more details in their paper or source code.</p>\n\n<p>Reference:</p>\n\n<ul>\n<li>Paper:&nbsp;<a href=\"https://arxiv.org/abs/1811.01910\">https://arxiv.org/abs/1811.01910</a></li>\n<li>Code:&nbsp;<a href=\"https://github.com/Wluper/edm\">https://github.com/Wluper/edm</a></li>\n</ul>",
      "rawMarkdown": "There are already some really good EDA kernels but I would like to add some extra information. A few weeks ago, I have read a super interesting paper from CoNLL 2018. The title is: Evolutionary Data Measures: Understanding the Difficulty of Text Classification Task.\n\nIn that paper, the authors investigated which characteristics of a dataset best determine the difficulties that dataset is for text classification task.\n\nThe result on this competition’s data is in this kernel (only used 2% of data): https://www.kaggle.com/shujian/test-the-difficulty-of-this-classification-tasks\n\n“Class Diversity”, “Max. Hellinger Similarity” and “Distinct Words : Total Words” are predicted as “GOOD”.\n\nThe overall “Difficulty” is also predicted as “GOOD”.\n\n“Mutual Information” is marked as “SOMEWHAT HIGH”, while “Class Imbalance” is marked as “HIGH”.\n\nThe authors provided some suggestions for imbalance:\n\n&gt; Class Imbalance can be addressed with data augmentation such as thesaurus based methods (Zhang et al., 2015) or word embedding perturbation (Zhang and Yang, 2018). Under- and oversampling can also be utilised (Chawla et al., 2002) or more data gathered. Another option is transfer learning where knowledge from high data domains can be transferred to those with little data (Jaech et al., 2016).\n\nIf you have interest, feel free to read more details in their paper or source code.\n\nReference:\n\n - Paper:&nbsp;https://arxiv.org/abs/1811.01910\n - Code:&nbsp;https://github.com/Wluper/edm",
      "votes": 35
    },
    {
      "id": 426933,
      "postDate": "2018-11-24T06:56:47.480Z",
      "content": "<p>Good work, data imbalance is really a big problem. I have experimented with 8W per component of normal problems and then trained with false questions. But the effect is not very good, have you done any related operations?</p>",
      "rawMarkdown": "Good work, data imbalance is really a big problem. I have experimented with 8W per component of normal problems and then trained with false questions. But the effect is not very good, have you done any related operations?",
      "votes": 3
    },
    {
      "id": 429897,
      "postDate": "2018-11-29T14:19:38.263Z",
      "content": "<p>Thanks for sharing! Really valuable information.</p>",
      "rawMarkdown": "Thanks for sharing! Really valuable information.",
      "votes": 1
    },
    {
      "id": 427046,
      "postDate": "2018-11-24T12:21:46.387Z",
      "content": "<p>Interesting work! Thank for sharing! It will surely come in handy with the Quora competition.</p>",
      "rawMarkdown": "Interesting work! Thank for sharing! It will surely come in handy with the Quora competition.",
      "votes": 1
    },
    {
      "id": 427003,
      "postDate": "2018-11-24T10:32:29.007Z",
      "content": "<p>That's really interesting, I'll definitely use that with my future work. \nThanks for sharing !</p>",
      "rawMarkdown": "That's really interesting, I'll definitely use that with my future work. \nThanks for sharing !",
      "votes": 2
    },
    {
      "id": 429556,
      "postDate": "2018-11-29T02:32:02.313Z",
      "content": "<p>Has data augmentation with synonym insertion worked well for anyone?</p>",
      "rawMarkdown": "Has data augmentation with synonym insertion worked well for anyone?",
      "replies": [
        {
          "id": 429722,
          "postDate": "2018-11-29T08:47:25.527Z",
          "content": "<p>For me, no.</p>",
          "rawMarkdown": "For me, no.",
          "votes": 1
        }
      ]
    },
    {
      "id": 429241,
      "postDate": "2018-11-28T15:14:47.257Z",
      "content": "<p>Thanks for sharing <a href=\"/shujian\">@shujian</a>, this is super cool !</p>",
      "rawMarkdown": "Thanks for sharing @shujian, this is super cool !"
    },
    {
      "id": 428905,
      "postDate": "2018-11-28T03:49:23.157Z",
      "content": "<p>Thanks, I appreciate being pointed in the direction of papers which in turn leads to a journey of discovery from reading the paper and investigating the more interesting cited work. </p>",
      "rawMarkdown": "Thanks, I appreciate being pointed in the direction of papers which in turn leads to a journey of discovery from reading the paper and investigating the more interesting cited work. "
    },
    {
      "id": 429596,
      "postDate": "2018-11-29T03:55:16.763Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    },
    {
      "id": 430668,
      "postDate": "2018-11-30T20:21:00.510Z",
      "content": "<p>Thanks, this is good.</p>",
      "rawMarkdown": "Thanks, this is good.",
      "votes": 1
    },
    {
      "id": 427092,
      "postDate": "2018-11-24T14:50:41.383Z",
      "content": "<p>thank's for sharing</p>",
      "rawMarkdown": "thank's for sharing",
      "votes": 1
    },
    {
      "id": 429069,
      "postDate": "2018-11-28T09:47:56.960Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 428891,
      "postDate": "2018-11-28T03:27:31.927Z",
      "content": "<p>thank's for sharing</p>",
      "rawMarkdown": "thank's for sharing"
    }
  ],
  "comments": [
    {
      "id": 426933,
      "author_name": "Bai",
      "author_url": "",
      "post_date": "2018-11-24T06:56:47.480000",
      "content": "<p>Good work, data imbalance is really a big problem. I have experimented with 8W per component of normal problems and then trained with false questions. But the effect is not very good, have you done any related operations?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 429897,
      "author_name": "ssh",
      "author_url": "",
      "post_date": "2018-11-29T14:19:38.263000",
      "content": "<p>Thanks for sharing! Really valuable information.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 427046,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2018-11-24T12:21:46.387000",
      "content": "<p>Interesting work! Thank for sharing! It will surely come in handy with the Quora competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 427003,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2018-11-24T10:32:29.007000",
      "content": "<p>That's really interesting, I'll definitely use that with my future work. \nThanks for sharing !</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 429556,
      "author_name": "Matthew Anderson",
      "author_url": "",
      "post_date": "2018-11-29T02:32:02.313000",
      "content": "<p>Has data augmentation with synonym insertion worked well for anyone?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 429722,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-11-29T08:47:25.527000",
          "content": "<p>For me, no.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 429241,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-11-28T15:14:47.257000",
      "content": "<p>Thanks for sharing <a href=\"/shujian\">@shujian</a>, this is super cool !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 428905,
      "author_name": "Joseph Romani",
      "author_url": "",
      "post_date": "2018-11-28T03:49:23.157000",
      "content": "<p>Thanks, I appreciate being pointed in the direction of papers which in turn leads to a journey of discovery from reading the paper and investigating the more interesting cited work. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 429596,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-29T03:55:16.763000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 430668,
      "author_name": "Aizaz Ali",
      "author_url": "",
      "post_date": "2018-11-30T20:21:00.510000",
      "content": "<p>Thanks, this is good.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 427092,
      "author_name": "youness",
      "author_url": "",
      "post_date": "2018-11-24T14:50:41.383000",
      "content": "<p>thank's for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 429069,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-11-28T09:47:56.960000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 428891,
      "author_name": "bbwen",
      "author_url": "",
      "post_date": "2018-11-28T03:27:31.927000",
      "content": "<p>thank's for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "426884": "There are already some really good EDA kernels but I would like to add some extra information. A few weeks ago, I have read a super interesting paper from CoNLL 2018. The title is: Evolutionary Data Measures: Understanding the Difficulty of Text Classification Task.\n\nIn that paper, the authors investigated which characteristics of a dataset best determine the difficulties that dataset is for text classification task.\n\nThe result on this competition’s data is in this kernel (only used 2% of data): https://www.kaggle.com/shujian/test-the-difficulty-of-this-classification-tasks\n\n“Class Diversity”, “Max. Hellinger Similarity” and “Distinct Words : Total Words” are predicted as “GOOD”.\n\nThe overall “Difficulty” is also predicted as “GOOD”.\n\n“Mutual Information” is marked as “SOMEWHAT HIGH”, while “Class Imbalance” is marked as “HIGH”.\n\nThe authors provided some suggestions for imbalance:\n\n&gt; Class Imbalance can be addressed with data augmentation such as thesaurus based methods (Zhang et al., 2015) or word embedding perturbation (Zhang and Yang, 2018). Under- and oversampling can also be utilised (Chawla et al., 2002) or more data gathered. Another option is transfer learning where knowledge from high data domains can be transferred to those with little data (Jaech et al., 2016).\n\nIf you have interest, feel free to read more details in their paper or source code.\n\nReference:\n\n - Paper:&nbsp;https://arxiv.org/abs/1811.01910\n - Code:&nbsp;https://github.com/Wluper/edm",
    "426933": "Good work, data imbalance is really a big problem. I have experimented with 8W per component of normal problems and then trained with false questions. But the effect is not very good, have you done any related operations?",
    "429897": "Thanks for sharing! Really valuable information.",
    "427046": "Interesting work! Thank for sharing! It will surely come in handy with the Quora competition.",
    "427003": "That's really interesting, I'll definitely use that with my future work. \nThanks for sharing !",
    "429556": "Has data augmentation with synonym insertion worked well for anyone?",
    "429241": "Thanks for sharing @shujian, this is super cool !",
    "428905": "Thanks, I appreciate being pointed in the direction of papers which in turn leads to a journey of discovery from reading the paper and investigating the more interesting cited work. ",
    "429596": "",
    "430668": "Thanks, this is good.",
    "427092": "thank's for sharing",
    "429069": "Thanks for sharing.",
    "428891": "thank's for sharing"
  }
}