{
  "id": 78759,
  "title": "Character level only model",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78759",
  "author_name": "Hamish",
  "post_date": "2019-01-27T16:14:56.078000",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>It turns out you can do surprisingly well (I mean not great, I got ~0.63 but better than you might initially think) using a character only model: <a href=\"https://www.kaggle.com/hamishdickson/char-level-only-lb0-630\">https://www.kaggle.com/hamishdickson/char-level-only-lb0-630</a></p>\n\n<p>My intuition around this is it's learning things like capitalisation and punctuation - which we know is important in how quora decide if something is sincere, but I've noticed a lot of people are inadvertently lowering/stripping in public kernels</p>",
  "messages": [
    {
      "id": 462080,
      "postDate": "2019-01-27T16:14:56.080Z",
      "content": "<p>It turns out you can do surprisingly well (I mean not great, I got ~0.63 but better than you might initially think) using a character only model: <a href=\"https://www.kaggle.com/hamishdickson/char-level-only-lb0-630\">https://www.kaggle.com/hamishdickson/char-level-only-lb0-630</a></p>\n\n<p>My intuition around this is it's learning things like capitalisation and punctuation - which we know is important in how quora decide if something is sincere, but I've noticed a lot of people are inadvertently lowering/stripping in public kernels</p>",
      "rawMarkdown": "It turns out you can do surprisingly well (I mean not great, I got ~0.63 but better than you might initially think) using a character only model: https://www.kaggle.com/hamishdickson/char-level-only-lb0-630\n\nMy intuition around this is it's learning things like capitalisation and punctuation - which we know is important in how quora decide if something is sincere, but I've noticed a lot of people are inadvertently lowering/stripping in public kernels",
      "votes": 3
    },
    {
      "id": 462099,
      "postDate": "2019-01-27T17:05:10.007Z",
      "content": "<p>A great point! I've not entered this competition - so I say this fully intending that it's taken with more than a grain or two of salt - but I have noticed in reading quite a few kernels and blog posts, that sometimes people head straight to the cleaning of data without considering whether or not they're throwing away useful information.</p>",
      "rawMarkdown": "A great point! I've not entered this competition - so I say this fully intending that it's taken with more than a grain or two of salt - but I have noticed in reading quite a few kernels and blog posts, that sometimes people head straight to the cleaning of data without considering whether or not they're throwing away useful information.",
      "votes": 1
    },
    {
      "id": 462380,
      "postDate": "2019-01-28T08:10:54.160Z",
      "content": "<p>I made several attempts with char level model and got similar score (F1 around 0.62-0.63). But I thought it's not good results :) \nHowever I haven't tried to use it with other models yet, may be it can add some value in ensemble.</p>",
      "rawMarkdown": "I made several attempts with char level model and got similar score (F1 around 0.62-0.63). But I thought it's not good results :) \nHowever I haven't tried to use it with other models yet, may be it can add some value in ensemble."
    },
    {
      "id": 462674,
      "postDate": "2019-01-28T17:44:59.330Z",
      "content": "<p>I tried too, characters (a-z 0-9 white-space) -&gt; Embedding -&gt; 1D CNN, got f1 score 0.63 at most :-(\nNeeded 20 - 30 epochs and &gt; 8000 min to get to this score (local hold-out set), so I did not proceed, but it was fun :-)</p>",
      "rawMarkdown": "I tried too, characters (a-z 0-9 white-space) -&gt; Embedding -&gt; 1D CNN, got f1 score 0.63 at most :-(\nNeeded 20 - 30 epochs and &gt; 8000 min to get to this score (local hold-out set), so I did not proceed, but it was fun :-)",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 462686,
          "postDate": "2019-01-28T18:15:18.307Z",
          "content": "<p>Oh interesting you got basically the same results with a CNN as I got with an RNN. The time was a killer for me too :(</p>",
          "rawMarkdown": "Oh interesting you got basically the same results with a CNN as I got with an RNN. The time was a killer for me too :("
        },
        {
          "id": 463584,
          "postDate": "2019-01-30T08:37:01.973Z",
          "content": "<p>If you're interested I managed to get this down to 10 epochs and ~2800s .. which is probably still way too slow to use, but at least I got to the bad result quicker 🙃</p>",
          "rawMarkdown": "If you're interested I managed to get this down to 10 epochs and ~2800s .. which is probably still way too slow to use, but at least I got to the bad result quicker 🙃"
        },
        {
          "id": 464148,
          "postDate": "2019-01-31T08:58:40.300Z",
          "content": "<p>Interesting, so you get f1 = 0.63 in 2800 sec, this is an improvement, it shows character level has potential.</p>",
          "rawMarkdown": "Interesting, so you get f1 = 0.63 in 2800 sec, this is an improvement, it shows character level has potential.",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 462099,
      "author_name": "ChrisBow",
      "author_url": "",
      "post_date": "2019-01-27T17:05:10.007000",
      "content": "<p>A great point! I've not entered this competition - so I say this fully intending that it's taken with more than a grain or two of salt - but I have noticed in reading quite a few kernels and blog posts, that sometimes people head straight to the cleaning of data without considering whether or not they're throwing away useful information.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 462380,
      "author_name": "antklen",
      "author_url": "",
      "post_date": "2019-01-28T08:10:54.160000",
      "content": "<p>I made several attempts with char level model and got similar score (F1 around 0.62-0.63). But I thought it's not good results :) \nHowever I haven't tried to use it with other models yet, may be it can add some value in ensemble.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 462674,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-28T17:44:59.330000",
      "content": "<p>I tried too, characters (a-z 0-9 white-space) -&gt; Embedding -&gt; 1D CNN, got f1 score 0.63 at most :-(\nNeeded 20 - 30 epochs and &gt; 8000 min to get to this score (local hold-out set), so I did not proceed, but it was fun :-)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 462686,
          "author_name": "Hamish",
          "author_url": "",
          "post_date": "2019-01-28T18:15:18.307000",
          "content": "<p>Oh interesting you got basically the same results with a CNN as I got with an RNN. The time was a killer for me too :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 463584,
          "author_name": "Hamish",
          "author_url": "",
          "post_date": "2019-01-30T08:37:01.973000",
          "content": "<p>If you're interested I managed to get this down to 10 epochs and ~2800s .. which is probably still way too slow to use, but at least I got to the bad result quicker 🙃</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 464148,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-31T08:58:40.300000",
          "content": "<p>Interesting, so you get f1 = 0.63 in 2800 sec, this is an improvement, it shows character level has potential.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "462080": "It turns out you can do surprisingly well (I mean not great, I got ~0.63 but better than you might initially think) using a character only model: https://www.kaggle.com/hamishdickson/char-level-only-lb0-630\n\nMy intuition around this is it's learning things like capitalisation and punctuation - which we know is important in how quora decide if something is sincere, but I've noticed a lot of people are inadvertently lowering/stripping in public kernels",
    "462099": "A great point! I've not entered this competition - so I say this fully intending that it's taken with more than a grain or two of salt - but I have noticed in reading quite a few kernels and blog posts, that sometimes people head straight to the cleaning of data without considering whether or not they're throwing away useful information.",
    "462380": "I made several attempts with char level model and got similar score (F1 around 0.62-0.63). But I thought it's not good results :) \nHowever I haven't tried to use it with other models yet, may be it can add some value in ensemble.",
    "462674": "I tried too, characters (a-z 0-9 white-space) -&gt; Embedding -&gt; 1D CNN, got f1 score 0.63 at most :-(\nNeeded 20 - 30 epochs and &gt; 8000 min to get to this score (local hold-out set), so I did not proceed, but it was fun :-)"
  }
}