{
  "id": 65172,
  "title": "Notes from my submission",
  "url": "/competitions/avito-demand-prediction/discussion/65172",
  "author_name": "",
  "post_date": "2018-09-07T01:30:46.612214Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>So this is just some notes about my participation, primarily for my own reference, however anyone is welcome to read or give feedback.</p>\n\n<p>Due to hardware limitations, I did not work with the image data at all, but focused on the text of the titles and descriptions. </p>\n\n<p>Most of the time I spent on this challenge involved trying to engineer features, since the previous competition I entered that seemed like where my models struggled. Since I started with a public model that had many count based features, I decided to try to focus on text processing. I did add a few count based features.</p>\n\n<p>For text processing I used the python library nltk to tag words as various parts of speech, the library natively supported both English and Russian, so it seemed like a good fit for the competition. Specific tags didn't seem to be useful, and many of the tags that I thought might be useful were not. Any tagging related to tense or plurality just created overfitting, but the percentage of nouns, adjectives, and verbs all worked decently. </p>\n\n<p>My final model got some lift from the features I created, although after reviewing the successful submissions, it seemed like none of them had a similar approach, so I'm guessing the reason I only saw a small lift was because there simply wasn't enough value in word tagging (for this competition). I'd imagine that with more advanced features derived from my features, that would possibly capture more of the nature of an advertisement, more lift would be achieved. </p>",
  "messages": [
    {
      "id": "382725",
      "postDate": "09/07/2018 01:30:46",
      "content": "<p>So this is just some notes about my participation, primarily for my own reference, however anyone is welcome to read or give feedback.</p>\n\n<p>Due to hardware limitations, I did not work with the image data at all, but focused on the text of the titles and descriptions. </p>\n\n<p>Most of the time I spent on this challenge involved trying to engineer features, since the previous competition I entered that seemed like where my models struggled. Since I started with a public model that had many count based features, I decided to try to focus on text processing. I did add a few count based features.</p>\n\n<p>For text processing I used the python library nltk to tag words as various parts of speech, the library natively supported both English and Russian, so it seemed like a good fit for the competition. Specific tags didn't seem to be useful, and many of the tags that I thought might be useful were not. Any tagging related to tense or plurality just created overfitting, but the percentage of nouns, adjectives, and verbs all worked decently. </p>\n\n<p>My final model got some lift from the features I created, although after reviewing the successful submissions, it seemed like none of them had a similar approach, so I'm guessing the reason I only saw a small lift was because there simply wasn't enough value in word tagging (for this competition). I'd imagine that with more advanced features derived from my features, that would possibly capture more of the nature of an advertisement, more lift would be achieved. </p>",
      "rawMarkdown": "So this is just some notes about my participation, primarily for my own reference, however anyone is welcome to read or give feedback.\n\nDue to hardware limitations, I did not work with the image data at all, but focused on the text of the titles and descriptions. \n\nMost of the time I spent on this challenge involved trying to engineer features, since the previous competition I entered that seemed like where my models struggled. Since I started with a public model that had many count based features, I decided to try to focus on text processing. I did add a few count based features.\n\nFor text processing I used the python library nltk to tag words as various parts of speech, the library natively supported both English and Russian, so it seemed like a good fit for the competition. Specific tags didn't seem to be useful, and many of the tags that I thought might be useful were not. Any tagging related to tense or plurality just created overfitting, but the percentage of nouns, adjectives, and verbs all worked decently. \n\nMy final model got some lift from the features I created, although after reviewing the successful submissions, it seemed like none of them had a similar approach, so I'm guessing the reason I only saw a small lift was because there simply wasn't enough value in word tagging (for this competition). I'd imagine that with more advanced features derived from my features, that would possibly capture more of the nature of an advertisement, more lift would be achieved.",
      "votes": null
    },
    {
      "id": "382786",
      "postDate": "09/07/2018 05:20:26",
      "content": "<p>Good work</p>",
      "rawMarkdown": "Good work",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 382786,
      "author_name": "ebbygorg",
      "author_url": "",
      "post_date": "09/07/2018 05:20:26",
      "content": "<p>Good work</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "382725": "So this is just some notes about my participation, primarily for my own reference, however anyone is welcome to read or give feedback.\n\nDue to hardware limitations, I did not work with the image data at all, but focused on the text of the titles and descriptions. \n\nMost of the time I spent on this challenge involved trying to engineer features, since the previous competition I entered that seemed like where my models struggled. Since I started with a public model that had many count based features, I decided to try to focus on text processing. I did add a few count based features.\n\nFor text processing I used the python library nltk to tag words as various parts of speech, the library natively supported both English and Russian, so it seemed like a good fit for the competition. Specific tags didn't seem to be useful, and many of the tags that I thought might be useful were not. Any tagging related to tense or plurality just created overfitting, but the percentage of nouns, adjectives, and verbs all worked decently. \n\nMy final model got some lift from the features I created, although after reviewing the successful submissions, it seemed like none of them had a similar approach, so I'm guessing the reason I only saw a small lift was because there simply wasn't enough value in word tagging (for this competition). I'd imagine that with more advanced features derived from my features, that would possibly capture more of the nature of an advertisement, more lift would be achieved.",
    "382786": "Good work"
  },
  "source": "meta"
}