{
  "id": 519468,
  "title": "Sharing my mistakes and learnings",
  "url": "/competitions/uspto-explainable-ai/discussion/519468",
  "author_name": "",
  "post_date": "2024-07-11T09:01:35.539893300Z",
  "votes": 20,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In the first month of the competition I had more than 50 failed submissions (still more than 45 failed after rescoring), so just sharing what I know since then.</p>\n<h6>About the datasets</h6>\n<ul>\n<li>All publication numbers have metadata, but not all publication numbers have patent data. Do be careful of this if you are fetching the raw data in your submission notebook. Here is a notebook finding such patents: <a href=\"https://www.kaggle.com/code/dilliontan/uspto-metadata-raw-data-check\" target=\"_blank\">https://www.kaggle.com/code/dilliontan/uspto-metadata-raw-data-check</a></li>\n<li>There are duplicate rows in nearest_neighbors.csv, you may want to sample with this in mind (for your validation index)</li>\n<li>The earliest neighbor patent for target patents &gt;= 1975 is in 1836</li>\n</ul>\n<h6>About train index, validation index</h6>\n<ul>\n<li>2500-3000 rows of nearest_neighbors.csv is sufficient for the validation index and can be created with the standard Kaggle kernel. For 4000 rows it will exceed the file storage limit during index creation but we can get around that by saving to a temp directory and then copying the final index to the output directory. Beyond that it will exceed the available RAM during index creation. But index size above 3000 is actually not useful in my opinion because it makes the term weights too conservative as compared to the test set size.</li>\n<li>I didn't find the train data useful but here's a notebook collecting all train data, also for Polars testing purpose: <a href=\"https://www.kaggle.com/dilliontan/uspto-train-data-loading-and-verification\" target=\"_blank\">https://www.kaggle.com/dilliontan/uspto-train-data-loading-and-verification</a></li>\n</ul>\n<h6>About specific data fields and operators</h6>\n<ul>\n<li>Wildcard works even for cpc (e.g. you can query cpc:G01N33/*), and can be applied to every query while still remaining under submission qualifying time. But in my case I found it was detrimental to the score as the precision loss is greater than the recall gain.</li>\n<li>Default operator for the Whoosh index is AND, so conjunctions can save on 1 token</li>\n<li>Default parser for the Whoosh index is multifield parser for title, abstract, claims, description and cpc, which means query terms without field specifier will search across all fields, but which also means multi token terms without field specifier will fail because cpc is keyword</li>\n</ul>\n<h6>About performance</h6>\n<ul>\n<li>In my current solution I use json for flat dictionaries and Polars + parquet for everything else. Pickle is too memory intensive.</li>\n<li>For Polars:<ul>\n<li>Always load as LazyFrame and select columns unless the parquet file is small. </li>\n<li>Avoid reshaping large dataframes or columns with list values especially with operations like melt and explode</li>\n<li>Avoid Object datatype for dataframe values, especially nested or dictionaries with variable keys</li>\n<li>For row selection, prefer joins instead of filter</li>\n<li>Avoid row based iteration, especially with .rows() (.iter_rows() is much more performant although still not ideal)</li>\n<li>We can batch prefetch data needed for the test publication numbers</li></ul></li>\n<li>For csv:<ul>\n<li>Loading directly using csv.reader and iterate is sufficiently performant</li>\n<li>Do not read_csv into Polars and then iterate</li></ul></li>\n<li>General:<ul>\n<li>Do not gc.collect() in row-based iteration, it will exceed the submission scoring time</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "2916801",
      "postDate": "07/11/2024 09:01:35",
      "content": "<p>In the first month of the competition I had more than 50 failed submissions (still more than 45 failed after rescoring), so just sharing what I know since then.</p>\n<h6>About the datasets</h6>\n<ul>\n<li>All publication numbers have metadata, but not all publication numbers have patent data. Do be careful of this if you are fetching the raw data in your submission notebook. Here is a notebook finding such patents: <a href=\"https://www.kaggle.com/code/dilliontan/uspto-metadata-raw-data-check\" target=\"_blank\">https://www.kaggle.com/code/dilliontan/uspto-metadata-raw-data-check</a></li>\n<li>There are duplicate rows in nearest_neighbors.csv, you may want to sample with this in mind (for your validation index)</li>\n<li>The earliest neighbor patent for target patents &gt;= 1975 is in 1836</li>\n</ul>\n<h6>About train index, validation index</h6>\n<ul>\n<li>2500-3000 rows of nearest_neighbors.csv is sufficient for the validation index and can be created with the standard Kaggle kernel. For 4000 rows it will exceed the file storage limit during index creation but we can get around that by saving to a temp directory and then copying the final index to the output directory. Beyond that it will exceed the available RAM during index creation. But index size above 3000 is actually not useful in my opinion because it makes the term weights too conservative as compared to the test set size.</li>\n<li>I didn't find the train data useful but here's a notebook collecting all train data, also for Polars testing purpose: <a href=\"https://www.kaggle.com/dilliontan/uspto-train-data-loading-and-verification\" target=\"_blank\">https://www.kaggle.com/dilliontan/uspto-train-data-loading-and-verification</a></li>\n</ul>\n<h6>About specific data fields and operators</h6>\n<ul>\n<li>Wildcard works even for cpc (e.g. you can query cpc:G01N33/*), and can be applied to every query while still remaining under submission qualifying time. But in my case I found it was detrimental to the score as the precision loss is greater than the recall gain.</li>\n<li>Default operator for the Whoosh index is AND, so conjunctions can save on 1 token</li>\n<li>Default parser for the Whoosh index is multifield parser for title, abstract, claims, description and cpc, which means query terms without field specifier will search across all fields, but which also means multi token terms without field specifier will fail because cpc is keyword</li>\n</ul>\n<h6>About performance</h6>\n<ul>\n<li>In my current solution I use json for flat dictionaries and Polars + parquet for everything else. Pickle is too memory intensive.</li>\n<li>For Polars:<ul>\n<li>Always load as LazyFrame and select columns unless the parquet file is small. </li>\n<li>Avoid reshaping large dataframes or columns with list values especially with operations like melt and explode</li>\n<li>Avoid Object datatype for dataframe values, especially nested or dictionaries with variable keys</li>\n<li>For row selection, prefer joins instead of filter</li>\n<li>Avoid row based iteration, especially with .rows() (.iter_rows() is much more performant although still not ideal)</li>\n<li>We can batch prefetch data needed for the test publication numbers</li></ul></li>\n<li>For csv:<ul>\n<li>Loading directly using csv.reader and iterate is sufficiently performant</li>\n<li>Do not read_csv into Polars and then iterate</li></ul></li>\n<li>General:<ul>\n<li>Do not gc.collect() in row-based iteration, it will exceed the submission scoring time</li></ul></li>\n</ul>",
      "rawMarkdown": "In the first month of the competition I had more than 50 failed submissions (still more than 45 failed after rescoring), so just sharing what I know since then.\n\n###### About the datasets\n- All publication numbers have metadata, but not all publication numbers have patent data. Do be careful of this if you are fetching the raw data in your submission notebook. Here is a notebook finding such patents: https://www.kaggle.com/code/dilliontan/uspto-metadata-raw-data-check\n- There are duplicate rows in nearest_neighbors.csv, you may want to sample with this in mind (for your validation index)\n- The earliest neighbor patent for target patents >= 1975 is in 1836\n\n###### About train index, validation index\n- 2500-3000 rows of nearest_neighbors.csv is sufficient for the validation index and can be created with the standard Kaggle kernel. For 4000 rows it will exceed the file storage limit during index creation but we can get around that by saving to a temp directory and then copying the final index to the output directory. Beyond that it will exceed the available RAM during index creation. But index size above 3000 is actually not useful in my opinion because it makes the term weights too conservative as compared to the test set size.\n- I didn't find the train data useful but here's a notebook collecting all train data, also for Polars testing purpose: https://www.kaggle.com/dilliontan/uspto-train-data-loading-and-verification\n\n###### About specific data fields and operators\n- Wildcard works even for cpc (e.g. you can query cpc:G01N33/*), and can be applied to every query while still remaining under submission qualifying time. But in my case I found it was detrimental to the score as the precision loss is greater than the recall gain.\n- Default operator for the Whoosh index is AND, so conjunctions can save on 1 token\n- Default parser for the Whoosh index is multifield parser for title, abstract, claims, description and cpc, which means query terms without field specifier will search across all fields, but which also means multi token terms without field specifier will fail because cpc is keyword\n\n###### About performance\n- In my current solution I use json for flat dictionaries and Polars + parquet for everything else. Pickle is too memory intensive.\n- For Polars:\n  - Always load as LazyFrame and select columns unless the parquet file is small. \n  - Avoid reshaping large dataframes or columns with list values especially with operations like melt and explode\n  - Avoid Object datatype for dataframe values, especially nested or dictionaries with variable keys\n  - For row selection, prefer joins instead of filter\n  - Avoid row based iteration, especially with .rows() (.iter_rows() is much more performant although still not ideal)\n  - We can batch prefetch data needed for the test publication numbers\n- For csv:\n  - Loading directly using csv.reader and iterate is sufficiently performant\n  - Do not read_csv into Polars and then iterate\n- General:\n  - Do not gc.collect() in row-based iteration, it will exceed the submission scoring time",
      "votes": null
    },
    {
      "id": "2935242",
      "postDate": "07/25/2024 05:19:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dilliontan\" target=\"_blank\">@dilliontan</a>,</p>\n<p>Your comment on the query on another post really helped me fix the scoring error—thank you so much. And congrats on your medal! It's a bummer you didn't place higher, though. Now that the competition is over, I wanted to ask you about something I've been curious about regarding Polars. Every time I try to read through the bulk of claims and descriptions, the memory crashed. Can I ask how you dealt with this?</p>",
      "rawMarkdown": "Hi @dilliontan,\n\nYour comment on the query on another post really helped me fix the scoring error—thank you so much. And congrats on your medal! It's a bummer you didn't place higher, though. Now that the competition is over, I wanted to ask you about something I've been curious about regarding Polars. Every time I try to read through the bulk of claims and descriptions, the memory crashed. Can I ask how you dealt with this?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2935242,
      "author_name": "gowillgo",
      "author_url": "",
      "post_date": "07/25/2024 05:19:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dilliontan\" target=\"_blank\">@dilliontan</a>,</p>\n<p>Your comment on the query on another post really helped me fix the scoring error—thank you so much. And congrats on your medal! It's a bummer you didn't place higher, though. Now that the competition is over, I wanted to ask you about something I've been curious about regarding Polars. Every time I try to read through the bulk of claims and descriptions, the memory crashed. Can I ask how you dealt with this?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2916801": "In the first month of the competition I had more than 50 failed submissions (still more than 45 failed after rescoring), so just sharing what I know since then.\n\n###### About the datasets\n- All publication numbers have metadata, but not all publication numbers have patent data. Do be careful of this if you are fetching the raw data in your submission notebook. Here is a notebook finding such patents: https://www.kaggle.com/code/dilliontan/uspto-metadata-raw-data-check\n- There are duplicate rows in nearest_neighbors.csv, you may want to sample with this in mind (for your validation index)\n- The earliest neighbor patent for target patents >= 1975 is in 1836\n\n###### About train index, validation index\n- 2500-3000 rows of nearest_neighbors.csv is sufficient for the validation index and can be created with the standard Kaggle kernel. For 4000 rows it will exceed the file storage limit during index creation but we can get around that by saving to a temp directory and then copying the final index to the output directory. Beyond that it will exceed the available RAM during index creation. But index size above 3000 is actually not useful in my opinion because it makes the term weights too conservative as compared to the test set size.\n- I didn't find the train data useful but here's a notebook collecting all train data, also for Polars testing purpose: https://www.kaggle.com/dilliontan/uspto-train-data-loading-and-verification\n\n###### About specific data fields and operators\n- Wildcard works even for cpc (e.g. you can query cpc:G01N33/*), and can be applied to every query while still remaining under submission qualifying time. But in my case I found it was detrimental to the score as the precision loss is greater than the recall gain.\n- Default operator for the Whoosh index is AND, so conjunctions can save on 1 token\n- Default parser for the Whoosh index is multifield parser for title, abstract, claims, description and cpc, which means query terms without field specifier will search across all fields, but which also means multi token terms without field specifier will fail because cpc is keyword\n\n###### About performance\n- In my current solution I use json for flat dictionaries and Polars + parquet for everything else. Pickle is too memory intensive.\n- For Polars:\n  - Always load as LazyFrame and select columns unless the parquet file is small. \n  - Avoid reshaping large dataframes or columns with list values especially with operations like melt and explode\n  - Avoid Object datatype for dataframe values, especially nested or dictionaries with variable keys\n  - For row selection, prefer joins instead of filter\n  - Avoid row based iteration, especially with .rows() (.iter_rows() is much more performant although still not ideal)\n  - We can batch prefetch data needed for the test publication numbers\n- For csv:\n  - Loading directly using csv.reader and iterate is sufficiently performant\n  - Do not read_csv into Polars and then iterate\n- General:\n  - Do not gc.collect() in row-based iteration, it will exceed the submission scoring time",
    "2935242": "Hi @dilliontan,\n\nYour comment on the query on another post really helped me fix the scoring error—thank you so much. And congrats on your medal! It's a bummer you didn't place higher, though. Now that the competition is over, I wanted to ask you about something I've been curious about regarding Polars. Every time I try to read through the bulk of claims and descriptions, the memory crashed. Can I ask how you dealt with this?"
  },
  "source": "meta"
}