Systems and methods for identifying documents based on citation history
Abstract
Systems, methods, and computer-executable instructions for identifying a document are described. A method includes receiving a query from a graphical user interface having one or more concepts, normalizing a set of terms or concepts in the query to create a normalized query, comparing the normalized query to a set of document centric concept profiles associated with a set of documents in a corpus, where each document centric concept includes a plurality of concepts and at least one reference value for each concept, where the reference value is calculated by tabulating the number of times a document associated with one of the document centric concept profiles is cited by a citing instance for the concept, and surfacing a document from the corpus with the highest reference value for the concept.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A system to generate document centric concept profiles for cited documents for use by search engines in assessing a corpus of documents based on reasons for citation, the system comprising:
a processing device; and a non-transitory processor-readable storage medium including one or more programming instructions that, when executed, cause the processing device to:
extract each citing instance from each citing document of a corpus of documents, wherein extracting each citing instance includes extracting citing text for each citing instance within each citing document, each citing text including at least one of a portion of text before the respective citing instance or a portion of text after the respective citing instance, wherein each citing text is indicative of a reason for citing a cited document of the corpus of documents;
identify one or more than one key concept from each citing text by comparing each citing text to a key concept list;
generate a document centric concept profile for each cited document by:
mapping each key concept identified from one or more than one citing document to each cited document; and
calculating a reference value for each mapped key concept; and
store each document centric concept profile in association with its respective cited document for use by a search engine in assessing the corpus of documents based on reasons for citation.
22 . The system of claim 21 , wherein extracting each citing instance further includes extracting an identifier associated with the cited document for each respective citing instance.
23 . The system of claim 21 , wherein the one or more programming instructions, when executed, further cause the processing device to store each generated document centric concept profile as metadata internally within its respective cited document or externally in a database.
24 . The system of claim 21 , wherein the one or more programming instructions, when executed, further cause the processing device to:
generate the key concept list by mining the corpus of documents to determine:
one or more than one concept associated with a definition within a standard resource; or
one or more than one concept having statistical significance within the corpus of documents.
25 . The system of claim 21 , wherein the key concept list comprises one or more than one legal concept.
26 . The system of claim 21 , wherein calculating the reference value for each mapped key concept includes counting a number of times the cited document has been cited for each respective mapped key concept.
27 . The system of claim 21 , wherein the one or more program instructions, when executed, further cause the processing device to:
adjust one or more than one calculated reference value by:
determining a number of key concepts mapped to each cited document, and decreasing the one or more than one calculated reference value if its respective cited document has been cited for a relatively higher first number of key concepts, or increasing the one or more than one calculated reference value if its respective cited document has been cited for a relatively lower second number of key concepts;
determining, for each cited document, whether any mapped key concept has been identified from a seminal citing document, and increasing the one or more than one calculated reference value if its corresponding mapped key concept has been identified from the seminal citing document;
determining, for each cited document, whether any mapped key concept has been identified via a number of different citing instances within a particular citing document, and increasing the one or more than one calculated reference value, based on the number of different citing instances, if its corresponding mapped key concept has been identified via the number of different citing instances within that particular citing document; or
determining, for each cited document, whether any mapped key concept has been identified via a number of different citing instances within a plurality of citing documents, and increasing the one or more than one calculated reference value, based on the number of different citing instances, if its corresponding mapped key concept has been identified via the number of different citing instances within the plurality of citing documents.
28 . A method to generate document centric concept profiles for cited documents for use by search engines in assessing a corpus of documents based on reasons for citation, the method comprising:
extracting, by a processing device, each citing instance from each citing document of a corpus of documents, wherein extracting each citing instance includes extracting citing text for each citing instance within each citing document, each citing text including at least one of a portion of text before the respective citing instance or a portion of text after the respective citing instance, wherein each citing text is indicative of a reason for citing a cited document of the corpus of documents; identifying, by the processing device, one or more than one key concept from each citing text by comparing each citing text to a key concept list; generating, by the processing device, a document centric concept profile for each cited document by:
mapping each key concept identified from one or more than one citing document to each cited document; and
calculating a reference value for each mapped key concept; and
storing, by the processing device, each document centric concept profile in association with its respective cited document for use by a search engine in assessing the corpus of documents based on reasons for citation.
29 . The method of claim 28 , wherein extracting each citing instance further includes extracting an identifier associated with the cited document for each respective citing instance.
30 . The method of claim 28 , wherein storing each document centric concept profile comprises storing each generated document centric concept profile as metadata internally within its respective cited document or externally in a database.
31 . The method of claim 28 , further comprising:
generating, by the processing device, the key concept list by mining the corpus of documents to determine:
one or more than one concept associated with a definition within a standard resource; or
one or more than one concept having statistical significance within the corpus of documents.
32 . The method of claim 28 , wherein the key concept list comprises one or more than one legal concept.
33 . The method of claim 28 , wherein calculating the reference value for each mapped key concept includes counting a number of times the cited document has been cited for each respective mapped key concept.
34 . The method of claim 28 , further comprising:
adjusting, by the processing device, one or more than one calculated reference value by:
determining a number of key concepts mapped to each cited document, and decreasing the one or more than one calculated reference value if its respective cited document has been cited for a relatively higher first number of key concepts, or increasing the one or more than one calculated reference value if its respective cited document has been cited for a relatively lower second number of key concepts;
determining, for each cited document, whether any mapped key concept has been identified from a seminal citing document, and increasing the one or more than one calculated reference value if its corresponding mapped key concept has been identified from the seminal citing document;
determining, for each cited document, whether any mapped key concept has been identified via a number of different citing instances within a particular citing document, and increasing the one or more than one calculated reference value, based on the number of different citing instances, if its corresponding mapped key concept has been identified via the number of different citing instances within that particular citing document; or
determining, for each cited document, whether any mapped key concept has been identified via a number of different citing instances within a plurality of citing documents, and increasing the one or more than one calculated reference value, based on the number of different citing instances, if its corresponding mapped key concept has been identified via the number of different citing instances within the plurality of citing documents.
35 . A non-transitory computer-readable memory comprising computer-executable instructions for execution by a computer machine to generate document centric concept profiles for cited documents for use by search engines in assessing a corpus of documents based on reasons for citation, the computer-executable instructions, when executed, cause the computer machine to:
extract each citing instance from each citing document of a corpus of documents, wherein extracting each citing instance includes extracting citing text for each citing instance within each citing document, each citing text including at least one of a portion of text before the respective citing instance or a portion of text after the respective citing instance, wherein each citing text is indicative of a reason for citing a cited document of the corpus of documents; identify one or more than one key concept from each citing text by comparing each citing text to a key concept list; generate a document centric concept profile for each cited document by:
mapping each key concept identified from one or more than one citing document to each cited document; and
calculating a reference value for each mapped key concept; and
store each document centric concept profile in association with its respective cited document for use by a search engine in assessing the corpus of documents based on reasons for citation.
36 . The non-transitory computer-readable memory of claim 35 , wherein extracting each citing instance further includes extracting an identifier associated with the cited document for each respective citing instance.
37 . The non-transitory computer-readable memory of claim 35 , wherein storing each document centric concept profile comprises storing each generated document centric concept profile as metadata internally within its respective cited document or externally in a database.
38 . The non-transitory computer-readable memory of claim 35 , wherein the computer-executable instructions, when executed, further cause the computer machine to:
generate the key concepts list by mining the corpus of documents to determine:
one or more than one concept associated with a definition within a standard resource; or
one or more than one concept having statistical significance within the corpus of documents.
39 . The non-transitory computer-readable memory of claim 35 , wherein calculating the reference value for each mapped key concept includes counting a number of times the cited document has been cited for each respective mapped key concept.
40 . The non-transitory computer-readable memory of claim 35 , wherein the computer-executable instructions, when executed, further cause the computer machine to:
adjust one or more than one calculated reference value by:
determining a number of key concepts mapped to each cited document, and decreasing the one or more than one calculated reference value if its respective cited document has been cited for a relatively higher first number of key concepts, or increasing the one or more than one calculated reference value if its respective cited document has been cited for a relatively lower second number of key concepts;
determining, for each cited document, whether any mapped key concept has been identified from a seminal citing document, and increasing the one or more than one calculated reference value if its corresponding mapped key concept has been identified from the seminal citing document;
determining, for each cited document, whether any mapped key concept has been identified via a number of different citing instances within a particular citing document, and increasing the one or more than one calculated reference value, based on the number of different citing instances, if its corresponding mapped key concept has been identified via the number of different citing instances within that particular citing document; or
determining, for each cited document, whether any mapped key concept has been identified via a number of different citing instances within a plurality of citing documents, and increasing the one or more than one calculated reference value, based on the number of different citing instances, if its corresponding mapped key concept has been identified via the number of different citing instances within the plurality of citing documents.Join the waitlist — get patent alerts
Track US2019310988A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.